Guardians of the AI Galaxy: Provenance, Pedigree and Zero Trust | RSAC Virtual 2024
In an era where AI feels like a tidal wave poised to revolutionize every corner of the tech landscape, it’s easy to feel like we need to reinvent the wheel within our organizations. But what if the key to embracing AI isn’t a sweeping overhaul but rather the foundations we already have? This session invites you to a journey that demystifies integrating AI into your organization by establishing the provenance and pedigree of your AI systems and applying zero-trust principles to their authentication and authorization.
Whether you’re engineering solutions, crafting code as a developer or architecting robust systems, this session speaks your language. It’s designed for software architects, engineers and technologists—those on the front lines of implementing and securing AI systems.
Transcript
Three people log into a corporate video chat ready for their daily team meeting. The first a compliance manager looks concerned. The second a data scientist, of course they're juggling three laptops.
And the third, a project manager is just trying to figure out how to turn off that filter that makes 'em look like a potato. The compliance manager says, Hey, I have a serious question. Can anyone prove to me that a particular piece of data was it used in the a application that you release?
You recently deployed the data scientist without missing a beat. Replies, that's like asking if I can prove that I didn't use my neighbor's wifi for our last internet call. Technically possible, but you're not gonna like digging through the logs, the project manager still looking like a potato chimes in.
And here I was thinking the hardest thing to prove was that I'm actually not a vegetable. Everyone laughs and the meeting moves on. But deep down, they all know the compliance manager is about to make them audit every piece of data from last year.
I jokes aside, we can feel the uncomfortable feeling all these people and the potential pain they're about to experience in a vast majority of companies. Right now, this is or will be the big question with your AI solution. So let's face eight AI aside.
Some organizations may not be able to answer this question to begin with. Besides taking Valium before conversations like this, how can we cope with this question? How can we turn this from Yes, but here are all the blockers to that question into yes.
And here's how we can verify for you right now. Answering this question was part of a hackathon. Some folks now are involved in this previous year.
The purpose of this hackathon was to operational a operationalize AI for the enterprise. My focus was the governance and security aspects. There were some other brilliant minds involved in this whole hackathon, although other, there were about three specific people I need to highlight who also helped shape these ideas and thoughts we'll cover today.
First, Steve McGill and I, we focused on security and safety of LLMs. Second Alexander and I focused on control, provenance, and pedigree. And since the hackathon, Joseph Enoch and I have been continuing to expand upon the idea, uh, definitely gotta thank John Wills and the text Textron team for the opportunity to do this.
And there's a really cool video of this whole hackathon on Textron's, uh, Textron's website. Um, along with this presentation, quick intro to myself. Um, I'm Bill and I focus on binding the gap between governance and technology.
Uh, most of my focus is showing people how to use technology to solve old world governance problems so they can do more with what they have. The five main areas I create solutions for and research are governance, engineering, autonomous assurance, cyber safety versus cybersecurity. And where does AI fit and where does it not fit.
Um, I argue that using your data to either train or be consumed by machine learning algorithm or existing AI product is really no different than what most people probably should be doing today. So I wanna walk through a bit of this argument in three steps. First, today I'm gonna describe the problem and why it's a problem.
Second, I'm gonna talk about the first principles to address this problem. And third, we'll talk about an architecture you can build and imp you build INTA implementation for, uh, that will encompass these principles and address this problem. The general problem we face is assurance.
Just proving that we're doing what we say we're supposed to be doing or proving assurance is proving that we're doing what we said we would do. A lot of what we would do is in there. Many see this as proving to some external or internal third party that we're compliant.
I think, and most importantly, we should view assurance as proving to ourselves that we're really not fooling or lying to ourselves about what we're doing. The most harmful thing to do is really to believe the lie you tell yourself. So there are three questions to ask ourselves to prevent us from believing some erroneous in our voices.
The first question, can we use this data for something? This is a question of control. Control, as many of you're familiar with is about ensuring an action, uh, can or cannot be taken.
When it comes to our AI and data, we need to have some clear understanding of two things. First, the sum and the sum thing. We need to know who is taking the action, the someone.
And we need to understand what the context of the action is. The sum thing, an example of control would be an answer to the question. Can the call center representative use the response, uh, to a customer from our internal LLM?
Questions of control are binary. They're either yes or they're either no. Control becomes key with ai, given.
AI allows for significant degrees of efficiency for information exchange as compared to the traditional ways of having somebody just simply email search for or interpret other information. So in essence, as we all know with AI and automation in general, we can efficiently do more bad things. Now, this situation of efficiently doing bad or wrong things is what controls attempt to mitigate.
The second question is, what's the question of origin for this data? This is a question of provenance. Provenance is commonly misconstrued with its gross relation to pedigree.
Now, provenance is the origin of something. Knowing the origin tends to be the first question to assess the quality of an item. So for example, imagine an analyst is pulling financial data from a shared company database for a report.
This database, this database aggregates data from many sources, although there is no tracking from where the systems that his data elements were pulled from. Therefore, there's no provenance for each data source or data element. So without a clear trail of breadcrumbs, it's challenging to verify the data's authenticity or accuracy.
And now our third and final question is, is this the right type of data to use? Now this is a question of pedigree. Data pedigree is not widely discussed as far as I'm aware, although it may be the root cause for many data and information issues most organizations faced pedigree is the ancestry or the record of purity for some piece of data.
Pedigree is important because data information are changed and mutated over time. Pedigree helps us assess data quality to ensure that the changes within the data lineage are the expected type and or good changes. So for example, imagine that same employee extracting some customer data from their CRM system.
This CRM is considered the system of record for customer information. That's the data's origin. It's provenance and the path that traveled through various systems is well documented.
However, multiple departments have edited the customer information over time without clarity of what's changed, why, and by whom. This is the pedigree. And now it becomes suspect.
Despite knowing where the data come from, inconsistencies and alterations could lead to the question about its accuracy and or reliability. So being able to answer these three questions, the questions of control, provenance, and pedigree are what we're going to are is how we're going to address the compliance question of prove to me this data is used or not used in your AI application. So now that we know the problem we're attempting to solve, let's cover two guiding principles that will help us handle a myriad of questions that arise during a solution implementation for the problems of control, provenance, and pedigree.
The first principle is cyber safety, security and cyber safety. Security specifically cybersecurity is talked about a lot and rightfully so. Uh, so the word security is used as broad to cover many terms to protect someone something, uh, from energy, um, from inj injury based cyber system.
Um, I'd argue we need to separate the overall security idea into two specific terms, cyber safety and cybersecurity. So what are the differences between these two ideas? Well, let's first talk about cybersecurity securities, we all know is protecting your cyber systems and the folks using them from a malicious actor with the 100% secure cyber system, and I know there's no such thing as 100% secure, no malicious actor can make the system cause injury to the user.
The focus on security is about bodyguarding against either known, which is finite or unknown, infinite threats. Now, cyber safety safety is not about protecting your system from a malicious actor. Safety is about ensuring that the users of your cyber systems are not injured while using the system just under normal operating conditions.
Now, safety is about having a clear path for the user to follow when there is an issue that could cause injury, uh, that would happen during normal operations. Safety is about having succinct and clear guidance for how you should use something during those operating conditions. Outta curiosity, do you recall the Japanese airline incident early this year?
Uh, that threw Japanese air flight, uh, JL 5 1 6 into a crash. The passengers of this airplane collided with another small airplane on the runway. All 367 passengers and 12 crew members escaped before the plane was completely ablaze.
Now a plane crash is a probable non-normal operating condition. The procedures and protocols in place by the Japanese airline where there's safety mechanisms that reduce the risk of bodily injury. This was not a security incident or a threat, this was a safety incident.
Now, similar incidents like this occur with our technology systems, or they catch fire when some non-malicious thing bumps into them. Now, what procedures and protocols do you have in place when safety situations arise in your tech systems that are not security based? Cyber safety should be a first class consideration for system design and operation, just like the main feature of any of your systems or any other types of behaviors you expect from your system.
So hopefully that difference is clear. Now, the second principle I wanna cover is what I refer to as the trust map. It's inspired by knowledge mapping, graph mathematics, and it's implied to the underlying data information and other inputs into an AI solution.
This principle is important because it answers two types of foundational questions that make up any type of governance or compliance, inquiry depth and breadth based questions. So depth questions, depth answers to depth based questions provide insights into how a specific piece of information from the bottom up or top down is used. So example, depth based questions are, can you tell me if your model was trained on some specific piece of data?
Can you tell me if your model is retrieving this specific piece of data? Or show me how you removed this piece of data from your model or the data stores. Now on the other side, we have breath based questions.
Broad answers provide insights that go into the scope or the total usage of a piece of data. So imagine, so examples of breath based questions are, um, of all your models, which ones use llama as a base model, uh, of all your models, which ones have been trained on this specific piece of data? And prove to me that this piece of data is not part of model A, B, C, or D, but is part of model E, F, and G.
Now let's go ahead and sum up these two guiding principles real quickly. Cyber safety helps you scope what you control for regarding normal and probable non-normal operations. The trust map allows you to verify composition of your models such that you can validate and verify your meeting of governance or compliance requirement from top to bottom, but also side to side.
Now that we've covered the problems and principles to address these problems, let's talk about an architecture. And an architecture is simply just a set of constraints for a solution implementation that it must abide by. Um, recently Joseph and I have called this the neural gatekeeper architecture.
Now, quickly before we dive into the architecture, let's recap with some common problem and common use cases as to to help understand where and how the this solves the problem. So for example, your AI powered app could have a component such as the model that is trained on restricted data. Now, what types of behaviors should you expect that user receives if they are or are not allowed to ask a question or make an action that consumes this piece of data?
Now, what happens when your AI powered application contains some type of outta date knowledge or information? You know, this type of behavior we should expect is that the user receives a message as if the knowledge doesn't exist. There may also be cases where you'd like to let the user know the response provided to them is in fact out of date, either which way the behavior you prefer is that there is an expectation that the user does not receive an answer or is restricted to action they can take given the outdated information.
All of this behavior is powered by the trust maps. In the trust map. Each thing within your system has a set of constraints that start with the data source itself.
These constraints are treated as authorizations for every transaction. So if you think about these from the core bottom, what they are is fundamentally authorizations that are composed over time. 'cause not one model's made up one piece of data or one question or one query recall involves one piece of data, one model, it involves a composition.
So using the trust map, you can answer those breadth and depth based questions, especially around control and authorization is, should this person or could this person ask this question and or receive an answer to the question asked based upon the data? Now this is where the gatekeeper comes into play. This component is responsible for four primarily main things.
The ingestion of data, ensuring the appropriate authorization and constraint Metadata is assigned at the time of data ingestion, adding the ingestion data and assets to the trust map and persisting the data into data source designed for retrieval by models or storage for the future of the training. One thing to point here around the gatekeeper is we refer to some things as sort of the golden data stores. It shares similar mo shares similar concepts with things like data lakes.
And the idea is your external data source and data and data elements are put in. But the gatekeeper does not pull from this specific source of origin. It pulls from its gold data stores because that's where everything is managed and that's where the maps and the trust map is kept.
So there's a couple of use cases that we can talk through on this, on how this works. First, let's talk through on consuming assets or data assets. Uh, this in fact is the defacto use case, which ensures that the consumption of any asset, whether it is a model, a piece of data or sort of how the model interacts with a piece of data or something like a uh, a vector data store, um, meets the trust map.
Uh, the consumption can be done by any actor, whether it's human and or machine. So very common primary use case. The second one is adding assets.
Assets. So the idea behind this is the gatekeeper isn't just about moving your data from one spot to another. The gatekeeper is about once you bring the data in, it has a role of ensuring that it's somewhere that's stored that can be retrieved, and it's in an area that is, um, that may be, uh, that may be itself controls.
So you can have smaller golden data stores with different constraints around the data stores themselves that represent the aggregate or the composition of the, the, the constraints of the, the actual individual elements of the data. But the idea behind the gatekeeper is to be the orchestrator of this full process when you say, Hey, take this piece of data, give it constraint one, two, or three. So these types of people can see it, these types of people can't, whatever it may be.
And put it into the golden data store, such that when we go to do things like either a query through a, an existing model or an application or as we'll talk through the next, in the next, uh, use case where we actually wanna build a model of this data, um, we can apply these types of controls that, uh, that, that, that, that relate to provenance pedigree. And of course, uh, so, uh, governance controls. Now the third use case here, let's think about an ML lops pipeline allowing something like an ML lops pipeline to request the data given its constraints.
So the idea here is that you're not pulling from the original source of the data. You're usually the gatekeeper and saying, Hey, provide me this list of a thousand pieces of data elements. And oh, by the way, I want to use some other model as a, as an intermediate model if that's the situation here.
Now, the gatekeeper, it does a couple things here. It records what data is sent to the pipeline. It also is applying the authorization controls if that person can have that data.
So if I send a list of a thousand piece of data elements or a thousand PDFs, whatever it may be, and let's say I, I don't have access to 50 of 'em, well, the gatekeeper would tell me, Hey, I'll give you these, uh, 950, but these other 50 I can't give you because based upon whatever it is, you don't have access to it. Now, the beauty behind something like this is that we can go further down on multiple types of dimensions. Imagine you have constraints around usage and or other types of constraints that you can describe that data.
So maybe I do have access to it since I'm an employee, but if I'm gonna be using it to answer an external facing question, I shouldn't be using it in that situation. So as we start to think about the, the multiple paradigms for how data can be used, especially as it gets baked into a model or consumed through something like a vector data store, there is a, there, there's a huge dimensionality to that problem. And the idea here behind the gatekeeper is to be able to structure that dimensionality such that it's simple for a user to request and then be responded back with either what they are or a not to do.
Um, as the ML model builds the other, uh, the other, uh, specific feature and value of this use case is building in providence and pedigree of that model back into the, to the, to the trust map. So inside the trust map, you can have all your pieces of data. Now once this is built into that model, that model can fundamentally be fingerprinted, restored back into your golden data stores.
And then from there, as you use the model, now you have an existing structure. So when somebody asks, Hey, what data is in this? What was used to train this?
Any of those types of questions, you can go straight to the trust map and say, here's the composition of everything that was used to train. Whether it's in, uh, whether it's the just a source data, it's intermediate or even used to test the model. So as you start to think about anything that can come into contact with the model or your ai, your system from a data, that's what the trust map there is to do.
Now the main inspiration, uh, for the narrow gatekeeper is the zero trust architecture, specifically the NIST 802 0 7 0 trust architecture. Now, this zero trust architecture consists of three main logical components. First, it consists of a policy engine.
Now this component's responsible for the ultimate decision to grant a resource, uh, for a given subject. So for example, a piece of data that you may request a resource is a piece of data. The given subject is you a policy administrator.
This component is responsible for establishing and or shutting down the communication path between the subject and the resource. So therefore, if I request a PDF, I don't have access to the, the policy administrator is the control aspect that basically says, Nope, can't send that piece of data out. And then third, the policy enforcement point.
Uh, this system's responsible for enabling monitoring and eventually terminating all connections between the subject, um, and an enterprise resource. And so this policy enforcement point is what sits in between and fundamentally what the, the gatekeeper is. Now the, the, the, the neuro gatekeeper is all of these together, but if you look at it from a big macro perspective, really is biggest value is this policy enforcement point.
So let's look at a highly simplified view of all the three zero trust components. What's some additional sub components? The purpose of these components is to enforce the constraints on ingestion and consumption.
The ingestion constraints are to ensure that the proper metadata is available and stored in relation to the data that it represents In the trust map, uh, the consumption constraints are allowed, uh, allow one to enforce, uh, access based upon the following questions. For example, can the data source be consumed by this model? And then can a user use this type of data source for of, for example, uh, from this model.
Now here's an example of how the neural gatekeeper uses the zero trust principles, both at runtime and from a build time perspective. The goal of this architecture is to provide an answer for a cradle to grave questions such as you have the top-down questions. Is this model compliance?
It's a very broad and detailed question to answer bottoms up questions such as what models have been trained on this specific piece of data depth based questions. You know, was this model trained on a specific piece of data? So as you can see from the bottoms up versus the depth, what models, a scope of models versus what specific one.
And then you have the breadth, you know, what models we're trained on this data on. With that being said, this is what we've come forward with as a bit of a, uh, we'll call it a bit of an open source architecture, although there is not, uh, no, no formal source code yet around it. But once to introduce this out to everybody as we start to think about going forward into organizations, especially around the problem with AI and managing the usage of AI based upon that data.
'cause to a large degree, I would argue it's not the art, the, the AI applications, the model themselves. It's what the models represent that we're concerned about. And so ensuring that representation is consummate or congruent with our expectations of our organization or whatever the expectations may be, is going to be key.
And of course, through assurance and ultimately some forms of compliance, and this is where the governance model comes in, is you need to be able to answer those questions. So today, in the neural gatekeeper, this is the architecture for how we begin to answer those questions. If you have any questions about what you've seen today, always feel free to reach out to me and have some conversations.
And thank you for your time.