AI Integration Without Upheaval with Bill Bensing at Techstrong Con 2024
It’s easy to feel like we need to reinvent the wheel within our organizations to keep up in an era where AI feels like a tidal wave poised to revolutionize every corner of the tech landscape. But what if the key to embracing AI isn’t a sweeping overhaul but the foundations we already have? This session invites you to a journey that demystifies integrating AI into your organization without turning your operational world upside down.
Whether you’re steering the ship as a CEO, crafting code as a developer, or architecting solutions, this session speaks your language. It’s designed for a broad audience – from the guardians of governance and management mavens to the visionaries of the C-suite, Digital CxO leaders, and IT all-rounders.
We’ll explore why embracing AI doesn’t mean bidding farewell to your current operation strategies. Instead, we’ll dive into how the pillars of good data governance—think data provenance, pedigree, and the classic duo of authentication and authorization—are surprisingly your best allies in welcoming AI into your realm.
Transcript
All right, let's do something real quick. Let's start off with a bit of a joke. Three people log into a corporate video chat ready for the daily team meeting.
The first is a compliance manager. They look concerned. The second is a data scientist and they're juggling three laptops.
And the third is a project manager and they're just trying to figure out how to turn off the filter that makes 'em look like a potato. The compliance manager says, Hey, I have a serious question. Can anyone prove that this particular piece of data wasn't used in our AI's training or in our rag implementations, the data scientist, without missing a beat replies, that's like asking if I can prove I didn't use my neighbor's wifi for last internet call.
Technically possible, but you're not gonna like digging through the logs and the project manager stills a potato chimes in. Hey. And here I was thinking that the hardest thing to prove was that I'm actually not a vegetable.
Everyone laughs and the meeting moves on. But deep down, they all know the compliance manager is about to ask them to audit every piece of data from last year. So jokes aside, we can tell and we can feel how comfortable they're probably feeling as well.
And we can feel these pain points. And a vast majority of companies right now, this is probably the big question when it comes to some of your AI solutions and your governance, and let's face it a i aside, some organizations may not be able to answer these questions to begin with for basic data. Besides taking Valium before conversations like this, how could we make ourselves more comfortable with questions like this?
And how could we turn this from a yes? But here are all the blockers into a yes and here's how we can verify you for you right now. Answering these questions was part of a hackathon that some folks and I were involved in this last year.
Now, the purpose of this hackathon was to operationalize AI for the enterprise. Now my focus was the governance and security aspects. There were other brilliant minds involved in this whole hackathon, although there were three people that I needed to shout out, shout out to, and highlight who have helped shape these ideas.
So for the first one is Steve McGill. Him and I on the first day focused on security and safety of LLMs. On the second day, Alex Ner and I focused on control, provenance and pedigree.
And then since then, Joseph Enoch and I have been continuing to expand upon these ideas. I definitely have to thank John Willis and the whole Techstrong team for the opportunity to be part of this. And if you're interested, they do have a video up on their website where you can see the whole three days shoved into about an hour.
So quick introduction to myself. I'm Bill and I focus on binding the gap between governance and technology. Most of my focus is showing people how to use tech to solve old world governance problems so they can do more with what they have.
The five main areas, or really four main areas where I create solutions and research for are a new concept I call governance engineering. And another new one I'm walking on call autonomous assurance. This idea of cyber safety versus cybersecurity.
And also where AI does not fit in an organization. 'cause believe it or not, it probably shouldn't be in a couple places. So I wanna argue that using your data to either train or be consumed by machine learning or existing AI product is really no different than most of what are you doing today.
And I'm gonna walk you through this argument in three steps. First, we're gonna talk a bit about the problem and why it's a problem. Second, we'll talk about some first principles to address this problem.
And then finally, we're gonna go talk about an architecture that you can build an implementation against to encompass these principles to address these problems. Now, the problem we face in general is proving that we're doing what we say we're doing. Ah, that's a hard, that's a big word.
That's a big thing to say. Prove to me that you're doing what you say you're doing. Many see this as proving to some external or other internal party as compliance.
Now, moreover, I like to think of this as, I like to think of another view that you know, it's less proving to somebody else and more proving to yourself that you're doing what you say you're doing, really, because you know it's you lying to yourself. That's really the most harmful lie that you can believe. So as a lot of people talk about this as compliance and proving to somebody else, really as we go through this a day, this is setting up a system so you can prove to yourself that you're being safe with what you're doing.
Now, there are three key questions to ask ourselves to prevent us from believing our erroneous inner voices. The first question, can someone use this data for something? This is a control question.
Now, controls, as many of you're familiar with, are about ensuring an action is or isn't taken. So when it comes to our AI and data, we need to have a clear understanding of two things. First, who's taking the action that's the sub.
And second, what is the context of the action? That something, an example control could be an answer to the question, such as, can the call center representative using a response to the customer from our internal LLM to other types of co control questions are binary. There are either yes or no.
So just like that one, can the call center representative use the response, uh, to a customer From our intel? LM is that yes or no? Control becomes key with AI as it allows for significant degree of efficiency for information exchange as compared to traditional ways of someone having to email search for info and interpret it for somebody else.
In essence, as we all know with ai, we can efficiently do more bad things. And this is the question that a control attempts to mitigate. Now our second question, what's the origin of this data?
This is a question of provenance. Provenance is commonly misconstrued with its crows relative pedigree, but provenance is the origin of something. Knowing the origin tends to be one of the first questions to assess a quality of an item.
So for example, imagine an analyst pulling financial data from a shared company database for a report. This database aggregates data from any sources, although there is no tracking from the systems where they're pulling each data element, therefore there's no providence for the data source. So without a clear trail of breadcrumbs, it's challenging to verify the report's authentic authenticity.
'cause you can't revalidate the actual underlying data's authenticity or accuracy. Now our third question, is this the right type of data to be used? This is a question of pedigree.
Now, data pedigree isn't broadly discussed as far as I'm aware, and not specifically using those terms, although I would argue it's the root cause for many data and information organization information issues that an organization faces. So pedigree is the ancestry or the record of purity for some piece of data. Now, pedigree is important as data or information is important because it shows that data or information has mutated over time.
So for example, imagine an employee extracting customer data from their company's CRM, their customer relationship management system. This CRM is considered a system of record for the customer's information. The data's origin ass provenance and the path that traveled through to the various systems are well documented.
However, multiple departments have edited the customer's information over time without clarity of what has changed, why it's changed and by whom changed it. Therefore, the pedigree of the data now becomes suspect despite knowing where the data comes from. Inconsistencies and alterations could lead to questions about the data's accuracy and reliability.
So being able to answer these three questions are what we wanna answer for our compliance question. Prove to me that this data is approved for use. So now that we know the problem we're attempting to solve, let's cover three guiding principles that give us the ability to determine how to handle others' questions that arise during an implementation of a solution for this problem.
The first principle, I'm gonna call cyber safety. Now, security, cybersecurity and cyber safety. Cybersecurity is talk a lot.
Security and cybersecurity are talk a lot and, and rightfully so. I don't think anybody's really talking about cyber safety just yet. The word security is used as a broad term to cover many things to protect someone from, um, to protect their system from or from injury.
I'd argue we need to separate the overall security idea into two specific terms. And let's talk about the idea of these terms, which is cyber safety and cybersecurity. So cybersecurity security is about protecting your cyber systems from folks that are using your cyber systems from other malicious actors.
Does that make sense? A malicious actor is going to take advantage of your system and cause injury or harm to somebody else. Now, although we know security can't be 100%, if there was a 100% secure cyber system, no malicious app actor could use that system to cause injury to another user.
The focus on security is really about body guarding against either known, so like finite or unknown infinite threats. So it's a bodyguard against threats. Now let's talk about cyber safety.
On the other hand, safety is not about protecting your system from a malicious actor. Safety is ensuring that all users of your cyber systems are not injured while the system is under normal or non-normal operating conditions. Safety is about having a clear path for a user to follow when there are either normal or non-normal errors happening.
And it's about having a succinct and clear guidance for how you should act during a probable non-normal operating condition. So let's talk about a recent example. Um, maybe you do or don't recall, but earlier this year in 2024, the Japanese airline, uh, had a, there was an accident.
There was actually a crash. It was a GAL 516. Um, this is a passeng airline and it was on the runway and come collided with a smaller aircraft.
What was really cool and interesting was all 367 passengers and all 12 crew members escaped from the plane before it completely caught a blaze no injuries. A plane crash is the example of a highly probable non-normal operating condition. It's not normal for planes to crash, but it is probable it has happened and it will, there's a probability that will continue to happen.
Now, the procedures and protocols in place by the Japanese airline are their safety mechanisms to reduce bodily injury. So this was not a security incident, this was a safety incident. So cyber safety should be a first class consideration for system design and operation, just like the main features of the system.
So hopefully this is clear and I'll ask you that. Is it clear to you the difference between cyber safety and cybersecurity now to make it easier to understand your machine learning and your AI and, and things are happening in training? Um, I wanna compare it to teaching children.
I, I think this is key, and I've used this analogy a lot 'cause a lot of people, um, they, they look at AI and machine learning as this, this very complicated and, uh, extremely complex. And while a lot of it is, it could be, it could be boiled down to sort of teaching a child. So when you teach a child, you give them data and information and you reinforce it and or they reinforce it over and over across time.
That's how they learn. Now the downside is when this information that you're trying to reinforce becomes obsolete or is now wrong for some reason. So in general, it can be hard to consciously forget something.
But what's interesting and different with AI is with ai you can actually forget things. Now, effective AI needs information and data that is current, accurate and complete. So ensuring information is accurate and complete is done through provenance and pedigree.
Now, ensuring information is current, especially on an ongoing basis, can be a problem. Now I've go ahead and borrowed the term repression to describe this principle. Now, repression isn't specifically a technology term, it's a psychological term.
So in psychology, repression is the mechanism that blocks a memory. Generally repression is negatively associated with memories in psychology, although that negative connotation, um, is the same. When I talk about the relation to AI and machine learning, it's when, say it's not the same, it's there's really no negative connotation of repression in ai.
Uh, but at as time and context changes, the information provided by AI will most likely need to change or may need to change. Now, repression is the ability to remove information or data that is not current, it's obsolete or has become erroneous. So with AI and some of the complicated tech around it, many people forget that unlike a human with ai, you can systematically repress data and information.
So think about how many times something has changed at a company you're at and you were probably told to stop doing something or to do it a different way and not just you, your team and teams around you. I can almost guarantee there was at least one email an announcement made or somebody held some type of change management seminar. Um, and it did not just happen once.
It probably happened more than once. Now ask yourself, how long did it take for people to stop or change? I'm betting it took a while.
That's assuming if they ever changed at all, they did not repress that information. So while the efficiency of bad responses from AI is significantly higher than it is with a real human, so it's easier to create bad responses with AI quickly as it is with a real human over time, you can effectively repress information from AI to ensure that there's no erroneous information in there such that you stop making it quicker. Now, let's get to our third and last principle, but this is one I refer to as the map.
It is inspired by knowledge mapping and graph mathematics. It applied to the underlying data and information and other inputs of an AI solution. Thus, principle's important because it provides the foundation for answers to two types of questions that make up the sa They make up any type of governance and compliance question.
These two types of questions are depth and breadth based questions. So let's dig in each of these deep questions, deep questions answer, uh, they they answer our provide insights into how a specific piece of information is used at the bottom. So almost like a bit of a bottoms up things like can you tell me how your model and can you tell me if your model was trained on this specific piece of data?
Or can you tell me if your model is retrieving or your AI solution is retrieving this specific piece of data? Or show me that you've removed this piece of data from your model or your data stores if it's not current, accurate or complete. Your next ones are broad questions.
These are the breadth questions. A broad question asks where a piece of information is used over a swath. And it could be bottoms up or top down, but over a swath of different items.
So if you think about it, you can ask like, Hey, are all of my models based on llama or are all of my models of all of my models, which ones have been trained on either this specific piece of data or these sets of data? And then prove to me that the data I have is either not part of model A, B, C, or D or it is part of model E, F, and G. So when you start looking and think about governance and compliance questions, they generally boil down to these two types of questions, a depth based question, show me about this one thing from top to bottom or bottom to top, or breath based question, show me about these, this, how these things are used across the swath of a, of a problem or a situation.
So let's go ahead and sum this up right now, this little section. So cyber safety helps you scope what you need to control for regarding normal and probable non-normal operations. So I wanna say that one more time.
Normal and probable non-normal operating conditions. And so in a probable non-normal operating condition. So for example, like our, like our crash, we know a crash is probable, it's not the normal operating condition of the airplane, but what we need to do is we need to ensure there is safety conditions around it such that they had protocols, for example, in the Japanese flight, they had protocols and processes to get people off quickly before anybody could have bodily injury.
Now, repression, repression, ensurers the pedigree of the model or the data by removing our forgetting what it's not either current, accurate or complete. So I think this clear, this is, this is key because we start to think about pedigree and pedigrees around the quality. This is interesting because now we can start to remove things that were there as things change to ensure that we keep quality over time.
And the third one is the map. It allows you to create verification of the composition of your models and or your whole AI system such that you can validate. So once you verify it's in there, you can now validate that you are or are not meeting some level of governance or compliance requirement.
So now that we've covered this problem, these problems and uh, principles to address these problems, let's go ahead and talk about, uh, architecture. Now architecture is not specific implementation details. Um, like specific technologies, I I'll go to my grave arguing that one architecture really the sets of constraints of a solution that we'll use to define how an implementation implementation is achieved.
So Joseph and I, who I talked about earlier, uh, recently, we named this architecture the neural gatekeeper. Now the neural gatekeeper to a large degree is an extension of the NIST 802 0 7 0 trust architecture. There are three main components to the zero trust architecture and we'll go by each one real quickly.
So the first one is the policy engine. This component is responsible for the ultimate decision to grant access to a resource for some type of subject. So what the policy engine does is given context and given the request, it says yes or no, you're allowed to either do or have what you need or not what you need, what you're asking for, you need asking for a different the policy administrator.
Now this is a component that's responsible for establishing or shutting down the communication paths between the subject of resource. So while the policy engine says yes or no, you may or may not have this, the policy administrator looks at what the policy engine determines and then says, nope, can't have it. And, and actually physically is that stop.
So I think that's key as you look at these two separate components is you have one component that determines and the other component that is the actor that creates the action of either allowing to continue the communication and or stopping. Now the last one is the policy enforcement point. And what this does, this the system that's responsible for monitoring or enabling the monitoring and individual termination of the connection between the subject and the enterprise resource.
So everything flows through the policy enforcement point and it communicates with the policy administrator and or the policy and it communicates with policy administrator to determine should this person, this untrusted subject coming possibly from this untrusted source, can they have access to our enterprise resource? So there's a couple other components around there, but those are the most basic components that there are. Now, uh, there's one piece though that's missing from this 800, two a seven, um, that we need for the neural gatekeeper and we, we call it right now the gateway.
Uh, this component is responsible for a couple things. It's responsible for ingesting data, ensuring the appropriate metadata is clued included with the actual ingested data. It adds data information to the map.
So the knowledge map of where it came from, from providence and to help establish pedigree. And it persists the data into data stores designed by the for retrieval by the models or for whatever it needs to be. The key here is if you wanna start thinking a bit like around ETL, we're not going back to the data sources, the original data sources.
We're actually creating a copy through the gateway and storing that such that we can add to and or remove that data makes it easier to manage. That's the responsibility of the gateway. So let's look at this highly simplified view of all the components, um, and some additional sub components.
Now, the purpose of these four components, the policy enforcement point, the policy administrator, the policy engine and the gateway is to enforce constraints on the ingestion of data or the consumption of the output of an AI system. The ingestion constraints are there to ensure that the proper metadata is available and stored in relation to the data that it represents. Now, the consumption constraints are there to enforce that, uh, to, to enforce and assess really two following questions.
Can the data be consumed by this request or say it's this model? And can the user use the data of this model? So there's an interesting thing here as we start to think about AI to AI models.
It is a bucket of trained knowledge. It has been trained on data, has access to other data which may have not been trained on. So to a large degree you have to ask this question and it comes down to an, to a to a authorization question.
Is this user given the context authorized to use sort of this body of knowledge with this external data and these multiple factors? And so by going and looking at and using this simple type of, uh, this architectural approach, what you're really doing is being able to provide ways to constrain the consumption of those resources that that question based upon whether that person given that context is allowed to either consume the data. So for example, let's say they can have use the model, but they can't have access to the data.
Well then of course they shouldn't be able to have a response to the request or let's say have access to the data. But the model they're trying to consume through, they don't have access to that. Or another option is for the context they're consuming.
Say the data's internal financial data, the model's a specialized model for internal financials, but they're asking a question to provide a response to some external party. Well, given that context and they shouldn't be allowed to consumer be generated a a response. And so as you start to look at what this, the melding the, uh, NIST 800, 2 0 7, the zero trust architecture with the gateway, this is what we're trying to deploy and we're trying to accomplish.
Uh, here's a bit of an example from both runtime and build time perspectives when it comes to compliance. Now the goal here is to provide a cradle to grave answer, such as you can have those top down questions. Is this model compliant?
You have bottoms up questions. You know what models have been trained on this data? You have depth questions.
Was this model changed on this specific piece of data? Or you have breadth questions? Were these models trained on this data?
So once you start to get to this, we start to think about sort where the map comes in and repression. The big key point and aspect here looking at build time runtime is we look at, we look at compliance and compliance over time is as we built something is a compliant, and then if our context of compliance changes, we can identify of what's been built, what do we need to change? So with that being said, that is an introduction to the neural gateway and the concept that we came out of the hackathon.
I hope this is all insightful and information for you. This is applied very mini as places or this can be applied very mini as, uh, very mini places, various, very, very minis. Did I just create a new word?
Various places you can apply this. It still uses a lot of standard concepts. That's why my argument is as you move into the AI age, managing data does not need to be different than how you're doing it now.
You just need to make sure that what you're doing now is effective and contains that solutions to the problems as discussed here. And these guiding principles, if you don't have the solutions, should be able to help you. Again, I'm Bill Bening.
If you had any questions, feel free to reach out to me and let's chat. Thank you very much.

