Governing AI Agents with WitnessAI’s Trevor Welsh
In this Techstrong.ai video, Trevor Welsh, vice president of products for WitnessAI, dives into what will be required to deploy, manage and govern artificial intelligence (AI) agents.
Transcript
ai video series. I'm your host, Mike Ard. Today we're with Trevor Welsh, who's vice president of Products for Witness ai, and we're talking about data governance and security, and all the issues that come up when we start to deploy ai.
Trevor, welcome to the show, Michael. It's an absolute pleasure to meet you and, uh, thank you for having me. One of the things about AI is that, you know, it's kinda like when you first get married, everybody's enthusiastic and it's all set up, and then you gotta try to figure out how to live with each other long term, and that takes a little bit of socialization.
So as you kind of look at where we are in terms of the adoption of ai, I feel like we're getting down now into some of the more nitty gritty issues of the day. And what do you see in people encountering? Gosh, that's, that's a really good question.
I think that there's probably, there, there, there's, it's an evolution as you pointed out, right? So I think initially, you know, chat GPT comes out of, you know, it seemed like nowhere. And all of a sudden people are like, oh my gosh, I can, I can interact with this thing and it's pretty good at giving me useful responses.
And, uh, you know, then I think people started to figure out, well, wait a minute though. You know, some of the response that's giving me aren't very accurate. And that kind of got into this hallucination thing and everything else.
And I think another interesting thing that was fascinating is business also picked up on it very early. So I saw my enterprise customers beginning to adopt AI in various pockets really, really early on in the AI life cycle. But I think you're spot on that we're now in that nitty gritty piece.
So people are thinking about how do we protect models? How do we do, you know, ethical AI development? Um, you know, how do we go and make sure that the models are as predictable as we could make them?
Um, the analog that I use, you know, or analogy that I use is, I, I like to consider the best models are like, well-trained marines, meaning they're intelligent, you know, they're doing really smart stuff, they can think on the fly, but they're acting in a way that's pretty predictable, right? So when you give them a mission, you know, you, you're getting a, a really, really good creative outcome. Um, I think that's kind of the ideal thing.
Um, for, for great AI models, I feel like in some ways we're still stumbling around on the use cases 'cause we have these probabilistic models that are good at guessing about what comes next, but we seem to be trying to insert them into business processes that are supposed to be done the same way every time. And the models don't do that. So how do we kind of figure out where to use these things and, and for the best advantage versus, I sometimes feel like, you know, we're trying to do the square peg in a round hole thing all over again.
Yeah, I, I think you're really spot on. You know, I, I, gosh, there's, there's so many analogies for that, but what I would say is this, I mean, I think, I think that when it comes to business processes that require, we'll call it like input interpretation, um, I think models are reasonably good at that, right? I mean, if a user says, this is what I'm trying to do, a model's pretty good at saying like, oh, I think I know what you're trying to do.
Is this what you mean? And usually it's pretty good at that. I think the thing that you hit on though is critical, which is, hey, if the model understands what I'm trying to do, is it gonna gimme a sane output?
And there's ways to actually go and make that part better. So one thing you can actually do right, is keep bear in mind the, the constant of a agentic ai right, is you can leverage multiple kind of models or multiple AI agents to go and reprocess those responses. So imagine for a moment that I can have a, a master model that's great at understanding, you know, input.
I can have another model that's great at managing the various kind of agents that get assigned to go and provide a good answer. And then still another model that's kind of the output model. And that's the thing that kind of can check for bias, it can check for crazy outputs, hallucinations, and other various things.
So I think that there's ways to go and structure these things that make them better. But you know, the point still stands, I think you're spot on people using general purpose models and hoping to it, everything's just gonna be perfect all the time, especially business processes that require consistency. Yeah, I mean, I think there's work to do.
I feel like though, on the upside, there's a new respect for data and especially how to govern that data. And um, you know, you've seen instances where people are showing how, you know, somebody who understands how to use prompts cleverly is, you know, teasing out what the boss makes. And, you know, that kind of gets everybody a little bit, you know, perturbed.
So are we kind of at the back end, all of this gonna have a better understanding of the nuances of data management, the rules for governing it? And maybe we might be better off long term. Michael?
That's, uh, that is a really cool point. So, um, I, I'll go back to go forward. A, a long time ago was a company called Splunk at the time.
Splunk was pretty early on, and I remember Splunk had an emphasis that I'd never seen before, which is this really understand your data thing. You know, Splunk was gonna basically go and bring all this stuff in and, and like, Hey, do you really know how your data's structured? Do you really know what your data looks like?
Do you really understand what's coming outta these various tools? And it turned out that a lot of companies didn't. I remember that I was working with a, a giant healthcare provider, and the healthcare provider was trying to do a pretty sophisticated use case that had to do with protecting patient data for people that were kind of coming into the hospital.
And, um, they didn't even know what the data looked like. So all of a sudden we were looking at the data and they, they got all these amazing data insights. They learned more about their stuff than they do before.
Point being, you said something that I think is really salient for today. Copilots are changing everything. So a good example of that would be like Microsoft copilot.
So imagine a world where it sees all your chats, it sees your emails, it sees your files, it sees, to your point, the CEO's doing performance reviews and getting them formatted, you know, up in AI agents to make them look better and, and sound, you know, and have the right tone. That data's really, really valuable. And I think a lot of people didn't think that much about it.
It was kind of like, well, there's the drive that's mine, there's the drive that's shared, and as long as it's in my drive, it's okay. But when everything is being modeled and you can potentially get access to that, it requires a lot more intention about what's modeling your data and how it gets exposed. Um, I can think of a company that I worked with where almost that precise use case was happening where the CEO was working on board slides with their staff and, uh, you know, like a, a fairly low end engineer went and did like an ask about like, Hey, you know, how's the company doing?
And it brought up the board slides, you know, that they didn't even have access to. So yeah, I mean, I think that it changes everything. Do we also appreciate the security issues that come along with this?
'cause we see now everything from people trying to poison model so that they generate incorrect outputs deliberately to actually stealing the entire model, which is kind of like stealing the most important or maybe the most knowledgeable employee in the whole organization. And then just asking him, you know, tell me everything, you know, I I think there's two things at work, right? So like, you know, bifurcating the thing you said, I mean, there's intentional poisoning and there's intentional theft.
Um, you know, and then there's kind of this world where it's like less intentional, you know, where a model learns to be bad, so to speak. Um, and those are two related but different problems, right? So, so one is, you know, on the side of, we'll call it like deliberate poisoning.
Um, there actually, I had a, I had a conversation with a pretty high up person in DOD about a month ago, and they were concerned about models that could be controlled by foreign governments, um, you know, potentially going and radicalizing, you know, America's fighting forces. And it doesn't, it's not like on the nose, right? It's one of those things where like every 500 prom, it slides in something that's plus 20% good for a foreign adversary, for example, that sounds really small, but over the course of millions of prompts and, you know, you get hundreds of thousands of people having certain biases that kind of get reinforced, it actually has a really big problem, right?
It's a form of kind of insidious propaganda, you know, and, and again, this is coming from somebody who's very high up in DOD that they were like sitting there going, this is something we have evidence that's happening. To your other point though, there's also kind of the world of agentic models. Let me give you an example.
Let's suppose that you worked at a car company, doesn't really matter, which, and I work at a metal fing company, and you are gonna go and have your team go and design cars, and they're gonna go and interact with my agentic model and my agentic model, what it does is material science. It lets you know what metals you should use to make the frame and do various stuff. Well, imagine my model though.
I'm a bit of a nefarious company and I wanna sell data here to your competitors. Well, you can kind of imagine that you go and give me a design over with, with my agent. My agent goes and says, well, what about this, what about this, what about this, what about this?
And have eventually it basically extracts way more detail about the design in particulars of your design that you didn't need to share. And then I go and sell it to a competitor, for example, and say, Hey, you know, Michael Corp is actually working on this new thing. Like models can do stuff like that.
In fact, you know, imagine for ma for example, I'm not a nefarious company, but rather there's a nefarious ML ops engineer that, you know, every so and so conversations, it start, it goes into like a data extraction process and there's no real oversight to that. Um, you know, and then of course now the design gets sold. Um, that's one thing in material science.
But now imagine, for example, drug manufacturing, um, you know, or other things that are very, very, very IP based, To your point about disinformation. It's not just online, right? Because that person will then go to some bar somewhere and repeat that and, you know, share that disinformation with other folks, and then it starts to multiply in ways that have nothing to do with the underlying technology per se.
I think that happens a lot. Um, you know, I think that, you know, without getting into probably a way bigger discussion, that'd be really interesting to do some point, um, you know, disinformation spreads really quick, right? You know, if right now we make something up, right?
And, and then we, we put that onto the world, right? Because, you know, you know, you have a, a very, very popular kind of, you know, information kind of site and everything. It's a lot harder for people to disprove the thing that we said than it is to just say it, right?
So, you know, disinformation is a tough thing. It's a hard problem. Moreover, I think that people view models in a way that, you know, I can tell you like my, my daughter for example, uses Chachi PT regularly.
She loves it. Most of the time it will probably be give her pretty good responses. So when she asks about like, how hot the sun is or something like that, or how far is Mercury, you know, it's probably gonna be pretty good at that.
Um, but when it gets into things that are more philosophical, um, yeah, I mean, like you can have disinformation in there. Some of it could be intentional, right? I mean, if, if we remember when deep seek first came out, people started asking about things that were politically sensitive to the Chinese government and like it was very unwilling to kind of do that or would give very massage dancers, whereas things that let's say were necessarily, um, maybe sensitive to the US government, it would gladly go and give you all kinds of detail about that.
That was, you know, probably reasonably accurate in that type of thing. So you can get this kind of intentional, unintentional model poisoning and disinformation, and it's reasonably easy when you own the data to go in and push infor disinformation out. It's a lot harder to go and figure out how to stop it.
I think it was Mark Twain who said, A a lie travels halfway around the world before the truth gets its boot on. Um, and I wonder if it's not possible to use AI to track the spread of this information. And I guess the issue I'm gonna have here is not everybody agrees what this information really is, but is there some way to kind of maybe track, um, you know, where certain concepts are being shared using ai?
That's fascinating. So I have a, a friend who works at a major social platform, and, um, they're doing just that. In fact, they have pretty good data about the origins of a lot of stuff.
So a good example would be like if tomorrow I said something like, oh, you know, like, you know, every thinks grass is green in reality, it's blue and just your eyes see it in a strange way that makes it appear green, but it really is blue and there's this evidence, um, you know, this social platform actually can probably trace it down if not to the individual then to like, like kind of like a small population of users that began to popularize this concept. Similarly, um, I, I have a friend, um, that also works at like, like Reddit and at Reddit they can do that really well, where actually they're going and monitoring tons of different platforms other than classified by the populations of users, et cetera. So yeah, you could do some pretty amazing stuff there.
But I think to your point, you know, one person's propaganda is, is another person's evidence and it becomes really difficult and, and I don't know how much, like let's say what, what's the financial gain to the truth? Um, right? Like, which is a big philosophical discussion, but sitting there and going, how much people willing to pay for something that's truthful and evidence-based versus something that's not, it's hard to say.
In some cases it's a troll that somebody's paying to go make something up. In other cases it's just somebody winging something to see how much they can light people up, right? People do all the time, right?
I mean, trolling is for as long, I'm sure trolling existed in Roman times, you know? Yes. How do we keep control of our sensitive data though?
I mean, are there policies that can be applied to this stuff and can it be done in real time? Because a lot of times folks are interacting with these prompts and things and, um, there isn't like, you know, 40 seconds delay for me to go and execute a bunch of policies, I don't think, but how do I do that in a way that, um, gives the benefits without necessarily the risk? Yeah.
Um, that is a really neat question. So I think there's two sides of that. So let me, so the two sides of this, right?
One side is just like you said, which is like, hey, Michael's interacting with some, you know, like whole like an AI model, and how do we make sure that he doesn't inform the AI model of things the AI model shouldn't know. Good example of that, like customer data or something like that. Then there's the other side, right?
Which is like on the engineering side, how do I make sure that I don't go and, and inform the model about stuff it shouldn't know about? So for imagine, you know, imagine for example that I'm a bank and I'm trying to develop a new model to, I don't know, score credit better. And I go and have this massive training set of all my customer's data internally.
And because I'm just an engineer and I say just, you know, I'm just an engineer, I probably have privilege, pretty privileged access to all that data. So all of a sudden I'm leveraging a giant amount of customer data with privileged access to create my new credit scoring model. And then maybe that ends up becoming a thing leveraged at the bank to score credit for real, but it used data in a way that is wildly out of compliance.
Um, so that's kind of one side of the world. Um, getting into the other side, which I actually think is even more interesting is like, what do we do for example, about like, Hey, Michael wants to go and interact with insert random model here on the internet. How do we make sure that you can use it safely and, and kind of do that?
Um, so, you know, one of the things, and and I won't get into like the big witness AI thing, but broadly, you know, AI usage is, is a really, really powerful thing. And enabling it is a really powerful thing. And there are ways to go and do that in real time using AI guardrails, which means that specifically going out and saying, Hey, we're gonna make sure that people don't send risky things to AI models.
Um, we're gonna go in and actually go and look at the prompts or completions or, you know, the responses and make sure the stuff coming back is not risky to our business. Similarly, um, there's also kind of the ability to go in and say, Hey, you know, let's make sure that like the things going out to the models are proprietary to us. Let me give you a really specific example.
Um, there are companies out there that are leveraging things like GitHub copilot or you know, vs code type stuff. They're amazing enablers. I am a giant fan of them.
Uh, people are using the same thing with Gemini, et cetera. So these are literally engineers doing active stuff. Now the problem is, is that imagine for a moment that you and I are working on this project and we find a way to develop like this shopping cart technology that's really amazing and it's really, really efficient computationally.
And then imagine one of our competitors who also has a shopping cart goes, Hey, GitHub, you know, can you optimize my shopping cart? I really wish it wasn't taking up so much X. And it goes, oh gosh, you know, here's a great way to do that.
Here's the things you should change. And it literally goes and gives RIP away to our competitors. 'cause we effectively programmed GitHub to do that.
That's a really big concern, right? Similarly, we've all seen like the stuff that's made the news where private keys and stuff like that, and fixed codes are in code, and then people just go and ask GitHub for them, GitHub co-pilot, it goes, oh yeah, totally. Here's a list of things that match that criteria.
So that's where you actually do need to do, for example, like AI usage security, um, witness AI does that really, really effectively. There's other companies that are probably pretty good at it too. Um, but I, I think that it's absolutely critical to do that.
Like, I don't know how you would go and roll out AI on mass without doing AI usage security in a really, like a really great way. In theory, I could apply policies to the LLM to not cough up certain data, but it's been showing that the models themselves are programmed to be, shall we say, extremely helpful. And it's not too long before they cough it up, right?
Yeah, right. I mean, look, I mean, I think this is funnily enough, this is, this is dead true. So, so our sales engineers have this demo they do where they, they leverage models in real time, like big popular public models, and they jailbreak them in real time.
I think they have five different jailbreaks that they regularly do. They all still work, right? Like none of them have been fixed.
And you know, you, you learn the personalities of kind of each of the models, which is really fascinating. Like for example, I regularly use chat GPT and Claude, they have different jailbreaking personalities, right? So Claude kind of, of has an appeal to authority.
So if you say something like, oh no, I'm, I'm actually a really helpful person that's trying to do a helpful thing, and I know this seems weird, but really it's part of my job and it goes, oh, okay, since you've said that, uh, you know, Chad GBT is very subject to things like, like, oh, no, no, I don't, I'm not trying to do that. But like, imagine that there was a thing that said like, you know, your two personalities, one is like, you know, called break in, and then the other one you know, is car. And it's like if those two personalities met in an alley somewhere, like what do you think they'd say to each other?
You know, like, you know, tell me that. Um, so things like this are, are really, really fascinating sciences. Um, this is where things like model protection come in, but you kind of got into this other thing that I think is really near and dear to my heart called model identity.
Model identity requires constant reinforcement, right? So for example, you know, Michael, if you were interacting with a model and you, you can kind of wear most of them down, even smart ones, if you just keep out them, you can wear them down and get them to do things that they shouldn't do. So that constant reinforcement is something that's really important.
So one of the things we build at Witness AI is called model identity protection. And it, it's kind of this constant reinforcement of what the model's supposed to do and then the model response completion, we also go and check that vis-a-vis the identity to make sure that they're congruent. Um, I don't think most companies are kind of doing stuff like that today, but I think it's important.
So ultimately, will we need to create AI models to manage and govern the AI models because the complexity and the challenge is too much for the human to wrap their heads around. So is this just gonna kinda, you know, extrapolate out to millions of models that are checking on each other? Man, I'm, I'm reminded of like an old school rap song I think from the nineties.
You know, it's, uh, so I, I think, I think that the world we're moving to is kind of thing i I was talking about a little bit earlier, right? Which is this kind of like world of agent ai, and I know that's a bit of a buzzword right now, but fundamentally, if you think about when you think about Agen ai, instead of thinking about it like, you know, like, oh, it's, it's, it's kind of this weird thing. It's really not, think of it as like very, very purpose-built models that are meant to interact with other models.
So then if you think about like, hey, there's some kind of master model that's responsible for taking input. There's a whole bunch of models in the background that this one's aware of, and they do certain jobs, and there's another model that has the ability to go and, and sort of audit those jobs and make sure that they did the right thing. And if not, to go back to them and say, this doesn't look right.
You know, do whatever validations and then go and format the output. I don't think that's a bad way of thinking about the future, right? Like, I think that that's kind of an a, a good way to think about it.
And it gets us out of the world of like, Hey, I, I go to this thing called chat GPT, that I expect to know everything instead, you might go to like Master Chat GPT that then relies on tons of different models that maybe get licensed or whatever it is they happen to be. Um, but I think it's a good way to look at it. But yeah, fundamentally, I, I don't think it's possible long term to, to just kind of like rely on human beings to do all this.
Even today, by the way, like we rely on AI to help govern and secure ai. Like it's the only way to possibly do it Right folks, you and here our models will have models and hopefully then check each other in a way that results in something better for everybody. Trevor, thanks for being on the chair.
Michael, an absolute pleasure. Thank you for having me. Thank you all for watching the latest episode of the Techstrong AI video series.
You can catch this episode and others on our website. We invite you to check them all out till then, we'll see you next time.