Cybersecurity Threats in AI Operationalization with Melissa Ruzzi
Melissa Ruzzi, director of artificial intelligence (AI) for AppOmni, dives into the cybersecurity threats organizations need to keep in mind as they look to operationalize AI technologies.
Transcript
This is Textron tv. Hey guys, thanks to the throw. We're here with Melissa Zi, who's director of AI for App Omni.
And we're talking about some of the inherent security risks that go along with AI ops because, well, we might be automating things, but maybe we're not thinking it all the way through as much as we should be. Hey Melissa, welcome to the show. Hey, thanks for having me.
Everybody's kinda looking into AI ops and folks who are at different levels of progression, no doubt. But from your perspective, when it comes to security, what do we tend to overlook? I think the one thing that we forget is that the core of machine learning, ai, anything is data, right?
We tend to focus a lot on the models, uh, even when people are talking about LLMs right now as well, right? So we focus so much on the models and we forget that it's like, hey, there is no ai, there is no machine learning without data. And actually the most precious part of AI is the data.
So the first thing that we tend to overlook is how, how much security we need to put around data. So when we hear about data breaches, everyone's aware of that and it seems that we tend to forget that that is also connected to lops. So one of the big problems on ML lops continue to be security breaches around data.
Either injecting data to change the model or extracting data through the model in different perspectives. And another thing that we tend to overlook is that we think like, hey, if it's running in the cloud, we're all safe. Well, it is safe as long as we put the correct configurations in place, right?
So if you put the correct configurations in place, your cloud provider will assure that it's safe, but he can't assure that whoever is doing that is doing the proper configurations. So the two big things is around the data and the second is around configuration. And then the configuration goes into everything that is running in the cloud from the data to the model itself.
Third part that we look into the machine learning models itself. That's the other thing that sometimes we tend to overlook. We want to do things too fast.
So sometimes people want to reuse models, they're out there that someone else developed and it's very important to know, okay, what are the other libraries, the other pre-trained models that the company's using to know that they are also safe? Are the bad guys targeting these AIOps platforms and and going after this data? I mean, I think part of the assumption is that, you know, these are too sophisticated for them to actually crack.
Yeah, and I think that's the part that we overlook or a lot of people overlook into because they focus on the machine learning part and may happen that you put a lot of safety around the model itself, but not the data because the data is being moved around, right? So we have kind of like, let's say the raw data that needs to be treated in order to go into the model and then the model is using it to then give some output. So there is data moving around there.
So there's a lot of little places or different places that the attacker can go into, almost including, and unfortunately we know about that, that it happened is like reverse engineering based on the output of the machine learning model or through an LLM, you can try to extract the data that was used by the model to give that answer. So there's different ways that the attacker can try to get in to get a hand on either the model itself or the data and then to either change the model. So that's one type of attack that happened I think with Tesla that some hackers hacking to the Tesla.
So the autopilot of Tesla will start doing things that was not supposed to do long time ago, um, has been fixed since then. And the others were to extract the underlying data that's being used by the model. Is this two separate motions then?
Do I need to figure out how to secure the model and then also secure the data or can I unify that? You can unify that. Um, there's two big important points to think about it, right?
But if you're running everything within the same cloud provider for example, is the same approach. You can think about the model as being just different type of data and take the same approach into it. And the idea is really what's going out to the world.
Every time that you have an interface to the world, you need to be careful of that. You know, we know about that even outside machine learning that with misconfiguration you can get even internal data exposed to the normal world than anyone can go and find it. It's not different with the machine learning.
So you can unify that for sure, but you cannot forget that there are two different fronts or two different things to look into. How do I kind of discover this stuff? Because I think part of the issue is the cyber criminals, they're getting smarter and they, I guess they don't just smash and grab anymore.
They kind of break in and linger and hang out for a long time and then I don't really notice them. Yeah, and that's a um, a big problem actually. We do that a lot.
That app Omni is about the SAS configuration, right? It doesn't take uh, too much to click maybe a wrong button and do some wrong configuration or misconfiguration to get your data records exposed. So it's the same approach into a lot of companies are using more and more SaaS and you can think about this machine learning AI pipeline as being another SaaS because you have data there, you are connecting to other parts depending what it's in purpose for.
So the core is like how can you extract one configuration, which we call that security posture, into how you have everything configured, and then are there any threat detection alerts happening around that? So to way, the way to think about the machine learning the mops is kind of like is another SaaS. You need to be aware of all the configurations around the whole pipeline, which includes model and data and treat that the same way that we treat pans SaaS security.
I don't think the folks building AI models as data scientists know all that much about security and maybe it's too much to ask them to know that. So do I need to go find cybersecurity specialists that know AI or where am I gonna find the skills and expertise to go address this? Yeah, I think that that's a very good question because if you frame the ai, the MO ops are being data, it falls under the same way that you treat any other SaaS, right?
I think the key to understand how to handle this is to really to look at the moops from a SaaS perspective. Hey, it's like having another SaaS in your environment. So a lot of the big companies that handle with a lot of SaaS, they are starting to look at AI and moops on that perspective.
So the key is that to not forget that the M mo ops is another security part and you take a look into and definitely yes, the same way that you look from a cybersecurity SaaS cybersecurity expertise on the SaaS, you have to look from that perspective as well into the AI, into the ML ops. So there is definitely this joint combination of teamwork between people that know SaaS security and cybersecurity with the people that are developing the ML ops. Are there different types of attacks that I gotta think through here?
'cause you know, you hear phrases like, uh, jailbreaking and poisoning of models are, are are these threats, you know, are there more and are they classified into what kind of buckets? Yeah, so if we go from a simple definition of the attacks, if I can put it that way, there's two types things coming in and things coming out, right? So there's that tax trying to JB break is data injection.
It could be LLM JB break, it could be even data injection for any other type of machine learning visor unsupervised that you can inject data into it. So there's one type of attack, you're injecting data to change the model and that has a huge impact depending on what the model is used for, right? So it can really, companies are using that to make decisions about their products or how to handle things.
It can have a huge impact to the company even economically, right? So that's one type of attack. It's different than a ransomware, but it can almost think as a ransomware, I know how the ransomware, they keep your files and then you can't have access to them anymore.
It's kind of that impact until you go and you change the data. So now your models are not working anymore. And that can have huge cascade problems depending on what you're using the models for.
Because let's say if you using the model to approve not a presu insurances in your company, now you may get all kinds of different things approved that you didn't want approved, right? And the economic impact of that is huge. So that's one type of attack.
The other type of attack is really trying to extract data. So if you think about all the ins and outs, there's all kinds of different techniques to try to get that, but really focus on this in and out is the core key to understand that what are all the different ways that data can get in? What are all the different ways that data could get out?
And at the end it goes to all those different techniques that attackers could use out there. Am I gonna need maybe an AI agent or some sort of external AI tool to verify or provide the guardrail for my AI model? And these things are gonna basically keep an eye on each other.
Um, a hundred percent. I think there's always good practice independent of attackers to keep watching the model. And there's different ways to do that depending on which type of AI machine learning model is being used, right?
So data drifting, for example, to watch like hazard data being changed tremendously, because if your data hasn't changed too much, the model won't change too much, right? And as soon as you see that, hey, there's some data drifting here, then it can go and take a look into why that's happening and find the root cause. Could be an attack, could be just, you know, it happens, right?
A benign change. So if you think about the whole pipeline, again from the data perspective, it's always from the data perspective to understand has the data changed? The other part is hyper perimeters and the parameters of the models, right?
There's another type of attack that if they come from the libraries directly from the models, they can change those. So any type of change and tracking that through the pipeline is what can help with that. Then there's a lot of information, there's a lot of data to deal with, right?
How can you handle all of that? That's where an LLM, for example, could help, right? So we have this direct ai, other machine learnings watching the machine learning and depending on what is finding, you could have a running on top of it to really kind of help, as will be a cybersecurity person and say, Hey, I saw this data drifting here and I noticed that the content of the data is this, that, that and that.
And normally you don't have, maybe this thing is happening. So definitely AI can help on top of watching all this. So ultimately, do you think that these issues are kinda holding up people in terms of their ability to operationalize ai?
I think we had a lot of enthusiasm a year or more ago, and maybe the hardcore realities of data management and security are kind of bigger than we anticipated. I think that if you start a development of AI already aware of those, it should not hinder it. I think the biggest problem is when you develop the whole pipeline, not thinking those things in consideration, and then when you're ready to go to production, you have to restructure a lot of things because you are taking a look.
I think that the whole key for it is going to think ahead of, and as you start developing ML ops, as you start the r and d part, the exploration part of the machine learning, if you're already thinking to how we're gonna put that in production, how we're gonna make sure that everything is safe, then it should not impact take into production at all. I think that's the idea, being as proactive as possible, already thinking to taking this to production, exposing it to the real world, and taking all the lessons learned around SaaS security, around how to secure the configurations and the data. I heard folks here, you heard it here, it's not necessarily rocket science, it's just that we don't know what we don't know.
And the problem with AI is the idiot tax is pretty high, but if you take a minute think about it, you can implement and get to something that well hopefully provide some sort of competitive advantage in the long run. Melissa, thanks being on the gym. Thanks for having me.
All right. And back to you guys in three.