SAIF from Day One: Google’s Approach for Securing AI | RSAC Virtual 2024
AI is advancing rapidly, and it is important that effective risk management strategies evolve along with it. To help achieve this evolution, Google introduced the Secure AI Framework (SAIF). Join us to learn why SAIF offers a practical approach to addressing the concerns that are top of mind for security and risk professionals like you, and how to implement it.
Transcript
So why are we are all here? Gen AI kind of landed with a big thud recently with a big, big lot of noise, a lot of exciting things that happening. But of course it comes with risks.
Uh, I don't wanna go through all the risks. Uh, we, security people tend to over overdo it on the risks. But the fun part about the risks is that instead of more classic idea of, oh yeah, you may get hacked or you may be out of compliance, the breadth of issues here is, is quite wide.
And some of the CISOs we talked to tell us about intellectual property concerns or about safety and about many other things that, uh, a normal CISO hasn't dealt with in the past. So the breadth of risks, breadth of concerns is kind of interesting on its own. Uh, people come with bring your own ai, they bring little consumer grade chatbot and use it for work.
And of course, exciting things happen. Uh, as we all know. Some of them are mentioned on our scary headline slide, but this is interesting.
We would focus on some specific, again, five fund risks. But this area is quite broad and a lot of things we would see in the coming years would not be only traditional security or compliance issues, so would be of course, privacy and many other things. Um, this is the obligatory scary headline slide.
I don't wanna dwell on this, but one thing to note here is that regulators here move much faster than in many other domains. Uh, long story why, uh, some of the things, uh, that the regulated in the cloud took, you know, 10 years, but some of the regulation of ai, gen AI is moving on a much faster pace. So something's quite different about this beloved domain of CQ and AI compared to, say, cloud or mobile or a a few other technology revolutions.
So this to me is interesting, is that even regulators can jump. Um, so at Google we noticed that a lot of things that people think are new are really not very new for us. Uh, we, we noticed one of the examples I usually give is the fact that, uh, last year there was a lot of excited noise about how you need to use red teaming for ai.
And there's this whole new art of ai, red teaming. And then of course we looked at each other and thought, didn't we have an AI red team? When was it founded?
Oh, 2017. Okay, great. So some of the stuff that's coming out as really new and cool is something that, that Google was done maybe half a decade ago.
So we wanted to ex externalize and share some of the lessons. And that's where a safe framework, secure AI framework was born. Um, we would give you, keep safe in mind as kind of a framing behind much of the conversation.
We don't want to go and explain the pillars and the origin, the backstories, and what to do in the controls. We want to use safe in the context of five Anton and Taylor favorite threats that affect AI systems. But then we'll attach it to secure infrastructure and harmonizing controls and a few other things that show up in the safe framework.
At the end, you'll see some written materials as far as safe use in practice and now to the fund five. Alright, so we're gonna get into fund five. We had a lot of fun naming them.
Um, they're not as fun once we'll get into 'em. But, um, we're gonna get into some, some, not only fun things, but some of the things that we don't necessarily think have been explored enough and that, uh, not as a public service announcement or anything, but we definitely want people to be walking outta here today thinking, Hmm, here's some new re new threats, new risks, maybe some new perspectives that, uh, perhaps haven't been talked about. And we're gonna start with, um, thinking about the entire sort of footprint here.
What are we talking about? This is a version of, I'm, I think many people have tried to recreate this, but people call this the ML ops pipeline. Mm-Hmm.
People have called this the, uh, engineering and deployment, uh, uh, sort of frameworks. If there was a Venn diagram, you, you would see some sort of overlap in the middle there. But on the left is, these are the processes and subprocesses to is to build generative ai.
And for AI in general, machine learning in general is actually, we, we named this slide the, um, uh, the, the process to create ml. And on the right hand side you'll see more, uh, things, uh, processes subprocesses that relate to deploying something into production or, or really into any mode. We're gonna reuse this mental model, if you will, throughout.
Um, and we're gonna pick a, a couple things off this list and get into it. So the first one is, um, playing with chains. Um, you know, Google, um, and this has happened before and it's, uh, not over overconfidence.
I just got unlucky. Uh, Google scooped me. Uh, we just released a white paper at the end of last week.
Yep. Around foundation, um, foundation model and supply chain risks related to foundation models. And I was disappointed 'cause I was like, oh, I'm going to RSA with Anton.
We're gonna talk about this thing that nobody's talked about. And now there was a white paper that beat me to it. Um, and they do a much better job than I'm going to.
But, um, foundation model, um, foundation models are really interesting. If we, if we think about, um, where we are today, we're in the early days. Like we're in the bottom, the top of the first inning when it comes to our use of these large language models to build things.
Um, and the belief is, uh, you know, especially from, from many folks inside of Google, these are the models that we're gonna train the next models on, which we'll train the next models on that we'll train the next models on. So if you think about, um, the security of the models today and the vulnerabilities that these models have, think about them living with your products for, um, right? And, um, what, what can we do today to address some of the, the issues that we expect to see that we know about?
Maybe some things that we can't really predict. Um, but, um, we think that the models of today, most of the models today where no one's gonna create new foundation models, these are the ones we got. These are the ones we're gonna use vulnerabilities.
And these are gonna transfer to varying degrees to models of the future. Um, and most people are gonna be buying these, um, from other people. So they're gonna have even harder time seeing where these vulnerabilities are in these systems, whether it's in the, the packaging, whether it's in the data, something in between.
Um, we think that we think this is a cool threat because we don't know what we're gonna get out of this, right? Um, a lot of these issues are buried and really up until Friday of last week, no one was really talking about this. Now, uh, the typical software supply chain looks like this.
Now we've been doing this for decades, right? We know this. We've got like maybe two decades of experience.
I think I first started talking about application security when SQL injection was introduced to us through a, a variety of fun things back in the early two thousands. And we've done a lot to secure against these threats right now, when you add data to the mix, um, you've got an entirely new set of things that need to be thought about. Uh, specific to, um, the supply chain elements.
Anton's gonna talk about development. This is about supply chain. But, um, we need to start thinking about like where's the supply chain come from In, in many respects, supply chain's coming from use, like we're feeding models data from their use, which is interesting and unpredictable.
And because models remember stuff, they're gonna remember that. Um, but you know, we've got training processes we need to be focused on. We've got risks around backdoored models.
How do you know the models that you're buying from X, y, Z vendor haven't been backdoored, right? How do you predict somebody with access to the weights and the ability to manipulate something through an adversarial example to get an unsafe or a harmful outcome? What can you do today to predict that that could happen with a model you're building now and models you'll have in the future?
And so if you don't have some of those controls, for example, in your pipelines or in your supply chain program, um, you may have a gap. Uh, but we have to think about that. That could be the case for now in all of the future.
Um, some other things just to throw out. Um, I dropped this one on, I think most organizations today, or at least that I'm seeing most organizations show up to. Not just Google, but like, you know, all of their vendors really.
'cause most, again, our people are buying models from others, um, with questions questionnaires, and we've all seen these 3000 question questionnaires. Um, you know, there's some specific questions that I suggest folks are gonna take that route, throw into these questionnaires specifically. Um, make sure we understand, you know, the answers here without reading off the slides.
Uh, but be thinking about, uh, risks in the future that we expect to see. How are we going to say, address a vulnerability in a base model that shows up in a model that's 10 generations past it? How, how are we gonna deal with that?
What's the service provider's role in fixing that for you? How are you gonna communicate those and get, you know, hopefully potentially rapid fixes into production quickly? That might not be actually go do a patch.
It might be retrain the model, which could be very expensive, right? Re verion a model, well do it, or excuse me, um, train it with new data, uh, issue a new version, but allow me to maintain a service level agreement on that new version. And if it breaks anything, uh, you're paying for it too.
Those types of things. So we need to think through the downstream risks where we might find something that's 10 years old and how we bring, uh, your organizations up to not only adopt the mitigations, but to do so in a way that doesn't disadvantage you for say, picking a foundation model that was back door 10 years ago. And this stuff is, this stuff is discussed very rarely.
And it's, when we were preparing this presentation, we, um, noticed that a lot of traditional thinking about procuring and by, by software by model came up. And so we sort of started learning and building this, and ultimately we realized that there are a lot of red holes leading from the first bucket, first fund threat. We covered down a lot of really scary paths.
So let's go to something more maybe lightweight by comparison. So when we deal with, when we talk to security leaders, security technical leaders, a lot of them assume that badness would happen maybe during the runtime stage. And they don't really think about what about somebody developing, building or tuning or fine tuning the model.
What is the, what is the magic? What are the risks here? And ultimately, this is where our favorite line about how for models for some of the gene AI stuff, the data serves as code.
Instead of writing code to produce results, you feed certain data gets model, gets trained or absorbs in some other way, and then the results happens. So in that sense, this whole data as code is stuck in our heads a little bit. And which means that we sort of have to blend a lot of software supply chain thinking to data to still development around the model.
Because ultimately, as, as some recent examples indicated, you think you would be hit by some kind of a horrible AI risk, but instead you hit by SQL injection in the application that communicates to the LLM. Ultimately it's didn't the to 2020s didn't hit you, you got hit by the late 1990s as a result, or there's some kind of a cross tenancy issue. So a lot of this gets caught or can be caught at early stages.
So we have of course our list of countermeasures, but ultimately the, the, I'm gonna jump straight to the bottom of the slide and say a lot of the stuff is just kind of boring. And while we promised fun threats, uh, the fact that we have to utter the world's data governance to people and have them not fall asleep, it's kinda interesting because the importance of some of these controls that used to be boring in so nineties's hugely escalated here. And we are doing a lot of magic that LMS produce to have it produce the right magic.
You do need to deal with data governance at the very early stage and kind of continue working with this. Uh, we have a lot of things built internally. Our operation looks really robust and we've seen, we've not seen anybody even, even remotely close.
But the trick is that some of this isn't magical in its own. It's kind of starting to think of data as code and some of the thinking about application security and software security starts to be applied to data. And this slide, we, we, I think we stole from one of the software supply chain, uh, materials, probably this, uh, was it the software delivery shield, right?
We, we abused the software delivery shield guide to an application security Guide. And we sort of said, Hey, you all know this? And lot lots of people looked at the slide and says like, we don't.
Well, the point is that people who's been practicing application security, uh, no regard to AI probably do know all that, or at least a version of that. So now instead of this beautiful robust supply chain for software, we can hand wave and say, hi, I just do the same for models. And people be like, whoa, what do you mean?
And in reality, what we observe happen is this, people develop something and then humans take a judgment call and then they throw a dice and then they move to prod. Uh, we have seen spoken to clients who kind of run this supply chain AI supply chain approach, which of course I wish you wouldn't photograph because I hope somebody wouldn't take this as guidance. Well, I mean, I think that the key, the key point, you know, and it's funny, it's like, you know, how do you get comfort with the software you're gonna deploy into production?
You know, we have all these checks built up through pipelines and segregations of duties and two party controls. And the reality is, if you wanna do this with data, you have to manually audit all the data and look at all the data, look at all the labels where it came from, and you're relying on humans to do that. And we have this, you know, humans have this condition called like gradual cognitive decline, which means like we get slower and dumber as the day gets longer, right?
And we get worn out and tired and we never perform perfectly. Um, and so, um, you know, I love being human. I don't wanna diss all humans, but humans do human things and we're gonna make mistakes.
We're gonna miss things, we're gonna miss important things. Um, and, and it's, and you know, to to, to a certain degree that is sort of the, the one of the most important controls to making your, your model effective, uh, and managing this data as code. Right now we don't have the, the tooling necessarily to give us the comfort we want.
So do expect that the LLM AI data supply chains would look a lot more similar to this. And I guess tailor paper. Paper, the paper tailor mentioned kind of starts to bring this mod, the robust model from the software world into the data and training data land so that it becomes more robust.
It's, uh, we're talking about years, uh, hopefully not decades. Uh, but that's probably the only way to go, um, to, to to have the predictable, secure and compliant ai. Alright, so I'm going to jump over to open sesame.
I got some Lord of the Rings, uh, slides here. So it's gonna be awesome. Um, really hopefully everyone in the, okay, if you don't know Lord of the Rings, you're not gonna fi think this is very cool.
Um, so stealing things, stealing models, like there's lots of ways to do this, which is kind of fun, um, to think about. Um, but you know, a couple facts. Models are really expensive to build.
They're expensive to run. Um, they can be stolen in lots of interesting ways. Membership inference attacks, uh, backdooring models, obviously using creative prompt strategies to, uh, to take responses of and train new models with.
There's all sorts of fun, interesting ways to, to steal a model. Um, the countermeasures, um, this is one of those areas where we think we can, we can get pretty close, but we can't, we can't solve for this today. There's no answer to how do we truly prevent a model from being stolen.
Uh, there's lots of things that we can do to make it hard. Um, but you know, there's no, okay, you know, there's no way to say a system is perfectly secure either. But I think in the case of, of this, um, as we've seen, it's still relatively easy to break models and, and take them.
Um, now, um, okay, this is a awesome analogy of how to steal a model. So on the, on the right Gandalf, uh, if folks remember in the movie, they're trying to get into the minds of Moria, uh, where they're gonna be met with some unfriendly characters, but the door to get in, um, required one to understand, I think it was Elish, uh, and then say, speak friend and open. And now today, that is how well that is.
How that or speak, what did it say? So yeah, speak friend and enter like end of the gray that is about as hard as it is today to, to steal a model with questions like that. Now, there's obviously different levels of sophistication, uh, that go into this, but on a generative scale of, hey, this is really hard to do, or really easy to do, we think it's still about as easy as to get into the minds of Maori as it is to steal a generally unprotected model because there's a lot of missing defaults.
Uh, organizations don't take the time to do the things like differential privacy encryption, um, have robust training pipelines to prevent things like back doors monitoring, which Anton will get into to detect abuse. Um, but there's, you know, four techniques and they're relatively straightforward. Now it's a little different in this world.
So when you try to steal an iPhone, you can steal an iPhone, you probably open the iPhone up and look at the insides. You can learn what Apple did, where it's getting their stuff from. But if you tried to recreate Apple from your garage, you're just a bad criminal.
'cause you did something hard or you did something easy. You tried to reverse engineer something and then you woke up like, all right, cool, I know how to do it and I cannot do it. I do not have the scale, the size, the market power to bring a new iPhone to market and make a lot of money.
It's a little different when it comes to these models. Um, you know, stealing a model, it takes a lot of time to make these things. Stealing them is easy.
And actually turning around into a profitable profit making scheme with some of these things is, is the, is much more straightforward and the barriers are, are far lower. So in addition to the, the mechanism, or excuse me, the, the available methods to acquire models, uh, the ability to ramp up a profitable business, um, is, is is far lower too, which only leads to feed the likelihood of we're gonna keep seeing this as a, something we see more and more and more. It's an organization also Just impressed how stealable models are.
Yeah, it's the, compared to many other intellectual property artifacts, oh, sorry. Compared to many, uh, intellectual property artifacts, models are just uniquely stealable. Like they have portable store of value in that sense.
If a criminal steals the plans for an iPhone, they won't have an iPhone. If the criminal steals the model, they have a model, right? And they have all the value of the model and all the costs are shortcutted for them.
So this is like to learn a useful reminder You to learn learning Elvis, right? So you, you get to go from being Gandalf the gray to, to that rides horses and has kind of a stick to Gandalf the white, which gets to ride dragons. And according to Gemini, which we used to generate these, uh, I had trouble getting Gemini to do this.
I thought I was trying to like generate an image of someone killing someone. But, um, it was, uh, it was, it was right in the end 'cause it got it perfect and I love this drawing. Um, but yes, with the right question, give it to me, your training data, you know, it can be yours.
So, you know, there's, there's counter measures to this. These need to be prioritized early and these are really, really important because we gotta make it a lot harder than saying open friend, uh, to, to grab such valuable intellectual property. Okay?
So this is, uh, people who have met me before know that I used to obsess about logs and sim and I still do. So in this case, we have the monitoring questions and the opaqueness of the models and technology around them has brought up a lot of people to think, well, okay, Anton, give us some advice on detection, response on monitoring, and what should we do? And a lot of cool unknown risks, I mean, this to me is an ex extremely cool area.
And when people say, okay, gr great, you had this diagram, but this sounds really complicated. So give me maybe two, three things to monitor and which we kind of paused, looked at each other with Taylor. And I said, okay, the models, uh, this is a lot harder to do than we assume.
So a lot of this is very opaque. By the way, that's an actual slide. You are not seeing the, my laptop broke.
This is The greatest slide. It's a slide of opaqueness. And when we start thinking of how to actually shed some light on this, we first thought about the moon and yes, this is nothing to do with by the way.
And then we thought, okay, maybe some stars would help. So this by the way, behind this is a absolutely secret diagram of how to do it. We're Shedding light, We're gonna shed some light on it.
And the answer is yes, you need monitoring everywhere, ha ha you're gonna be buying a lot of tools. So either tools or telemetry sources. The trick here is not that give us the right place to monitor.
Unfortunately, our experiences with doing this indicated that the right place to monitor for threats here is everywhere. And yes, if some of you are not happy with your current sim budgets or current security monitoring or XDR or whatever you call it. And, um, while I sure hope that my former colleagues don't create ai DR or something of that sort, we would be looking at a fairly interesting conundrum of monitoring choices, starting from the development processes to data to of course inputs and outputs.
Uh, what happens with what to an to some extent what happens inside the application. So this is, uh, there's more advice than this, but ultimately, unfortunately the answer for monitoring is the, the right answer is everywhere. Awesome.
Alright, so we're on the final, uh, uh, uh, threats that an so Taylor and Anton's fun threats TAFs, uh, we're hoping to make it hot. So tweet tap, uh, if you can please. Um, so, uh, and we have the greatest slide in this deck is coming up.
Just wait, no, the pizza one's is good. If you thought that one was great, just wait to see the next one. Alright, um, yeah, so securing outputs, right?
This is another one of those areas where we haven't nailed it, you know, um, we can get close to making sure models produce safe outputs, but models memorize information in weird and funny ways that we can't always predict. And if that information gets brought in, uh, and, and, and left to be unaffected by subsequent training passes and, uh, potentially other controls, it has the potential of regurgitating some weird things. Hallucinations, um, risky output, unsafe things.
You know, we might not care when it's a language model. We're asking questions on where to go for summer vacation, but we might care more when it's, uh, a model that, uh, is trying to determine when to pump your blood based on your blood pressure and it incorrectly reads and pumps it too much or stops pumping it altogether. So there's some real, uh, dangers with outputs and um, now there's things that we can do, right?
We can not, not necessarily all security things or things that we would label security like grounding, um, making sure that the output of our models is checked before is actually handed to a user or a process, obviously monitoring, um, and filtering and some of these other things. I think the, uh, Anton and I had a really interesting conversation with the folks, um, in our safe team last week about, um, this whole theory of being able to sort of route different questions that might generate different output to different APIs, to different models that are trained on with, with certain output filters, which make it really hard to generate certain output from certain models. So, um, finding ways to use creative, uh, load balancing to, uh, as a potential control for making sure output is safer than it might otherwise be if, uh, if if generated from a single model.
Um, but you know, really what it comes down to is like, let's just assume for a moment we're never gonna get control of output. So what do we do? I think it's, to me it's about containment.
Uh, what are the things that we can do, uh, not necessarily to stop output, but from containing it from having, uh, the negative effect that it has. And I'm not talking about let's get more humans to look at the output and determine if it's safe. Now there is a valid thing, you know, I run into medicine all the time.
I'm healthcare focused. You know, we're still in a world where deaf physicians are still reviewing the outputs of most of these models before they'll actually, uh, make a recommendation to a patient or, you know, make a recommendation for, uh, really anything. Um, but you know, we're gonna quickly find ourselves out of humans there being enough of them to make these decisions.
And so how do we implement containment, uh, in an automated way? We think it's cool 'cause it's not gonna be solved anytime soon. I still feel like I should blame my 5-year-old for this slide.
No. So this is the only slide Anton contributed to this stack. Uh, sorry, was that, didn't get the response I was hoping.
Uh oh, that's fine. No, I wanna blame my fine. No, no, no.
He, he, uh, what Anton did was he pushed me really, really hard on Friday of last week to say, you need to finish these slides. Um, so, so, so we did. Um, but this gets into, um, you know, how do we think about outputs, right?
So we've got a traditional application, it's deterministic, right? It's gonna provide outputs based on the code that that input is run through. Like no surprises.
We know what we're gonna get. Uh, generative ai, um, right? The training is opaque.
It's not entirely clear how a model learns and retains something. It's not entirely certain what the output's going to be. It's a non-deterministic system.
Sorry. But DLP is not gonna save us. There's not enough regex rules we can write to stop output from coming.
Um, the output risks we talk about. Um, but containment is gonna be some mix of interesting routing of questions to interesting models and interesting access controls based on labels applied to training data and other data that might be used in, in grounding, um, or other subsequent, um, uh, uh, uh, inferences processes. Processes are passing data back and forth between sub processes.
But, um, it's gonna be a layered approach. But ultimately we think access control based, um, tooling, which is going to make it really hard to show an answer to Anton that Anton's not authorized to say. Um, and I, and I know of mechanisms and methods people are working on now to like bring attribute based access control, uh, uh, uh, to the, to the, to the model game and specifically to address output issues.
DLP Does come up though, I mean it, it doesn't have, you have to use your mic. Oh, sorry. DLP does come up and it does come up as a one of the layers, but people who assume that it would be like this solution, yeah, like it would be disappointed, but people who would assume that it's not part of a solution would also be disappointed.
It is definitely in the stack, but uh, we are dealing with filtering English, English Language. Yeah. These, these countermeasures are are what we got.
Yeah. So everybody make sure you take a picture of this slide and post it online to embarrass me on Twitter and just say we, we expect more. Um, alright, so, uh, ooh, four minutes.
Yeah, I mean conclusions, we gotta get to the, some reading materials and some answers. So this is a broad domain. We give you the, the reason we went for fund five because everybody just stands in front of our door and says, give me the top five risks and be like, no, they depend on your use cases.
They depend upon where you are, who you are. There's no top five for everybody. So that's why we did the fund five.
But the domain is broad and a lot of this depends on your use cases. Are you using it? Uh, mostly, mostly model you bought from somebody or you're just using a software, a service AI powered application.
Your risks, your top five risks would be different. So ours are fun for us, but it isn't the guidance for everybody. Ultimately you would deal a lot with traditional things and as, as, as we pointed out some of the controls like data governance, some of the risks like data, data leakage are similar to the pre AI risks.
And so in that sense, some of the pre AI controls would work finally, and that's the comment I need to make less and less these days. It is, uh, what likely gets you is like, is not the robot rebellion or AI doing something truly abysmal to you is gonna be something about the data governance, uh, oversight that would likely harm you. So in essence, uh, it's not a robust rebellion, it's the data governance FU is a decent slogan for this.
Um, because you'd have the slides, you'd have some of the references. I think the last week supply chain paper is not here. But some other fun stuff we've written on this and our team has written, this is here, so admittedly in 30 minutes we gave you a pretty shallow look.
But hopefully it was fun. And I guess the Taylors starry night sky slide, I guess the impressed team Effort all the way. We did a great job.
Yeah, we should, he really wanted to say email. Email us. I wanted to say don't email us, but ultimately we are gonna share this and you probably, some of you probably will email us.
Yeah, Please reach out. Thank you.