Generative AI and Security: What’s Next? | RSAC Virtual 2024
There’s a lot of noise in the news about where AI security is going and the implications over the next few years. Why don’t we talk with the people who are actually working on, and with, the major AI platforms and see what their vision is for the future? In this quickly paced Q&A, we’ll talk with Jason Clinton (CISO at Anthropic) and Joshua Saxe (Meta) about how they see all this playing out. You won’t be able to get any closer to the source than that.
Transcript
Yeah. So I had that slide and, um, you know, basically I would add that, you know, based on all the information that I have available to me, I don't see this slowing down. So, um, you know, one of the, one of the things folks out there will, uh, will, will trot out as a reason to not be, uh, concerned, or that the models may be slowing down, as they say.
We've already trained on all of the available, uh, information on the internet and like, you know, where is more, more training data gonna be coming from? And to that, I would say, like, all you need to do is like, look at AlphaGo as inspiration for where things are going. Um, for those of you who don't remember, uh, AlphaGo was, uh, an effort to, uh, train like the most powerful, um, AI model for playing the game of Go, which is like potentially the most complicated game, um, uh, in history.
And so the way that it worked, the way that they achieved a breakthrough there was to have it play itself over and over and over again millions of times. Um, and so there's potentially this breakthrough in tokens and training data, which you can have just simply by having models interact with other models and sort of work through the reasoning and on an objective scale, which is, you know, the alpha goes metric is whether or not you won the game. Um, and of course, if you have models interacting with each other and other models rating those interactions, and they rate the outcome as positive or, or the solution is achieved in a sort of a solution space, then then the training tokens are just generated, uh, in that space.
And so, yeah. David, I want to jump in with your, your thoughts on that. When we're talking about how Go was basically training itself, where do you see the various disparate models training each other and where that's gonna go?
Well, I, I, Jess, we got that going. Got this going. Yeah, there we go.
Yeah. Um, well, I, I think that, you know, speaking as a science fiction author, I, I think that the range of potential ideas that can then be weighed by these entities, uh, is not limited to reality. Of course, we know this from the hallucinations, from their outputs, but when they are competing with each other, that adds, um, and interacting with each other and notice the, the, uh, go, um, model that Jason mentioned that's within a single Mind, we are all complicated inside.
There's, it's been proved we're very, very competitive within our sub selves. So I spoke of individuation, but we are then distill what's going on inside this as, as the Go program did Alpha Go, did, uh, into a, a front that interact faces with others. So there are layer layers and layers and layers.
It, it may sound weird, but the internet of existing websites, facts and all of that is just the first layer. Some of you have heard of Cantor, the mathematician who had talked about layers of infinity. There's the infinity of all, uh, integer between all integer is, is the, um, infinity of rational numbers between all the rational numbers is an infinity.
Each of the rational numbers is an infinity of irrational numbers. So we're talking about levels of complexity here that will feed these AI programs for a very, very long time. Yeah, that's right, Josh, to follow up on that part of the discussion.
Thanks. One of the things I think would be interesting to, to investigate here is are the major AI platforms going to be checking each other for accuracy and credibility? So will the, will the major will OpenAI, will meta the, all the major ones, will they be a way to build security into the various systems by the systems checking each other?
Yeah, that's a good question. Um, yeah, I, I don't know. Well, so I think that, so I think we're all learning from each other.
Um, right now. This is, hopefully this is relevant to your question. Like, you know, so, um, like you've seen, you've seen the AI world close down a lot over the last five years or so, and, um, you know, I would say five years ago, and it was extremely open.
So the, the, the major companies, even the for-profit companies tended to publish everything that they did or everything of value. And, um, you know, there's less collaboration. Um, but, um, there are sort of rallying points around like benchmarks and this sort of thing.
Um, and also I think the way that we're doing safety, like, like a lot of what Jason was saying in his presentation around the way, um, I forget the category. I think it is, is it a l Yeah, a SL like the safety level. I mean, I think we probably speak this a similar language between the, between the major companies and also, um, share some of the same evaluations.
Um, so I think, I think we check each other with respect to how well we're doing against these public benchmarks, which have become sort of a common language for, but I don't think, I think to more directly answer, I mean, yeah, I, I am not seeing a, a trend towards the models checking each other. I mean, who, who knows that could happen in the future, but it's not, it's over the horizon if that's gonna happen. I Think there's two small ways that it's organically happening.
Um, not in a big way. Uh, one is, uh, I don't know if y'all have seen this, but there's this, uh, sort of emergent, um, red teaming. So, uh, ai, red teaming.
I don't think we've, we've defined that this is the idea of trying to break a model and sort of sussing out what, how it breaks. Um, you can now use a large language model to automate the generation of one of these, uh, these red teaming prompts. And so, uh, it turns out that sometimes actually using an, uh, a, a, uh, language model that is not the same as the one that you're trying to attack is a little bit more effective at that.
So that's kind of a, a small way this is happening. Um, the other thing is, um, and I, I didn't really have a chance, I don't think any of us had a chance to talk about this when we were talking about the different safety barriers. But if you're deploying a safety barrier around a model, like, uh, and especially if you're a customer and you're doing a deployment, you have the opportunity to pick any models you want.
Like, um, it's important to pick a model that doesn't correlatively fail in the exact same way. So if you are, if you have a large language model in front of your large language model, and they're both exactly the same family, and they both are subject to exactly the same jail breaks, like what have, what have you gained from, from that, right? So that, that's an important principle in this general ecosystem of the, the safety filters.
I just wanted say one thing quickly to, um, that the, just the thought that I had when you were talking before Jason, about people, um, about this debate over whether or not skilling laws will hit a wall or whatever, um, I guess it just made me think, so hearing you talk about that, like, to me, that debate, well, first of all, I don't think, yeah, I agree with you. I don't see any evidence they're gonna hit, hit a wall. Um, um, I think like it's just as interesting to talk about in all cases.
Like, like, okay fi so I, I don't know what's at stake in in, in saying we're gonna get to a GI and by 20, you know, 40 or in five years. Like, I think what's, I think it's more important to think about what we should do now and what the algorithm should be for addressing AI safety. And like to me, like we're, our basic situation is we're in a cloud of uncertainty around, um, how fast progress is gonna go.
Like, yes, you can predict this is a little jargony, but you can predict the air, you can predict validation loss based on the scaling laws, right? Yeah. You can predict the error rates of your model in predicting like the next token based on scaling laws.
But, um, but we're totally surprised by the emergence of capabilities, right? Yeah. Like scaling laws don't predict when we're gonna be able to solve the exploit challenges that Sahan and I talked about just now, right?
They, they predict incredibly accurately next token prediction, but not accurate. Like we, but we don't, they don't, they don't give us that much of a guide to knowing when there'll be some sort of capability breakout. Um, and so, but, but I think so, but, so I think we need kind of an algorithm around like sensing the risk environment and responding.
Yeah. Like it's like we need like a OODA loop as an industry, you know, and, um, exactly like that's very tight, you know, and allows us to, to sort of sense the risk and respond to them quickly. And I think there's like a lot of debate about when we're gonna have like, human level intelligence and all that, that I think that's somewhat independent of just having a tight ood loop around safety.
Yeah. Agreed. Um, I would just, I would just follow up on that and say, um, completely in alignment with your, your direction here.
Even even Zuck, uh, said on the Doores podcast. Like at some point, um, if there is a dangerous capability in a model that made out trains, like yeah, there's gonna be a conversation about whether it makes sense to release that. Um, so it's good to hear that, uh, sense of caution.
And I agree. I'd like from a, from a, from a policy perspective, I would really like to see, and, and Anthropic would like to see like standard industry benchmarks of evaluating whether or not we've crossed some threshold and not being run by the companies, uh, themselves, but actually being done by outside auditors. Um, this works in the financial industry already auditors in the financial sector just like audit the books of companies and people are expected to follow the rules.
And I can imagine a world where, um, you have auditors for AI safety who are running standard benchmarks and sort of, you know, do the crash dummy test, uh, in the same way that we do with, with cars vehicle safety before products are released where we just know, you know, there's no risk with this model. It's fine. And we, and we can put it out the door.
Yeah. Let me extrapolate from what you just said, Jason, and that is that we have national test facilities and, um, the red teaming, um, exception that Amit talked about in policy, I think is terribly important. It gives us a huge advantage Yeah.
Um, uh, for our civilization because we can, um, uh, have less bureaucratic control from above. And yet I wonder if maybe just like we did with nuclear weapons and other things, we should have a, some one or more test facilities that are isolated from everything Yeah. Out in some desert somewhere.
And you basically, you each company sets up their systems a copy of their systems, and then all the red teams and all the other companies are invited to, um, to yeah. Come and audit competitively instead of from above by an agency. And that way it can't, whatever happens to it can be evaluated in isolation and is not Yeah.
Could, can't infect that company's original model. Yeah, it's a good idea. I, uh, there's several other things in this space that I think are super important for government to play a role in and across the world, not just in the us for example, the national research cloud ideas going through, through DC a bit right now.
And that would give academics and the broader community an opportunity to perform safety evaluations at a, in an environment, uh, funded by, funded by, uh, on a grant basis, uh, federal, federal dollars. So there's, there's an opportunity for that. And then that would, that could lead to these much more security conscious, isolated engagement environments that, that you're talking about.
Expanding on that idea, keeping on that tangent, one of the things you and I have personally talked about is model stealing. Yeah. And the idea, I, I've actually seen that I have a podcast called AI of the Law in you, and one of the cases is somebody stealing the actual dataset from another company and then extrapolating the model from the dataset.
Yeah. Distillation attacks isn't the other word for this. Um, so if you have a model deployed on, uh, the open internet and you're not monitoring the rate at which someone is, is querying the model, they can take from you a, a great deal of the model capability in a specific area.
Like let's say, for example, that your model is really good at, um, um, uh, rating, uh, you know, comments as toxic or something, the small, small feature that we're talking about here. Not, not a general intelligence thing. If you had a million samples of, of, of your model, uh, performing judgment on a, on a million different kinds of cases, and you let somebody capture all of that, that's a training data set for their model, their bespoke model that does exactly that one thing.
Um, so this is, this is a general pattern of abuse that we're seeing now, and it's, it's, it's been concerning. Um, this is by the way, publicly, uh, widely known. 5 being released as a chat bot.
Um, and that became a training data corpus for all kinds of, uh, lower, uh, lower lower rank adaptation, uh, fine tuning things that people did on open weights models. And so, um, it's definitely a concern people need to be aware of when doing these deployments. I guess he covered that pretty well.
Alright, good. The, I I wanna relate back to, to one of the, uh, podcasts that I did, um, about Air Canada. Did you guys follow about the chatbot in Air Canada?
Yeah, yeah. Uh, David, can you give us just a brief summary or should I do that and then you jump in? Oh, Well, I'm not an expert on the Air Canada.
Oh, that's okay. I, I'll give you the scenario and then we can talk through it. Air Canada has a chatbot on their site, and what happened is somebody who needed to go to a funeral talk to the chatbot and said, can I get my money back?
And the chat bot said, yeah, you got 90 days. Don't worry about it. You're shaking your head.
You read, you read it. Right. And so what had happened is that Air Canada's defense in court was, we're not responsible for what our chatbot said.
It's its own individual entity comments, thought It was worth a try. As I told you, there are these hyper intelligent, um, extremely agile verbal creatures that already exist, and they'll try anything including hallucinatory, and they're called lawyers. I think.
I think when we look, we look at this case, we're looking at a case of, um, uh, needing to understand the operating environment of a deployment. And this is, this goes to, this goes to, uh, Meta's presentation, and it goes to Matt's presentation. There's so much that we need to think about when a model's being put in an environment where it has an opportunity to make a decision.
And this is a case where the company did not constrain the out the parameters of, of what was allowed. And this is just another example of where we need this, these safety barriers on the output side to ensure that it conforms to the, the constraints of the company business use case. Good.
Yeah. One aspect to this is that, uh, obviously in that case a common sense, a source of common sense as a backup, um, is necessary if you're giving advice, advice to people that they are going to rely upon for real life decisions. Now, one way to do that is the way there used to be backup drivers in these Waymo cars, but there aren't anymore because companies try to reduce the most expensive thing, which is human beings on their payroll.
Um, now backing up a, an AI system or an LLM system with a completely different LLM system that can flag a human being to say, does this make sense? This is a layered approach, and it actually mimics some aspects of how our minds work. Yeah.
I love that framing. Um, oh, go ahead. Oh, well, I was gonna say something different.
You, you, you go first. Yeah. Yeah.
I guess I, I love that you brought up robotics because it's a, it's a really good, important principle for us to remember. Um, sometimes the consequences of even the wrong movement in a robotics context can be injury or death. And, um, if you have a large language model running a, um, a robot, um, I don't recommend this right now, but if you did, um, then, then, you know, there should be another model monitoring that, that robot's movements in the world and sort of like stopping or avoiding harm wherever possible, even if the core model says that that's the action that should be taken.
Yeah. Yeah. No, I was just gonna say about it's Air Canada, I think I knew the least about, I just knew this was a chat bot that went off the rails.
Um, yeah, I mean, I guess one reaction I have to that is I think that incident gets at like a, a deep problem in what's going on in AI right now, which is like, um, we don't know. Okay. So I think that when people encounter these technologies, at least for many of us, myself included, like that, well, when chat GPT came out, I think there was a wave of euphoria that came over the tech industry, right?
It felt like we'd, we'd sort of discovered, and I think this is true, like a new pillar, like, uh, you know, like of the way technology, like a new pillar of technology, right? That is just gonna become ubiquitous and could be applied everywhere. Um, but then there are these subtle limitations that sometimes block entire application areas of the technology right now, right?
Um, and we found this over and over in, in security at meta. I mean, we, um, like when the sort of LLM revolution started a couple years ago, um, we made long lists of thing of areas that, in that, you know, of manual processes we were gonna automate with LLMs. Um, I mean, uh, I'm just thinking back to that list and the number of Xs through this use cases actually didn't work out when we tried to apply them, is enormous.
It far outweighs the, you know, the success stories, uh, that we've had. Um, I mean, it could be that a customer service bot, I mean, so I, I've heard other stories of customer service bots, like a, you know, for some car manufacturer and it goes off and starts recommending the competitors, right? Or whatever, this sort of thing.
Um, I mean there's just like, there's an optical illusion that happens when you start using these technologies, um, that dissipates when you, when you try to actually realize business value in many of the cases, not all of the cases, I think, um, and there are just open like this problem of guard railing and LLMI mean, we're working on that, right? And we showed that in our presentation. Um, but this isn't like an open scientific problem, right?
Like we don't, we don't know how to control these models with the level of granularity we need for many use cases. Um, and, um, so I think, I think, you know, anyways, that sort of thing is a canary and a coal mine for me, you know, around like, what's gonna happen. Like, I, I'm a believer in AI and what I do for a living, right?
But I think, you know, there are gonna be ebbs and flows to, to sort of the level of excitement and, you know, and as we encounter more and more of these blockers, it makes me wonder, you know, sort of where's the eeb start in this particular cycle? I'm going to, um, go one more question before lunch, 'cause I know we all need to take a break here. I think one of the things, and again, this is because you and I have discussed the most, Jason, one of the things that we need to keep foremost in mind is bias and ethics.
And this is different than hallucination. I've tried to have, I've had people talk about that in the same breath. And those are two different things that what we're talking about is bias and ethics within the system itself.
Jason, you wanna start that off? Yeah. Yeah.
This is a super important topic and it, um, it doesn't get a lot of airtime at a conference like this, but it does get airtime, um, in the, in the public discourse quite a bit. And sometimes it's pitched as an either or thing, like you're only doing safety, um, at the expense of bias or vice versa. And that's not even, that's not true.
It's not, it's a false dichotomy doing, doing bias, uh, training as part of that RLHF that we talked about this morning. And having robust data sets for that and, and looking at the ethical implications of AI deployment are part, um, parcel of every one of the models that our, that our companies are putting out there. So, um, there's, there's amazing, uh, data sets out there that can be acquired as some of them are made by, by nist.
Um, so you've got, you've got the opportunity to address this in a systematic way. And I think some of the other ethics things that come in are, are concerns around power and, uh, access to models. And this is where I think like the national research cloud thing that I mentioned earlier is super important so that everyone, uh, across the public sector has the opportunity to engage with this, uh, frontier technology in a way that makes this a much more accountable democracy.
Yeah. I, I, I want to speak up also on the question of, um, personal property rights as far as your information is concerned now in the transparent society. And ever since I went to early RSAs, way back in the last century, um, there's been the discussion of, of privacy.
I am of the opinion that, that reciprocally accountable enforcement of your right not to be harmed is more important than your right to prevent anybody from knowing things about you. And therefore, if you have knowledge in a generally transparent world, you are more safe. So I lean towards light, I lean towards things being open, but having said that, uh, I'm a member of some communities that are quite incensed about the possibility that all their ip, either novels and books and things like that, or information about you being used in training sets without your permission now that one of the reflexes is to ban that and it can't be done.
But one of the things that Jaron Lanier talked about decades ago is that you don't have a right to keep other people from knowing things, but you do have a right to some interest in the benefits from the uses that they made of your knowledge. But in order for vast training sets to distribute to be accountable, the results of grass training sets the ais that use them, uh, for the benefit of these castles, for the benefit of these companies, in order for them to be able to distribute the benefit that they got some profit that they got to the, the, the people who contributed to the training sets you need, not just micro transactions, you need nano transactions, possibly even PICO transactions. And this is an area that of AI development that I think is, is underdeveloped and getting under use of resources because past efforts at Microtransactions failed.
I think I know why, I think I know why a, a microtransaction system, uh, micropayment system can work very well and very easily, but in the long run, if you're gonna be talking about the rights people have to their information, uh, and yet not clog the system with lawsuits, I think you need some way for them to benefit from these, their participating in these training sets with their blood type, with their genetic information, with whatever's being fed in, they, there should be some kind of an accounting system that they benefit from the use of all this information. Okay. Josh, can you get the last word, Josh?
If you got something? I Tend to do that. Um, no, that's just a, that's a fascinating idea and, um, I think there's, I think an outstanding research problem is how you attribute back to the contributors.
Um, so like right now we don't really know, so if we use an AI image generator to generate an image, um, I know, I think the jury is still out on how you attribute that back to the, the, to, to the training data that actually contributed to it. And It's impossible now. Yeah, It's, I mean, we, we just don't, like, that's a, you know, this is sort of research problem.
I'd love to see lots of academic literature on, you know, 'cause I don't think it's totally intractable. Oh, I think it's probably, there's probably always some uncertainty, like I think probably can't solve absolute, but I think, I think there probably, we probably could anyways. I think it's, I think it would be interesting to see more work in this, in this area Claude was doing, or was it, uh, perplexity or Claude that was actually giving the, the resources that had used for the inference.
Oh, Citations. Yeah. Yeah, yeah.
That's true. I mean, we're out of time, I guess, but it doesn't, they don't entirely solve this problem, I think by, by, by, by giving citations, although that's relevant. I agree.
Good. Yeah. Will you say thank you, please?
I, I enjoyed that.