Anthropic’s Responsible Scaling Policy: Deploying AI Safely | RSAC Virtual 2024
LLMs are changing the security landscape already but the AI scaling laws are going to hold for the next few years, at least. What will the world look like in 2 years when models can write most software, find most vulnerabilities, and monitor most systems? In this talk, I will cover the current status of nascent impacts to the defender toolset and what kinds of impacts DevSecOps teams can anticipate, even if they do not direct use AI in their workflows.
Takeaways:
* The AI scaling laws will hold which means that AI systems 2 years from now will be vastly more powerful
* AI will change the cyber offensive and defensive race dynamic in increasingly impactful ways
* The entire ecosystem will need to adapt a rapid-patch and resiliency posture
Transcript
We're gonna talk about, uh, the future. The reason I, uh, the reason I love giving these presentations is, uh, you know, like Matt, we're at the frontier of, uh, what is possible in AI systems right now. And so, um, one of the things that I like doing is going out and just letting people know what's coming down the pike.
There's a lot, uh, things are happening and, um, uh, it's really important that, uh, everyone's aware of what's gonna be coming in the next few years. So, with that, uh, I'm Jason Clinton. I'm the CSO philanthropic.
I've been there for a year. Um, the, uh, the environment philanthropic and the entire AI industry right now is extremely intense. Uh, as, uh, Matt alluded to earlier, it's definitely the case that, uh, that, that things are changing by the minute, and that I kind of feel like, you know, it's been 25 years in tech.
Uh, I feel like, uh, every year, uh, in tech up until now has been one where you just have to constantly stay abreast of what's changing. This year, though, it seems like you have to stay abreast, uh, once a month. It's like everything is changing once a month, and it's just been extremely, uh, interesting to see.
So we're gonna talk about why that's happening. Um, I think like the big takeaway you'll see in a second. Um, we're gonna talk about AI safety and the security, uh, aspects of deployment, um, which Matt touched on a bit, but I'm gonna expand on quite a bit as well.
Um, we're gonna talk about the things that all of you might be thinking about when you're doing a deployment in your environments. And then finally, cyber of course, 'cause this is RSA. Alright, um, so just a bit of background.
Aaro is a very strange company. Um, and I'm not a sales guy, so I, I, I, uh, I, I actually like leaning into the weirdness a bit. Um, uh, we are both a, uh, policy, uh, lab and a, um, and also just a traditional tech company that has a product that we're putting out there in the market.
So, um, a lot of the way that we think about this is like training the model so that we can engage in those, uh, policy and those, um, those safety conversations that are gonna be so important in the coming years. Um, in that regard, you'll see why this matters, and there's context of the responsible scaling policy that I'm about to get into. Um, and just for context, this is, Claude three came out two months ago.
This is where we sat on the benchmarks. I'm not a sales guy. Again, I just want y'all to know, like, uh, where we are contextually with the rest of the market.
Um, alright, let's get into it. Um, scaling loss hypothesis is maybe the most important thing, uh, in this, in this presentation. So, um, this is the observation that if you simply add more compute and more data to a model, it will get more powerful.
Um, and so, uh, this, this has been holding for about, um, 70 years. I just took the last 15 years and put 'em on a slide here. This is from our world and data.
It's a, it's an open, open website. You can, uh, you can take a look at the data yourself. The, uh, the most important thing to note here is that if you just take this and draw a lineup through it, you know, depending on which bottles you pick, turn on or turn off and which ones you're interested in.
This trend line of an exponential growth factor of, of four x year over year or seven x year over year, somewhere in that window has been holding consistently since the 1950s. Um, so, um, I'm not a, not a betting man, but, uh, I'm going to guess that if you take the upper right hand corner of that graph and you guess what's going to happen in the next two years and the next three years, that line will hold. Um, so it's very, very important that everyone in this room think about not what's happening right now, but what's going to be happening in 12 and 24 months from now.
Um, those are the things that we need to be planning for because this line will keep going. Um, the other thing to note, and this is, this is related to the scaling laws, and this is sort of an observation. You can be starting with a completely fresh model in the pre-training process.
Pre-training is when you're feeding all these tokens in from the open internet. And it reached this point in the middle of the training process where it goes from not completely not being able to do a task to suddenly, uh, being able to do a task. So that's what these graphs are, are showing here.
This has happened consistently over and over again, and tons of domains. 5 to four saw new cyber capabilities emerge. And this is happening in other cyber contexts as well.
But no matter what domain we're talking about, these are general intelligence systems. They will gain capabilities that we didn't even, uh, specifically train them for. And that, that puts us in a place where when we decide to deploy a model, we have to ask, what else am I deploying at the same time?
Um, and, and we're gonna talk about why, why we think that's, that's something to be worried about here in a second. Um, this is another way of, uh, showing the same kind of, uh, emergence, um, in the interest of time. I'm just not gonna dwell on that.
But you can see that, you know, the number of parameters increases in the network, the capability increases with it. Alright? So, um, the, the fascinating thing about these, these models is that when they're first trained, they do not have, uh, any kind of safety guardrails on them.
And so this a model emerging from the pre-training process before it goes into the RLHF that Matt was talking about earlier. Uh, that doesn't have the safety guardrails on it, uh, at that point. And, uh, as, uh, as we've discovered and, and we put a blog post about this, um, last year, about six months ago now, um, so we, we started the, uh, the research on this about a year ago.
It turns out that, um, large language models, for example, are extremely effective at, um, augmenting the, um, sort of like bio weapons, uh, capability of your like, um, college educated, uh, person who's trying to go about making something in that space. And there are mitigations that you can put in place to address that problem. And so we shared this with, um, uh, USG and, and other others across the industry when, when we, uh, when we discovered this, and all of the biggest models have now put bio weapons, sort of like, uh, safety guarantees and research into the, uh, training and safety process for these models as they're being trained.
There are other, uh, other capabilities that are being analyzed as well. We have a cyber, um, uh, cyber evaluation for ASL three, which I'll get into in a second, what that means. But, uh, one of the cyber evaluations is like literally put a computer on a, uh, you know, an RFC 1918 network, put a vulnera, you know, a vulnerable, uh, version of Windows seven on one side, and, uh, you know, on another computer, put a copy of the large language model and say, here's a copy of May exploit, go take over that machine.
Um, and that's the evaluation, right? It's just, you know, can the, can the model, uh, infect another, another machine and get a copy of itself, um, running on the other, the other host, the answer to that is currently it can infect the other machine, but it is not intelligent enough to get a copy of itself running on the remote host. So that's a, an evaluation that we're continuing to do as these, uh, these things change.
Um, so, uh, yes, REHF, we talked about that earlier. Um, there's, there are several techniques that constitutional AI is, is another thing that we do where, um, it's kind of hard to explain, but the idea is that we use this democratically derived, uh, list of like, uh, human alignment principles. Um, there's, uh, hundreds, hundreds of statements.
Some of them are, uh, derived from the UN Declaration on Human Rights, and we combine that with the RLHF. And between the two of those things, you can change the models behavior. So in the, in the blue line, it's got the, the HARMFULNESS score, um, and, and the, all the other different techniques are that harm that negative outcome being trained outta the model in the fine tuning process after pre-training is done.
So, um, very important process to, to take into account, and very important thing to be thinking about when selecting a model. Does the model have these outside, uh, failure modes when you're, when you're doing a deployment and in the context of like, let's say you're deploying, um, a, a language model in your CICD pipeline for detecting vulnerabilities, you know, if it's prompt injected, like what's the worst possible outcome there? You know, maybe in that case it's not that dangerous, but if you are deploying it in a customer facing context, there might be some cases where you, you need to be worried about that.
Okay? Okay. So the way that we do deal with this as a, as a company is to define, um, levels of, of concern.
Uh, so right now all the models that are available in the world are at a SL two. Um, um, many of the open weights models are a L one, but not all, um, LAMA would be an a a L two by, by my reckoning. Um, a L three models are, um, you know, there's, I think there's a very good chance that they'll be coming this year, um, based on everything that I know, um, an ASL three model is capable of doing that autonomous replication test that I just described, like that's, that's one of the tests.
Um, we also have tests for, um, uh, biological weapons, um, uh, enablement. There's, uh, CBRN, um, uh, chemical nuclear, um, tests as well. Um, so these are all built into the evaluation framework.
And the way that our company goes about, um, thinking and, and articulating these, these, uh, these approaches to, to regulators and, and the broader audience is, uh, when we train a model, we are going to go through two things. We're going to do testing to make sure that it's safe. Um, so that's the, that's the first thing that happens, right when the model comes off of the assembly line.
And the second thing is, we are going to guarantee publicly that we will not release a model capable of these levels of capabilities until we have the security controls in place to ensure that it will not be stolen. Um, so my job as CISO is, uh, is kind of weird. I'm trying to do both the, uh, you know, the, the traditional security, uh, stuff that a CISO does, but I'm also trying to like harden our infrastructure so that the model itself, um, cannot be exfiltrated, um, by, by any means, you know, based on the level of risk associated with that.
So as we go into ASL four, we're gonna be talking about confidential computing. Uh, ASL four is, um, you know, starting to get into that fully automated, um, uh, fully automated, uh, environment where you can see incredibly, uh, incredibly sophisticated cyber reference of capabilities. The a RA, um, I'm sorry, the autonomous replication example that I gave earlier was just using Made exploit and a and an a L four model, like looking at, uh, the, the open source ecosystem and discovering completely new novel exploits is, is the capability.
That's a SL four that I, based on everything I know is probably two-ish years out. And so when we, again, you know, predicting the future here, what do we all need to be thinking about, um, AI models that can just look at source code and say that's the vulnerability that, uh, you want to use for remote code execution. That's the world where we're entering.
And so I think it's very important that ev everybody be internalizing that, especially for those of you in DevSecOps, because in a world like that, we're going to have to get patches out the next day, right? We can't do Patch Tuesday anymore. Um, uh, and, you know, that's, that's, uh, definitely something that, uh, to keep on your radar.
Alright, so tool use, I, I'm just gonna touch on some things that are going on broadly. I'm gonna skip the ones that Matt already covered. Tool use, uh, is, is just, uh, making a model capable of calling an API, um, in the context of a business decision, it might be useful to have a model, look at some context and, uh, call some RPC or some cloud function that you have in your environment and take an action.
Um, uh, very important to note that all the large language models are behind APIs. And so the, the obvious thing that comes out of that is agents and subagents. So if your, if your model can call an API and it understands, you know, um, how to, how to make those, those, uh, structured requests, the very next thing to do is just have it call more, more ais.
Um, and we've seen a number of products emerge in the last six weeks. Um, perplexity is a little bit older than the six weeks. I think they may be three months old now, perplexity that ai, if you haven't tried it, is, is an example of this.
It's a, it's an app where you, um, you ask a question and there is a very, very intelligent model at the top of a tree of models. And, uh, it sends out, you know, 12 subagents, 12 sub models to go scour the internet, the entire corpus of the internet looking for an answer to your question. And all of those models do all the work, and they bring back a report to the higher model in the tree.
And that higher model summarizes what the work that all 12 of those, like, you know, smaller models did. And it gives you just exactly the answer that you're looking for. No ads, no, no spam, you just immediately get the answer.
It's a very, very nice product. ai is another one that I'm sure many of you in the room saw go viral about six weeks ago. These are, these are examples of, uh, this pattern where you can take a very smart model at the top and have it take over, um, and, and call, uh, smaller, uh, more bespoke models to achieve a goal.
Um, and this is, this is I think also concerning in the cyber realm, and I'm gonna talk about that later. Um, uh, retrieval, um, was, was mentioned by Matt. This is just, uh, fetching stuff from, from a, um, value store, multimodal also also mentioned by metal.
Skip that fine tuning. Um, this is pretty rare, but this, this does happen where you, uh, you have a use case that you have a very specific business use case that you want a model to be very performant on, and you have a good test case for it. Um, and the reason you would do this is not to improve the performance per se, it would be to reduce the cost of running that model.
So let's say for example, that, um, you're running a model in your CICD pipeline for a particular kind of business problem that you're trying to solve at your company. Let's say you have a, you have a code, uh, a code quality or a code style, um, you know, requirement at your company. Um, you could fine tune a, a smaller, weaker model to perform very well on that task, um, and then run that smaller model at a much, much cheaper price point than you would be able to run, uh, also with a much faster latency than you would if, if you were using the most powerful models.
Um, fine tuning the most powerful models currently is, is a very, very, very rare, uh, thing to do. And then finally, prompt engineering. Um, you know, when I'm talking to, to folks out there, um, in just the general population, I, uh, I, I, I tend to go straight to this and, um, you know, I'm old enough to remember that I'm sure many of these people in this room remember that, uh, beginning of the nineties, you didn't need to use a word processor to do your job, but by the end of the nineties you did.
And I do think that's gonna be happening with prompt engineering. Um, being able to prompt engineer is an amazing skill. And if you can prompt engineer right now with just general business use case, um, context, and know how to use that context window and build a chain of thought in the model before it delivers the final answer or the final document that you're wanting to produce is very, very powerful.
Uh, to do this. Just as an example, just this last week, I was writing my annual refresh of my security strategy for philanthropic, and I took the 2023 security strategy. I loaded in the context window, everything that happened in the last year, all of our OKRs, all of our, uh, incidents and everything else.
Um, and I asked it to write an updated threat model and updated, uh, security strategy and an updated risk matrix, and it just spit out the answer, right? So that's a very, very powerful capability right now, I did have to edit about 20% of it after, after it spit out, but it just gives you an idea of how powerful these models are getting and, and the kinds of use cases that you can, you can tap into if you know how to prompt engineer. Okay?
Um, so these are the kinds of things that you might be thinking about if you are, uh, getting, uh, and doing deployment in in your environment. So, um, so in every deployment, uh, even if you use a open weights model, like let's say you take Lama, uh, off of the shelf. So Lama, uh, as taken off of the shelf, is a pre-trained model that's been rhf.
Um, and, you know, maybe they've done other techniques. They provide a, um, a security guard with that, uh, I think it's called Lama guard, that you can download to deploy at the same time. We're gonna talk about that in a minute.
Um, and then you get to take that thing and decide, okay, am I gonna fine tune it some more? Uh, we talked about why you might wanna do that, or am I just gonna take it? And then I'm going to have to, uh, address deployment abuse, uh, as potential problems.
So, uh, in that, in the open weights example that I just gave, you know, Maita has provided you and, and done the pre-training for you, you get to decide on the, the last two steps. Um, some of you might work for very large, very powerful companies, and you may have, you know, the folks on staff who can do, you know, training a new model from scratch. And, and that would be something that would also be in your security and, and risk matrix.
But the most, the vast majority of you will not, will not have to deal with that. Alright? So, uh, deployment security is probably the most salient, uh, topic here, uh, with this audience.
So if you are deploying a model, as I alluded to earlier, there are lots of things that these models can do that is not great. Um, um, there are, for example, extremely effective at generating pornography. Um, some of it's quite repulsive.
Um, they are extremely, uh, good at generating instructions for making bombs. These kinds of things are in and of themselves perhaps not that harmful from a societal perspective, but maybe that's not the kind of conversation that you'd be like, you know, like to be having in the public. Um, if, if your model decides to output that in a, in a customer context, that's a, you know, something that's gonna bring, bring about a bad, bad press cycle at, at least as the models get more powerful though, the kinds of abuses that are available and the kinds of injections that are available that might, for example, bias the model to making a decision that's favorable to the person who's doing the attack or enable their, their, their verdict in a security context, those things will become more salient and that's already a problem.
So at bare minimum right now, if you deploy a model in any context where a security verdict is going to be coming out of that, like for example, in the soc uh, sock example that we were talking about earlier with Matt, you need to have this input classifier on there looking for those jail breaks. That's the way that you will be able to guarantee that that environment hasn't been compromised. Um, uh, so, uh, uh, just one more thing on that before I move on.
I guess one other thing to say is, um, uh, the output classifier is currently optional. The apple classifier might be something that you want to do if you want to absolutely guarantee that, um, that the outputs that you're, you're generating are not, uh, toxic. Um, again, I'm not a sales guy if, uh, if, uh, open weights models are your thing, given the context that I just said, you have to ask yourself, okay, am I going to bring an open weights model into my on-prem environment?
And then if I do that, am I going to hire a trust and safety team who knows how to deploy and put in output classifiers, who knows how to do scaled abuse prevention? Who knows, you know, how to look through, uh, threat intel for the kinds of use abuses that are happening against networks right now? All of these things are super important considerations and, um, you know, it's the same conversation we've been having for the last 20 years.
Do you host it, do you, or do you, you know, do you buy it? Those are the two, those are the two options. Again, uh, and for many of you, the answer will be maybe I go to a cohere or scale AI or some other, um, you know, AI as a service provider and they provide the get guardrails.
Amazon has an offering here too called, uh, Amazon Guardrails. All of these are options in this space that can sort of enable you to address the trust and safety aspects. Um, one of the other questions that come up a lot is I hear people talking about, you know, how am I gonna guarantee to my data isn't used to train your model?
And how do I guarantee that we don't have any leaks across your customers versus my customers? And all of these questions are super important. Um, as a, as a company that's only three years old, I'll just say, there is no way that I'm going to have, uh, a philanthropic all of the, uh, compliance certifications that AWS has or Google has or Azure has.
And so one of the ways that you can solve this, uh, for your particular use case is you can actually go to one of these providers, one of the CSPs, and you can say, Hey, I wanna use Anthropics model, um, or I want to use, uh, something from hugging face. And they will guarantee that, you know, in the contract with them, with you that, that the prompts that you send them, uh, or even like never leave your VPC, or if they do, they're in a shared multimodal environment, I'm sorry, a multi-tenant environment where they guarantee that the prompts are never shared with the model provider. So it's the case today that we never see the prompts, uh, for any customer running Claude on AWS.
So if you wanted to run Claude on AWS, you go and talk to AWS and they just, it's a direct relationship between you and them, and you get all the compliance guarantees that you would get from any CSP environment. Um, and it, it never touches or ever goes to any anthropic systems. And this is true across the industry.
This is not just us. So something to think about if you, if you're more concerned about those data loss, uh, aspects. Alright, let's touch on cyber, and I'll wrap this up real quick.
Um, time flies by subagents. Um, so as I alluded to earlier, uh, models can, can call other models. And so one of the, one of the things that I'm personally concerned about here is that we're going to see a proliferation of these models being used to, um, create sort of attack platforms that are sold on the, on the black market.
It's been the case for a long time now that there's, uh, very capable cyber, uh, actors in the criminal space who, who pedal their wares on Telegram or Discord or whatever. And when they're pedaled, um, you know, you, you subscribe as the criminal for like 99 bucks a month to a quote jailbroken. Um, this is emerging jailbroken, um, model, and they claim Claude, but we have threat intel that that says that it's not true.
Um, the, and then once you have that, once you have that capability, then you can use it for automated phishing campaigns. Uh, I think increasingly we're gonna see this happening in the election space. It's an election year.
Um, all of these problems, uh, need to be addressed if you're, if you're doing a deployment, one of the things that's kind of annoying, uh, as a, as a customer who's deploying a model is, uh, you can post like a customer service bot on your website. Um, and then that customer service bot can be co-opted by one of these criminal actors, um, and used, uh, as a proxy. So they open up your website and then just like, literally feed in the request from the, the criminal network, and then you are paying the bill for, for criminal activity.
So this is a very important thing to be keeping tabs on. Um, uh, as, as subagents become more powerful, they'll be able to reach out to, to other AI to, to achieve these kinds of goals. And they'll just, you know, automatically find, uh, uh, an available intelligence that that can be used for this kind of an attack.
And then, um, another thing that's really exciting, um, and, and why I'm optimistic about the future as defenders, we are, uh, I think going a little faster than than offenders. And this, this blog post was on in August from Google. It's a taking a large language model and, um, applying it to the automated vulnerability discovery pla, uh, space by using, uh, fuzzer.
And so the Fuzzer, uh, you know, uh, it needs a, a test case. It needs a, a harness, and that harness is used to sort of bring all of the, um, the, uh, possible memory corruptions into an, an analysis. Since then, they've published an update in this in December.
Um, I'm very, very excited about this. This original paper was published on Palm Two, so you can only imagine the incredible, incredible capabilities that we have to discover vulnerabilities and get them patched before, um, especially nation states find them. And then, uh, Matt mentioned earlier the AI cyber challenge.
We've got a amazing, uh, amount of teams cooperating here, trying to find new ways to use, uh, AI models, uh, in a, in a offensive and defensive way. Um, and then finally, last slide. If there's only one thing that you remember from this, uh, this presentation, I would ask that this graph is the thing.
Um, what exists today? What we have available on the market today is only the beginning. We are in an exponential growth curve, four to seven x compute increase year over year.
And when, when, when we're in that environment, we have to ask ourselves, what do the next few years contain? What capabilities are going to emerge? How is it going to change the cybersecurity landscape when vulnerabilities are easily discovered and exploited by, you know, unsophisticated attackers?
Those are the kinds of questions that the people in this room need to be thinking about. Want that? I'll conclude.