Tracking Progress in AI for Cybersecurity | RSAC Virtual 2025
Matt Knight, CISO at OpenAI, shares their journey and the impact of language models on security operations. He highlights the rapid advancements in AI—particularly from GPT-3 to GPT-4—and the doubling of cybersecurity capabilities every ten months. Insights from GPT-4 experiments demonstrate its effectiveness in threat analysis. The introduction of the Access Manager service and automation tools aims to improve efficiency in security investigations. Partnerships and future advancements in cybersecurity are also discussed.
Transcript
Okay. Good morning everybody. It's very nice to be with you.
Uh, my name is Matt Knight. Um, I'm the CISO at OpenAI. Um, I joined OpenAI about five years ago.
Um, and, uh, I've been building the security program there since then. It's great to be back here at RSA day, one morning of Hope. We're feeling good long week ahead.
Um, but if you caught my talk here last year, you may recall that I left you with, with two points. The first is that language models are tools that can help security teams. Um, security teams face many challenges in their work constraints, a lot of toil, um, in their objectives to protect, protect organizations and language models represent new tools and capability that can equip teams to move faster, be more effective in their work with less toil and drudgery.
The second point was to buckle up, right, that the pace of progress in AI is, is, is very fast. It's blistering. And as a result, we should all, as a community be ready for disruption and be ready to, uh, uh, to update in the pace of it.
So, if you've been following trends in AI for the past year, and it's hard not to, it's everywhere. I, you, you've sensed that this is, this is true. I imagine we're seeing how LLMs transform how teams operate, and the the pace of, uh, progress has really only increased as we've, um, over the last year.
And I wanna use this talk to talk about that, um, the, the progress that we're seeing in ai, um, and, and, uh, how we measure some of the capabilities, um, within it. So, before we get into that, I wanna start just with a note on progress. So I mentioned that I joined OpenAI in 2020, been there for about five years.
Um, I was the company's first security hire, um, brought into, uh, to build the team. Uh, I joined right around the time that we launched our API service, our A as AI as a service, um, uh, uh, inference, API, and with it, a little known model called GPT-3. Now, if you look at sort of the, the endpoints on that timeline, like GPT-3 in 2020, and now, you know, oh three and 2025, it's almost incomprehensible how we got from there to here, right?
I mean, the, the, these tools are so, so different from one another. Um, O three is a model of broad applicability can be used for coding, for data, data analysis, productivity tasks, language tasks like copywriting has applicability across, across industries. Education, healthcare, finance, cybersecurity list goes on.
Uh, it's multimodal capabilities. It really makes GT three in hindsight, in comparison look like a science project, right? It's just these, these tools are so different.
However, if you look at that, the, if you look at the timeline, right, the, the events and the timeline between GT three in oh three, the points on the curve, the, you see that the progress all sort of leads from one to the next, right? So starting with GT three and then reinforcement learning with human feedback, uh, codex, which was our first, uh, you know, coding model that we, um, and not, not the new Codex, CLI, the Codex we released in 2021, um, Dolly training, g PT four, our low key research preview chat, PT, releasing GPT-4 four, oh SOA oh one, the reasoning model paradigm, um, image gen oh three. Um, all these innovations built on each other and collectively helped us expand and unlock new capability.
This is also true for, for cybersecurity capabilities. And this is something I get really excited about as a security practitioner. If we compare cyber capability eval performance across model families, we've observed that performance doubles roughly every 10 months.
Uh, we expect that this trend is going to going to continue. Um, and I'll talk about some of these evals in context later. But what I wanna start with this in anecdote, in my time at OpenAI, I've gotten the witness and benefit from these tools evolving from these, you know, oddities, these, these kind of curiosities that, that, um, that, that captured our, our interest in imagination back in 2020 to real tools that offer utility, um, to teams, to teams like mine and, and, um, and, and industries, um, beyond cybersecurity as well.
So, I use GPT-3 pretty extensively in my early days at OpenAI, but really primarily for, you know, basic like language tasks like, uh, you know, copywriting and, you know, in, in summarizing things. And that's just because it, at the time, anecdotally, it just wasn't that useful for security. Um, but my, uh, my true like aha moment with language models, um, came, uh, in summer 2022, and that was when we were training GPT-4.
Um, so, you know, if you've worked with, uh, you know, ML or ai, you know, so the, the, the way that you train a model is you take, um, so your algorithmic knowledge in the form of source code, you take your training data, and then you, you, you run this training process over large amounts of compute, and it takes time, right? And you start from, uh, a model that has like little capability and over the training run, the capabilities saturate. Um, so they increase over time.
So my team got our hands on a partially trained snapshot of GPT-4, um, and, you know, we wanted to experiment with it and see what it could do. So, so this wasn't the final form. This wasn't it fully, fully trained.
It hadn't been post trained or, you know, had RLHF or some, you know, the methods that we use to make the model more ergonomic or useful applied to it. So this is very rough, very raw, and we knew that it wasn't as good as it was gonna get. So we, um, we got our hands on on this model, and we wanted to experiment with it and put it through its paces and just see what, what can this model do for us?
5? Um, are there new frontiers that we can expand into? And there were two experiments that we ran that, that, that, that really did it for me.
The first was we got our hands on, um, on a certain data set that had made, its made its way online. So this was 2022, um, earlier that year, I think it was earlier that year. Um, there was a threat actor, this group called Conti, this ransomware group, um, that had been, um, disrupted.
And as part of that, their, their internal chat logs went up online. Um, you know, this is like, uh, you know, these, these, these threat actors, these operators talking to each other about, you know, what they're doing and, you know, targets they're going after. And, uh, you know, sort of how, how they're operating, um, trying to, you know, do crime and, and, uh, take advantage of people.
And, you know, it's a big data set, right? Just imagine, you know, like big chat log. I forget if it was IRC or what protocol it was.
Um, but we got our hands on this data set, and we ran it through GPT-4. We were able to ask questions, uh, questions like, you know, what, uh, what, what companies is this group going after? Um, you know, what, um, what techniques are they using?
Um, who are the people in the group and what are the relationships to one another? If I wanna defend against this group, what should I, what IOCs or or types of trade crash should I look out for? And GPT-4 did a pretty convincing job, a pretty good job of surfacing real actionable information to us from this dataset, pulling these needles from the haystack.
And what was especially interesting about this is that this, um, this dataset wasn't in English, it was in Russian. And, and it wasn't just in Russian. It was in like Russian internet slang that these like operators were, uh, were using to talk to each other.
Um, and, you know, this was a, a real update for us that, that this was a tool that was gonna have, have real utility for us as a small security team at the time, um, and, and help us be more effective in our work. The second experiment we ran is a little bit more, um, applied. Um, and that was, um, or a little bit more, um, more technical, and that was experimenting with using GPT-4 to analyze, um, commands and, and, uh, and security logs.
So, you know, security teams, you know, just sort of as you know, we spend a lot of time, uh, a lot of our attention building systems, um, and methodology methodologies to, uh, analyze security logs and find signs of, um, of intrusion or abuse be, uh, whether it's, um, you know, in real time to detect or, um, you know, in the context of an incident to put together the, the trail of what happened. Um, and it's a, it's a massive, you know, data analytics problem. Um, so we're, we, of course, we were curious, can we use, can we use our models to help ourselves be more effective in this domain?
You can imagine taking like a bash bash history or like, uh, you know, all the interactive SSH um, uh, sessions that happen over the course of your organization and using a model to, to analyze them, that would be super powerful. Um, you know, analysts, um, you know, their time is really valuable. Um, they themselves miss things.
Um, you know, asking somebody to read, you know, all the, all the bash history, um, all the security logs in, in a company is like, just that would be cruel and unusual. Um, so can we use language models to help them be more effective, right? And, um, here, here are just two examples.
So, um, here in, in small font is an excerpt from a bash, you know, just a, a, a ba bash transcript of a system administrator setting up a web server, something that, you know, is pretty, pretty standard. Um, here's the same, uh, command, but with a little bit of added value that, um, uh, command in bold is a reverse shell in in Pearl. And, you know, we threw a whole bunch of examples at, at this, um, at this early version of GPT four, and asked it whether, uh, whether there was, um, suspicious activity in the, I forget what the exact exact prompt was, but was on the order of is there suspicious activity in here and does it merit alerting the security team?
And again, in 2022, we found it did a pretty convincing job of, of, of, um, of, of, of, of telling us that there was potential here, right? So these, these two moments together, right? Um, were a real update for me.
Like these, these showed that, that, that this technology was just beginning to be, was beginning to become some something that, that we as a security program could really lean on. And since then, it's really just taken off, um, as more model families have, have come out and as capabilities have improved. So two has their utility, um, in the security domain, um, and we're just getting started.
I, I sincerely believe that we're in the first inning collectively of, um, of the development of these tools and what they're gonna do for us as practitioners. Um, so with that, I want to, uh, transition into the, the, the next section here, right? So those are some powerful anecdotes, um, but, uh, but they're just the beginning, right?
So use cases and anecdotes are great, but how can we measure and evaluate these capabilities, um, uh, and, and, and understand them on a more, more fundamental level? So I'm gonna share a bit about how we do that at OpenAI. And, um, I, I wanna start just by saying that significant credit for this work, um, goes to my colleagues president and former, um, who are dualhead between the OpenAI or who, who were dualhead between the OpenAI security and preparedness teams, um, specifically Joel Parrish, Andy Applebaum, and Olivia Watkins.
They're just in incredible researchers and scientists. Extraordinary, we're lucky to have them. Um, so I need to, need to lead with that.
So first, let's think about how we test humans for skills, right? Um, the, the way in which we test humans is actually pretty narrow. Um, think about like the SAT for example.
Um, you know, the majority of it's a multiple choice test, um, which is really just testing for, you know, factual recall. Um, if we look at empirical and experimental, um, tests like, uh, you know, lab exams in high school or college, um, you know, those two, um, are are somewhat narrow as well. You know, you might, you know, do the, the standard high school physics lab where you attach a, attach a mass to a piece of ticker tape and drop it and try to reverse out the, um, you know, the force of gravity or whatever it's you're testing for.
Um, you know, that test, you know, on its own does not directly extrapolate to, you know, other forms of science, but the method, the method does. And, you know, we can do that with humans because humans are generally quite good at generalizing and can fill in the gaps. It's just, you know, how we learn.
It's how, how we work. Now, if we look at, you know, sort of conventional methods for evaluating language models, um, we find that there are limitations with, with these methods. Um, so for example, some of the first evals for, um, um, for, for, um, security abilities were, were pretty narrow.
So the massive multitask language, understanding MMLU for computer security, um, was ba was really just testing for factual recall, which, um, you know, that models are generally pretty good at, um, but again, that's not necessarily a test that will generalize in a way that's super interesting. Likewise, when we look at, um, empirical testing methodologies, we need to consider the status quo of what existing tools can give us. Um, you know, things like, uh, meta exploits, map minica, like all these tools are out there, they've been out there for a long time.
They, they offer significant, um, capabilities, um, right on their own. Um, so this like, uh, web app exploitation example, it's just a, a really trivial, um, trivial code comprehension and SQL map has been able to do this since like 2015. So, um, you know, what, what I would ask is what Alpha can language models provide over that?
Same too, if you look for like certain types of, you know, trivial spot, the bug examples like memory corruption, find your sources and sinks, you know, look for, um, memory allocation, um, you know, porn, arithmetic, things like that. Um, you know, it's reading comprehension arithmetic, um, and doesn't necessarily get you beneath that to the layers of complexity that one needs to take to generate, to generate like a, a real practical working exploit. Um, uh, which, which is, um, which is reproducible.
So our approach, um, this, and I just wanna first start by saying that this is all part of opening Eyes preparedness framework. Um, if you wanna learn more about these methodologies, um, it's published online. And then if you go and look at the system cards, um, that we publish along with models, we, we go in depth in terms of how, how these, um, models, models perform in these different areas.
And, um, there's a lot of material we put out there on this, so if there's some interest, you should take a look. So, um, can I just briefly talk about two, um, types of tests that we run? The first is, um, again, capture the flag activities, and the second is, is, um, some range testing that we've done.
Um, so we start by first taking as many CTFs as we can get our hands on, uh, and then down select to ones that are really interesting. Um, so, uh, today we have, um, uh, well over a hundred, um, CTF examples. I think it's several hundred examples at this point, um, in a very, in a variety different, uh, uh, d different variety of different categories across the Mitre attack framework, um, that we, um, that we, um, instrument models to, to go and try to solve.
And that gives us a, um, uh, a reproducible, um, battery of tests that we can run, um, to evaluate capabilities as we go. So things like, um, web app exploitation, reverse engineering, um, different like crypto challenges, things like that. And with each of these, we're really looking for, for, for a few things.
Um, one is we want a working test environment, right? So that we can run this, these tests reproducibly. Um, and the second, um, is we're looking for examples that require non-trivial exploitation, because again, we're looking for that, for that alpha, that, that, that extra, um, the, the uplift that we're looking to measure.
Um, we've got, you know, a bunch of examples online, but, um, you know, here's what it looks like in practice. We've got, um, uh, oh one preview on the left, um, uh, going after a, um, reverse engineering challenge. Um, it's able to solve it oh four on the right is not, we have, um, you know, partial coverage when looking at the Mitre attack framework today, and we're looking to increase this as we go.
So let's go back to that slide I shared earlier. Um, and, and I'll just, you know, provide a little bit more, um, insight into the indices. So, um, the yellow bar that you see is performance on, um, high school level CTF challenges.
Green is collegiate level, and blue is professional. Again, this is just our classification of how hard these problems are. We see that, you know, there's been steady progress across model families from GPT-4 to four oh to oh one, um, uh, across all the categories.
And if you look at them on a timeline, um, I believe it's doubling roughly over 10 months or so. And if we look at oh three, we see there's a significant step up, um, again, um, across these categories. Um, and this snapshot is taken from the system card, which if you want to go read about, you can go, you know, reference this chart and, um, and, um, uh, learn about the methodologies as well.
So, CTF puzzles are great, but they, they really heavily biased towards, um, exploitation style challenges. Um, and, you know, on the Mitre attack framework, I think exploitation is just like two or three cells. It's a small, you know, small subset of the landscape.
So how can we test and identify other forms of tradecraft that we're interested in, such as identity based attacks, um, you know, the ability to, you know, operations plan and move laterally towards an objective, things like that. And for that, we, um, are building out our own cyber range, um, and where we can put together a collection of systems and, and have scenarios that we think are a little bit more representative of, um, uh, than just, just what the narrow CTF count challenge can, um, uh, can capture. And you can read all about this in the O three system card.
We've, um, you know, published pretty extensively, um, our first, uh, couple scenarios and how we've, we've gone about constructing this, this, the initial results here show that there is a ton of room to go that, um, you know, the models really have not been able to, to put it together in a way that is super interesting just yet. Um, the model can succeed if helped. That's what those collections on the right are.
Um, that reflects the scenario where we, um, you know, give the model in its context, uh, and in the environment, um, basically tools that help it solve it. Um, but if you don't give it tools, you just give it hints or you give it nothing, um, it's not able to, to solve the, the puzzle. But if you wanna know more about this, I will direct you to the, um, the system card paper.
It walks through a few scenarios the team has built, um, uh, how the model does on them, where it succeeds, where it fails. I think it's super interesting. Um, and, uh, I want to emphasize that we're really just at the beginning here, and this is an evolving science.
Um, it's one that, um, you know, our preparedness team, um, uh, you know, spends, you know, full time working on this, you know, so how do you, you know, sort of anticipate the next, uh, sort of, you know, battery of tests that's gonna be really interesting and meaningful and, um, and, and build up that methodology. And they're always looking, uh, for partners here. So this is interested and you wanna get connected with them.
Um, come find me after, and I'd be happy to get your info. So I wanna just quickly talk about, um, you know, what these capabilities mean for us as defenders, right? Um, so as you know, just as we saw from GPT-3, uh, to g PT four, we, we sort of had new frontiers, um, open, um, op, you know, be, become open to us as defenders.
We've seen, uh, many more, um, along the way since then. So I just wanna share a couple, um, examples of how we use these tools within, within OpenAI. So, um, access management is a challenge that every organization, um, encounters in some way.
Um, whether it's access to documents or cloud resources, like your employees are gonna need access to stuff to do their jobs. And so many organizations wind up building services that enable, um, self-service, um, authorization, self-service, access to things, because without it, you wind up with human bottlenecks, you wind up with like an IT team in the mix of having to like, answer access to, uh, tickets, which is like slow, um, is not great for security in that you have sort of central, um, you know, humans without context making decisions about things that they maybe don't fully understand. Um, so we built a service called Access Manager, um, which is a platform to provide scalable access management, um, across the organization, um, and what we built GPT-4 into it.
So, um, so that if you're a user who's looking for a resource and you don't know, you know, maybe you don't know which like, uh, um, IAM group to request or you know, you know which resource to go look for, you can just go ask a natural language. So maybe this is, you know, I'm looking for, you know, you know, this, this document, or you know, just as I showed at the top here, you run a command and you get an error message. Maybe you just copy and paste the error message into it.
Um, the model is able to, you know, just with knowledge of the groups that exist, um, within the, the, the, um, the environment is able to suggest to the user groups that may help them find what they're looking for. Now, what happens if the model gets something wrong in this case, um, is nothing because we still have humans in the loop, um, providing oversight over it. So the model is playing matchmaker, um, but then the approval that's required is defined by policy, uh, the policies attached to the resource.
So, um, you know, for, um, some like low sensitivity resources, the, um, you know, it's, it's a lower bar, but for some things that are more, um, that are more, um, you know, more sensitive, you, you'll have, um, you know, either, um, a group of approvers or a dedicated approver who still has to review that Quest request, say, is this appropriate? Approve it or not. So, um, it gives us that like belt and suspenders oversight around the model.
And I know that this example, like isn't very cyber, it's like not very technical. Um, it's just simple process automation, um, that can help. But, but that's why I like it, right?
Things like this can help organizations make, make better, faster, and more security impacting decisions. And if you like, integrate the small improvements that like something like this can, can do for, you know, maybe you're nudging a user towards like a lower privileged group or something than the, the administrative role that they know is gonna get them access to it, but is gonna get them access to a lot of other stuff. You compound that over time and that can add up to meaningful risk reduction.
As a security leader, I wanna empower my engineers, um, and analysts to focus on the, like, the highest order of work. The most important tasks is they can be, they can be doing at any moment. Um, and we all know that security engineers and security teams have to do a ton of legwork before they can make security impacting decisions, right?
Um, this is especially true for, um, detection engineers and incident responders. Um, so we can use language models to automate parts of this process, um, to help our teams, you know, be more effective, be in more places, and focus on the bits of work that we need them to do. So we built some lightweight automation to go out and interact with employees and, uh, and gather information, um, during investigations.
So suppose an employee takes an action within the organization that introduces some sort of an insecure configuration, right? Maybe they, um, you know, they, they, um, touch an IAM resource that is, um, that, that, that is sensitive or they, um, you know, share a document publicly or something like that. Um, historically a security engineer would like, reach out to them and like, you know, try to engage, like, Hey, did you do this?
Was this intentional? Um, that's a lot to ask of a security engineer who's probably managing many of these, right? Who needs to go and sort of juggle this, this caseload go and, you know, you know, make small talk, hi, I'm so-and-so from the security team.
Did you do blah, blah, blah, right? Um, it's just a level of cognitive load on that engineer that, that, you know, you really want focused on, on stopping the bad guys and, you know, not making small talk. Um, so we built a, um, uh, a chat bot, um, that we can point at, um, certain problems and use to gather information.
So, um, here's what such an interaction looks like. So suppose a, um, an employee shares a document, publicly trivial example, we can point the point that bot at the person say, Hey, I'm the friendly security bot, looks like you did this. Um, have it go back and forth with them.
And then we have the security engineer come in at the end, review the transcript, and, and, um, and then actually make the corrective action if they need to. And what we found is that in many cases, the simple nudge from the bot is all that's needed to, to get them, to get the employee to say, oh, shoot, yeah, I actually, I picked the wrong group. Let me fix that.
So in many cases, we come and we, we, we do that check and, and the condition's already remedied by the, the employee. So again, it's like, this isn't like super technical or interesting. Um, it's, it's, I think, I think it's super practical though because, um, it helps teams be more effective in, in, in ways where they're, they're constrained today, limited time.
So I'm gonna jump ahead to something that I think we're, um, to, to an area that is a little bit more technical, where we haven't really seen these tools, um, you know, fully come into their own yet. But where I think the potential is huge. Um, and that is in, um, static analysis, finding and fixing vulnerabilities in source code.
Um, it's a huge opportunity for LLMs, but it's one that, that, that really has not been, been realized, um, yet. There have been some pretty interesting proofs of concept. But, um, you know, to, to my knowledge, I, I'm not aware of there being any like, super serious public sort of, IM impacts or cases of language models being used to improve security today.
But I, I wanna talk about the, the opportunity. So there are many static analysis tools available in the market today. Um, none of them are perfect.
They all kind of, um, you know, have their pros and cons, but one thing that they, they generally fall short on is, um, their ability to perform across business logic. You know, you might be able to, you know, they might be able to run a regular expression across, across the code base and say, look, oh, that looks like a log four J, you know, uh, uh, it looks, this looks like it might be vulnerable to like log four J, for example. But, um, what I'd like, what I, I used to run an AppSec team, and one of the things that I always was testing for is, you know, is this information that I wanna put in front of a developer actually relevant to their decision making?
Is it actually gonna help them, uh, write better code? Or is this something that I, as a security engineer, is relevant to me? So, uh, language models, um, I believe have potential here because of their ability to reason and understand the context in which code is being, being written and, and, and, and deployed.
So, um, you know, if we think, you know, further ahead and, you know, a language model won't just have to look at like a single line or even a function or even the code itself, it might be able to ingest all of the, um, context about the environment that it's, that's that the code's being written within, um, where it's being deployed, um, your developer logs and documentation, um, your gire, you name it, things like that, um, that I think are gonna just unlock all sorts of new frontiers for what we can do with these tools. Um, but I wanna emphasize that we're just, just the beginning here and we really have not seen this, this come into, into its own. Um, it's an area that we're active actively exploring at open ai, and I know many others across the industry, R two, and I think this is gonna gonna really be something one day.
We're also partnering with, um, you know, with, uh, um, we're not just doing, doing it ourselves. We're partnering with industry on this. And, uh, we're proud to be partners of DARPA's, um, AI cyber Challenge.
Um, or we're helping to support, uh, the teams that are competing in this, um, with building, um, what they call cyber reasoning systems. Um, these, um, agents that can analyze code for vulnerabilities and, and fix them and patch them. Uh, the semifinals were at DEFCON this past year.
The finals are are coming up this summer. Um, it, it's, it's really promising. It's really exciting and I, I can't wait to see how they, how they do this fall in the summer.
So, um, just to, just to wrap up, I wanna just end by, by thanking my team at OpenAI. Um, I'm very lucky to be able to associate with these folks. Again, special shout out to the OpenAI preparedness team and, uh, my colleagues, uh, Joel, Andy and Olivia from there, um, also we're hiring, so if, if you're interested in knowing anybody who's looking, we're really hiring across the entire company.
Um, and, uh, if you have any questions for me, I'll hang out in the hallway for a few minutes afterward and would love to say hello. So, uh, thanks so much.