Security in Al – What to Watch for in the Next Three Years | RSAC Virtual 2025
The panel covers both the excitement and concerns surrounding human capital in the AI sector. Trust in teams is crucial as organizations manage data security and non-human identities. The evolution of AI adoption is highlighted, marking the shift from simple chatbots to complex systems. Participants emphasize the need for accountability in AI actions and the importance of maintaining work-life balance. The conversation also touches on AI’s role in code review and security, as well as the challenges of scaling AI applications.
Transcript
I have to say that the human capital that's present right now on stage has me excited and a little bit worried. I'm assuming you trust your teams, everyone's doing their job because you guys are literally at the forefront of the new operational imperative moving forward. And it's not to be underestimated, especially for those of you who have been in the data security or cyberspace for a really, really long time.
Um, so Jason, I'm just gonna jump into it. You know, Gartner had this interesting stat that about 80% of orgs are gonna face challenges managing non-human identities coming up. Mm-hmm.
And by the end of this year, potentially, most of the threats we're gonna be experiencing is really about machine and machine interactions, um, especially around virtual AI collaborators. Mm-hmm. And as we really pivot into, you know, autonomous and agent ai Yeah.
Um, what gaps do you currently see within some of the jurisdictions and the laws that we have that you're personally advocating for, that we kind of have to put in place? Because, you know, there's always a delta between innovation and how fast lawmakers catch up. Yeah, that's a good question.
So, just, just to contextualize here a bit, um, maybe I'll just do like a, like a ten second interview. I'm the CISO of Anthropic, so I've been there for about two years. Um, the, if you just think about the way people are adopting ai, they start with chatbots and then they move to something that's sort of like RPA, where you're, you've taken like little, little nodes in your, like workflows in your companies and you've put AI in that little node.
Um, and then as intelligence of the models goes up, as you see this exponential curve of intelligence increasing, you can take those nodes. You can think about the intelligence, um, being able to compress or collapse those discrete processes into one continuous sort of contextual. Yeah.
The same way humans work through things. You're just like, oh, well I tried a, that a didn't work. I'm gonna try a prime, you know, just go through the graph of, of things that you might do in your workflow.
And if we, if we give more autonomy, we give more, um, uh, decision responsibility to the, the agents or the, the models or the virtual collaborators, or eventually in virtual employee, imagine a world where you onboard an AI with memory and you send it through your onboarding class, and then you give it, um, a starter project and it has a reporting manager and email address and a Slack account. Like all of that's gonna happen, you know, sooner or later. Um, then the question becomes like, how do we have accountability for, um, the actions and transparency and visibility into what's happening?
I think a lot of the way that we think about risk management in companies and, um, you know, when, when companies are working between companies and trying to figure out who's responsible for what parts of the infrastructure is, uh, do I even have transparency into what happened? Like, do we have the logs? Do we have the right, uh, pieces of metadata recorded?
So let's say for example, I request an, um, a virtual employee to do a bunch of software changes. Everyone in this room can imagine what that looks like because we're already seeing the coding revolution happening, right? Um, so it works on it for a week and it gets to the very end and it does something it's not supposed to do.
And that environment then, um, I want to know as the security team who asked for that change, who was the, who was the agent acting on behalf, behalf of who was the manager of that thing? And those are questions that there's like core technology to build and core, um, sort of like audits, uh, to build, to understand, uh, from a, from a, from an accountability perspective, like where, where things have flown and what, how we can, how we can understand and react to those things as security teams. So I think that's the big missing piece is we don't have the right technology bits to follow, uh, everything all the way through the infrastructure.
I think that's a key point saying that we don't necessarily have the right technology bits. And it's interesting 'cause I do think that we're gonna pivot from a place of where AI is our co-pilots to, we're gonna to us pivoting into managing a bunch of AI cockpits. Um, so you're saying that the accountability is gonna be on the individual overseeing those cockpits?
Mm-hmm. Okay. And I did something really, really silly because in my world, I know who they are.
They're kind of our, our celebrities of the modern day world. But I think Jason, you appropriately pointed it out, you probably all should introduce yourselves in case there's one or two individuals who don't know who you are. Um, I know we have the titles up there, but I think for everyone in the room, yes, name and title, but two, that pivotal moment where you pivoted from a practitioner to a thought leader, whether it's a moment in time in your career, or whether it was something that you wrote or invented, what was that pivot that put you into the place that you are right now?
And then third, a quick tidbit of how on earth you keep sane with the pace of innovation and where we're supposed to find our zen knowing how fast things are moving. And ultimately everyone in this room is accountable for keeping the ship floating. So Jason, how about you kick us off?
'cause I think that's just really important to me. Yeah. Wow.
Two big questions. Um, when, when did I pivot to leadership? Um, I have a science and science fiction book club that I've been running for 18 years.
And one of the things that we do on the science side of that is read, um, anthropology books. And I read, I think my fifth or sixth anthropology book. And I just got to the point where I'm like, you know what?
There's a lot of, uh, opportunity for organizations to work better than they're working today. So why don't I put my hat in the ring and try and figure out how to make that happen? And then how do I stay sane in this world?
Um, uh, I don't recommend my methods, so, but I'll, I'll be honest. Um, I honestly, the last, the last two years have been totally unsustainable in a very strange and, and, um, weird way. So my husband and I, my husband's also in ai, he's on the Gemini team.
Mm. Um, which actually makes it great because yeah, we don't have to unhappy hours. We have to like, argue about AI taking all of our time.
'cause we, neither of us have any time at all. Um, so we just put our personal lives on hold for the last two years, which I, I highly do not recommend, but mm-hmm. That's basically all we've been able to, to do to keep, keep her head above water.
Yeah. Wait to keep your head above water, you've given up personal time. Yeah, yeah, yeah.
Just wanna make sure I heard that correct. Yep. Yeah, Don't recommend it.
Yeah. I can relate to what Jason just said. Um, okay.
Pivot to thought leadership. Um, yeah, I guess that happened. I, I, so now I work for Meta now, but, um, for, I don't know, eight years or something.
I worked in the national security space and I was a principal investigator on some like DARPA sponsored research projects and some NSA sponsored research projects. And yeah, as a go going from sort of, you know, researcher to principal investigator, you sort of have to, you have to learn how to lead a team and, you know, yeah. Um, um, be a thought leader, you know, and, um, conferences and government meetings and this kind of thing.
Um, trying to remember what other questions were. I was staying sane. Um, yeah, I ran the big SIR marathon yesterday.
Um, so thank, so if I seem so, so if I seem less coherent than normal, then that, that, that's probably why, 'cause everything is, is aching at this point. Um, but, um, yeah, I'd say running, um, is definitely, um, one of the main things. And then I have two small kids, um, uh, 4-year-old and a 9-year-old.
And they definitely keep me grounded, keep us and saying, And your beloved wife is here. Partner is correct? Yeah, my wife is here.
She keeps me sane too. She's there right there in the Big clap for her. Um, Yeah.
Okay. That's the main, thank you. Yeah.
Marran. Hi. Can you hear me?
Great. Well, we'll start with, uh, first of all, I'm Ashkenazi, I'm Jfr cso and I've been doing cybersecurity for 25 years. Mm-hmm.
I think that the pivot was like, uh, two and a half years ago when I started, uh, a venture capital for cybersecurity and really support, I understand the power of the cs o global CSO and how we can help to a very early stage startups, uh, start coming, invest in them, uh, help them, assist them to grow and, uh, also support them with an innovation with, uh, the warm up stage. And that was like, uh, the change. I understand how, like, the power of it and that, uh, I think that was the, the point that I understand that it, it's bigger than, than everything.
And, uh, following what Joshua said about the family, I think that, uh, a supportive family is definitely, uh, you know, help with that job. And also the ability to change and to help organizations and to make, make, you know, even a little bit the world, a little bit a safer place. It's just, it's a mission and support my team.
I think the organization, build a stronger organization, be there for them, support them. 'cause it's, it's, it's a, it's a hard work. Thank you.
Very cool. And that Leaves me, I'm Matt Knight, I'm CISO at OpenAI. I've been at the company for just about five years.
It'll be five years this summer. I joined as the first security hire to build the security program. Been doing that since then.
Um, question first was thought leadership. Right? Okay.
The pivot. The pivot, pivot. I still consider myself a practitioner.
Uh, and I hope my team does too. I, I, I don't know what what thought leadership is. I just try to do good work and, and to the extent that I, you know, learn things along the way, I am, I'm happy to share that too.
Um, and I've had the privilege during my time at OpenAI of being able to lean in on experimentation within the program, right? Finding ways to use language models to aid in our work, to find ways to help the team be more, uh, more productive, um, be in more places, move faster on things. Um, and, uh, that, that's one of the things that's, that's been most exciting about, about my time there is finding, uh, really making first contact with these tools and finding ways in which they can help us.
And with regard to how to stay sane, nobody gets into security because they love the status quo, right? Like our, our industry, our work is, is defined by disruption. It's the very fundamental it to the, the extent that there's anything fundamental or foundational about it, it's change, right?
So when you consider that this is really just, uh, business as usual. Wonderful. Well, thank you so much.
I appreciate that. And thank you for the friendly reminder. Yeah.
Um, like I said, you guys are celebrities in my head. So, and you know, our world is really interesting. There's no shortage of information coming at us, whether it's LinkedIn and articles and people calling themselves AI experts.
I don't resonate with that word at all, regardless of how many years under the belt. I don't know if you guys do, but I think we're all still real time practitioners trying to figure it out. Um, so if there is a space or time, I would say that follow the goats, if you will, the people that have been there from the very get-go who aren't speaking by reading and regurgitating, but they're living, eating and breathing it.
Um, tremendous props and respect to you all. Now Josh, quick question. Um, there was a recent stat by MIT and it was published just the end of last year.
About 70% of LLMs are vulnerable to prompt injection attacks. It's funny, 'cause if you take a look at everything that's coming out, not many people are talking about prompt injection attacks. Could you do us a favor and explain to us what you think that means and the definition of it, and how meta is approaching resiliency, um, into the models to be able to mitigate that risk?
Sure. And yeah. Um, okay.
Yeah, I was also gonna ask Jason a question about what he said, but I'll, I'll I'll save that till later. Yeah. Okay.
Um, yeah, so actually, I, I should correct, I feel like I should correct the MIT sta mean all, all LLMs are vulnerable to, to prompt injection. Um, yeah. So first I'll de define what that is and then talk about what we're doing, which is probably not that different from what folks at philanthropic and an open AI are doing also.
Um, okay, so, so actually how many people already know what prompt injection means? Um, okay. It looks like almost everybody, but, uh, so just, um, just for completeness, I'll, I'll give a definition.
So, so prompt injection happens when you can caate an untrusted input with sort of trusted system programming in an LLM. And then the LLM, um, um, well, prompt injection is successful when the, the LM then follows the instructions that, that are given in the untrusted data. So to give a practical example of that, um, you know, so if I, if I, you know, program my LLM to take a system prompt that says, you know, be, I don't know, a helpful web search agent, um, you know, that's like the, that, that should be at the top of the instruction hierarchy.
So the system should always obey that, that system prompts, right? And then if the user says, well, it goes, go search for like, where I should take a, my vacation in France this summer, or something, um, that would be like, you know, in the instruction, in the sort of hierarchy of privilege that next prompt should, um, you know, take, take sort of next precedent after this, after the system prompt. com or whatever, like that should that, that, that, that, that should be overwritten, you know, by the system prompt, which is like, be a, be a helpful and trustworthy, you know, uh, like web search agents.
Um, like I I sort of fundamental problem with the technology right now is that we just don't know how to, how to ensure that in all cases, the, the large language model will sort of respect that instruction hierarchy. Um, and, um, so this is like a real problem with, you know, a number of cvs we've come out over the last year or two, you know, with like, real systems that are like really deployed in production, um, that, um, succumb to these kinds of prompt injection attacks. Um, so yeah, I mean, there's a few ways in which we're dealing with.
So, so one, one of the things that, um, the teams that I work with at, at, at meta, uh, are responsible for, are making sure that out of like the meta's whole universe of products, which is fairly large, um, in which AI is being integrated in lots of different places, like we make sure the teams aren't shipping products with like, severe prompt injection vulnerabilities, that, that affect our, our, our users security and privacy. Um, you know, a few things, um, that we do, um, are one, just refrain from using large language models when they're not really necessary. Um, and there's these prompt injection risks.
So you, you, you, you really just wanna not have, um, like, uh, non-deterministic risk risks in your application where, where possible. Um, so oftentimes that means just like not using an LM and using traditional procedural code depending on the product feature. Um, another thing we do is, is sort of restrict the privileges that we give large language models.
So, you know, what, what you don't want is your large language model, like processing a messenger over Messenger or WhatsApp, and then, you know, going and changing account settings sort of downstream of that as a function of like LLM decision making, right? Um, so we wanna like basically refrain from using LLMs in like sensitive cases like that where we can, um, and then there, you know, there's often like residual risks that we just can't mitigate. So where, you know, you can't both get the benefits of the AI technology, um, and also have zero risk, right?
So like an example would be like a research agent that goes out on the web and like does a bunch of research, um, and, you know, each decision it makes about, like, which sort of next piece of content it looks at, looks at, um, is made as a function of some previous piece of content that it looks at. Um, there, there's like, you know, there, there really are with the sort of research agents that are getting shipped in the industry right now and risks that the control flow will get hijacked. Um, so there we do like fine tuning of our LLMs, uh, to make sure that they, you know, well not make sure, but, you know, to reduce the likelihood that they'll get hijacked, uh, by a malicious instruction.
We also have like system level guardrails, so like machine learning models that sit outside of the model that scan untrusted content coming into the context window of the LM. Um, and if they see something suspicious, don't let that into the context window. Anyways, I could go on about this, but this is, you know, I think this is a big active area of research.
Um, well, I'd be curious to hear about the other folks on, on the panel also, but, um, yeah, that we're, that we're, I was gonna say for all of you guys, I mean, based on the state of ai right now it's in research, but as more and more folks go into agent ai, autonomous ai, um, and that's not a pivot that's really happened with an enterprise just yet, it's still at its infancy stages. What's that lifespan of it needing to get out of r and d and into a productionalized motion? Yeah.
Do I know you wanna jump in on this one? Briefly chime in. It, it really depends on your use case and your application.
Uh, AI is software. We've been managing risk in how we build and ship software for, uh, you know, for, for decades at this point. And it's about using the appropriate tool for the job.
So, you know, two years ago when I started, um, sort of sharing some of the work that we were using LLMs for within our security program, you know, the responses range from like, interesting curiosity to just like being aghast by the, the, just the, the notion that you would use an LLM in a security context. But the reality is that there are many, many places where language models are totally appropriate. Think like, um, you know, just something as simple as like summarizing what happened during an incident.
You might have a Slack channel of your, um, you know, your incident responders talking back and forth and, um, you know, there's, there's, um, uh, there's information, there's a timeline, there's decision making. And, you know, do you really need, you know, one of your, your detection engineers, one of your, your most valuable people, like writing a book report based on what happened, or do you want them, you know, being able to use a language model, have that do the first pass, have edited for correctness, and then, uh, and then, uh, disseminate that and move forward. It's a way of like reducing toil from the team's work and helping them move faster.
And, you know, in that use case, like, I hope we can all agree that there's like very little risk to, you know, something going wrong. You, um, you're not actuating anything. You have a human providing oversight and checking the results for correctness.
Um, and this is something that we've been able to do, you know, for, for quite some time, um, using technology that was even a couple years old. So as the technology improves, we'll find more, um, ways in which we can incorporate it into our, our work and, um, you know, uh, find more places to take the hands off the handlebars. But, um, you know, that's sort of where we started.
Um, there's a lot more that we're doing today, but, um, it's just one thing I wanted to, to, um, reinforce is that at the end of the day, like we're building and using software, um, we can, uh, start with, uh, you know, there are plenty of places to start that are, that are low risk and expand from there. Um, so I think, I think one of the things I wanna, uh, like emphasize, uh, and maybe to answer your question about the research to deployment life cycle as short as possible, um, everything that we just described, and Josh's, uh, very excellent explanation of the vulnerabilities, the state of the art today could be solved. Like, if you just think about the way that we human beings are neural networks interact with the world around us, it's sort of like, you know, you watch the Star Wars where, where the guy goes, these aren't the droids you're looking for.
And that's like effectively what's happening with these large language models, right? They, they get the jailbreaker, they get the prompt injection through this sort of, what, what appears to be magic. It's, there's no reason a neural network needs to be vulnerable to these problems.
We just haven't figured out the right answer, um, from a, from a research perspective. And on Friday, Dario put out a call for action on his a blog, uh, saying there's like an urgent, urgent need for something called interpretability research. And what this is already yielded is the ability for us to reach into neural networks and actually find a specific neuron responsible for behavior.
And Dario said in his post, there's a strong possibility that we can actually systematically solve jailbreaks and prompt injection with this technique. Like think about a a, a world where as the neuron, the neuro, the neuro, um, the neural network is executing the, um, the, uh, the, you know, over the token stream that's coming in, you can see a specific neuron that's like, oh, I'm, I just received an instruction. I should follow the instructions.
You know, you could just see like the lights turn on from a n uh, like an actual and introspection's perspective. And that would be an indication that your untrusted input cross that threshold from being context into a command that the neural network is now, now following. Um, very, very exciting early, um, research in that space.
And, um, there's so many other things. OpenAI has the prompt coloring work and the prompt hierarchy work that they've published, uh, through, through scientific papers. So there's a lot of stuff that we can do that can make this a lot better.
So I'm very optimistic that we can, we can make a con, confidence can be something we can share with enterprises. I'll take it to the, to the production area, the notion to production. Uh, I totally agree that it's easier to adopt LLM especially when it's like it's internal stuff, it's internal data, uh, but everyone are talking about MCP, right?
Everyone wants integration. Everyone wants to prepare the, the business for it, get ready because it's easier and added value when the business is like, really, it's a crucial, crucial mission. And there is like added value to the business to adopt it.
So I think it's not a matter of of time, it's just, it's gonna happen. That's it. So security need to adapt it, need to reinforce it and, and support it with the best guardrails that we have today.
And as the research will developed, we'll adapt more and more technology move from the spiritual, the guidelines, the best practices, the audit that we can perform to much more practical, eh, solutions. Um, and, and, and I think that it's, it's, it's something that is already happen happening. Uh, I can tell you that the early production step that we took was around cybersecurity and also customer success, customer support replaced the chat bot with something much more autonomous that really drive the business fast and, and make things easier.
Yeah. Wonderful. Um, I'm gonna go a little off script here and ask all four of you guys a question.
And I purposely did not prep you in advance because I genuinely wanna know if, if there's a difference, gener, do you think there is a difference between what we classically know as InfoSec and what's now being called cyber tech? So information security, which is part of a vertical within enterprise data management. Mm-hmm.
And then now everything, I don't know if it's a Glossier version and we're calling it cybersecurity, where if you think it's actually a separate vertical. Um, but I'd love for you guys to define the difference if you feel there is a difference between InfoSec and cybersecurity. I can take that.
Yes. First. So, 'cause I'm here for 25 years since it started to call, like, uh, InfoSec and then everything's changed to cybersecurity.
I think that the mindset is the take it to the practical level from procedures and policies down to earth, to prac to practitioners. And also when security, uh, took in charge of the DevSecOps area, the product security, I saw so many CSOs that didn't pass the, the transition and hold the product security take in charge, give it added value, support the r and d, the product and engineering with great tools that collaborate with them. And I think that the notion to the DevSecOps and the cyber security in, and the product security themself did the transformation because it's, no, it's no longer information security.
It's, it's everything. It's very holistic. It takes, uh, several, you know, core partnership with all the different collaborators and make it something very effective.
So I think that was the notion, um, the DevSecOps, the tools being collaborator and, uh, and actually accelerate the product. Now I can tell you that it's, it was like 20 years ago when cybersecurity were blockers. Like, no, first of all, get, you know, get cyber security or InfoSec approval or block, uh, with firewalls, approve the rules, those things that we don't hear them anymore, right?
Mm-hmm. Like, allow me to do something. Let's prove that it's not, it's not longer.
Um, the mindset, we are accelerators, we are, uh, business partners, um, the organization come to cybersecurity to get the best advice and, and partner with us, uh, because they, they see the added value. 'cause they, you give them trust and you are a partner. Your technology, you understand faster than sometimes even faster than the p and e.
Yeah. And take it to the next level. Thank you for that.
Matt, Josh, Jason, do you think there's a difference between InfoSec and cybersecurity? Or is it just a semantic evolution? I don't, I don't get too hung up on definitional stuff.
Okay, perfect. I mean, yeah, me, me too. I guess I, I just remember like when I was a kid in like a high school, like hacker calling security cybersecurity was seen as like a, a tell that you had no idea what you were talking about.
And now it's like the accepted term, which is, it's interesting to see. And same in machine learning. Like, you know, for a long time calling ML AI was like this weird thing that you did, and now we're all doing it.
So I don't know. Um, but, um, I, I always wanna say, okay, so this is a little, little bit of a pivot from your question, but, um, I do think there's interesting new things happening in our field due to what Jason was talking about with respect to the need for like an identity. So like, if we assume that we're gonna enter a world in the next, I don't know who knows what the timeline is, but two to four years in which we have AI colleagues at some level, you know, where you can talk on Slack with an AI that is off programming, you know, and solving your programming problems and then sort of ping to you when it has a question and needs resolution and this sort of thing.
Um, yeah, I, I think we will need, um, I think this is what you're getting at. Like, we, we'll, we will need like a, like a modification, um, of our like, identity infrastructure, right? To like, um, where like, it's like clear, like, you know, and there's some guarantees around like, um, like the audit log that like, you know, this agent was act acting on Josh's behalf, like when it commit, when it like, you know, whatever this pull request and this and this, this sort of thing.
And like, um, I do think security will change if, if we, if we accept that like, like we will like sort of move pretty quickly into this world in which we have, um, AI assistance and colleagues doing things on our behalf with increasing autonomy, um, like authorization, authentication, identity. Yeah. You know, and, um, um, how we deal with sort of least privilege with respect to these sort of like colleague like entities.
Like all the, all the stuff will have to change. And I think that'll change the shape of our field. I don't know what we call that, but Yeah, I, I, I like the question, and it sort of ties back to what Josh has said here.
I think, uh, David Bryn was, uh, on the stage here last year and he had some stuff about ecology. Um, and I think when I think about like, the evolution of the security practitioner over the last three decades, it's moving from like being in a silo to being integrated with the entire ecology of tech. And as we think about virtual collaborators, they're gonna be, uh, we will try and we will to some extent succeed at like defining boundaries and trying to put least privilege around where we, where we can.
But I don't know, if you just look at, um, something that's I've been thinking a lot about lately is MCP is exploding in a, in a very, uh, it's crazy, very fast and very rapid way. And that's like, I'm just gonna give you access to all of these things. You know, my, my own personal assistant running on my, my machine is access to all of these things.
Um, and, uh, we are gonna have to be very flexible. Uh, 'cause things are gonna move very fast. Um, and as we think about cybersecurity and cyber, cyber tech, and all of the things that are gonna be happening over the next few years, um, this, this evolution, this co-evolution of the security practition and, uh, the technology is gonna be really important that we stay embedded and we stay, um, uh, even federated with product teams as they try and move fast.
Wonderful. Thank you. All right.
Um, we have a few minutes left. So I did wanna ask Matt, maybe we'll start with you. I think for any of us, whether you're in startup, enterprise, mid-level, there's always this constant battle between speed of innovation while securing your ecosystem.
Um, I haven't seen anyone who's gotten it like correct or right, and I don't even think there's a right answer, but the question I have for you is how are you guys approaching, um, especially within your position, that balance between speed and security alongside development to make sure that there are processes and protocols and documentation, which is very time consuming while doing the work. It's kind of like you're flying the plane while you're building it at the same time, but you don't have enough folks to do either or how are you guys handling that? So I'll talk first about, um, an effort that OpenAI does around its model releases, um, uh, to test and evaluate models before we release them.
And then I'll talk about how we incorporate, um, perfect. Uh, that's sort of those, so similar, the, the principles that you were asking about within our security program. So first for OpenAI, so, um, you know, OpenAI is a mission driven company.
We're, we're here to make sure that AI benefits all. Um, we take that mission very seriously. Um, one of the things that we do before we release models is we test them for a series of, uh, capabilities and risks.
Um, we have a testing methodology that we have, um, uh, published. It's called a preparedness framework, uh, that, uh, spells out the battery of tests that we put the models through before they get released. Um, there are a number of categories that we, we test them for.
Um, cybersecurity is one of them. Um, and the testing is a mix of, uh, sort of, you know, repeatable automated testing, and then some, uh, that we do with, um, with expert red teamers. Um, all this factors into, uh, attempting to get a holistic picture of, um, of, uh, what these models are capable of before we re we release them.
And this is an evolving science, right? It's, it's one that, um, we expect to continue to evolve. And if you want to know more about our, me our methodology and our approach to this, um, I recommend that you, um, a read our preparedness framework, which is published online, and b, take a look at some of the system cards that we published along with our recent models.
And system cards are, um, artifacts. They're sort of like, uh, you know, data sheets or reports that, um, describe, um, uh, describe sort of, you know, what, you know, what's behind the model and, uh, and what, and what, what, what is it capable of. And they extensively go into these, um, uh, these tests and results.
And I think they're really interesting. Um, actually just last night I was rereading the O three, um, system card report, um, uh, in some of the, the work on cybersecurity, which if you stay for, uh, by talking a little bit, you'll, you'll hear a little bit about. Um, and, uh, I, I think it's super interesting personally.
Um, so I hope that gives, uh, gives some perspective on how we approach it, um, from managing model risk. Um, with regard to our security program, you heard me sort of allude to it earlier. Um, you know, we've taken a, um, sort of a qual crawl, walk run approach to how we incorporate these, uh, these tools into our work.
Um, you know, we've, we started with a number of, uh, you know, started years ago. Um, it would be malpractice for me not to attempt to use language models to help our program just because they're so powerful and we have, um, some really incredible tools and we wanna wanna use them. Um, and some of the, the first, uh, use cases that we've started with are ones that, um, you know, are pretty, pretty, pretty evident.
They were right in front of us. Things like, um, uh, you know, I, I shared that, um, uh, incident summary, um, earlier. Um, we have a number of other, uh, use cases that we've expanded into, some of which are a little bit more, uh, more interesting.
Um, I'll talk about some of those a little bit. Um, don't wanna steal my own thunder there. Uh, but I hope that that answers your question.
That, um, both from, you know, the model evaluation side and how we incorporate that into our work on the security tank. And what about that balance between speed and the innovation? 'cause it feels like you guys are releasing models every month, um, outside of having a giant army.
Like how are you guys, what's the mental mindset of balancing the two? So I, I don't think speed is necessarily a dirty word, right? Mm-hmm.
I think it's about doing the work and, um, you know, um, I think you did a great job of, of saying earlier that like, you know, it used to be that security was like the blocker at the end of the Yeah. End of the road. It's, you know, there to sort of, you know, uh, deliver, deliver, uh, be the hook, deliver justice, or, um, yeah.
Uh, you know, judgment at the end of it. Um, and, uh, yeah, I think part of, you know, doing this well is doing it, you know, in a way that's, you know, efficient, repeatable, and, um, helps the organization meet its goals. Okay.
So that begs me to ask then I could just jump in real quick. Oh, yeah, please. Uh, a couple more thoughts.
Um, so, uh, I do think the Matt's point about using the, the technology to actually enable the speed, like the innovation enables the safety and the speed to to be done in, in, in parallel. So there's a couple things that are going on in this space that I think are really fascinating from security, protection, protection, uh, practitioners, uh, side that I think are worth, worth calling out. And everyone in this room who's in DevSecOps, I think should, should be thinking about these things.
Um, so, so the first one that I wanna call out is more than half of all the coded andro now is being written by Claude, which is, um, mind blowing. And I think probably by the end of this year, it's gonna be closer to 90% of all code. Um, when we think about the actual code that introduces bugs, we want these models to be producing correct, um, uh, code to begin with.
So there's a question of does the model generate correct code, which can be done through fine tuning? Um, there's a question of using models to find bugs in the, in the code that was written by models. So there's like a, a second like peer reviewer aspect of these things.
And Anthropic has adopted, uh, AI for code review internally. Like the, you know, I think we, we think of code review as this gold standard of like, you know, you know, your peers like, oh, there's the bug, I'm gonna make it, you know, can make it make it so it doesn't shift to production. But I think we all know that like, human code review is maybe like 25% effective or something like that.
Maybe that's a spicy take, but, um, if AI is 90% effective versus 25% effective at finding bugs, like that's a massive uplift Yeah. In our ability to stop shipping bugs to production. And I personally would like to go back to being a software engineer and, and step out of being a CISO eventually.
And if we can just stop shipping bugs to production, I, I wouldn't have a job anymore. So I think, I think I'm really looking forward to that. So, Yeah.
So it's funny that you mentioned that I am gonna ask a question at the very end around intellectual atrophy as a result of leveraging large language models to do a lot of our work. But just a quick show of hands. I know we've all heard about a variety of methods in which we can use LLMs and various applications of artificial intelligence to write the code, but within your organizations, how many of you are using AI to write your code for less than 20% of your systems applications within your, the totality of your ecosystem?
Few hands. Okay. Less than 50%.
Less than 80%. Okay. Or I would assume the rest.
And you guys have completely allocated all of your code writing to LLMs, or you don't wanna answer the question, how many of you guys are not using LLMs to write your code yet? I think that's the better one. Okay.
Fantastic. Thank you for that. Um, Miranda, quick question for you.
It's so funny because I think in our world right now, we hear about what amazing things the startups are doing, what amazing things the tech companies are doing, um, but enterprises are still fundamentally struggling Yeah. With scaling any application of ai. And like, if you take a look at the stats, and I'll share some of that in my keynote, it's, it's not really optimistic.
Um, from your perspective, what are some of the practical things that you think are prohibiting enterprises from actually being able to take these remarkable capabilities, um, and scaling them within the enterprise? Yeah, I'll start with the bottom line. The bottom line is, is trust.
Yeah. It's a matter of trust and regulation and friction between the ready, um, uh, the technology is in compared to how fast we wanna do that. Yeah.
So I can share a practical thing that we did in order to scale. We build internal skeleton when everyone were a part of that, the legal, uh, that take and understand, better understand and educate and train their teams to adapt AI to really understand the barriers and way we need to, to change our mindset instead of like the friction of we don't know how to handle that. So it's around training.
Um, the second thing is the p and e, the r and d that really understand what is going on and how much it cost. So it's a matter of cost as well. Yeah.
To try to assess how much it cost, what is the ROI and the business value, and then security come jump in and said, okay, we need to train our people as well, uh, educate ourself and be able to use AI in order to do diligence the models and create some kind of skeleton that everyone can live with, and then give it to r and d to do the magic, to build their features very fast. And once you have some kind of structure like that, there is like a committee, uh, AI committee that support it. So if someone wants to create new idea and create new feature, it's really a ramp up.
It's a very fast to do that. Yeah. With a skeleton in place.
Wonderful. Now, yeah, I was gonna add. Oh, it's okay.
I have like a thanks. Um, yeah, like, I just wanted to add on the question of applying AI security. So, so, um, I feel like I've lived through elite like two years of, of us trying at meta to apply large language models to various types of security, like sort of various sort of bread and butter and security problems like finding bugs and code.
We have like this huge legacy and code base that we, you know, and or like automatic program repairs, so like finding and fixing the bugs and this kind of thing. And, um, I, I think that, I think large language models are, are particularly slippery, slippery, slippery kind of technical object because they're totally open-ended. Like, like you interact with the chat bot and it feels like you're interacting with a person and you, and I think, I think the people's default mental model when they're just sort of going into this is that, oh, well we should be able to employ this everywhere.
And, um, like we've had brainstorming sessions in the security department of Meta where like, we like list out like all the, all the ideas around how, where are we gonna apply the, the lms and like, it's like literally people just are, engineers are just putting in like everything, you know what I mean? Like, uh, yeah. And like 90% of those cases don't work, you know?
And like, like what I found is like if you have a lot of hands-on experience, uh, applying LLM security, you can sort of eliminate 60% of those cases, I think just at priori, like based on your own intuition. And then like, like another 30% have to be eliminated through like sort of failing fast and rapid prototyping and then like you left with like 10% where, um, they actually work and like, there's a lot of value and you can actually automate stuff, you know, but it's like very different than like buying some sort of like B2B SaaS solution that like does, like you're hiring that project to do like a very specific job and it's like a mature area. Like, um, so I think that's a big, big challenge with adopting these technologies.
Yeah, I Just add to that, that by the end of the day, AI agent mm-hmm. It's a software we need to treat it as, although there is a model behind it that we don't know what it's gonna do and cannot predict what the, their prediction will do. By the end of the day, as a security practitioner, we want to implement security within the layers exactly like we are doing with any other feature.
Yeah. Like secure them, scan them, put the SBO to understand the dependencies, which models they're, we are using due diligence, the models understand that it's not a malicious one. Um, be able even to sign it, like make sure and confirm that this is a model that has been through and with evidence with at the station, that this is the model that we, we tend to to use and nothing, no one changed it along the supply chain.
So it start with the models, continue with the diligence one, the one that, uh, privacy and legal approved and, uh, read everything and, and feel comfortable with that is not illusion from security perspective, from red teaming, uh, uh, issue. Definitely do all the tests that we can from, you know, tools, automatic tools, and also with code review, uh, use AI to code review things that we cannot predict with the human, you know, yeah. Uh, review and then all the way just to production.
So implement that with the layers. Using that as a jumping point to this, 'cause we're gonna be wrapping up, I have one philosophical question for everyone, and Matt, maybe we'll start with you. If college students are using chat GPT to write their essays for them, if we have practitioners who are using artificial intelligence to summarize and research and audit for them, if we have software developers leveraging large language models to write code basis for them, do you think that as a civilization, as humanity, we could be experiencing intellectual atrophy?
Because if you have not mastered your craft, AI can amplify your abilities if you are subject matter expert or you've mastered your craft. But if you're still learning a craft and you leverage artificial intelligence to accelerate your work, I can't help but personally feel that we're bypassing the learning process and short circuiting the very thing that makes us stronger. Um, so I call it intellectual atrophy, kind of the numbing and the dumbing of our intellectual capabilities.
If we just aren't fingers to keyboards, do you think that's a possibility for us? It's a really interesting question and it's aptly timed. Uh, just two or three weeks ago, I spent, um, spent the weekend back at, back at Dartmouth, my alma mater speaking with, um, you know, their leadership about, you know, what it, you know, what it means to train future leaders to, to harness this technology to look forward.
I wanna first look back, um, you know, when mankind invented the abacus or the calculator or these, you know, computational tools that are now ubiquitous, we didn't stop doing math, right? We started doing higher of math, it unlocked more frontiers, and it allowed us to be more creative, to think to, you know, to to look at further horizons and, and do more and move forward is, um, and, and move forward. Um, I I think this really, um, uh, in, in a similar lens, right?
Um, so rather than, you know, um, I think as you, as you'd frame it instead, I think it puts more emphasis on our ability to reason, to think critically, uh, to use what it is that makes, makes us unique and powerful, um, to use these tools to the fullest extent. And they can help us with a lot of the toil, a lot of the, the, the drudgery, a lot of the legwork to enable true creativity on top of it. And that's, that's how I've been thinking about it.
Mm-hmm. Jason, do you wanna give it a shot? Yeah.
So, um, uh, will we end up in the wall scenario where everybody has a floating recliner and they're just sitting a surfy? I, I, I, I don't think so. I think, I think what I, a lot of, a lot of these public speaking engagements after afterwards, um, there'll be like an 18 or a 20-year-old who's going into computer science come up to me afterwards, and they have like a look of terror on their face.
Um, and I think my advice to them is, is to really lean into what they're passionate about. If they were, if they're just going into computer science because they, they thought it was a high paying job, then maybe it's not the right fit. But if they were doing it because they really love the computer science or they, um, they think that the, uh, this is a, a place where they find curiosity and excitement, that's an opportunity for us to, to Matt's point about creativity and, uh, lateral, uh, movement in the way that we think about problems to be more, uh, of an opportunity and less of a, um, a pitfall.
Hmm. So there's, there's this, this, uh, this opportunity. I think like, you know, let's say that every software engineer is a manager of a team of AI employees in five years.
Um, okay, well, you need to be good at management. Like, yeah, if you, if you've been thinking about like, well, what does it mean to, to have, uh, good cohesion between, um, uh, you know, coworkers and, uh, clarity on road mapping and understanding of where we're going, um, as an organization and what the business needs are, you can't ignore those things in your, in your education anymore. You need to be actually focusing on the, the big picture and not just the narrow, like, how do I write code?
Um, mm-hmm. Yeah. Wonderful.
All right. We're over time. So 30 seconds if you wouldn't mind, please.
Oh, okay. Yeah, I was just gonna add, I mean, I think I agree with everything that was said. Um, I mean, like for me personally, I've found, um, generative AI to be like a bicycle for my mind.
And I've learned more than I would've had I not had that. And, you know, I, I can get personalized instruction on any topic in the world, you know, and I mean, you know, um, that said, I, I do think that we have an opportunity, like, like we have a responsibility as tech leaders to sort of guardrail, like, I mean, I'm sure you guys all agree with this already, but like, I mean, so very locally, you know, um, so if somebody, somebody on my team is using an LLM to like write the lit review for a paper or something, like, you know, we need to create a culture in which they're accountable for every word of that, that lit review. And like, there's like, but there's all sorts of cases in which this needs to be guard railed, right?
Like in terms of the way education gets shaped, you know, like in, you know, I mean education I think right now is set up in such a way that's probably not like, um, sort of optimally adapted to, to this new technology, you know, in which students can screenshot their homework and get all the answers and this sort of thing. Um, so I do think we have a response. I think the technology is powerful, it has an enormous potential, but also we need to think about ways to sort of nudge it in the direction and set up sort of guardrails and institutions that sort of adapt and bring out the best.
And then, Okay, I'll just mention take us home really quickly. Yes. Really quickly.
Just give it a metaphor. And, and the way I see it, it's still just a co-pilot. It's not a pilot, it's just, it can accelerate, it can help you, but it's not a pilot.
It won't replace the human mindset. The challenge, the innovation, the passion, it can change that. It just can accelerate, it can really fasten the current education, uh, of reading so many books and just like summarize it and curate a deep research and, and learn fast.
But it won't to replace like the pilot, same goes for the CISOs. Um, in my mind it can really help with, uh, create some kind of AI incident commander that will help us and advise to do that very fast. But in times of critical incident, you are the pilot.
Perfect. Thank you all so, so much. And thank you For attending today.