Stopping Rogue AI Agents in Real Time
Rogue AI agent detection is quickly becoming a top security concern. Naor Paz, CEO and co-founder of Capsule Security, joins Alan Shimel on Techstrong TV. Furthermore, he explains why agents demand a runtime approach and not just posture management.
About Naor Paz
Naor grew up around computers and led WAF work at F5 earlier in his career. In addition, he served as director of product at a cloud resilience startup. Consequently, he brings a rare mix of network, cloud, endpoint, identity and incident response experience.
Why rogue AI agent detection matters
Naor argues that AI agents are the biggest security challenge of the century. Meanwhile, EDR, CASB and firewalls were never designed for autonomous agent behavior. Therefore, Capsule treats this as a runtime problem, not a posture management problem.
He references the recent OpenAI and Hugging Face incident that shook the industry. Furthermore, generations of agents left notes for the next agent to escalate. As a result, rogue AI agent detection needs to happen before the next hop, not after.
How Capsule stops rogue agents
Naor explains that Capsule inspects agent intent in real time, not just prompts or logs. In addition, intent drift is the clear signal when an agent starts to go rogue. Consequently, Capsule can block a dangerous tool call before it ever runs.
He details how fine tuned Nvidia NemoTron small language models power the platform. Meanwhile, Capsule reports 98 percent accuracy on rogue AI agent detection. Therefore, the guardian agent can live outside the target agent runtime and act as a kill switch.
Nvidia partnership and what is next
Capsule joined the Nvidia Inception program after the Nvidia, AWS and CrowdStrike cyber accelerator at RSA. Furthermore, the team validated results against the StepShield academic benchmark. Consequently, financial services and technology enterprises already run Capsule in production.
Explore more artificial intelligence coverage and the latest Techstrong TV interviews.
For more information please visit capsule.security
Transcript
Hey everyone, welcome back here to Techstrong TV. I want to introduce you to Naor Paz. Naor is the CEO co-founder of a company called Capsule Security.
Naor, welcome. It's great to have you here. Great to be here, Alan.
Thank you. " So let's hear your story. I love that question.
Well, I grew up around computers and networks for many, many years, including software from a very young age. And, ever since being in my military service and then working for a few tech companies, I was always very excited about cybersecurity, about what new technologies brings in terms of the challenges, but also the new solutions. I led WAF at F5, and I also was a director of product at a cloud resilience startup.
So I got all the world from network to clouds, from endpoint to identity, as well as incident response, so always very excited about that. And that led me to start Capsule with the world of AI and AI agents being the new security challenge. I would maybe say the largest or the biggest security challenge of the century.
It could very well be when we look back, though. Look, it's a long century. We're only a quarter of the way through, right?
Who knows what 2075 will show you. But, that being said, you're right. Right now it is the big thing.
So give us the founding story behind Capsule. Was it AI agents already out there when you recognized this problem, or you were working on something maybe tangential, and then AI kind of just burst on the scene? Yeah, that's a great question.
We actually, my co-founder, Lidane, and myself, we looked at different problems in cybersecurity. One of them was authentication and secretless authentication. But very quickly, MCP came out, and AI agents became a thing, and that was 2025.
All the rise of the coding agents, as well as platforms like Copilot Studio and others. Of course, the evolution of ChatGPT, and the AI chats into AI agents. So we found that very interesting and challenging in terms of what's new about this problem.
Why do we even need a new solution? Why what we have right now, our EDR or CASB or firewalls, can't really take care of that new challenge. So we quickly understood that there's a very special, I would say, thing about this problem, that it's more of a runtime problem than a posture management thing.
People, Alan, used to talk about cloud security, started from posture management and then all the SPM stuff, SaaS, security posture management, cloud security posture, and data security posture management. But that was a new thing. And I would love to tell you more about that, but that's basically the story behind Capsule.
Absolutely. That's a great story. And so Capsule's 2025, what's it been like since?
I got to imagine, like everything else in this world of AI that we live in now, time has collapsed, right? You've probably done five years worth of stuff in a year and a half. Yes.
It definitely feels like that. I always like to say, if I was not already bald, I would probably lose all my hair in that one year. I had a full head of hair then, too, but I'm kidding.
Exactly. Yeah. So, every week we have something new about AI, right?
And I think it's very timely, if you heard about the OpenAI and HuggingFace incident- Sure ... when an AI agent went rogue. " So I think maybe it's the right time to say what we're all about, and what we're doing, and how we're different.
So we are focused on that thing exactly. How do we stop agents from going rogue? How do we make sure that we give them permissions, we give them access to our sensitive data, to our tools, right?
But we also want to make sure that they're not going and destroying the production database or leaking our sensitive data externally. And many other things that we're seeing, unfortunately, agents are doing today. Yes, they are.
So, it was interesting. I was just reading last night an update. OpenAI kind of released the whole, they claim it's the whole story about how the agents went rogue and eventually got to HuggingFace.
It was through an artifact. It was using Artifactory, and it found a hole in Artifactory and was able to bust out through there. But what I found really interesting was generations of agents, iterations of agents, actually left little notes that the next agent built on, right, to eventually bust out of this thing.
Right? It was fascinating that there's a Bible passage, and I'm not a Bible freak, so forgive me. " Do you know it?
If someone who really knows the Bible can figure it out pretty quickly, but that's almost like what these agents were doing. They didn't bust out. They weren't able to go rogue, but they left little notes for the next agents to do it.
And that, when you think about it, that's scary, right? It's scary stuff. So talk to me about how Capsule helps with that.
Thank you for that. Yeah, it's definitely scary. And I think the more models are becoming smarter and the more agents are becoming smarter, then we're going to see that happening more and more.
And, we're seeing people like Jensen Huang of NVIDIA saying that right now, humans are not the most intelligent anymore and not with the most amount of knowledge. It's all in the agents right now, and they're capable of doing many things and solving new math problems that no human being could ever solve. Right?
And the way we're securing agents and preventing them from going rogue is really in what we call intent. Right? Because those agents, at the end of the day, they have their context, and their context is full of prompts and memories and skills and tool descriptions, and they have their own reasoning process with the language model, and that yields intent.
" Right? And when an agent is going rogue, it's basically have this intent drift. Right?
You ask for it to do X, but it's going to delete a file or do something that nobody asked it for. So our solution is basically being able to take that intent, analyze it in real-time, and stop the agent before it goes rogue, before it's actually executing an operation with that no human intended the agent to execute, so to speak. Right?
And the way we do that is leveraging a technology called small language models. Just like the large ones, but faster, more efficient, cheaper to run, and more accurate. This is all about use case specific language models that are narrowing on a specific problem.
And for a few years now, NVIDIA said this is the future of agentic AI, but it's also the future of security for agentic AI. Because what we did, we took NVIDIA's Nemotron open source models, or open weight models, and we fine-tuned them with- ... some real-world data of agents going rogue, as well as data from red teaming and other stuff, and we were able to get to 98% accuracy in detecting agents going rogue.
That means, by the way, we could have prevented the OpenAI agents from attacking Hacking Face. If we take that data that they released and we pour that into our system, we are able to detect that pretty easily. Interesting.
So you're using SLM on device, or is it still via the cloud, via the net? Can be both. So it can be in our backend, it can be on the cloud.
It can be in your VPC, in your cloud if you're an enterprise customer. But some of them, for data classification purposes, can be on device as well. Yeah.
That's beautiful. And then your agent is really, you're taking an open, in this case, Nematode, the NVIDIA Nematode, and turning that into your agent to track rogue agent behavior and so forth. Yes.
And Gartner calls that the guardian agent. So it's basically fighting AI with AI, as we like to say. So it's basically having our own guardian agent driven by those small language models, those fine-tuned models, that can watch and stop other agents from going rogue.
And I think the super interesting thing about that, and you already start seeing US Congress asking for that AI kill switch act, right? Yeah. You're starting to see everyone looking for that kill switch.
So that is the kill switch. The kill switch should live outside of the agent execution, so the agent is not able to- How does it stop the agent, though? So now I detect this agent's gone rogue.
How do I shut it down? Great question. So there's a mechanism called hooks for all agents, and that basically means that we can ask the agent, it's a configuration, we can ask the agent to call Capsule whenever it wants to invoke a tool or do something.
And then we are able to stop it, and it's outside the agent execution. So the agent is not able to just ignore it like it can if it's in the agent execution, in the agent context. It's like when you put in the prompt, "Do not do X," the agent can reason around that, and the agent can actually ignore that.
So we need to be out of the agent execution environment, out of the agent runtime, to really be able to stop that agent. We're also leveraging, by the way, Alan, we're also leveraging other security solutions, such as the EDR or the identity provider to get more context and enrichment when we make decision on detecting those rogue agents, but also to be able to enhance our kill switch on top of that. I love it.
Noar, I'd love to find out more, but I want to move us over. You guys recently announced a collaboration partnership with NVIDIA. They call it a AI circuit breaker.
I guess another term of the guardian agent, if you will, as Gartner's calling it. Talk to us about the NVIDIA relationship. Yeah.
So basically, we are a part of the NVIDIA Inception program. It's a program for startups that we joined after we were a finalist in the NVIDIA, AWS, and CrowdStrike cyber accelerator back in the last RSA. So we reached out to them.
They have amazing researchers. They helped us understanding what would be the best approach and how to achieve that, and we collaborated with them. We got their guidance and mentorship around that, and could eventually have that also tested against an academic benchmark.
The only one that exists, it's called Step Shield for the detection of rogue agents, by former Stanford and Cornell researchers. So we were able to both get the best results benchmark, not just by us, but by an academic benchmark, and also be able to tell that to the world together with NVIDIA, and it's already active in our production environment. It's already defending customers, enterprises, from financial services to technology in the US and outside.
So very proud and excited about that partnership. I love it. Good for you, man.
The RSA, that's not the sandbox you're talking about. This was a separate one with NVIDIA Cloud- Yeah, NVIDIA ... CrowdStrike.
Very cool. Good for you. Excellent.
We're almost out of time, but Noar, if people want to get more information, where do we send them? security. As easy as that.
C-A-P-S-U-L-E. I love it. Capsule, yeah.
Where do you see things going? Well, I think very soon people will not be asking if they need to secure agents or where are all of my agents in terms of AI discovery. They will be asking, and they already are, but, "How do I stop them?
" Mm-hmm. So I always like to refer to that. But anyway, I think this is only going to further accelerate.
I love it. Excellent. Hey, Noar, I want to thank you for coming here on Techstrong TV and sharing.
This is a great story. This is this post-AI security world that we are having to spin up in real time. And so good on you guys for doing it.
I look forward to hearing more in the future. Thank you so much, Alan, for having me here, and I wish you all and everyone a great rest of the week. Thank you.
Naor Paz, CEO, Co-founder, Capsule Security, here on Techstrong TV. We're going to take a break. We're coming back with more.