Red and Blue with Vanessa Sauter and Brennan Lodge – The AI Security Edge EP2
Vanessa Sauter and Brennan Lodge join Caroline Wong for an intense discussion of red (offensive) and blue (defensive) capabilities and how artificial intelligence is changing each. Vanessa ponders the ability for artificial intelligence to generate more secure code than a human can, and stresses the importance of red teaming and manual penetration testing when it comes to the use of LLMs. Brennan describes how artificial intelligence can be used to manage governance, risk and compliance requirements and frameworks.
Transcript
Welcome to the AI Security Edge. I'm your host, Caroline Wong, and I'm so excited to introduce one of the newer tech strong TV podcasts, text Strong TV features, your favorite video series, industry thought leader commentary and analyst research on DevOps, security Cloud native and Digital transformation in a podcast format. We all know that AI is revolutionizing cybersecurity, both as a weapon for attackers and as a shield for defenders.
The AI security edge dives deep into the evolving cyber battlefield where AI driven threats, challenge traditional defenses and cutting edge AI solutions offer new ways to fight back. Our podcast explores real world case studies, expert insights and practical strategies for building cyber resilience in an AI powered world. Whether you are a security leader, practitioner, or AI enthusiast, we are so glad you're here with us today.
I'm joined by my good friends and colleagues, Vanessa and Brennan. Vanessa Solder helps enterprises secure gen AI functions at scale. She is an expert at threat modeling and attacking rag based l LLMs.
She loves to talk about AI governance and privacy assessments. Vanessa has led hundreds of security and privacy reviews for customers at a leading artificial intelligence company. She has personally pentest dozens of enterprises and launched hundreds of bug bounty programs.
She speaks at conferences like B Science sf, the Diana Initiative and Developer Week, and you can also find her writing in Forbes Law Fair and Dark Reading. Vanessa, welcome. Thanks so much, Caroline.
It's great to be here. Brennan Lodge is a cybersecurity expert with over 15 years of experience in data science, ai, and threat defense. As founder of BE Logic Inc and a professor at NYU, he has led AI centric cybersecurity solutions at top financial institutions, including HSBC and Goldman Sachs.
Brennan is known for pioneering the use of retrieval, augmented generation rag models in cybersecurity, streamlining, threat detection, and enhancing regulatory compliance. He is a technical advisor and an award AI researcher driving innovation in cyber defense Strategies across multiple industries. Renin welcome.
Thank you, an honor to be here. Thanks, Caroline. So the first topic that I want to hear about from each of you is can you tell us about your experience with ai, Both Personally and professionally?
Vanessa, let's start with you. Yeah, sure. So, so right now I am the principal solutions architect at a company called Prom fu.
Uh, we are an open source company that's developing an l lumm evaluation and red teaming tool. So what that means is we help users and we help companies, um, sort of take their LLM development applications to the next level in terms of quality, um, and output. And then we take it a step further and do red teaming, um, further applications or their foundation models, which we'll assess, uh, technical security risks that align with O osp, um, any type of privacy issues that might come into play.
And then this really cool nebulous area around trust, safety and AI codes of conduct, which is what happens when, um, foundation models or LM applications behave in erroneous or, or harmful ways, which we can always talk about more in the, um, you know, the topic around insecure code and what that means. Um, so, so, you know, my professionally, my, my role right now is really helping accelerate that deployment, um, and engaging with our community of more than 60,000 developers who are helping, um, use this tool and contribute to making it better, um, and really driving the red teaming forward. Um, and then, you know, on the side, and previously I've done a lot of bug bounty hing for LMF applications, um, pen testing, and then in general, just sort of being on top of the, the AI security governance part of it.
Um, we use AI for, for everything. Um, you know, we use cursor to help accelerate and ship code faster. Um, you know, personally, perplexity has replaced Google for me.
I almost never use Google anymore. Um, you know, I use chat GPT pretty much every hour on the hour for some random task or the other. Um, you know, we're, we're working on building agents for a lot of the work that we're doing.
So anywhere that that AI can, can play a role and be a companion in my professional life and in my personal life, you know, we, we, we use it safely and we use it, um, to move faster. So Yeah, thanks for having me. And you know, Vanessa, uh, the, the red teamer and me, the, the blue teamer here, hopefully there won't be too much conflict in the, the podcast today, Caroline.
Uh, but it's, it's always awesome, you know, hearing the two different perspectives, right? So, and that's where I come from in my AI use. It's been an uphill battle, I'll say till, you know, since 2015 when I was building out, yeah, remember the, the terms data science and big data, uh, solutions, um, and building them to be realistic.
And it's still a, an uphill battle with ai. Um, you know, the, the hardcore security defenders are a bit reluctant on the, the use cases. And, you know, I continue to, to prove them wrong.
That's why I'm here, right? Um, and, and rag, right, really powerful, uh, solution and good for, you know, the CISOs and defenders in that you can manage your own destiny, have an on-prem, uh, solution of, of rag, and, you know, build out fine tune, customize your, your own models without having to use the, the APIs from, from OpenAI. So on the more on the professional side, uh, have started, you know, some consulting and then also a, a product that uses rag, um, specializing in the the GRC use case of, uh, gap analysis between policies.
There's hundreds of, of regulations, uh, across the world that, um, relate to information security. Uh, consumer protection data privacy laws are only growing, um, getting more, you know, um, you know, verbiage with them. And, you know, having to understand and incorporate those regs to make sure you're abiding by the law is getting tougher and tougher, uh, within the, the GRC front.
So, really excited to, to get going and, and talking more about AI offense and defense. Awesome. Awesome.
I just wanna take a moment and say, I'm so thankful that the three of us get to be together in this moment, um, because things are changing so incredibly quickly. You know, if we had gotten together a year or two ago, our conversation would've been entirely different, and I'm sure that things are gonna be even that much more different six months from now. Um, you know, on the AI security edge, we are really kind of exploring from a cybersecurity and resilience perspective, how is artificial intelligence helping attackers, and how is artificial intelligence helping defenders?
I'd love to start on the attacker side and get your perspectives on what is the difference between being an attacker today with all of this ai, everything, um, and how has that really shifted? Yeah, I'm, I'm happy to, to dive into this, um, a little bit. So, so there's a couple of ways that we can actually, um, quantify this.
Um, and that's what's really fascinating right now is we're seeing a lot of research come out around, um, how can, can LLMs at the foundation level, right? So the models themselves, how can they, um, produce, uh, first of all, can they produce insecure code, which is a risk for, for attackers? And, and, and how are we shipping code, um, when developers are, are using ai, um, how do we know that they're not shipping code that isn't secure that can be exploited by attacker?
So there's sort of benchmarks that we're, we're looking at right now in terms of, you know, what is the security of the code and the quality of the code that that LMS can produce. Uh, and then there's the second level that, that we're starting to see a lot of of research coming in, and that is around, um, how effective are LLMs, again, at the foundation level at, um, identifying exploits, exploits in zero day vulnerabilities. Um, and so that is the, the efficacy, and there's different benchmarks that are being developed around how LLMs can actually take information from their environment, um, and subsequently be able to craft, um, highly accurate exploits that attackers can use for, um, for, for, for, you know, attacking different enterprises or different types of infrastructure.
Um, so, so that's unique. And then there's a another level that we're seeing, which is this agentic, um, action. So, so what we're seeing is, again, you know, the base level is, um, are we producing code that is insecure?
The second level is, um, can, can attackers use LLMs as part of their arsenal in terms of, um, exploiting, uh, vulnerabilities in the wild? And then the third is, um, can attackers actually develop age agentic systems where, um, they, they actually can behave autonomously where there's not, there's not always a human in the loop. Um, and they're able to, uh, use LLMs both to identify, um, and craft exploits, but then also take it a step further, um, and exploit them without having to constantly have feedback, uh, or have a human take action in that.
So, so right now we're, we're really starting to quantify the first two, uh, the third really becoming more of an issue. Uh, we're starting to see this a little bit in the bug bounty space actually, and this, it's a really good indicator in terms of sort of what the trend is, where there's a lot more age agent activity happening, um, with public bug bounty program about policies, um, you know, in the bug bounty space in terms of how do we accommodate that. Um, but, but that's really the, the main concern that we're seeing.
And, and obviously as LMS grow in capacity, um, and their reasoning improves, the ag agentic capabilities improve, um, that will continue to pose more of a risk for enterprises on the, on the blue team side. And, you know, whatever your, your response is initially from a SecOps perspective, it'll have to be an eighth or a 10th, um, of what, what the response time is now. So, um, just things are changing.
They're moving really fast and, you know, we just have to prepare accordingly. Yeah, Yeah, I totally agree. Uh, I see it as a, you know, double-edged sword, um, one for the defenders, you know, it can help with, with automation maybe getting quicker to identify and synthesize, you know, the huge number of, uh, attacks and alerts that they have to, to go through.
Um, and, you know, the double-edged sword, right? So on the, the attacker side, um, making it easier to, to automate craft, you know, new zero days, right? And I think the overall theme is just creativeness that, you know, we as humans are leveraging AI to do and to, you know, think differently.
That means attack differently, defend differently. Um, you know, the, the research, you know, is still out there, you know, at first, um, you know, I thought it was more geared, you know, when the models came out, you know, two years ago with the open ai, uh, announcement that it was, you know, 90%, Hey, this is gonna be, you know, good for, um, the, the offense. But now with, you know, the transparency of the, the models, uh, you know, articulating how the model is thinking and processing the, the information that's really helpful for the defenders, um, especially the, the junior analysts that need that upleveling, uh, you know, a tie to raise all boats here, the boats being the, the junior analysts, um, you know, in college, uh, folks coming into the cybersecurity industry that need that, you know, level up in, you know, experience that they don't have, and understanding, you know, what this type of attack is, what this, you know, vulnerability is and why it matters, right?
That's always been really hard to determine, uh, on the, the blue team where you get, you know, CBEs and, and notices of, hey, the CBSS score of nine point a. Yeah, it's concerning, but why does it matter to me? Now, the, the models can, can really understand the environments, um, the, the assets that you have, you know, the, the types of vulnerabilities that that impact you.
Um, and then yeah, more zero days, right? Uh, for the attackers. So fighting fire with fire, right?
Hopefully, you know, it, it helps and there's less attacks, but they're always gonna be there. Wow. I mean, I think fundamentally one of the questions that I wonder about is what produces more secure or less secure code a human or ai, you know?
And I think one of the kind of, you know, similar but not exactly, uh, the same things that I, that I wonder about is, you know, the other day I was in San Francisco and I was riding in a Waymo vehicle, um, and I actually loved it. Um, but my husband, for example, feels very, very uncomfortable, uh, in a vehicle with no driver. Um, and I happen to think, okay, uh, I might actually trust kind of a smart car that is not drunk, is not texting, is not, um, you know, potentially gonna have some sort of severe health issue, you know, unpredictably.
Um, but what do you, what do you each think about that? Um, this question of what produces more secure or less secure code a human or ai? So, um, I think that, that anything that produces code is gonna make errors, and that's, that's sort of the baseline that we should always operate under humans will, will inevitably produce and secure code.
That's why we, we have pen testing. That's why we have software development life cycles. That's why we have bug bounty and change management.
Um, the interesting question that, that I'm starting to see is actually not human versus ai. It, it's actually what ai, um, so, so with the, the, the open source community, one of the things that's very interesting is that you can take open source models and you can fine tune them according to your use cases. So on hugging face, there's hundreds and hundreds of these models out there, many of them can, can create code.
Um, there's actually research coming out that we're seeing where you can actually train these models to have back doors in them. So, uh, if they're starting triggers, um, that, that a developer's using and they're using a, a poisoned open source model, um, that model will actually produce, uh, intentionally insecure code. Um, they can subsequently, uh, cause and exploit within, within the environment.
So, so what's really important is actually the, the due diligence aspect of what models you're using when you're producing code, um, and using buttoned up models or doing very, very intense red teaming on the, on the open source models that you're using. Um, obviously if you, if you opt for a closed source model, um, like, uh, you know, open AI models, um, Claude, um, you know, they have very, very tight security and trust and, and safety regulations. Um, you know, llama has exceptional, um, uh, trust and safety reports that that, that they publish.
But you have to be extremely cautious about the, the type of model that you're using. Um, and then, you know, the recommendation that I always have is, um, continuing to have defense in depth. So always assume that, um, that there will be insecure code that will be published or it could be shipped, um, and then incorporating additional policies into that.
So, so I don't think there is such a thing right now as, as human versus ai, the assumption is that they partner together. Um, and, and what happens is you, you create at an organization level the different layers, um, for, for defense, defense in depth, whether that's for, uh, you know, a separate agent that reviews the code. So you actually have an LM as a judge as part of what's being shipped, um, having human reviewers, um, having change management processes.
So there's different layers that you can get into, uh, but, but inevitably, you know, things are gonna happen and you just kind of have to prepare for it. Yeah. Yeah.
And it is good. We've got the Vanessa's of the world, uh, you know, pen testing, right? And, and red teaming, and it is that human factor that I think you're always gonna need, uh, in cybersecurity and in in ai, right?
You think back to the, you know, car manufacturing process, right? By, you know, Henry Ford and, you know, putting the, the car together, right? There were no safety belts, there were, were no airbags, right?
But it took, you know, a a series of unfortunate events of car crashes, you know, to, to implement those, you know, safety precautions that we now have. And, you know, similarly with, with the ai, it's still greenfield, right? And it's gonna take the, the red teamers and the pen testers to experience those, you know, outer edge, you know, type of experiences with the, the model to then put some safety precautions in place.
And even with the manufacturing process that is, you know, fully automated that, you know, Elon Musk wants to do with his, his car manufacturers, it, I still think can't be done. You're gonna need that human in the loop to identify those defects, right? That are produced by the, the machines, the manufacturing process, and then rinse and repeat with refining, right?
The, the tough part is, you know, most of the AI models are, are black box, and even the, the scientists and, you know, professors like me still don't have no freaking clue what's going on behind the scenes as to how they're, they're working, right? Um, what we do know, and I think we need to be better at in the industry is the data that we use to create the models. Um, we could do that in the enterprise.
We could do that with, you know, RAG and the, the custom models that we're building the data that is held within the, the Vector database. Um, and there are some really good tools for visualization and, and transparency, and then the, the pen testing, right? Testing for, you know, racist remarks, uh, that may come out from, uh, the, the model and, you know, everything else that may impact or security concerns or giving up intellectual property, right?
It's, it's a lot of work that us as cybersecurity practitioners are, you know, back to the, the drawing board, but hey, that's, you know, why they pay us the big bucks in cybersecurity, right? That's right. You know, you two 20 minutes of talking with you two feels like it flies by in an instant.
Um, I know that in my brain there are just like sparks going off and I know that for the next like 24, 48 hours, I'm just gonna be like thinking about our conversation. Um, and I really hope that is the same for our audience today. Um, thank you so much for joining us today on the AI Security Edge.
Um, make sure that you get a chance to, uh, check out our previous episode, uh, with Daniel Mesler. Uh, and we have a lot of fun episodes coming up for you. Um, if you loved today's podcast, don't forget that Techstrong TV has a bunch of other podcasts that are also super great.
Um, thank you Vanessa. Thank you Brennan, for being here with me today, and thanks so much, uh, to our viewers and listeners. We'll see you next time.
