Vulnerability Management in the Age of AI and OSS | RSAC Virtual 2024
This session will discuss the rise of AI and LLMs and the role of open source software (OSS) in this evolving space. It will cover the convergence of OSS and vulnerability management and software supply chain security. It will cover the nuances and complexities of OSS security and vulnerability management as well as the OSS Top 10 Risks list, and resources such as the OWASP AI Security Checklist and Guidance.
– OSS is increasingly powering modern software including AI
– While incredible and valuable, OSS has some security and software supply chain considerations to be accounted for
– Vulnerability management is a complex topic and, if not approached correctly, can create tremendous toil for engineering and development peers of security practitioners
Transcript
Hello, my name's Chris Hughes. I'm the Chief Security Advisor at Indoor Labs and also the president at Acquia. A little bit about my background.
I've been in the cyberspace just about 20 years. I was active duty Air Force. Uh, I've spent some time as a federal government employee with a couple different agencies, the Navy and the FedRAMP team, and also have worked around, you know, both the federal and defense market quite a bit, as well as in the commercial space with tech startups doing some advisory work and software security and things like that.
Uh, so today I'll be giving a talk titled Vulnerability Management, the Age of AI and Open Source software. Uh, so open source software to start with. You know, there's a lot of good, we have a thriving ecosystem.
You can see this kind of eyesore map I've put up here, uh, from CNCF, the Cloud Native Computing Foundation. It's a thriving ecosystem. As you can see, there's a lot going on, and there's a lot of reason for the open source software adoption that accelerates innovation, uh, can save.
It has a robust community of contributors, you know, people creating projects and contributing to projects. It allows cross organization organizational collaboration, and it can speed up things like time to market, save on, uh, costs of research development timelines and, and investments and so on. Uh, there's some metrics around that.
You know, it says 97% of organizations are using open source software. Uh, 77% intend to increase their use of open source software. Many are looking now to sponsor projects, you know, financially contribute to projects, and the people who are maintaining these projects that they rely on, that the whole internet and the modern digital landscape we live in relies on.
And we see a lot of uptick in adoption, particularly around open source software for DevOps and cloud native ci cd tooling. And now, as we'll talk about later in the talk here, we're seeing explosive growth in the AI ecosystem as well. That said, it's not all good.
There are some negative drawbacks, or at least concerns or risks to be concerned with when it comes into open source software in the ecosystem itself. Uh, as I talked about, you know, 60 to 80% of modern software code bases, for example, are comprised of open source software as a metric put out there by organization known as the Linux, uh, foundation. But that said, software supply chain attacks are on the rise.
Everything from, you know, code Cove and, and SolarWinds XE utilities you don't see is pictured here, but it just keeps coming and coming and coming. And many of these projects are ultimately supported by unpaid volunteers. Uh, we saw this recently with X XE utilities.
Essentially, the maintainer was overwhelmed. You know, they were doing this in their free time, kind of a hobbyist, and it, it was socially engineered to take advantage of that reality and kind of, uh, try to insert malicious backdoor, for example. Uh, so a lot of these, you know, people who are maintaining these things that the modern internet, the modern digital landscape relies on, are doing this in their free time or out of goodwill, or just out of a passion for software development.
Uh, they didn't intend to put this out there to be relied on by everybody or necessarily respond to things like service level agreements or, uh, you know, request to, you know, add additional features or functionality or remediate vulnerabilities. And malicious actors are taking advantage of this reality. You can see the kind of attacks you by year reported.
It continues to climb. Uh, organizations, uh, you know, malicious actors, I should say, have realized rather than targeting one organization, uh, uh, you know, uh, for example, they can target, target, a widely used open source software component is used by thousands of organizations around the world and different proprietary products, or even other open source software, uh, libraries and components. And then it can have a massive downstream impact across the ecosystem.
And most organizations simply don't have good visibility of what open source software components they're using. Uh, they simply haven't been keeping a good inventory of this. And this is despite the fact that we know that software asset inventory is a security best practice that's been around for decades at this point.
In things like the sans or CIS critical security control lists, uh, many organizations can tell you what software, what applications they might be running. Many organizations don't have a great ability to look across their enterprise and tell you what libraries, what components they're running, what transitive dependencies they might have, and it's being taken advantage of right now by attackers. Uh, so there's been expansive growth of open source software.
I gave out some metrics earlier. I wanna dive into that a little bit earlier, a little bit more, I should say. As we talked about almost a hundred percent, 96% of total code base.
This is a report from Synopsis, their open source software security risk analysis report. 96% of code bases contain open source. Uh, almost 80% of the actual code base itself is open source software.
You know, we're kind of just sticking that together, adding some glue in there, and some proprietary first party code. Uh, and many of these things have concerns, whether it's, uh, you know, vulnerabilities or they haven't even been maintained in multiple years. You can see the mean age of vulnerabilities, for example, is nearly three years old.
Uh, if you think about, you know, zero days exploitation timelines from attackers, uh, there's a lot. It's a lot of time for, uh, malicious act to look at a widely used component and figure out how to exploit it, how to take advantage of that. And, uh, you know, another, uh, key factor that we need to look at from the risk perspective is there's a high bus factor, as it's called.
These metrics are truly astounding. You look at this and you understand that 25% of all open source software projects, for example, have one contributor com, uh, you know, committing code to the, to the library, to the component to the project, and then 94% of 'em have 10 or fewer. Uh, so 94% of all open source software in the entire ecosystem have 10 or few people contributing to it.
Uh, so every, you know, for every kind of widely used, thriving, you know, uh, open source software project like the Linux kernel or Kubernetes, et cetera, there's many, many, many others that simply are just being maintained by one or a small group of individuals. And that's a risk for the ecosystem. If that individual decides to quit maintaining it, they don't respond, uh, quickly to say a vulnerability, or they are able to be socially engineered, like we talked about the XEU utilities case.
And if you look at these code bases, for example, almost 50% of their components in these code bases have had no new development in almost two years. So they, you have a lot of kind of, uh, legacy, uh, components and projects and libraries, you know, in modern code bases that just simply haven't been maintained, haven't been kept up to date and so on. And then, not only that, but if you look at this, there's, uh, actually in the, even the ones that are being maintained, it, you know, it says 1% of code bases success for risk components that were nearly 12 months behind, uh, in terms of code maintainer, updates and patches.
So even if there's an update or a patch, these components, these libraries within your applications are simply not being maintained, not being updated regularly. Uh, and also another thing that I, I wanna dive into here is what we call the open source software top 10 risk list. This was created by indoor labs and others, and contributed now and OAS project, for example.
And we often talk about vulnerabilities, you know, CVEs, for example, common vulnerabilities and enumerations. But these are lagging indicators of risk. That means we know there's a vulnerability, we know there's a problem.
And it's also, you know, we need to think about looking at leading indicators of risk. Is it maintained, for example, uh, is there a compromise occurring of a legitimate package? Do we have untracked dependencies that we aren't accounting for?
Is our license risks, uh, for example. And then also looking at, you know, kind of software bloat, you'll see number 10 on the list here is under our oversized dependencies. For example, if we have oversized dependencies pulling all kinds of libraries and points into our code base that simply don't need to be there, this is now an expansive, uh, an expanse of our attack surface.
Uh, it, it creates a situation where malicious actors can go in there and take advantage of this. Uh, so most organizations simply haven't had a good inventory, their open source software components and libraries that they're using, nor understand from a transparent transparency perspective, you know, what, uh, components, what libraries, you know, uh, uh, products that they're using from proprietary vendors have in their products. As you talk, as we saw before, even most modern code bases are overwhelmingly made up of open source software.
So that means that pro products that we're using from a third party, uh, have many third party components, third party libraries, and so on in them. And we just don't have a good level of transparency to understand what's in there on the risk, uh, to us because of that. Uh, so this is a great resource to start looking at some of the other, you know, risks aside from, uh, known vulnerabilities.
You know, things like unmaintained software, compromises of legitimate packages on track dependencies that we simply don't have a good, uh, inventory of, or if we have software boat bloat going on in these code bases and applications, for example. Uh, another thing I wanted to dive in here is the runaway open source software ecosystem and the explosive growth of vulnerability. These are two great images from my friend Josh Bresser works over at ancor.
He runs a pro, uh, a podcast called Open Source Software Security. I definitely recommend checking that out. And some of the, uh, the figures he puts here are just outright astounding.
It's a bit of an eyesore. Uh, but if you look at the left here, it shows you the number of packages published per month, you know, you know, going back some time. If you look at the last 15 years as he did here, there's over a hundred million packages published out in the open source software ecosystem, a hundred million.
And if you look at the right chart here, this is a NIST national vulnerability database that tracks, uh, CVEs or known vulnerabilities, for example. You can see we have explosive year over year growth when it comes to known vulnerabilities being put into the, uh, NIST national vulnerability database, or being published as CVEs. Uh, as you, as you can imagine, there's only about 250,000 ever, uh, you know, CVEs ever record in NIST NVD.
And it's safe to say that with a hundred million packages out there in the open source software ecosystem, and us using them in all sorts of places that we don't have good visibility or inventory of, there's likely many, many, many more of vulnerabilities that we simply don't know of yet, or haven't quite, you know, been discovered, reported, uh, you know, had A-A-C-N-A go and report 'EM as a CVE. And, and organizations simply don't have a good inventory of where these components are in their environments. What systems are running 'em, you know, where they may be vulnerable.
So the problem is just growing exponentially, and it's past the point where we can even try to, uh, to sustain the way we approach handling vulnerability management currently, and this is a, is this is becoming evident when you look at the state of vulnerability backlogs. You know, most organizations, for example, uh, 2023 saw record, 30,000 CVEs are known vulnerabilities reported to, to this national vulnerability database. Going back here, you can see on, on, on 2024, we're on pace to exponentially grow over that, significantly double digit figure growth year over year.
Again. So critical vulnerabilities, even in that year, we're up over 60% over the year prior. So if you look at things like the common vulnerability scoring system and how we remediate vulnerabilities, typically you'll see, you know, we gotta fix all criticals and highs before you can go to production based on CVSS scores, for example.
It's not sustainable to try to take that approach. 1 million for large, you know, complex enterprise organizations, for example. And then you see that the other organization here, CIA Institute, found that organizations typically have the capacity to only remediate one out of every 10 vulnerabilities in their environment in a given month.
So, new vulnerabilities keep coming in, as we talked about, uh, you know, exponential year over year, year over year growth, tens of thousands of CVEs being published. Organizations simply can't keep up. They're drowning in this technical debt.
The security technical debt and vulnerability backlogs keep growing and growing and growing. Uh, we don't have a situation where, you know, attackers are looking for a needle on a haystack. It's a, a, a stack of needles.
You know, they're essentially looking for a, a very easy to find targets because we have a massive amount of vulnerabilities that simply haven't been prioritized, triaged, you know, remediated. And they're just sitting there as part of our attack surface waiting to be exploited. Uh, it vulnerability management essentially becomes a sisyphean task of pushing the rock up the hill.
Anyone who has worked in this role knows how that feels. You just, you know, you patch new ones keep coming out, it's over and over. You can't keep up.
You're drowning, uh, as the EM image shows trying to kind of take water off a sinking chip. And we're also not doing ourselves any favor when we look at things, how we operate, when we operate in the, you know, the quote unquote DevSecOps, right? Uh, the goal of DevSecOps was to break down silos, uh, where security, you know, was gonna work in, in symbiosis, symbiosis with, uh, our development peers, our engineering peers had this great collaborative relationship.
But that's not how it's played out. That's simply not how it's worked out in reality. Uh, you know, we're typically throwing, uh, toil and, and, you know, kind of labor over the, over the fence to our development peers.
As you see in the image here, uh, we're, you know, we're dumping massive vulnerability backlogs on them, you know, spreadsheets with thousands of vulnerabilities from tools with little to no context. Uh, and we also, we often take this kind of guilty until proven innocent mindset. You know, they have to go through and justify why something's a false positive or why it's not a critical finding, or what are some compensating controls that are in place, uh, things like that through the security team.
And that doesn't make us any friends. That doesn't make us any, uh, you know, partners, uh, that, that are willing to work with us. And it often builds resentment and frustration from our development peers.
And we've shifted left. We've brought all the tools, all the acronym soup of SAS and das, and infrastructures, code scanning, and SCA, we've moved all the things into the pipeline. The problem is, a lot of these things are producing a ton of noise with no context.
For example, you know, I, I put out some great, uh, resources here. Things such as no exploitation using the cisa, uh, KEV, no exploited vulnerability catalog. We're looking at exploit pre or exploitation probability.
EPSS. Uh, is it likely to be exploited in the next 30 days, for example. Uh, and also exploitability, is it, is it reachable within the code base?
For example, is it actually called at runtime? What about the architecture? Is it internet facing?
You know, it, what is the, the context that can determine if it's actually even in a position to be exploited? We often don't do that, that hard work. We kind of just dump these massive lists on our development engineering peers with little to no context and tell them to figure it out.
You know, we're not really, uh, you know, breaking down silos and putting up guardrails. We're often just building gates and building frustration and, and resentment from our peers. Uh, and as I show on the image on the right here, another key aspect of that is reachability analysis.
Uh, maybe you have some transitive dependencies, for example. And you need to understand is the, the vulnerable component actually reachable? Because if it's not, you can drive down the noise and the to, and the frustration, our development peers significantly, which can help them focus on the actual, the actual, you know, most pressing critical risk of the organization.
And we have guidance to lower. We look at software, supply chain security, open source software, uh, software assurance, you know, whatever phrase you wanna take. Uh, just to throw a few examples out at you, we have the N Secure Software Development framework.
We have SALSA from Google. We have the NSA and CSA now putting out secure software supply chain guidance for developers. Oh, WASP has their guidance.
CNCF has their guidance. Uh, CSA has their secure by design publications, for example, we're, we have no shortage of, of, of guidance and best practices. I like to say we're best practices rich, but we're implementation poor.
Most organizations simply don't know where to start, how to make sense of all this different guidance, which one they should be focusing on. Which one do they need to actually meet? You know, do they have compliance and regulatory requirements forcing them to use one or the other where they get started?
Uh, security needs to come in and provide that clarity and then that guidance to our engineering development peers to help 'em get going on the right track, and not just bury them in, in a mountain of best practices. We need to move towards actual implementation if we wanna drive down risk and have some secure outcomes. Another factor that we'll be discussing in this talk is the rapid rise of gen AI and LLMs, uh, generative AI and large language models.
Uh, we've seen this is the fastest growing technology ever. You know, there's some great charts out there showing you that. When we look at previous technological waves, like the internet, uh, mobile devices, cloud computing, uh, gen, ai, a, ai and LLMs have grown faster than any previous technological wave.
The rapid adoption is truly astounding to see, uh, everyone's excited about it, you know, what can it do for the business? And, you know, I list some security use cases here, like, can it revolutionize the way we do soc, for example, instead, throwing more bodies at things and alerting and telemetry and findings and notifications. Can it address that problem?
Can it help us with code generation and scanning? Can it help us produce code, right? You see copilot and other things out here that are helping people produce code faster, uh, be more effective, be more productive.
And obviously, developers are incentivized to get, uh, new features out there as quickly as they can. And on our, on our side, on the security side, uh, can it help us perform things like malware analysis, which I'll be talking out, uh, talking about in a moment. Can it help us, you know, perform better?
Uh, scanning for code, uh, vulnerabilities like we just talked about in terms of reachability analysis and, you know, uh, vulnerability, exploitation, so on and so on. It's also having, uh, opportunities, you know, people are looking to see, can we use these technologies for things like incident response? Uh, rather than having a human looking for a nous activity, uh, activity through, you know, massive, uh, stack of, of logs and data security, data lakes, you know, you, cm, soc, you name it, uh, can help with, uh, those activities, software, supply chain, security.
Uh, does it have promise there? Compliance, we're seeing a lot of use of gen AI and LMS when it comes to con creating, creating, um, you know, compliance documentation, policies, processes, uh, running control statements of how you, uh, align with, you know, FedRAMP, nist, and hipaa, and high trusts, and you name it, the, the myriad of compliance requirements we have, and on and on and on. Uh, that said, you know, this technology is still being, uh, flushed out and still being matured.
You know, what can go wrong? As it turns out, there could be quite a few things that can go wrong. And we're still learning that as we go, uh, go about our adoption and evolution of using this technology.
So, another problem that we're gonna talk about is, you know, what exactly is open source ai? When we talk about open source software, and I, I shared this article here from the MIT, uh, technology review publication. It says that the technology industry can't really agree on what open source AI means.
And that's actually a problem. For example, uh, is it just the model? Does the license?
Is, is there licensing restrictions around how the model can be used by whom, under what context? And a lot goes into creating these AI models. It could be the underlying source code, it could be the trained model, for example, the training data that the model was trained on to be pre pre-processing code and pipelines that were part of that process of training a model.
And it could also be code governing that training process. And even more. So what exactly is open source ai, for example, is it just a model or is it all the other things that go into, uh, creating that model, operating that model, the infrastructure to, you know, under, uh, underpin the model and its operations, and you know, how the data was trained, what the training data set looks like, and so on.
So it's a lot of information that we need to take into context, determine, you know, what exactly is open source AI as an industry, how do we kind of define that as we move forward? So first off, I wanna talk about hugging face. This is, you know, if you're not familiar with this, this is the main place where developers are discovering and sharing open source software.
LLMs learn large language models. It's an incredible resource and incredible, uh, community of folks going out there, hosting, you know, models and data sets for the community. It allows users and organizations to each host models they can share with other folks, uh, put out there for the community use.
And each model is essentially a Git repository of model data and metadata about the model and so on. And, you know, they also provide, uh, example code on how you can deploy and run models in various cloud services, whether you're thinking about Amazon web services, Azure, you know, machine learning, and they even offer train, uh, their own services, for example, for deploying models such as spaces as they call it, which is free hosting, uh, for hosting some of the models, running some of the models and, and, you know, testing or production environment and so on. And, uh, you'll see why that's important, important here in the morning in a moment.
And they have leaderboards to track kind of model performance, what models are, you know, thinking about the GitHub star system, right? It has, uh, you know, a system out there to track model performance. What models are performing the best?
How's the community responding to them? You know, can you have people provide, uh, responsive feedback to, you know, encourage others to use these models and so on? And it has APIs for folks to upload models, and they have to go ahead and create an account to do that.
And just to show you the level of adoption and, and use that we've seen in hugging face, for example, it's kind of titled it's all about the lms. And it, you'll see why here in these figures, uh, there's over 470,000 hugging face models that they host on their site. There's also over 97,000 data sets that they host on the site with over a million model downloads every single day.
So just think about that, that rate of adoption, that rate of u rate of usage, uh, and how quickly that's grown, you know, in, in terms of our community of, of, you know, people tinkering with machine learning, gen, ai, lms, and how quickly that's growing. And it's good, but also it has some risk associated with it, which we'll hear about here in a moment. Uh, so if we think about this hucking face due to this massive adoption that they have in terms of the models, the data sets, uh, you know, all the users being able to contribute and, and provide resources to one another, it also now becomes this critical target in our software supply chain.
For example, in 2023, they had an incident where there were some exposed a, uh, API tokens. It could have potentially impacted 700 organizations, including some of the leading, you know, software and technology organizations out there. And it was over 1600, uh, hugging face tokens exposed as part of this incident.
And then we talked about spaces where people can go and host, you know, uh, models and deploy models, for example, uh, there were some unauthorized access, potentially exposing secrets. So they went and it, you know, kind of proactively, uh, you know, rotated secrets and recommended customers. Do the same users do the same?
And there was some great research put out there by Wiz, for example, showing that these, you know, these models could be used for cross tenant attacks. If you think about in the cloud context, we have a multi-tenant model in the cloud. And these cross tenant attacks can impact millions of private AI models and apps because of the numbers of users that're using them, the number of organizations that are using them.
So you might have a, an attack scenario like the one shown here in the image where you have some malicious access to a model or, or, or model poisoning as they call it. And it can impact thousands of, uh, downstream users or even millions of downstream users to the, due to the number of outsized, you know, outsized number of organizations that are using this service, uh, using this platform when it comes to open source software, collaboration, gen ai, uh, s and so on. Another big one that we have to mention here obviously is open ai.
It's sparked an entire ecosystem of open source software innovation. Uh, if you look at this here, after it's released in January 20th, 2023, there was over 600 new packages using this API, uh, just an NPM and PII alone, uh, you know, showing that uptick of, you know, uh, adoption and use of the, uh, OpenAI, API, for example, and then 276 existing packages add those APIs. Uh, API calls to their package, uh, to their projects, for example.
So we see a new wave of LLM powered applications. If you look at the graph, it's just a massive uptick of adoption. And that's all within less than a year.
You know, we're, we're a little bit over a year, I should say at this point. So you imagine that keeps growing exponentially as more people keep using it. More organizations take advantage of open API, uh, start using these services to power their modern applications for all the different use cases we talked about, whether it's code development, security, uh, use cases, chat bots, uh, business functionality, you know, all these type of things.
You know, it's gonna lead to exponential growth. And that has, and it's obviously good to drive innovation and acceleration in our ecosystem, but also has some risks associated with it as well. It's taking a look here.
We have a popular open source software machine learning framework known as TensorFlow. It's used by tons and tons of people. You can see it's, uh, you know, wildly popular.
And then we start to look at, you know, where are some of the concerns from a pedigree or providence, uh, perspective. It has 2,238 direct contributors. If you look at the third degree contributors, fifth degree contributors, seventh degree contributors, that number just keeps growing and growing and, and growing.
So you have all these contributors from all over the world contributing to these, uh, open source software, you know, machine learning libraries and projects. And that's great because it leads to innovation as we talked about, uh, innovation, you know, uh, new capabilities, new features, new functionality. But it also creates a problem in terms of, you know, where did this come from?
Who contributed to it? Uh, do they have malicious intent or are they, you know, kind of just a regular user that's trying to provide some, uh, functionality and value, uh, to the project? All that becomes incredibly complicated to look at when, you know, to, to unpack, when you look at the number of, uh, contributors that are, you know, to the nth degree of contributing to this wildly popular machine learning framework, notice tensor flow.
Uh, so let's take a look at some of the top 100 AI reposts, for example, on GitHub, you can see the, you know, how they're structured here and some of the findings. You know, we, we talked about, you know, uh, the number of dependencies and transitive dependencies, for example, the average number of dependencies in these projects is 208 when you look at direct and transit of dependencies. So now you have to look at those 208 dependencies and understand, are they vulnerable?
Do they have known vulnerabilities? Is it maintained by whom? How often is it updated?
Um, and do they have dependency, uh, bloat, for example, all the things we talked about in the open source software, top 10, all these factors come into play now. And if you look at this chart here on the, on the right, uh, 11 of these repositories have over 500 dependencies. Uh, and then you have to unre, you know, kind of unpack all those open source software, top 10 risks that we just talked about across that list of hundreds of dependencies.
Uh, and that's just for one repo on GitHub for an AI project, for example. So it's a lot of, you know, things to try to make, wrap your head around and make sense of. And continuing down that, you know, that path of known vulnerabilities and other risks, uh, 52% of the repos have one or more vulnerabilities in their dependencies.
So over half of 'em have a known vulnerability, let alone the other nine open source software risks that we talked about from the OSS top 10, 15% or more have 10 or more vulnerabilities. Uh, so that's the number of vulnerabilities. Like, are they known to be exploited?
Are they likely to be exploited? Are they reachable? All that comes into play in terms of the business criticality.
Uh, when you're looking at that, you know that doing that risk analysis of where's deployed in your environment, you know, uh, what systems does it impact? What type of data does, uh, those systems interact with or store? And then looking at your third parties, when your proxy that you're consuming, for example, do they run these components?
Are they addressing the vulnerabilities and risks that are in that environment or associated with these repositories and projects, for example? And so, if you're using any of these, you know, any of these, any of this code from any of these repositories, what your vulnerabilities really affect you, this is where we get back to things I talked about, like, uh, CS kev, EPSS, reachability analysis, business criticality of the system, for example, and the data sensitivity. Uh, these are all, you know, things that need to be considered when we start to go and push, you know, vulnerability scans out on our peers and, and try to minimize that toil, uh, that they're gonna have to deal with.
And, and the noise of just the, you know, the noise that these modern tools are, are producing that don't have context as they should. Uh, so another thing that we've seen in this, in industry and this ecosystem as gen AI and all them keep growing is, you know, AI coding tools. Are they a friend or are they a fo a lot of people are interested in, you know, for example, can AI coding tools be used to help us accelerate code development?
Uh, or can they be used conversely to help us perform things like malware scanning and analysis? Uh, if we think about vulnerability scanning of code and things like that. And it's obviously, it's, uh, you know, early in the adoption lifecycle, and the answer is it's too soon to tell.
And I'll talk about that. Why here. Uh, so LM assisted malware reviews, there's still ways to go.
We did some research here with our team, indoor labs and, and collaboration with some of our partners. We saw that prompt injection vulnerabilities must be addressed through pre-processing, for example. So removing con uh, renaming identifiers, things of that nature.
And it also makes it easier, you know, make it easier for attackers to blend malicious behavior into legitimate packages. Uh, so as we've started this study, you know, whether it can be used for malware reviews, which we obviously wanna accelerate our activity on the defensive side of cybersecurity, we're seeing high false positive rates and also false negatives due to inter procedural data flows, you know, across files, across packages, things like that. Uh, so the high false positive rate obviously creates, you know, more challenges around toil, frustration noise, uh, for folks that are reviewing these, you know, these findings.
And then false negatives, obviously that's very concerning, because we could have something that is malicious and, and not be picked up appropriately by these LMS and using that context. And so right now, the jury's still out. You know, LMS can assist with some, uh, code review and malware review of, of code, for example.
But we still need a human in the loop. We can't, we're not going anywhere quite yet. Another, uh, interesting topic, as I talked about, is gonna be used to aid exploitation.
There's a lot fears of, you know, AI is gonna, uh, just be unleashed on the ecosystem, and it's just gonna go, you know, uh, make all of our systems collapse, and it's gonna empower attackers to outpace us. And there's even some headlines, you know, kind of, uh, that are out there generating this fear. Uh, the one I point out here came from the register.
They, they went and took a look at a paper title, LLM Agents can autonomously exploit One Day Vulnerabilities. You see the headline, uh, just by simply reading the advisory. So reading the CVE, you know, reading a description of the CVE, things like that, reading its criticality, and they use open AI's, GPT-4 for this research.
And then they claimed in the research, uh, they was capable of exploiting 87% of critical CVEs, uh, not quite right. So we had someone that went and did some really thorough analysis, uh, in another paper that I'm citing here, titled No LLM Agents Cannot Autonomously Exploit One Day Vulnerabilities. And they went and debunked this point out that first it was done on a very, very small subset of simple, uh, uh, you know, code bases, for example, or, or vulnerabilities, I'm sorry.
And, and these vulnerabilities often had proof of exploits that existed. And, and, you know, were already out there and able to be used by these tools. And also, they often required assistance from humans, again, much like malware review, uh, to effectively reach their goal of, of, you know, exploding of vulnerability.
Uh, but that said, it does point to the reality that, you know, first off, you know, LMS and Gen AI have potential to help with malware analysis and code review, for example. Uh, if it continues to improve, which we largely suspect it will, but it also has the opportunity to accelerate exploitation of vulnerabilities, uh, for malicious actors. So it's important that we stay ahead of them in terms of being competent with these tools, these technologies in terms of how we adopt them, how we learn to work with them, integrate them into our organizations, how startup, uh, organizations go out and, uh, you know, inve make investments and bets on these technologies and try to prove that it can be done.
Uh, we have a lot of potential in both directions, which is both, uh, promising and concerning, obviously. Uh, so where do we go from here? We know that Gen AI and LMS have to potential both improve software security, but it also can accelerate malicious, uh, activity.
For example. Uh, we see, we know that developers, when we look at incentives, uh, you know, they're not incentivized to slow down. Uh, you know, kind of, we hear a lot about secure by design and things like that, but developers are primarily incentivized to move faster, get more features out, uh, to, you know, burn down that product backlog for the product manager, you know, and, and team, uh, they're not incentivized to go slower, you know, be more rigorous, uh, implement more governance, for example.
So if these tools offer, offer the por opportunity for them to produce code faster, be more productive, at least in, in their mind in terms of, uh, in their incentive structure, they're likely gonna do that. So it's very important that we as the security professionals do the same in terms of staying familiar with these tools, these technologies, looking to use them ourself, uh, to keep pace with our, our development peers. Uh, definitely not the right approach would be like, you know, outright banning these things, uh, trying to, you know, minimize their use.
Um, you know, it's just gonna lead to what we saw, saw with, you know, shadow IT or Shadow Cloud or shadow SaaS. Uh, they will find a way to use this. So it's very important that we work collaboratively with them hand in hand in hand, and build proficiencies, both as ourselves and the organization, uh, with our peers and using these tools and technologies.
And if we fail to learn from the lessons of the past, we'll, we'll all but ensure an insecure future. Uh, we've got be, uh, as I said, we've got plenty of best practices out there, but implementation needs to be part of that conversation. Uh, we need to keep, uh, you know, not just showing best practices, not just dropping documents on organizations and peers, but actually be there working with them hand in hand, showing them how to securely use these technologies, raising awareness about some of the risks, and just, you know, understand that we're part of this process too.
And a lot of this still needs to be flushed out and determined, and we won't be able to become competent with these technologies if we're not working with them ourselves. Uh, and rather than kind of, you know, using what security typically does, which is the office of no, or FUD or just, you know, kind of drumming up fear, uncertainty of doubt around the technologies. Uh, so that's it from my talk today.
Thank you so much for everyone who tuned in.