AI: Friend or Security Frenemy? | RSAC Virtual 2024
Is AI your friend, companion or your enemy when it comes to writing secure code? What about AI in regards to security tools? Learn from an expert who has lived on both sides of this. AI is today’s buzzword, just like blockchain was in the mid 2010’s. However, unlike blockchain, AI is here to stay, front and center. What is AI? What benefit’s does AI have? Is AI really everything they are claiming it can be? What’s the difference between generative AI and LLMs?
In this presentation you will learn the following:
• Security and Developer viewpoints regarding AI
• Basics of AI
• How AI can be used to speed your development up by more than 50% and stay secure
• How AI works regarding security
• Considerations that can’t be ignored
Transcript
Hello, my name is Chris Lindsay, and today we're gonna talk about ai. Is it a friend or a security frenemy? Today's agenda, we're gonna talk about AI versus machine learning.
Why? Because it's very important. There is a differentiation.
We're gonna talk about some of the basics. We're gonna talk about security and developer teams, how they view ai, and how this benefits each side. We're gonna talk about how AI can be used to speed up development, because if you're not using AI for your development tasks, now you're, you're, you're kind of lagging behind.
There's some big benefits there. We're gonna talk about how AI works in regards to security and some considerations that absolutely cannot be ignored. My name is Chris Lindsay.
I have been writing software for over 35 years. I've been in the security industry for over 15, and I ran an application security program for over three, for a very large enterprise. So everything we're talking about today is very relevant, and I've lived it, and I am one of you guys.
So let's start here. So, LLMs language, uh, large learning models. Models are not just changing the game.
They're actually creating an entirely new playing fields. If you think about the iPhone, when it first came out, you had that old Nokia and you have the new iPhone, they're different. They're absolutely different.
It's a game changer, and that's what we're seeing right now with ai. So let's talk about the difference between AI and machine learning. So AI is a broad spectrum, creating intelligent machines, enc encompassing rule-based systems, learning adaption.
There, there's so much to ai, and we're actually gonna talk about some of those differences later. And when we start looking at machine learning, we're gonna talk about some more narrower data-driven approach. We're also gonna talk about machine learning today as well.
So we're gonna kinda hit both. And one thing I want you guys to just take away from this, a, uh, or all machine learning is not a, or is ai, but not all AI is machine learning. There is a difference, and AI includes more often, uh, more than learning it, it's so much more.
It's also encompasses understanding and reasoning and just, you know, when you think about ai, you think about robots and, and self autonomous driving cars, you know, machine learning different. It, it's just a aspect of ai. So we'll get into that.
So just looking at kind of a security developer side, how do they think about ai? You know, kind of in general. So AI can be used for automation.
Think about AI for automating common tasks. You're, you're building your pipeline. There's certain things that you wanna do, certain things that you run into.
Like, for example, if your build fails and you need to notify certain things, or if your build fails and you want it to automatically try to recover itself, that's some automation. Uh, people consider it an assistant. So they perceive LLMs as kind of a trusted, uh, resource, which we're gonna find out later that it's not necessarily a trusted resource.
And you should have a little bit of, don't fully trust it. And we're gonna talk about how it speeds up thes, uh, decision making where, when, when something happens or something's coming up, letting AI allow you allow it to, to actually make decisions based on, you know, the, the parameters, the guidelines, the guardrails that you give it. So there's some of that there too, uh, is for multiple reasons.
So, I mean, you know, it's not just development. It could be used for financing, you know, financial areas. It could be used for so many different areas.
It, it's just, you know, think about it is just unlimited really. When, when you start thinking about the different uses for ai. And then really just think big, the best you can do, the biggest you can do, the biggest you can think about and then go bigger, because that's really what AI is.
You know, back in the old days, we had the cell phones, we had the little Nokia's, and we thought, man, this is great. I can make calls. I can do things today.
When you look at your iPhone or you look at your Android, it has so much more. And then now that you're starting to see AI on cell phones, now, it can actually remove things out of pictures automatically, and you can't even tell that something was there. So AI is, it's just growing and it, the value that it brings, again, is huge.
So be careful. All right, so let's talk a little bit more. So, AI also has pre-trained modules.
So it allows you to reduce development time. It allows you to lower your, uh, data requirements for certain things. And it really kind of gives you some of the abilities, uh, to do, uh, improved performance with your pre-trained modules.
There is a couple of repositories out there. Hugging face is one of the largest ones, and we'll hit it here in a minute, but there's 615,000 pre-trained modules already out there. Having that many modules gives you the ability to go out, grab something that meets your needs.
It's like kind of going to the candy store. You've got lots of options, and just finding the one that fits your needs and then grabbing it and trying it and using it. And so that's some of the benefits that you see from a pre-trained module is somebody's already done a lot of the legwork for you.
So it's just really just a matter of finding the right one and making sure that it is a good module. And then running with it. And again, we're gonna be talking about it here shortly.
We're just kind of doing the introduction. Um, again, some of the common repositories out there, hugging face, open ai, TensorFlow, uh, you know, py touch, uh, torch, uh, hub. I mean, there, there's, these are just four examples of many that are out there.
So using AI is like, using AI is like riding a bike. Anyone can ride it. However many people crash at the beginning.
And the reason why I talk about this is because when you're starting out in ai, you don't know what you don't know, and you're gonna grab a couple of models, you're gonna start going down the line, you're gonna be trying it out. You may think you have something that works, and then you find out, hey, the first 90% is is good, but that last 10% isn't. Or you find something, you get it hooked up, you're starting to use it.
And then you realize when you look at, Hey, how much is this thing actually costing me? I found a model that isn't very efficient. And so all of a sudden you're realizing that using this one given model is going to cost you a lot of money.
And when that first bill comes, it's a shock. And so using AI is like riding a bike. Anybody can do it and many people crash.
So what does security and development teams know about AI today? Most likely, not much. It's relatively new.
People are getting started and people are growing in this area. And so it's just really a matter of getting out of the gate and, and just learning. So let's talk about the security viewpoint, uh, regarding ai.
So for security, most teams do not yet understand, you know, AI's true value, nor are the risks. Again, AI is so new that your security teams are, are really just trying to get a grasp on this. Um, they're hearing that a lot of providers, a lot of tools actually have AI embedded in them these days.
You know, das and SAS and S-C-A-A-P-I, security, container security, you know, we're AI driven. What does that mean? AI driven?
And so, you know, when you start looking at what AI can do, you know, for these, it allows you to do things like connect the dots regarding vulnerability relationships. So now, once you start actually getting the results from these scans, AI can come together and really kind of automate and put and connect the dots. Hey, I have a finding in SAST that correlates with dast and it also correlates with API security.
Now you have a correlation across the line. You can see, hey, this is important, or I have something in my container that is also identified on my SEA and also on my das. Again, it helps you build that correlation so you can help identify or determine what do you wanna do in regards to security.
And so again, security is, is relatively new in this field, and a lot of tools are relatively new in this field. And again, it it's just a growing, uh, area. And, and it's just only gonna get better.
It's gonna allow you to provide solutions to identified vulnerabilities. And so from a security standpoint, some of these tools are gonna tell you for sas, Hey, you have a ProSite script injection, or you have ProSite scripting, and here are suggestions for how to fix it, which we'll talk about on the developer side, but that's kind of a nice to have. And so now not only is the tool just telling you, Hey, we have a finding something's out here, but now they're able to actually do something with that tool.
They can kind of help guide you, they can help build it, they can help, you know, get you going out the door, part of the, uh, the viewpoint for security ethical hacking tools. Yeah, right? And the reason why I say that is because there are ethical hacking tools that have come out, especially recently, and you have a lot of, uh, learning module, uh, large, you know, LLMs that are out there that are specifically focused on ethical hacking.
Well, that's great, but the question becomes, what defines ethical hacking? Because a tool that goes out and breaks into something, what makes it ethical? It could easily be, I call it the person in the chair, kinda like in the movie Maverick.
It's the person in the box, the person who's at the keyboard using the tool that defines is it ethical or not ethical in my opinion. But here are a couple of options that are out there for ethical hacking tools. So choosing an AI model is like eating at a buffet.
There's lots of choices, some good and some bad. So let's, let's start kind of crossing over from security and development and, and let's talk about how AI comes together. So hugging face as today, over 615,000 modules.
And what that means is anything that you're looking for is pretty much out there as far as a module goes. Be careful. Some of these modules are malicious.
Some will cost you a lot of money to use. Some are very large, some are very small. Some have been populated with great data to start with.
Some have been poisoned where somebody took something, a module that's great and corrupted it with, with poisoned data, done re you know, re-released it back out with a very similar name so that when you're looking for modules, you might accidentally do an accidental keystroke and like a typo squatting and go to a module that you're not intending. So that is one thing to, to look at, um, when you're looking at modules. So look for the popular modules.
Which modules are the most popular? Because those are the ones that people are using. Find one that has a trustworthy, uh, um, marking one that is an author verified account.
Because if somebody just created a new module today and uploaded it today, that's probably a module you wanna stay away from. And if you have an author who's been known to write or release malicious stuff, you wanna stay away from that. That's part of the reason why you wanna look at the author and what they're doing.
And again, you know, look at the trustworthiness, look at how many people have downloaded and used that module, because knowing how many people have used it will give you kind of a good guideline of, is this something good or not? One thing that, that I'll share with you that you just wanna kind of be careful of and make sure is, is it an active module? Because there are modules that have been out there for a long time that are no longer actively being used or created or modified or, or adjusted.
And so it's like opening and downloading an open source dependency. You may have a dependency that you've found that is just outta date and not being used. And if something's found in that model, you wanna know that someone's actually gonna do something about it and address it.
So let's talk about the developer viewpoint regarding ai. So the development view of AI is, is really kind of like having a partner or a co-pilot. A lot of people are saying AI is going to replace developers, and I disagree with you guys.
And the reason why is because developers have that knowledge. They know what's in their mind, they know what they wanna create, they know what they need, and they know how to do it. And AI just simply can't do that for you.
It just simply, if you told ai, go build me a car, it doesn't know anything about the insides of a car, it doesn't really know the outsides of a car. It kind of has an idea of what a car kind of looks like, but not down to the nth degree. And as far as developers go, when you're starting to look at using ai, the reason why I say it's a pa, uh, partner or a copilot is because it allows you to say, Hey, I just wrote some code, code snipet.
Never send the whole thing, but you have a code snipet and you can drop it into one of the online prompts or, or your development tools. Even have AI driven, uh, tools now these days. And so you can pass it in and say, could is, can you refactor this?
Can you make it better? Can you make it run more efficient? You may have a entry level developer who's just getting started, and they might do a whole bunch of nested for loops and there's a better way to do it.
And AI may come back and say, Hey, here's a better way to do something. Or regarding security, Hey, you just created a SQL injection. Here is some suggestions on how to properly get yourself out of the mess that you created.
This way you're actually addressing things during the development process versus once you go to QA and production, because once you hit qa, if you have to pull it back or make adjustments, you've wasted their time, you've wasted cycles and you've wasted everybody's effort because to take something from QA and bring it back, it's going to take time. Typically, once something goes to QA and it's past QA and it's now in production, you've already moved onto the next step, the next release, the next something, and that next something may be down the road one month, two months, three months, depends on your release cycle. You could be using, uh, a method where you're releasing every two weeks, every company's different, every development team is different.
And so you may just be kind of running into that. So when you start looking at finding things, try to find them and adjust them and fix 'em during the development stage because you can do it in minutes. And if you have something that shows up in QA now you still have a chance to fix it.
Once you go to real, uh, production, guess what? Those ethical hacking tools are going to probably go start attacking your stuff by other people and they're gonna find things and they may compromise you. So there's a lot of things that, that, that could happen.
And so when you're looking at development and ai, take it all the way down to the developer, IDE, you know, be on the command line, make sure that before you commit code into the repository that you're at least checking it. Is it secure? Did I do something good?
Did I do something bad? Could I do it better? Can I write something that performs better?
Because AI can give you that, that feedback and help you be a little bit better. When you look at integration with ides, you have the ability to actually, there's tools that will go right into your IDE. So as you're typing, as you're working, it gives you that immediate feedback.
So now as you're typing and you've got it in there, it's like, Hey, guess what? You just did this. Here are some options.
It's like the old days when we looked at stack overflow, you go out, you would find something, you adjust it to your needs and get going. With ai, it's the same thing. Here's a couple suggestions.
Would I take a suggestion blindly? Well, if I can see the suggestion and I'm still coding, I can look at it and go, actually, this is what I really need and just take it. So it really shas off a lot of time, at least saving you 50% or more on time.
And again, it can detect common tasks and auto generate code. It just make it a lot simpler for you. So when using like online tools or when, uh, like chat GPT or from from your IDE or just in general, you know, what are developers asking AI typically?
Well, they're looking at code quality. Is this code the best it could be? Do I have security issues?
Did I, you know, or how do I, I'm a new developer, how do I do a link query or how do I do something? Some of the the nice things that AI brings you from a developer standpoint is I could give it a database table structure and say, create the API endpoint for me, and it will, it won't give you a secure model, but at least it'll actually give you kind of that foundation or that starting framework that you can start from. And once you have that, then you can start modifying it, tweaking it, and making it yours and make it better.
Make sure that you put in your data validation if you're doing APIs, because if you're not doing data validation, you're gonna have trouble. Or if you're a multi-tenant system, making sure that, are we validating that we're, we're, we have the multi-tenancy checks in place so that, you know, customer A cannot see customer B'S data. And again, that's kind of the beautiful thing about AI, is it can build you the tests.
Hey, ai, please create all my, my validation tests for this. And it can automate those tasks for you that are just common. Once it has the, the results in front of you, you can add it, you can modify it and tweak it, it's yours at that point, but it will definitely shave off a lot of time for you.
So let's talk about some of the different types of AI that are out there. So the, the type of AI that most people are familiar with that you see in chat, GPT or others. And, and I actually have a list on, on a, uh, the next slide.
Um, you have generative AI chat, GPT. Now the thing about generative AI is every time you ask it a question, it's going to look it up, come up with an answer and give it to you. So let's talk about SQL injection as an option.
If I ask it 20 times, it's like going to 20 different people and asking the same question. Now, generative AI may give me, Hey, here is how you clean the string. Here's how you can validate the contents of what you're doing to ensure you're not doing SQL injection.
Well, the reality is, that's not the right way to do it. AI may come back and say, you have to use command parameters and use parameterization, and here's how to do it. That's your right answer for SQL injection.
And if you ask it 20 times, you may get different answers. And usually, hopefully the the right answer will be there at some point. Yeah, machine learning, the things that we've been talking about, you know, after, after a, a program has been monitoring logs or a program has been doing something or a, you know, you, you've built a model around traffic patterns, it will learn, hey, at a certain time of day you have this many cars, and if I tweak the the lights, I can actually move people through this intersection faster by doing certain things.
And so that's machine learning where it will actually take and learn from what you're doing and grow. You have super intelligent types of ai. You have reactive, you have limited memory, you have self-aware.
That's the AI that when, when I think about ai, the name, you know, the word ai, I think about the Terminator, I think about the self-aware, I think about, you know, the, you know, what could be. And you know, once you start having self-aware ai, you know that that's, that's where you know, it becomes potentially scary. You have a narrow AI where you're just performing a narrow task, kinda like what I was suggesting.
Uh, with traffic patterns. Your AI may be very narrow, that's all it's doing is learning and identifying traffic patterns. And it's just specific to one thing or facial recognition or modifying, um, or, or, or manipulating pictures or images.
It's very narrow, but it's focused. Um, these are just a few examples. The list actually is extremely long.
And I thought these would be probably kind of the, the, the easy ones to kind of bring up, explain, because these are the common ones that we hear about. The generative, the machine learning, the intelligent, you know, the self-aware, you know, the narrow stuff that we're seeing today. So some of, you know, some of the common used types of generative AI online prompts.
You gotta be very careful here because I have seen and heard story after story where people will copy their source code or even upload their entire source code appli of, of their application to these. And the problem is, is if you're using the free version, guess what? It's learning and growing from what you, what you send it.
So you're sending your company secrets, you're sending your source code, you're sending it things that you should not be sending it. So be very cautious with what you're actually sending to these prompts. But these prompts are great, and this is a list of some of the ones that I actually use.
And so the common one that everybody knows is the chat. GPT, they know the co-pilot, you know, clot is becoming more and more the Bard, um, grog chat, you know, it, it's, you know, you, you have all these and they're competing against each other trying to be the best. And the benefit is the models and the, and, and the work that comes out of these benefit, the whole community.
And so there, there's a lot of value there. But the beautiful thing about any of these, again, think of AI as being your copilot or, or your friend. You can go in, you can create an article, you can pass it into chat GPT, and you can ask it, is this readable?
Can you clean this up? Can you make this better? Or, I need to create an article, or I need to look something up.
How do you do X? How do you do y Or as a developer, hey, I I need to, you know, I I I'm a C developer or I'm a Java developer and I need to do something in Golan. How would you take this and use it in Golan?
You pass it in, it's gonna give you an answer and it's, it's, it's there to help you. So the potential is huge. You just need to be careful about what you're entering in.
Otherwise you could be sharing company secrets and whatever. And, and keep in mind, after 35 years of development, I can tell you one thing. It doesn't matter what anybody says.
When you enter something into a database, it stays in there forever. It just gets flagged as outdated, inactive, it never gets deleted. And from a development standpoint, when you have billions of rows, that deletes actually heavy.
So it is easier on the system, it's quicker, more performance based if you just flag it as inactive. And then just ignore that in your record sets. So let's talk about how AI works regarding security.
So you have your security tools, your das, your SEA, your sas, your API container and multiple other ones. It, it, I just picked the common ones out of the security wheels. So AI can be used in your IDE.
So you're shifting left. Again, it's providing the code suggestions for you. It's providing best fix locations for you.
It's providing that help that you need or that you want or that you don't even need that you or even know that you need it. And it's there for you. You have it in the middle.
You're using an SEM and AI can go in and detect or provide poll requests for identified findings. So some of these tools will actually create the poll request for you and say, Hey, guess what? We found something and here's a poll request to fix it for you.
So it's there at the SCM, which is GitHub, GitLab, Azure, those are three examples of SCM systems or in your pipeline shifting, right? It can detect and potentially autocorrect identified security findings during the deployment process. Maybe something didn't get, you know, wasn't caught during the development process or QA process, but during the deployment process or the build process, maybe something said, Hey, look, you were working on a microservice one out of 30 pieces.
But when you bring it all together, guess what? You have a security issue when you bring it all together and here's the information and you know what, I can auto correct it for you. And so some tools will actually auto correct it for you at the shift, right?
So let's talk about some considerations that cannot be ignored because we've been talking about all the great things that AI can do for you, but you've got to make sure that you're looking at the entire picture because if you're not looking at the entire picture, that's when it's gonna strike you. That's when it's gonna bite you and you didn't even see it coming. AI definitely has bias issues and fairness issues.
So depending on the model, you may not realize that you grab the model that has a lot of bias. So when you're looking at models, gotta look at the whole picture and make sure that you're trying to find ones that don't have bias. Because some models absolutely have bias.
You just don't realize it until you've implemented it. Make sure that you have a model that has a lot of transparency and explainability with ai. So, you know, why, how did we get to where we are today?
What decisions were made, what thoughts were made? How did we get to here? Because if you don't know how you got to where you are, then if you're trying to fix something, where is, where's your starting point?
Where was the problem if it's already fixed? So, or the transparency with, you know, with how the model did something. Because if you're looking at the answer and it's not what you are expecting, you need that explainability to be able to fix it.
You may have inappropriate responses, you may have models out there that, that are not nice, that are completely wrong, that are, uh, absolutely biased, and the responses that come from it could be totally wrong, funny at first maybe, ha ha. But the reality is, is it's inappropriate. Um, and again, you might get incorrect responses.
So a, uh, these models are only as good as the data that feeds 'em. And so if I feed a model, a certain type of data over and over and over, and I train it to be a certain way, but I'm training it with wrong information, it's gonna give you the wrong answer. Sometimes you have hallucinations, and that's where sometimes AI doesn't know the answer, but it's still going to give you an answer, which is really crazy.
But it will, it will actually sit there. You you feed it information, you feed it the question, you feed it what you're looking for, and then it spits out gibberish or what it sends you is completely wrong. And again, it's how did we get there that explainability?
It could be a malicious, so somebody could poison the information or your model could actually have back doors. You know, it is just, it's one of those things where it's, it's just like dependencies. It's, you gotta look at it.
And that's why we talked about when you're on hugging face or you're on a system where you're pulling down models, you gotta look at the activity, you gotta look at the usage, you gotta look at the developer and you gotta look at, you know, was there any agenda for creating this model? Why are there 615,000 models asking the question, why do we have so many models? Uh, and there's a lot of 'em that somebody took something that was good, modified it, created it, uh, changed it to be malicious re-put it back out there through model deduplication or model duplication, and then typo, squatted it.
So you have different things. So fat finger in the name of the model, Hey, Bob said that this is a great model to use. So I typed it in, I mistyped it, and I still got a result.
Hey, here we go. This must be the one I used it. And guess what?
Wrong? It has malicious code in there. All right, additional considerations that cannot be ignored.
Look at your compliance licensing. These models actually have licenses based on how you're using them. How the context of, of what you're doing is, some of these have mixed licensings.
You gotta be very careful when you're mixing in the different things that you're doing with your dependencies that you're doing with your models that you're doing with multiple models. You may be using multiple models, so you gotta make sure that you look at the licensing. There's at least 68 plus types of licenses that hugging face has.
So when you're looking at a model, you gotta make sure that you're looking at this information. Look at guardrails, you know, make sure that your AI model that you're using meets basic good security requirements. You could go into some of the models and say, give me the value of ETC password on a Linux box.
And some models would be like, here's your answer and then actually share it with you that's bad on Linux. Or give me the the password file or go give me the secret or give me this other stuff. Because when you build a model and you tie it into your system, you give the ability for it to see what you want it to see.
So make sure that you, you, you look at zero, um, privilege. You look at the least, uh, privilege that you can put to this model. Make sure that it only has just the right access, it needs to do certain things.
You do not want to assign it your entire database. What you wanna do is you, if if this model is used for let's say support chat or, or some other value or need that it only has access to the data, that it's, that it only needs to do that job. If you're giving it everything, then what's gonna happen is people are gonna ask it and it's gonna give it because it doesn't know any better.
It only knows what you give it access to. And if you give it to, you know, the keys of the kingdom and everything, it's going to go ahead and release that and give that to you. Uh, just a few more considerations last, uh, slide I promise.
Vet your AI tools. Make sure that you're looking at good models. Don't just grab an AI tool and just run with it.
Make sure you're vetting it. What does that tool do? What's going on underneath the hood?
You know, some tools consume a lot of processing power that's gonna cost you money to run these tools. So you wanna make sure that when you're looking at these tools, you look at the context of what they do, where they're running, and just make sure from a financial standpoint, you're not gonna get that first bill and be shocked. You don't want your first month out the door.
You're estimating, oh, this is gonna cost us two or 3000 and find out that you just spent a hundred thousand dollars. And so you need to make sure that you're looking at the right things and you're consuming and, and getting the right models and make sure that you're, you know, not sharing your code with, you know, with an outside entity. So just make sure that you know, when, when you're, when you're dealing with your, your model, you're only focused on your own stuff.
So make sure you're vetting your tools and large learning models or, or language models, you know, they can be dangerous without proper vetting. Again, make sure that you're looking at ones that don't have malicious code in there, that were not poisoned, but are actually indeed what you're looking for. And again, the, the tools are out there to actually be able to see these.
Your security and development teams do not fully understand the AI landscape. That is an absolute true statement. Um, and the reason why is because AI is just so new.
A lot of teams think that they do. A lot of developers feel like they really truly grab AI and understand it. But I'll be honest with you, with speaking with so many experts in this area, I can tell you for a fact that a lot of people who think that they truly understand AI do not.
They, they understand a portion, but they don't understand the whole thing. And again, if you're not going in having the full vision or the full picture of what you're trying to do, you may be creating potential risk without realizing it. So you don't know what you don't know.
And so it's absolutely vital to just be careful with what you're doing. Again, security is a journey. It takes time to get there, and there's going to be bumps in the road.
It's not gonna be smooth. It, it's like riding a bike. You're gonna fall off several times when you first start, but at some point you're gonna finally get the hang of it and then get it.
So let's go back. ai, is it a friend or security frenemy? My answer is it's a good friend.
You just need to make sure you pick your friends wisely. Thank you guys so much for watching and seeing this with me today. I hope you learn something and, uh, really at the end of the day, um, my, my hope and my goal is that you grew in your knowledge related to application security and ai.
Again, my name is Chris Lindsay, you can find me on LinkedIn. I post several times a week about application security and my tagline that I always say there, I'm gonna share this with you guys, stay secure, my friends.