How to Prepare and Manage Vulnerabilities in the Cloud with Ben Nicholson | SecOps Vision 2024
Does an on-premises security strategy translate to the cloud? What are the differences and how can you secure your cloud resources and prevent breaches? In this presentation, Palo Alto Networks‘ Ben Nicholson explains how to secure your cloud and how to prepare for any breaches that might occur.
Transcript
Hello everybody. My name is Ben Nicholson and I'll be talking about breaches today, how to prepare and how to manage vulnerabilities in the cloud. So, a little bit about myself.
I am, uh, Ben Nicholson. As I mentioned earlier. I'm a global practice leader at Palo Alto Networks.
I work in Prisma Cloud. I've been at Palo Alto Networks roughly about eight years or so. So today we're gonna talk about enterprise architecture and application framework, modern SOC and enterprise's, current state identity management and vendor consolidation, pre breach versus post breach activities, breach response attacks and protections and vulnerability management.
Okay, so if we look at the enterprise architecture or the standard enterprise architecture, obviously that's changed a lot over the last five to 10 years. Previously, you know, we were all working in a data center mode. Now that a lot of companies have migrated to the cloud environment, there's a lot of considerations as people have moved into the cloud.
Now, this is something I really want to talk about, and this is the overall application lifecycle. Why are people moving to the cloud? This is a conversation I have a lot, and as a part of my role, I get to go to a lot of different CISO roundtables and things like that.
And a conversation that comes up a lot is, why are we moving to the cloud? What are the reasons? And cost is very rarely ever the one that people are doing to move to the cloud.
People are not moving to the cloud because it's cheaper. The reason they're moving is because it's designed for rapid iteration and continuous application rollouts. They wanna move faster.
They want to take advantage of all these new great, uh, cloud capabilities. So if the reason why we're moving to the cloud is for rapid iteration, continuous application rollouts. One of the things that I found very interesting in the state of cloud native security report, it stated that 42% of customers were deploying code into production on a daily basis, while 72% were deploying code into production on a weekly basis.
So if we think about that, we have a large number of people that are moving to the cloud that are deploying code and updating applications on a daily and a weekly basis. What is that doing for our vulnerabilities? It's skyrocketing our vulnerabilities 'cause we're deploying code so quickly.
So now we're in the cloud and our vulnerabilities are much different. The way that we approach security has to be different in the cloud than it was in the data center. The way we approach it has to be different.
You know, when we talk about vulnerabilities and we talk about earlier in the application lifecycle, people have been writing bad code for a long time. People have been writing bad code, deploying applications, causing problems. We said, oh, this is vulnerable to ASQL injection attack.
We'll just put a firewall in front of it. But now, because we're designed for rapid iteration and continuous rollouts, the way that we were approaching it previously is not the same way we can approach it. Now, the way we approach it now is we need to get involved in the security of our code and build side or in front of our DevOps teams.
Now, from that perspective, when we say, okay, we know we need to get involved because we're having issues, we're having these skyrocketing vulnerabilities from the DevOps side of things, where is their concern? Lie? Well, from the DevOps teams that I've worked with, their concern lies in deploying applications quickly.
They wanna get things up and running. Their concern is not security. My concern is security.
So initially, you know, I, some of have a background. I was a head of a security team before I joined Palo Alto Networks. And as a part of that, I, I realized that this might be a problem.
And I put together this security meeting where I was like, we're all gonna meet together once every couple weeks and it's gonna be kumbaya. We're all gonna work together on security and we're gonna figure out all these security issues in the organization weren't gonna resolve 'em, right? But the reason why it didn't work is because I came in with a mindset that everybody in that room was gonna have that same security mindset and they were gonna look and that's where they were, where, where their goals were.
But that's where my goals were because I'm part of the security team, the dev team, their goals are rapid iteration or other teams are in there for securing their own budgets. When we're working with other teams and their goals are not security, we have to figure out what's important for them in order to be successful. So in other words, what I found is if I go to a dev team and I say, let's talk about your security review process.
Where are you currently doing a manual security review process and how, how much time is this taking When you get vulnerabilities, are they going back to you? Is a ticket being created? Is this wasting a bunch of your time?
Is this slowing you down? Because if it is, we can use a automated security review process and integrate security into code and build so that we can speed you up. So as a part of better security, we're also going to make you faster and help you with rapid iteration.
That's the best way in order to approach it so that we're figure out ways that we can work together that's beneficial for both parties. Now why am I saying all this? Why am I talking about vulnerabilities in the cloud and the proliferation of them and them skyrocketing?
The reason why I'm talking about it is because the good guides are losing right now what we're seeing is a large number of attacks in the cloud environment and a lot of different vectors. We're seeing CSCD pipelines being hit up a lot. Poison pipelines, hijacked pipelines.
We're talking about, um, you know, s three buckets. A lot of different things that are going on, ransomware that are going on in terms of both data center and cloud environments. So what's going on is modern.
So are being overwhelmed with alerts and they're not stopping enough cyber attacks. Going back to my original point. If we have a skyrocketing amount of vulnerabilities and then we just say, oh, we're gonna send them all to the soc.
The SOC doesn't have a magic button where it can just fix everything. 28% of alerts are being ignored. Less than 30% of SOC teams meet their goals for key metrics.
And the average days to identify and contain a data breach are 287 days. So it's not working the way that an average enterprise has a large number of assets and they're sending all these alerts over per day. It's not working.
In order to remediate this, we need to understand how do we reduce the amount of vulnerabilities? How do we target those vulnerabilities? How do we do vulnerability management, which I will get into later as well.
Additionally, in terms of issues that we're seeing in terms of cloud environments, it is, uh, cloud infrastructure entitlements is overwhelming. The number of machine identities have exceeded the number of human identities. So when we're looking at identity management and identity management has been an issue for a long time, but it's becoming even a bigger issue.
When we go to the cloud, 95% of identities used less than 3% of the permissions that are granted. Now, why does that happen? Why do we have so many of these overly permissive rules?
Well, when you're moving, one of the things that I found is that we're asking a lot for our security engineers as they're moving into cloud environments. Uh, my background was I was in networking, then I went to security, then cloud security, right? And early on in my career, specifically when I was learning how to do pitch firewalls or I was just learn understanding firewalls and what rules to be, to be put in place and things along, along, along those lines, to put myself into a situation like where we, which we expect our cloud security engineers to be in right now is a lot.
We're expecting them to understand networking, understand firewall and cloud security, understand identity management, understand all the stuff that's going on at A-W-S-G-C-P, maybe Oracle, I am, you know, there's, uh, uh, you know, Alibaba or whatever cloud infrastructure. You're using Azure. And then, and then be able to correlate all those informations and be effect, uh, correlate all that information and be effective.
It's a lot. And as a result of that, what winds up happening is I'm gonna go deploy something into AWS, how am I going to go deploy it? I'm gonna say, well, I'm gonna use a terraform script.
I'm gonna deploy it up into my AWS environment. And then I'm trying to understand the permissions that are required. What winds up happening is that I say I'm gonna select specific permissions and then it doesn't work.
And it says you need, it gives you like a random error that's like you permissions, blah, blah, blah, blah, blah. And you're like, okay. And then you wanna keep on adding more permissions to like a rollar or something like that.
And then it says, and then a lot of times people while use wild card characters, well winds up happening the next time they want to deploy something, they use the same rollar or they use additional permissions as, as a result, you wind up having these rollar that are the overly permission with a one to wild card characters. And you can go online and see all sorts of different hacks that, you know, address this type of thing. And in addition, when you're gonna go deploy, uh, I know a lot of times we have lab environments, people come to me and they say, oh, I need access to this lab.
And we're like, oh, what type of permission do you need access to in GCP? And it's like, well, I, I don't know. I need to deploy some instances.
I need this, I need everything. Right? So whats up happening is you start off small 'cause you wanna do lease privilege and people say, oh, I need access to this, I need access to that, I need access to that.
Not everybody has time to say all these specific permissions. And so they say, okay, we're gonna do this stuff overly permissive. And as a result of this, we have all these overly permissive roles and users in, in the cloud environments now, and this is creating a security risk.
So if we look at how do we resolve this, what we need to do is remove the permissions from all alarm that are not in use. If you scan your work cloud environment and look at permissions that haven't been used within three months, within six months, within nine months, you start removing those permissions that are not used. You can start reducing your re risk factor vendors and tools.
Consolidation is coming. So one of the things that I found early in my career as well is there was a lot of talk around defense and depth. Don't put all of your eggs in one basket.
Make sure you're getting all the security intelligence from all these different vendors, right? That's a result. What we have is 41% of organizations work with 10 or more cybersecurity vendors.
39 is a vendor average 31 tools is an average, right? That does not work and it does not work for a number of reasons. One of the reasons, what I found in my career is that you would say, okay, we're gonna get all these great tools.
We're gonna, you know, get in depth on here, here, here. Then the people on your team wind up believing the person. You know, Bob over there, he was really great at that tool.
Then he leaves and you're like, oh, what, what did that tool do? I don't know, Bob was great at it. And then, you know, it ends up getting added to another person's, you know, um, project list.
So then somebody who was managing three tools now is managing four or five tools and they wound up managing this tool. Well, this tool, well, this tool sort of, I know what's going on. This one, I don't, no idea what's going on there.
And so what winds up happening is you think, oh, I'm gonna get all this great information, but you want 'em not, that's not what winds up happening. And if I can look back on my career, I can say this a hundred percent that when I look back at all the different tools and how they weren't correlating data correctly, we would've been better off if we had just gone with a platform approach. Even though we had all these different vendors being like, oh, use our point of point of, you know, point to product and this is gonna give you this great information.
At the end of the day, if we had just done a platform approach and we had been able to see everything, a single tool, we trained everybody on that single tool, then we all would've been in a better place. Now I'm not saying one tool for everything, and that's just gonna fix all of your issues, but vendor consolidation and tool consolidation is where the industry's going right now. Every time I'm going to any event, everybody I'm talking to is how they're trying to consolidate how they're saying, okay, this isn't working and it never worked before, didn't work before, it definitely doesn't work in the cloud.
And that's where kind of the industry's going right now. And what are the cost of an action breaches are guaranteed. The number of percentage of organizations that attacked in the nine in the last year, 96%, 33% of security professionals experienced operational disruption as a negative consequence of a breach.
4 million, which I actually think is pretty low. Uh, where comparatively to where it really probably should be. So let's talk about breaches.
So when we talk about breaches and what happens during a breach, one thing we want to talk about first is pre-B breach activity. What should we do before a breach occurs? Because I've been a part of several different breaches when they occur and it becomes crazy town.
So you don't wanna be trying to figure out, oh, a breach occurred, what do we do? Now at that point, you really want us talk about pre-B breach activity. One thing I wanna talk about is normal activity can erase some evidence before triage collection occurs.
What does everybody do when they find that they have a virus on a machine? They say, oh, wipe the machine. If you have a hacker in your environment, did wiping the machine have any effect?
Probably not. So if you're thinking about what type of activity might cause problems that might erase evidence, I know earlier in my career I was at a company and we had a server that was sending, uh, FDP traffic to a, to a place that definitely shouldn't have been sending FDP traffic. And we're like, oh, there's this big problem.
As soon as they discovered somebody ran into the data center and they took the hard drive out of the machine, and it was like, okay, now. And we started talking about what did that do with the evidence? You know, oh, you can't touch this, this has to be gone to a third party.
It kind of basically destroyed all, all the evidence. So when you start talking about what type of behavior you should have before triage collection occurs, everybody has to be clear on what the rules are there. So when we talk about free breach, breach activity and post-breach activity, security audit and pen tests are going to be really important.
Security awareness training, security controls, and access controls. I know earlier in my career where we did a, we were doing data disaster recovery testing and we said, oh, this is what we're gonna do. We're gonna do X, Y, and Z.
We're gonna do all this great testing and we're gonna figure out what happens. Remember, we did the test, we broke more stuff than we thought was gonna happen. So if you haven't done the test ahead of time, you think, oh, this is the way it's gonna work.
And then what actually happens is not the way it works at all. So make sure that you're doing security audit pen tests. Understand what are your policies and procedures when, when breach activity occurs, activating the RIR plan, IR plan, isolate and contain.
Again, testing is gonna be really, really important. I've been a part of this as well where we say, oh, we're gonna isolate and contain all these things that we isolated them and then we couldn't get access to them anymore. And then we're like, oh, now what?
So again, and then we were like, where were these? You know, so, and testing these things all out. Activating RRR plan, isolate, contain understanding what are gonna be your processes.
Again, gathering and preserving evidence for legal and determining the root cause, fixing the immediate problem, coordinating with law enforcement, and again, con, providing continuous communication stakeholders and the post-incident review. Now, when it comes to post-breach activity, a thorough investigation and conducting a lessons learned, it's really, really important that when you do this, that we're not getting into the blame game, you know, of, you know, Bob had just done his job. Maybe we won't have to deal with all this in the first place.
A lot of times with a lessons learned where we get in, I've gone in front of boards before and talked about this is what happened, this is what happened, this is what happened. It's important to say, let's put together a remediation measures that are actually going to prevent similar breaches that are going to have a real impact. A lot of times people are searching for answers and so they start saying, oh, well let's do this, let's do that.
Let's additional these 10 additional policies and procedures. That's not really gonna fix anything. If you really talk about, okay, if we're going to implement these remediation procedures, fix the issue and then move forward.
That's the way to really kind of, uh, enforce these things. And then of course, updating the monitoring systems to catch similar attacks, updating the IR plan. And then of course we talk about breach assessment.
Breach preparedness, your breach response plan. First start off before anything occurs, identify the cross-functional stakeholders, including corporate communications, legal teams, and third parties. And then assign a timeline if, when a stakeholder should become involved in how they shouldn't be initially notified.
What you don't wanna have happen, and I've had this happen, is breach occurs. Everybody's trying to figure out who to talk to, when they should talk to what's going on. You wind up having like 10 people in your, in your cubicle, whatever, trying to be like, we should talk to this person.
Look for this person. You know, like, everybody's like, no, we're following the plan. These are the people that need to be communicated.
This is where we are. We're gonna step two, three. Uh, this particular timeline is when we're going to involve this particular team, have everything in order ahead of time.
If you try to try to figure out what the plan is after the breach, everybody's just gonna be in crazy town. Details of the information to be collected and shared by the security operations teams are to along are defined, along with the SecOps commander responsible for providing information to stakeholders. Of course, information about the frequency of updates, method of updates and communication processes are detailed.
And then training and policies are created to prevent leaks of breach detail beyond the breach response team. So again, breach response plan, activate the incident response team, isolate the affected systems, collect and preserve evidence, assess the scope and impact of the breach, notify relevant parties, contain or remediate the breach conduct. Post-incident reviews.
Update the incident response plan. That's your eight step breach response plan. Now, if you've ever been a part of one of these events where something happens, generally speaking, when you're looking and reviewing these type of API logs, there's not gonna be something that says breach occurred here.
Right? There's going to be, automation is required. You need to understand what happened, what was the event name, what services interacted with what date?
Again, timelines are really important. What date did it happen? How did they get on there in the first place?
What were the actions that they were taking? What happened during the breach? And then what happened after?
How did they move laterally after that? Understanding the entire part of what, what is occurring is gonna be important. So you can say, okay, these are the type of methods, methods that you're using.
So when we talk about remediation, remediation is not wipe the machine remediation is understanding what's going on and how they're interacting and how we can fix it. And again, constructing your adversary timeline, understanding where they're going, what they did, where they're going afterwards, and what type of attacks are being used on service Exploit vulnerabilities. Are they doing privileged escalations?
Are they doing lateral movement and container key container protections, exploit technique prevention, malware prevention, malware, uh, or extens of said malware prevention, malware, sandboxing, all that good stuff. Now, when we talk about vulnerability management, this is a really, really important part where we talk about, again, I was talking about this earlier on, but understanding where, so that we are not wasting time determining who, how, and what, who are the teams to be responsible receiving the alerts? What's the fixed process?
How will we prioritize each item? How do we scale, move forward towards infinite tools, building out for automation, iterating on improvement to the taking system? The thing with processes is that, if I think about it, anything that we, a lot of times early in my career, we were just always involved in this kind of, oh, there's an incident.
We gotta, we gotta, you know, we're working towards things all the time, but we're really focusing time on processes. I recommend at least spending one to two, you know, a couple hours per week, just focus on processes, focus on vulnerability management, making sure that that's perfected, that you have the right SLAs that's being assigned to the right teams. That they're actually resolving the vulnerabilities in a way that is, is effective for everybody.
So that across the board, you're, you're doing the right thing for your organization, that you're managing vulnerabilities in the right way. And the only way to do that is to take time towards it. Again, customer needs to see the way of data that integrates with their existing tools and processes to reduce alert fatigue and pre improve team cohesiveness and make the data gathering tooling more effective in the customer environment.
So thank you everybody for, for your time. I'm hoping this was really, really helpful. And yeah, thank you.





