Redefining Cloud Threat Detection with Anna Belak | SecOps Vision 2024
In today’s cloud-centric landscape, rapid and agile threat detection is paramount, with cloud attacks often occurring in less than 10 minutes. This necessitates a shift in security operations thinking, embracing the “distributed, immutable, ephemeral” mindset. This presentation introduces the 555 Benchmark, an innovative approach to cloud threat detection and incident response, with a goal of detecting signals in five seconds, triaging high-fidelity alerts in five minutes and responding within five minutes. Drawing insights from threat research conducted by Sysdig, Orca Security and CrowdStrike, we explore the urgency of this benchmark and its significance in securing cloud environments effectively.
In this presentation, ex-Gartner analyst Anna Belak, Sysdig’s director, Office of Cybersecurity Strategy, shares the 555 Benchmark framework and what key approaches to use.
Transcript
Hello, my name is Anna Beic and I am the director of the Office of Cybersecurity Strategy at Sys Digg, a cloud and container security company. My job is to bring thought leadership to the world and to help people innovate securely in the cloud. Today we're gonna talk specifically about cloud threat detection, but we're gonna start from a slightly unusual angle.
Did you know that 90% of commercial flight is automated? Some of you knew that, and the rest of you are never getting on a plane again. Well, that would be a mistake.
Uh, flying is actually very safe, much safer than driving and autopilot dates back to the 1930s. So we have been working on this problem for almost 100 years now. Early autopilot focused mostly on keeping the planes flying straight and level, so it wasn't so much of an autopilot and more of an assist for the manual pilot.
That was very much required. Then we added more and more sensors over the years. We started with gyroscopes and basic computers, uh, which led to a lot of innovation in the first military aircraft during and after World War ii.
And today, commercial airplanes are equipped with tons of sensors. We have added accelerometers, uh, GPS tracking systems, internal navigation systems, and then we have sensors for almost every component of the aircraft that plays any role in it. Staying a, a flight and being safe, like the temperature pressure altitude, the air speed, the fuel systems, the engine systems, all the systems have sensors that send data back to the some central analytic system that understands what to do should any of these parameters change.
There are also several layers of redundancy and safety. For example, all aircraft are equipped with at least two, sometimes three independent autopilot systems. And of course, they're vigorously tested and certified by the FAA.
Now, there are a couple things I didn't mention that are not automated. First of all, the communication with the air traffic controller can't be automated. The ATC tells the pilots when to go and where to where to go and when to go, and it notifies them if there's any change to the air traffic.
Now, the ATC theoretically could communicate with autopilot directly, but this is not legal because the human must be in the loop to ensure that they have received the instructions correctly and clearly pass them on to the plane and that the plane is behaving in the way that it should. The other two, uh, key elements of flight that are not automated are takeoff and landing. Now, lots of people think that takeoff and landing are automated because they are now so flawlessly smooth relative to how they used to be.
But that's not true. They're smooth because the pilots have gotten very good. Uh, according to the FAA regulation, again, takeoff and landing of commercial flights can't be automated.
It must be manually performed by the pilots on board. Now, this is not because it's technically infeasible. In fact, those of you who thought that takeoff and landing or automated are not completely off base, you could take off and land a plane automatically.
We just don't because on the one hand, the FAA regulation implies that it wouldn't be safe. But on the other hand, for the pilot himself, uh, manually controlling the plane during those stages is actually much easier than trying to oversee the automated system as it attempts to take off or land the plane by itself because there are so many dynamic elements that occur all the time. So, long story short, um, automation is amazing.
It empowers us to do all kinds of things that we couldn't do before or specifically to stop doing things that we didn't really wanna do before. Like the middle part of the, of the flight that is automated for the plane is the most boring part that requires the least skill and expertise and uh, kind of magic reaction time. But the takeoff and landing are the hard parts.
So the point is automation frees up our pilots to focus on the important pieces, like to deal with unexpected inputs and it can free up your talent to drive innovation as you move to the cloud. And we just got owned, yeah, that timer was a countdown. How long it takes, uh, us to execute.
Well, someone to execute, not me, 'cause I'm a good guy. Uh, the Scarlet Eel cyber attack, and I'll explain more in in more detail what Scarlet Eel exactly is. But the point is this is a very cloud specific attack.
It leverages the cloud. It only works in the cloud, and it is heavily automated. So it only takes three minutes and 42 seconds for an attacker to use scarlet Eel against you.
Now, the average cloud attack takes 10 minutes, uh, as we reported in our third report last summer. And that means some of them take much longer and some of them take much short, shorter, like scarlet eel. But 10 minutes is all it takes on average to act, to complete an attack.
I mean, this is from when you've been discovered to when there is impact to your environment. Um, now you might know from mandiant's report this year that the average dwell time for attackers is 16 days. Now that includes on-premise environments.
But the point is, if we're used to a dwell time of 16 days, then an attack that takes 10 minutes is crazy. Like we can't keep up with that. And this is a real challenge.
Now, additionally, we have the SEC disclosure rules that came out, uh, recently as the spring and that go into effect in fact, uh, in December. And the SEC disclosure timeline is four days. So that's within four days of you identifying that there is a material incident in your environment, you must disclose that to the SEC.
Now again, if 12th time is 16 days, so we don't know an attacker is in our environment until 16 days later, that four days start to look pretty scary. So what are we gonna do? Well, we're gonna defend at that speed, right?
Um, all that stuff I said earlier about planes was not just to waste four minutes of your time. One of the takeaways is that it took several decades for us to get there, right? The dream of automation dates back to the thirties.
Thirties, but it became a reality only in the eighties for, for airplanes. I mean, and then it became standard practice only in the 21st century. So we're talking the past two decades or so, and we're now seeing a very similar pattern with adoption of cloud.
First of all, it automation before cloud was really, really hard, right? Like systems were never designed to speak to each other. And so making them speak to each other was painful and ineffective.
And then cloud gave us this thing that the unified control plane and a set of well documented APIs that are designed exactly for that, for c for systems to speak to each other. Uh, in fact, if you were like me, uh, when you first discovered, you know, operating systems and data center, your question, your first question you asked was like, can we have a data center for the op, an operating system for the data center? Like can we just control the whole data center from some central brain?
And like now we do. That's what cloud is. It's amazing.
Um, but of course there's a catch, right? Actually there's three catches. One I already mentioned.
The bad guys are all over this, right? They have embraced the gift of cloud innovation. They have embraced delivering, uh, I mean we're delivering profoundly useful products to our customers and amazing experiences really fast.
That's a whole dream of, uh, you know, DevSecOps and cloud and, and fail fast. But the bad guys are able to use the same benefits to mine. Cryptocurrency really fast and steal data really fast.
So they're right there with us. Uh, the second catch is that a lot of us are moving to cloud with an on-prem mindset and a set of bad on-prem habits. And some of this is unavoidable because change is hard, especially culture change.
And then like refactoring your huge business applications is also hard. So you can't just snap your fingers and be like, oh, it's in the cloud now it's modern. So we get that.
But the more you resist, like the less you want to subscribe to the distributed, immutable, ephemeral paradigm of cloud, the more difficult it is for you to move at the speed the attackers are moving in the cloud. And the third one is, again, it's defending fast, right? So like we not only have to innovate fast, but we have to secure fast too.
And we often forget security as that last bolt-on piece of like, oh, it's just getting in my way. But the only way to move fast enough in cloud to keep up with the bad guys is to build security in. So, you know, shifts left.
I'm not actually gonna talk about that. But it's also to make sure that on the threat detection front, we can keep up with any attack that does hit our environment and block it long before uh, it gets anywhere close to a reasonable blast radius. In fact, one of our customers, uh, famously said, I don't want to know 15 minutes after I've been breached, right?
Like, that's not very helpful. I wanna know as soon as possible as early in that kill chain so that I can shut it down and contain the blast radius. So for this reason, sig and I, uh, set out to set a benchmark that will help you understand if you are in fact fast enough to defend yourself in the cloud.
So how does one create a benchmark? Well, uh, we have three key components here that come into play. Um, I used to work at Gartner, so my desire is to always get as much kind of like some objective real data from customers and end users as possible.
So that's what we did. We interviewed a bunch of customers. We actually worked with a bunch of industry analysts to get their input.
And then we went and looked at the threat research, like what attacks actually happen, how that affects our customers, and what our customers do in response to that. Uh, and so we designed from that data a very simple framework. It is called the 5 5 5 benchmark for cloud threat Detection and response.
And it is anchored very strongly in the threat research that shows that an average attack takes 10 minutes, hence 10 minutes to paint. So if you can detect a signal within five seconds of that signal occurring, so that can be any number of different signals, like it could be a log that comes outta an application or the cloud. It could be a system called, it could be something from the network traffic.
Uh, you gotta detect those signals within five seconds. And then if you can correlate those signals to each other within five minutes, so that means two independent signals occur. They may be meaningless in isolation, but if you have them together, suddenly they mean something and that's bad.
Uh, and you have to be able to do that as quickly as possible. You have five minutes to do that, and then you have five more minutes to initiate some kind of response. Now, a lot of this response can and should be automated, so you should be using the auto response actions that are available to you in cloud.
Um, but as with the pilots, like some of this response will have to be manual. So your most important, uh, component there is gonna be to make sure you can get all the data and then to deliver that data to the correct people so that they can then take the necessary potentially manual actions to, um, contain the attack. Now, our benchmark is designed specifically in a way that focuses on the challenge and the opportunity aspect, because our point is the cloud is different, right?
It's not like on-prem and because it's different that there are new challenges. So some things that come with cloud make it more complicated for us to respond fast or respond at all sometimes. Um, but some elements of cloud are opportunities, so they actually make it easier for us, or they create feasibility in places where previously they're were not possible or it was incredibly hard to do.
So we're gonna walk through these one by one. So the first piece, five seconds to tech threads and the text below, it tells you exactly what I mean. It says collect detection signals from the cloud service provider and cloud security tools within five seconds, uh, specifically to ensure visibility into ephemeral assets.
Now, if you don't have ephemeral assets, you're probably not clouding correctly. I would say, uh, most of us who are moving to cloud do have containers and lambda functions and other workloads that are short-lived, and those workloads are prone to attack. And we need to be able to know when something has happened, right?
So what is the challenge? The challenge is the bag are fast. So for us to detect within that scope of the attack completing in three minutes, 42 seconds or 10 minutes or what have you, we just have to move fast.
But most of the time we don't have the tools and we don't have all the data, right? If we just have logs, we're missing a lot of information. If we have legacy tools, we might have some data but not the rest, or we might not have it fast enough or we might not have the right context.
There's lots of things that might be, uh, missing. Uh, and then the automation for evil, like I said, all the stuff that is given to us by kind of the cloud providers and this new paradigm of op operating is really useful to the attackers and their full-time job is attacking. And our full-time job is something else.
And then we're also trying to defend ourselves, right? Um, on the other hand, the opportunity is that we now have access to a lot more data than we did before, right? So before, you know, you have to configure every single log source to get the logs.
Now you just have a pile of logs, you know, in the cloud that's just handed to you, right? Um, we have technologies like EBPF that let us get data from the system kernel. We have all kinds of visibility tech that we can use and often open source by the way.
So it's free, uh, that we can get the information out of the systems, um, hopefully within five seconds. And then we have automation for good, right? So all the automation capabilities like Terraform and CloudFormation and other kind of, uh, languages and formats for automating the provisioning of workloads, for automating the response actions for remediating infrastructure's, code manifests.
All these things are now, um, available and actually quite widely used by advanced teams. So Scarlet Eel I mentioned earlier, this is a sophisticated attack that are threat research team detected in the cloud. And what's curly Iel does, like it has a fairly complex kill chain.
This is why it's cool, right? It's very fast, but multiple things are occurring. Uh, vulnerability gets exploited, uh, and then a crypto monitor gets deployed.
But that's actually kind of a red herring because what happens next is the attacker steals some credentials and he uses those credentials to then move laterally and steal some data, right? And the, the reason this is interesting is that a lot of people will see a crypto mine and will shut it down, and you should do that, but sometimes it stops them from investigating further. So you need to collect all the signals and you need to make sure you're able to analyze those signals.
Otherwise, you might miss a lot of the kill chain of something like curly Ill. Now the second piece is five minutes to correlate and triage those signals that we just correct collected. And the point here is that we have to get them and connect them to each other very, very quickly.
But we also have to understand that we aren't gonna know necessarily which ones are good or bad immediately, right? So the challenge here is that we're gonna have many, many different services, right? Like we have data from lots of places.
I I kinda just said like, I need data from lots of places, but the point here is selecting which data is gonna be relevant is hard. And we have, we run the risk of overwhelming ourselves too much data and not enough signal as now we're just kind of drowning in noise, right? We might have a multi-cloud architecture.
Lots of people are in multiple clouds, whether by choice or by accident, mostly it's not by choice actually. Um, you might have to collect signals from different providers or different SaaS services or other, uh, sources of information that may not be just like in your collective choice, right? Um, and then lack of security context.
So lots of the data that's available to us is very useful, but it's not actually designed for security detection and response, right? It might have been designed for monitoring or troubleshooting or some other reason just auditing, right? So using that data for security is a good idea because it could have relevant information, but it's sometimes very hard to understand that data in context.
Now, on the opportunity side, again, we have API based access to all kinds of things. So we can just query an API and get stuff we want in seconds. And that's amazing.
Um, we do have identity as a key control. So this is really cool because on the one hand, uh, we're not so good at it yet, but we're working on it as an industry. Um, if you can track an identity across boundaries, right?
If you can see somebody using the same identity to do bad things in different contexts, then you can quickly connect those events to each other, which is much more difficult, um, in, in the, the old world. Um, and then we have artificial intelligence, right? So when we're talking about correlating signals that go together, um, it is kind of the golden age of ai.
And AI is the, the thing that is good at correlating things to go together. So hopefully, um, things that we previously like squint at dashboards and hope we can connect things or kind of guess in a bunch of if and statements and, and clunky rules. Now we can kind of leverage machine learning to accelerate the, um, innovation in the correlation of relevant signals.
So this example again, is showing scarlet el. And the point here is, first of all, that you have two different data sources that you needed, right? If for us to detect Scarlet El, we needed the cultural logs, happens to be in AWS, could be any cloud, um, and we needed a system called a machine learning now in isolation, right?
So each of these signals, if it's detected, has about this level of severity, right? So sumeral lateral movement, you see there's those are green because it's like there are legitimate reasons why that could be occurring, right? Reconnaissance is a little orange, but like things that look like reconnaissance could also be just activities that are performed by the user interface to kind of enumerate things for you, right?
And it comes to, minding is always bad, but it's kind of like not that interesting, right? Crypto mining by itself, like whatever, shut it down right? Now, the point is that when these things happen in a particular order or in connection with other things, then we should be able to presume that the combination of those signals is, uh, is worse than them in isolation, right?
So if for example, crypto mining is bad, okay, I shut it down, but assume role and reconnaissance together are already worse than either of them independently, right? And then if we add reaching out to an EC two metadata service to those two, now we're in the right, right? Like we're, oh, okay, we're recon maybe nothing.
Assume role, maybe nothing. Assume role metadata service, maybe something. And then when we add all these signals together, now we're in the deep red.
But like this is almost certainly bad. This is not normal activity. No cis admin should be doing this for sport.
Um, and so what, what the challenge here is though is that you want to be able to interrupt this skill chain as early as you can. So as soon as that is like right enough for you, which again, maybe up to your specific organization's risk appetite, you wanna shut it down because once a kill change is over, you can look back in time and say, oh yeah, this combination of things is terrible, which is like what our team did, right? Like this attack when they first saw it had never been seen before.
So they kind of retroactively went and studied it and found all this stuff, but you know, they're doing research and you're defending a real environment. So you wanna shut this down before the exfiltration of your intellectual property occurs, which is how this attack actually ends. The last part, and perhaps the most important part is the five minutes to initiate response.
And by that we mean that he used the flexibility of the cloud to initiate tactical response actions within five minutes of this high fidelity detection. So if we're able to correlate all the signals in the last part, uh, now we can take some action in response to those signals. And, uh, ideally as soon as possible, as soon as we suspect any weirdness, we can take some action to at least prevent, um, to control the blast radius, if not fully stop the attack.
And the challenges on this side come from huge complexity of environments, right? One of the cool things about cloud is it lets you create and scale things really, really quickly. And that's amazing, except that we take full advantage and now we have these huge, very complex environments and it's sometimes very hard to tell what all the signals mean or which things go with which other things and and so on.
So actually extracting information from those environments and then taking appropriate response actions on the correct scope is quite hard. Um, we also have ephemeral assets, like I mentioned earlier. So these things come up, they perform one task, they disappear, they might live four seconds.
The average container nowadays lives for, I wanna say less than five minutes. It's really short, right? So if you use the virtual machines that may be up for hours or days or weeks, um, that asset may be gone with all of the data associated with it.
So you need to be able to get the information, but then also to take actions on things that may not appear. So then you have to go remediate them somewhere else. You have to prevent them from being deployed not to shut them down, right?
And then we just see inadequate press and tooling, right? So cloud is still quite young and security in the cloud is still quite young. And so threat detection is analogously also young.
So all of the systems that we have built out for our on-premise SOC are not really yet ported to the cloud. So we are sort of in this phase of innovation and building these systems out and adapting them to the new mode of operation. But most of the tooling is just not really there, right?
So that will come, um, we're working on it. And in fact, SOAR systems are much more feasible now because like I mentioned earlier, things that were hard to automate because we didn't have the APIs and the control plane now make it much easier because we do. So, um, I'm sort of getting ahead of myself on the opportunity because we can do auto remediation, we can make pretty extensive and pretty cool playbooks, uh, in the cloud to take response actions without human intervention necessarily.
In, in many cases, um, the deployments are repeatable, right? So there are no more like handmade bespoke configurations, at least I hope not. Uh, because everything you have built is codified in hopefully cloud formation, terraform kind of templates that are hopefully stored in some version control repository and labeled so you know exactly what they're for.
And so for you to bring up a set of workloads or an environment that you had shut down or that may have been compromised, it should be just a click of a button essentially. And then that also gives us this built-in resilience. Like I said, you have, um, infinite compute available to you, so you can always provision another one.
And if you have workloads that are tainted for some reason, it's very easy to replace them, uh, because that's sort of, again, the whole point of cloud is to have resilience at scale for, um, anything and everything you choose to build there. And so this slide, uh, shows that same kill chain, right? We're gonna talk about what you could do at each stage to have either stop the kill chain or to at least contain the potential impact of this attack, um, based on what we saw.
So the exploit of vulnerability, like you're not gonna see that. You're gonna see the post exploitation activities. And so the first thing we actually see is, uh, reaching out to the EC metadata is C two meditative surveys.
Now, when you see this, you may just want to kill the container that tried to do that, right? Um, maybe that's normal activity and maybe it's not. But if you're feeling like it's a little suspicious, you could just kill the container, okay?
When you see assume roll, now this is one of those where this could be perfectly normal because lots of people are allowed to assume lots of roles for normal reasons. Um, I believe in this case, this was a machine role that was assumed, and this is kind of unusual, like normally you don't assume machine roles. Uh, but on the other hand, like if this is legitimate activity, you don't wanna shut it down.
You don't wanna just like disable that role because it might be necessary for certain activities. So what you can do is you can add restrictive policy to the role so that maybe that role can still be used in its normal scope, but this particular IP address that is associated with it right now maybe has limited scope or maybe can't use the role at all, okay? Reconnaissance.
Now, reconnaissance is often, like I said, normal because there may be things that look like reconnaissance that aren't, but here again, if you see somebody doing recon, you can reduce their permissions just in case, because usually if they're doing recon, they might be up to something no good later. So it's better to just prevent them from having the ability to do no good things later lateral movement. Um, this gets into quarantine.
So when you see lateral movement, it's, you're pretty certain that's bad. So you want to prevent that account from, um, having access to things maybe. So this is a potentially a quarantine situation.
And then crypto mining is always bad, so you kill a process. But again, uh, keep note of the fact that just a crypto mining by itself may not be the full scope of the attack, right? So if you see a crypto miner on a note, in fact, this is like one thing you can kind of look for.
If you see a crypto miner alongside other weird behaviors, it's almost certainly worth investigating and maybe, almost certainly worth quarantining that node or that account or whatever's going on. Uh, because as we showed in our first report, it costs you like $53 to mine, $1 a crypto coin. So $50 of AAWS cost.
Uh, so there's no legitimate reason to mine. So when you say mining alongside other weird stuff, maybe that's like really bad. But anyway, the point here is that you can actually take, uh, a series of automated response actions in this fairly complex kill chain that would either prevent or limit the amount of impact that this has on your environment.
And that's sort of the story. So, um, again, 5, 5, 5 benchmarks. So five seconds to detect signals, five minutes to correlate the signals together to get a high fidelity detection, and then five minutes to initiate response to, um, potentially very complex skill chain, somewhat automated, somewhat manual to make sure that you're able to prevent attacks before the attackers are able to complete them.
That's the game. And what you need, you need a cloud mindset. You want to embrace the immutable, distributed ephemeral paradigm.
There are actually many ways in which embracing that makes you more secure by default, if you will, right? Like if a workload is immutable, that means it shouldn't be changing ever. And so like if you're using a container and you know it's immutable and then you see that container drifting from what it's supposed to look like, you can kill that.
Like almost certainly that's bad because you know that it's immutable, right? So we get a lot of benefits like that from embracing the paradigm. Um, it also becomes a lot easier to bring those things back, right?
When the whole infrastructure and kinda workload configuration is programmatic. Then if you were to shut it all down, you can bring it back with just a few clicks, right? Like, yeah, it's gonna take a few minutes to provision all the infrastructure, but you know, it's not like back in the old days where you had to, you know, go rack new servers because your servers were dead, right?
So we just still have to keep track of downtime, right? Because downtime may be impacted business, like our company might be losing money for down, but in theory, in cloud downtime is much, uh, much less of an issue because we can always provision more, um, more systems to run those same workloads again. Uh, the other caveat out here is we require a platform approach.
Um, and that involves data ingestion and detection engineering. Now, I'm being kind of, uh, vague here on purpose because I don't wanna say that you need to buy a specific product or specific suite of products to accomplish this, but the reality is you need sort of like a central brain that is able to connect the pieces together and correlate them and allow you to actually operate this quickly. And some of you immediately thought like, oh, I have a sim, and like, yeah, you need a SIM and you probably still need a sim.
But in many cases, a SIM may not be fast enough because um, by the time your data got to sim, that attack may have actually completed. So you sort of need a system that is able to connect all the relevant pieces of information. And by the way, that includes your DevSecOps pipelines, right?
Which often are not in scope for SIM or traditional security operations. So you need to be able to connect all that information in real time. And then you also need to have a system of record where you can keep data longer, should you have to do long like investigations to look back more, uh, further in time.
Um, and then lastly, you need automation of enrichment and response actions. Uh, but it is more feasible than ever, right? So the ability to take a piece of data and connect it to some other piece of data or piece of asset context, right?
If I have something that's happening on a compute instance and I know that compute stances of vulnerability because I didn't patch it last week 'cause I made an exception, like that piece of information is super easy to connect automatically to that event information, right? If you have them all in, in some single, you know, pin of go glass, I'm sorry for saying that. Um, so the point is like a lot of the tools are right there in front of us and we can just grab them and connect them, and that does require work.
Uh, but on the bright side, like once we put in that work, then we can free up our talent to really drive innovation, right? So we can not just fly the planes or take off the landing. We can be like blue angels and really soar.
Uh, and I hope that all of your organizations are here to, um, soar and that you would much rather build all this automation right now so that your SOC teams and your, uh, security teams can actually do the hard work that has to be done manually. Uh, that's all I have for you today. I hope that was interesting.
The key idea of course is uh, we are aiming to secure every second at sig and we hope that you can secure every second, uh, in your infrastructure by trying to meet the 5, 5 5 benchmark for cloud threat detection and response. Uh, please go check out our virtual booth where you can get your very own copy of the benchmark, uh, in detail. com slash 5 5 5 4 more information.
Uh, thank you for listening. It's been a pleasure.





