Incident Management in Modern IT Environments – Cody Cornell, Swimlane
Cody Cornell, chief strategy officer for Swimlane, explains what makes incident management in modern IT environments so challenging.
Transcript
This is texturing TV. Hey guys. Thanks for the thrill.
We're here with Cody. Cornell is the chief strategy officer for swimley. We're talking about incident responsive management and how things are getting a little bit more complex in this exciting world of cyber security that we currently live in Cody.
Welcome the show. Thanks happy to be here. 15 is always been an issue whether you're working in traditional it or dealing with cybersecurity incidents, but it seems like it's becoming a much more pressing problem lately.
We're generating more alerts than ever. It seems like we have more things that are monitoring different things and that all sounds good, but sometimes it's too much of a good thing. So what is the best approach to Incident Management these days and cybersecurity and what should people be thinking about the kind of limit that the tea without compromising the quality of the capability?
Yeah, and I think it's something that a lot of organizations are looking for answers to right and in you obviously that the vendor Community is large and Broad and there's lots of different voices out there but to your point, you know, the amount of telemetry we're producing right now is unprecedented and it doesn't look like it's slowing down anytime soon. And I think most organizations have come to the kind of conclusion that it doesn't matter how many people I have there's still too much information to manage. I have to find ways of doing prioritization and Automation and things along those lines to do it.
So I think what we're seeing folks kind of adopt is, you know, obviously from our perspective being an automation company that they feel like that's a real tangible way to reduce the amount of work the teams have to do manually just the there's probably you know, 100 or 120 different use cases were organizations are doing you know, multi-step processes by hand be it in Excel or email or notifications or any ticketing system and they realize that's not a good use of those very kind of rare and Offensive resources and they're really looking at ways in which they can offload that work to different systems. And also feels like people have more Point Solutions than ever and each point solution generates more alerts and try not that you know one incident can have 500 different alerts get generated and we don't really have a way of kind of coalescing all those alerts in this one common thing that I can understand and act on so what is the problem with our alert systems today? And how do we streamline this whole thing?
Well, I mean, I don't know that there's something wrong with the alerting systems. I think it's just the nature of how attacks and data loss and all these things happen is that they're multifaceted right? I mean, you've probably see elements of data loss data exfiltration account compromised lateral movement all these things, but you see them in your Cloud logging you see them in your sim.
You see them in EDR you see you see them throughout all of your threat detection approaches and you know, thank goodness. They're working right that they're sending alerts, but generally It's associated with one set of activities or you know, a single activity. So I think what what organizations are trying to do is figure out how do they correlate that together?
They not correlation from a sense of how do I detect that behavior? But how do I correlate all these different alerts together into a single case so that I can work them cohesively and you know, we don't have a bunch of wasted effort. So what we're seeing is, you know people using things like what we call like alert correlation intelligence correlation.
That are kind of part of their case management system, but also kind of part of their their daily operating procedures is how do I actually look historically at all, the things that have happened in my environment and how much how many of those match how should that should be brought together? Are they associated with each other and the tricky part in that is that you know attackers are smart, right? They're trying to use deviations.
They're trying to manipulate their infrastructure and the techniques they use so they can't be correlated. So you have to bring in some things that you probably can't do as a human right things like fuzzy matching and things like that to correlate these things together so that people can you know, see that they're Associated but work them as a kind of a cohesive set of activities. The it environment itself is also becoming more Dynamic.
We see things like containers and serverless Computing Frameworks and multiple clouds. And can we keep Pace with all that and we're collecting more data than ever but I just have to wonder if with all these workloads so highly distributed you have multiple problems a day then that are unrelated to each other and you know, are we going to be overwhelmed by that as well? I think teams are overwhelmed and I think that kind of the nature of how we manage infrastructure the expansion that infrastructure you mentioned, you know, we've all been through mobile and virtualization but now serverless and containers and Edge compute and all these other elements are obviously driving a lot more data from a logging infrastructure perspective.
There's monitoring tools for each one of these they're producing things that we have to respond to. So I think folks have to find a way to keep up. They don't really have a choice and and what this is doing is really changing not only the amount of information they have to process but also the way that they have to remediate you know, historically we thought about Remediation in the sense of I have to disabled user account or isolate a workstation or a server but those things are physical devices.
I can do that on a network remediation is much move much more into kind of the infrastructure change management process that is much more real time from you know, Devar devops SRE, you know, get up perspective and moving your remediation thought process. To you know, make sure that you have connectivity and and access to those and the permissions to do things but also understand how that's going to affect your infrastructure if you're making automated changes in real time. So I think it's both a ingest problem.
But also our mediation problem that that folks have to Define answers for and part of that is how do you reduce the amount of things you have to do by hand. But also how do you make sure that you're getting high fidelity information you're responding to so that's both a tuning of threat detection, but also a automation of you know detection of bad behavior. Is too much of our focus on the incident as an event to be remediated versus an ounce of prevention and I'm asking the question because for example, we see tons of misconfigurations that are out there and that becomes the fundamental original sin as it were that gets exploited by somebody right?
Maybe you know is the definition of an incident need to change when it's the actual configuration is misconfigured. That's the incident and we need to get in front of that versus the actual Cyber attack that and inevitably inside. You know, I totally agree and I I think you see people trying to make the shift but I think it's also a hard shift right is the incident the fact that someone exploited a misconfiguration that caused data exfiltration or what should the incident been on the initial Miss misconfiguration.
I think it's kind of to the point of the question and I think yes and I think folks are thinking about that and how do I take programs that are you know, if it's system hardening vulnerability assessment data loss prevention threat hunting all these programs that you know over time. Maybe have been you know, hey, I run these quarterly or weekly or at some time interval and move them to continuous activities to be proactive so that when I do see that something is misconfigured. I actually proactively trying to reduce that kind of risk or that that threat exposure there and that moment as opposed to waiting for it to be exploited to take or take response.
Right? So, I think that desire to move to becoming proactive is is really where people are trying to go. But the the boat anchor in that is that it's a lot of work.
So again, you know, obviously we have our our bias here at swimline because we think about automation all day every day, but the way that folks get in front of that and kind of relieve themselves of all that work that they could never get to the system their backlog or sits on the back burner is to find the things that can be automated. It can be checked continuously that don't require someone to go and do that by hand so that you don't end up in just what you said right that that configuration that actually ultimately becomes exploited. Where does this incident response capability lie within organizations?
Because in theory maybe it's part of a devops workflow. But that seems to be focused more on the before things are deployed and maybe more updated but then there's a whole security function that focuses on things and runtimes and production environments. But where does this incident response responsibility actually my or should reside.
Yeah, I I think It's a general consensus that like security is a team sport, right? You can have a head coach. You can have players that work in particular positions, but it doesn't work if it's not working in concert.
So, you know, if you don't have deep relationships integrated processes, you know stand-ups with your you know. Peers in you know infrastructure in devops and it and audit compliance privacy. You name it?
It's really not gonna be effective. There's there's no way you can do the job. Well, I don't know if you can win the game to you to continue the analogy, but I think there's you can't you have to have someone that's quarterbacking it you have to have someone who's understands it who is going to make sure that the activities are coordinated.
The communication is seamless but there's no one group or individual that can make make the program successful if you're within an organization of any size because there's just there's too much like you said, there's too many things going on. They're living too many different places. So you have to have cohesion across your security and infrastructure and clouds and all of these things.
If not, you're really gonna be you know, you can't move the most important metrics around me time to respond meantime, you know, you know resolve restoration all those things. How do we embed that muscle memory into an organization? I remember talking to one fell and I see you'd asked me some question about some issue and I said, well, you know, I gave him what I understood to be their resolution.
And then I asked him I said, you know, how come it is you're dealing with this, you know, there's a volumes of literature on this subject everywhere and he looked at me and he said son is very hard to think about fire prevention when you're holding on to a 10 inch hose for dear life. So is the problem is is that it seems like we're so caught up in that bailing wire and holding everything together that we can't think through. What is the right process or best practices for an incident and kind of have that built into more Collective thinking Yeah.
I think it's both in kind of a organizational strategy that has to be put in place but also like tactical things that have to be done, right so I think you have to make space for Preparation. Right if you look at any of the Frameworks for incident response or alert management or Security in general. There's always the That portion right?
So if you're talking this there's you know, you know analysis detection analysis and response and recovery, but the beginning in the end of that is, you know preparation. Like how are you prepared for this? And how do you get better at over time?
I think as an organization you have to Fine time to do that. If not, you're always going to be caught in the hamster wheel of responding but the other side of that the Tactical thing that you can do there is what what is the thing that your team is spending the most time on today that is actually not providing a lot of value that is very very repetitive that you don't need to do by hand anymore. And and by tackling that one thing that low-hanging fruit you can free up time to do these other more strategic activities of preparation and kind of post incident activity like Lessons Learned and Retros and things like that that I think are really really important, but you have to you have to make the time either through you know, Priority prioritizing it or you have to reduce the amount of daily operational work to give yourself time and I think you can't probably do it one way or the other.
It's probably a combination of both. We hear a lot about AI do you think AI will save us from ourselves someday? I think that the Problem with security and AI especially as like an incident responder or as a security operations team is that you know AI models work really well Machine learning models work really well when things are very very predictable so they can build very good training sets.
So they can they can do these things. And yes, there are things around outlier detection and and things like that, but the human element of security is always going to be a chess match between you know, what do we think? Somebody's gonna do what do they have in place?
And how do I circumvent that and those models will end up being very very complex specially, you know post threat detection and I think that there is a future for AI obviously we're working on that in internally really around what where is machine learning applicable and what we're doing and actually will help our customers and not be just kind of like a marketing buzzword, but I do think it has a place in the future but I do think that people have to be realistic about what kind of do for them now and how to plan for it down the road. All right, folks. I heard it here first whether you're camping in the woods or running an IT shop.
There's no substitute for being prepared. Hey, Cody. Thanks for being on the show.
Thanks, Michael. Really appreciate it. Back to you guys in the studio.