It’s all About Scale: Designing Your SOC and IR Capabilities for the Long Term with Andy Ellis | SecOps Vision 2024
The structure of a process or organization at small scale is often about generalization – everyone has to be able to do anything – but at large scale, you want to design for specialist roles. Learn the essential items to scale, which ones to specialize in and how to outsource unimportant (to you) parts.
Takeaways:
* It’s okay to have a hero-oriented small-scale SOC/IR program, but that can’t scale.
* How to build robust IR processes that handle both trivial and toxic incidents.
* How to staff a scaling SOC organization.
Transcript
Good morning, or maybe good afternoon, good evening, or good night depending on when you are in the world. I'm Andy Ellis, operating partner at Weill Ventures. And today I want to talk a little bit about the scale of designing your SOC and your instant response teams for the long term to look at sort of what we're getting out of our soc, what we're trying to build for, and to recognize that what you might do is a small scale program will be very different than what you do as a large scale program.
And that's not actually a problem for that small scale program. It, it should be a little bit different. So, you know, when we think about what is, what is an incident and what does incident response look like, you know, in the beginning, organizations basically are just cruising along.
They're totally happy. And then something bad happens, right? You get an incident, you lightning strikes, boom fire, we've gotta go solve this problem.
And this is sort of the origin of incident response. In most organizations, very, very rarely does someone's create an incident response function before there's an incident that you need to solve it for. So we're gonna look at sort of how we evolve.
But usually the first thing that happens is like some people go and they like have to go put out the fire and they say, oh, maybe we should have buckets. And you start to organically build out your instant response capabilities and especially those reactions after the fact. And maybe if you have enough fire start, you're like, oh, we should, we should have an arson investigator or some form of fire department going around trying to reduce the flare up.
So think about that as, you know, safety and resilience organizations trying to keep incidents from happening. But here's sort of the theory. Often what people think is like incident response starts at really easy.
It's like you've got this easy bake oven that can, you know, you can do everything even if you're just a little kid. Um, and that's not really the reality. You know, almost anybody who's been in instant response knows that, that it looks very, very different.
In fact, in practice, you know, you're basically sort of running around trying to stick your fingers in dikes and be like, oh, I gotta stop a problem here. And as soon as you stop one problem, you see another problem, right? And you really don't fix problems, you get them to tolerable.
So we put band-aids over broken pans of glass and often try to tell other people what's going on. We're often working sort of this inscrutable fashion. You know, incident responders are very used to their family seeing them at a computer all the time because unfortunately incidents don't usually stick to, you know, the, the right time of day.
It's not like that you, you work from 10 to seven and there are no incidents outside that. And of course incident responders think of themselves as heroes. And that's really is what organizations build for is when we think about the capability maturity model, if you're familiar with that, you know, level one, you know, the very non repeatable, hero driven processes is where incident response generally starts from is we take superheroes and we say, Hey, just keep, keep doing this.
Keep holding the line, get more fingers to put into the dikes. Um, and like more band-aids to slap on things. But as long as you can keep the company from failing, we're happy with that.
And you know, that works for a while, but at some point it's too much for this one person to handle and we're going to need to scale it. So let's think about the dimensions that we're going to scale on and what we're going to invest on. You know, the first one, very obviously people, you do not solve incidents without people at some point.
Um, you know, much as we would love to see be a beautiful automation that just dealt with all of our incidents for us. First of all, almost nobody will trust automatic remediation if the incident tempo, right? 'cause something was working and now it broke, let's go fix it, is a very scary proposition because it broken in unusual fashion, right?
We think about the technology dimension, like what systems do we buy to make it easier for us to do incident response that we can detect what's happening, we can solve these problems, right? That's what we're gonna be looking for here. And then of course there's the dimension of process, which is how formalized you know, is this incident response program.
And when I look at process, I like to think that the very first piece of process that almost everybody puts in is incident categorization. How severe is this? And I like to think that there's basically, you know, three types of severity.
There's toxic, you must deal with this problem if you don't deal with this problem, the company's gonna have a really bad day. There is trivial, which is eh, if it doesn't go well, it doesn't go well. And then there's in between it, there's intermediate sometimes also could be interesting, the things that we don't know, it's trivial, but we don't yet have evidence that it's toxic.
So we're gonna need to do some investigation, right? So that's the first piece of process that many people will put in is this step where you categorize an incident so you can communicate like what people have to do. And then of course some way to share that communication.
So when people first start, this is often what you have, right? Your people is one person, one defender, usually not their full-time job. This might be a security professional, it might be an SRE might be somebody off of your help desk.
You know, often it's someone who has very good operational experience. You ability to put literally your fingers on a keyboard and solve problems. Their first process is usually entirely built on top of email.
I send people email, I ask them to do things and they solve it. And our technology is usually a command shell, right? We have access to systems that we're gonna go collect data from and like this works.
And honestly for many organizations, like if this is you and you're able to solve incidents within your capacity with these tools, great, you're doing fine. Like pay attention to the growth of your organization. You know, as incidents become more frequent, like what's working, what's not working here?
You know, maybe you'll scale up different pieces of these. Let's talk about how we do scale those up. You know?
So one evolution that often happens, you know, on the people side is we just start adding more bodies. And there's a challenge here, which is if you're just adding bodies to an incident response team where you say, oh, we're gonna create an incident response team. And so if a person who's full-time job is incident response, well if you haven't really scaled up your process or your technology yet, what you're trying to do is really clone this person, which is always a bad idea.
The first person you have incident res doing incident response will always be your most versatile and best person. They can solve any problem, they can coordinate, they get things fixed, okay? If they can't do all of these, you replace them with another first person until you have this person who is successful.
And organizations will say, well why don't you just hire somebody and teach them to be like this? Well first of all, as long as this person's around, it's hard for anybody else to become as good as that person. They're really just not going to succeed at, you know, learning all the things because there's always someone to fall back on.
But more importantly, this doesn't actually scale well, both from a training perspective, it is in fact really hard to scale all of these people. Uh, but it also doesn't scale well because at some point they all share the same limitations. You know, think about incidents.
You know, we start out and an incident is basically this one person drives an incident to conclusion and when they're done with an incident, they move on to the next one. But as soon as they move on to an incident, they often forget about those old incidents. I mean, they're in the back of their head, but from a drive a prob a process forward, like, oh, they'll, they'll forget about it.
Maybe they come across it in their inbox at some point and say, oh, I gotta fix this. But maybe somebody who's more process oriented coming in here, you know, they might not have as deep technical capabilities at finding weird and unusual problems is where you're need going to need to invest in because they're gonna take and solve more problems for you. They might not identify more problems, right?
And so maybe one of these people is focused on solving or you, how do I remediate? One is focused on identification. And so as you're gonna evolve this people organization, you want to think about better specialization, right?
The first person is the biggest generalist you have. But as you go along that more specialization, you want people who basically get hired to take work off of this person and to identify what this person is dropping. And those are the two capabilities you're aiming for is how do we not build clones of this person but take parts of this person and replace what they do?
And often this will make this person very happy if you're the first incident responder and you don't have to deal with the things you used to deal with that are easy for you, you know? But maybe there, there were, were hard for these people, but they can get better, right? This gives you free time to go solve other problems.
So as we think about how we're gonna evolve people, we wanna understand that it also does need to evolve with process and technology. Like I'm showing this as if we could just just grow people out, this would be a bad idea. If you did not change your process for your technology but just hired more people, you are wasting these people's energy because some of that investment should go to making their job easier.
But let's look at like what this evolution might look like from a technology perspective. You know, we started out where we go have to hunt it down artifacts. Well, over time we probably evolved to something like logs and to sim, you know, and maybe we're looking at the next generation of SOC platforms, right?
And the goal here is to say, well let's pull all of this data into one spot and maybe let's start to do some analysis, right? And the goal here is like, well the more that the sim can do to correlate, to figure out what's happened and the more that our sock platform can connect the dots for us so we can see what unusual things are happening, well that's less that our people like exactly what just happened. They need to have this knowledge.
So they need to move up in lockstep. 'cause what you're really doing with these technologies is not getting ahead of your people. What you're doing is moving in parallel with the people that the more the technology can reduce the work this person has to do to get to a conclusion so they can get to the conclusion quickly, right?
But if, if it's just getting to a conclusion that is inexplicable to the people who are trying to drive the process, then the less likely that gets paid attention to. Now another important thing to understand is like this is a lot of data, right? It's one thing when you're trying to go out to a bunch of machines and gr all of their logs.
But as you start to pull all of these logs together into one place, this becomes a substantially expensive investment to make to understand like where is all of your data? How is it all being aggregated? What is being collected?
But at the core like this, this technology goal here is to take what your people are doing and offload them so they can do more and actually to help drive your process forward. 'cause oftentimes what processes are doing is sort of mimicking like what the people are doing and what the technology is doing. So that's what we wanna look at the, what the technology does for us.
Now finally, we're gonna wanna look at process. Now most people when they say process, they see something like this, oh, I just grabbed Jira and I slap it in. And let's be honest, like what that is is that's slightly better than email, right?
We've taken an email, we've turned it into a ticket, something that we're tracking, whether it's in Jira or in RT or honestly in Salesforce I've actually seen people implement incident response programs in Salesforce. What you're really just doing is you're taking this, this construct, there was an email and just sticking it into a database with some level of persistence and you're just sort of moving it forward. And the challenge is that misses the, the important thing of process, which is process design, right?
And I like to think about as you think about incidents, um, we talked earlier about severity, but then let's talk about phases. Because phases also drive incidents. And to me, incidents have several phases, right?
Phase zero before the incident actually happens, but you're on a collision course for it happening happens a lot more in the safety world than the security world. Like it's hard to say, oh this person's about to break into our systems, you know, but it is sometimes easy to say, oh, we have systems that are overloaded that are about to fail, right? Phase zero.
Phase one is you're like in this serious problem. Um, we need to remediate it. Something is broken.
Phase two is often we've, we've dealt with the immediate break, we've bandaged over it. It's at, we're at risk of it recurring, but it's not currently happening, right? We've moved forward and now we're sort of in this cleanup phase.
And then phase three it's like we've restored systems to normal order. And if you're really professional, you have a phase four, which is we have learned from this incident and we have brought those learnings back into our design program. And think about that as just a process that has nothing to do with like, go fix this, right?
Go fix. This is all inside phase one. But this process that your people can then drive that might pull in data from technology but helps you understand how you're getting better and better over time.
You are less likely to suffer from one incident, you know, happening to you multiple times. Like that's the heart of good process. Maybe you're doing process mining to understand where you have incidents that, you know, maybe they were in that intermediate state, they clearly weren't toxic.
Maybe they actually were trivial but they weren't handled at all or they got stuck in some state. 'cause often when people think about process, they think about it transactionally, right? We're used to looking at tickets and operations teams, ticket comes up, you try to close it, you move on, right?
And so just, just a straight transaction, you're not thinking about it from overall perspective, did it get dealt with in the right way? And I think as users we've often seen the, you file a ticket and the ops team that looks at it says, wow, this is a hard one to solve right now, so I'm just gonna close it as won't fix, right? That's an example of a process that is failing because it's built on top of a transactional system instead of on a process system.
And so a big piece of process as you go to scale out what you're doing is understanding all of the ways the process could fail so that you can start to fix those. So you can build in controls to make sure that whatever you've built the process for is actually providing value. 'cause imagine that if your technology detected someone one dwelling, but they hadn't taken anything yet and someone said, oh, well let's go investigate and take a look at this.
And then your process didn't have an action for someone to do so it gets forgotten about. And then two years later it turns out, oh hey, like that thing that we saw that turned out to be the major breach that you know now makes our companies like very visible in the headlines, right? That's a, a process failure more than a decision failure over here.
Although the decision failure is not recognizing that you don't have a process for solving a problem like that for paying attention and investigating. If you only have a choice of cleanup or forget, you should always clean up. And that's where good process design can really come in for you.
So let's just quickly look at like, sort of what these lessons are. And so first we think about people like those heroes that you have in the early days. The more process you have, the more uncomfortable they're going to get.
Uh, I had somebody who worked for me who used to say that heroes can't survive in a level three program. And they're pretty much right. Like if you have somebody who likes to just come in, solve the problem and get out, and then you say, oh, by the way, before you do everything, you need to open up the ticketing system.
You need to document here, you need to document there. They're gonna be overly constrained by that. And so that's not a problem on their part.
That's just a mismatch between the maturity of your process and the ways in which your people want to operate. So go find systems that are out still in level one and let them go deal with that. Let them go deal with a different problem.
While you hiring people who want to work in that level three or level four system, they're like, no, no, I'm completely happy with, I've got a checklist, I follow my checklist. Um, I don't do anything without documenting what I do. Right?
Those are great to have once you have that process, once you have that technology. But until you have that, don't hire those people either. So you really have to think about when you're hiring heroes, when you're hiring people who are gonna blaze a trail for others to follow.
When you're gonna hire people who will follow process, when you'll hire people who will focus on the continuous improvement for your process, because you have to make sure you have the right people and then the more people you have, the more specialization you're gonna want to have. Now that specialization can be at the organizational level. Maybe you have incident, you know, responders, an incident response team that all they're doing is coordination that they're the experts in, look, whenever there's an incident, we come in, we'll run the incident, but we're always gonna have a technical person doing forensics and analysis for us.
And that's an okay kind of specialization to have. 'cause it turns out that often your technical leads don't want to be your process leads. And so one way to deal with this hero problem is to split that role.
And that lets you take your, your technical heroes and have them last a little bit longer. Somebody else will manage the process. Maybe you want to have people who are experts in your privacy.
So whenever there's an incident that touches on privacy, they can come in and help people understand like, how bad is this? What can we, what can't we do when we think about our processes? You know, we should just recognize in the early days what our processes look like.
Is ticket resolution, like there was a fire, we're gonna deal with the fire when the fire is out. Cleanup is not really our problem, but mature processes need to notice when there are deviations needs to monitor for them and also need to start to expand into recovery and say it's not just good enough to put out fires. Our job is to help build new systems after the fact so that we don't have to deal with these, you know, fires continuously happening so that we get back to normal recovery.
And just because you have a process tool that you buy from someone else doesn't necessarily make it the replacement for process design. In fact, probably one of the biggest challenges many companies have when they hire a tool or buy a tool, you know, usually hire a tool, um, when you buy a tool that has process baked into it, is you just take whatever process is there and you try to force that into your organization. And if your organization is uncomfortable with process, that's gonna be a problem for you.
You're not gonna be really able to successfully roll out that process and make it work for you because you don't even understand the process and neither do the recipients. And then when you think about technology, the more time that your people are spending doing something, obviously the more value that you get out of the technology. Now that makes sense from just a an economics perspective.
If I'm spending five people pulling logs and I have a system that pulls logs, I save five people. But qualitatively that also matters. The more you understand a problem inside your organization, the better you're gonna be able to use a technology to solve that problem at scale.
If you don't know how to pull logs and you just buy a system, you say, oh, I got this sim and it will just aggregate all my logs. Well, if you didn't know where your logs were, the SIM isn't necessarily going to just go find them for you. Like you need to build those connectors and you know, go do that integration.
So the more you have done in advance of buying a technology to solve the problem that the technology will solve, the more value you will get out of that technology. And also, a lot of people like to buy ai. AI is the big buzzword.
I gotta address it a little bit here and say, oh, we'll use AI to solve the unsolvable. I really wouldn't aim for that as your first goal. You should recognize that like 99% of the incidents that will come across your desk, you know, from toxic to trivial, are known repeatable problems.
You've seen these before your analysts see them. The more you can focus your systems on dealing with those problems, the things that are understandable, that are predictable, the more value you're gonna get. And when something weird happens, sure, if you had a, you know, machine learning system that could say this was weird, that's great.
But that's really where you wanna get your humans involved because those are opportunities for you to better understand your environment, to better understand your adversaries, to see what's happening, and then to have technology augment and help those humans. So focus your technology first on the repeatable, scalable problems, then worry about the anomalous problems, but leave those for your humans as you do your transitions. So thank you.
I hope this was really valuable for you in understanding you how to approach your scaling of your SOC and your incident response programs. If you have any questions for me, you can find me on Twitter or LinkedIn as CSO Andy.





