Service Level Objectives – Kit Merker, Nobl9
Kit Merker, chief growth officer for Nobl9, explains why attaining and maintaining service level objectives (SLOs) has become more critical in the wake of the company receiving an additional $15.8M in funding.
Transcript
This is Textron TV. Hey guys. Thanks for the throw.
We're here with kit merkers Chief growth officer for Noble 9. 8 million dollars in funding and we're gonna talk about what the plan is for that along with observability and slo's okay, welcome to the show. Hey, thanks for having me good to see you.
See you guys. This is not the first one you guys have raised and I guess the question that immediately comes up is you know, I guess after everybody gets a virtual beer. What are you planning to do?
And what comes next? yeah, we you know, we've been around since 2019 and We've raised some funding before from battery and CRV and some other investors. She's been great.
I think what's different for us now is we've really hit our stride in terms of customer adoption. And so everything that we're doing now in terms of growing the business to help customers and help them standardize on on slos. We've added service now and Cisco both the Cisco Investments as investors in Noble mind.
And so the focus for us is continue to build the best SLO product we can make sure we support our customers and make sure that they're getting the best, you know, uptime and service and everything else and then, you know continue to educate the market about slos and continue to invest in different open source and Community projects to make sure that people know how to use slos and how to make and service level objectives and how they can impact the reliability and velocity of their organization. Yeah, you guys have been talking about slos otherwise known as service level objectives, but the concept at least been around since I don't know maybe the first Mainframe, but your approach is SLO is code essentially and we're trying to make all this programmable. So can you explain what the difference is for the uninitiating?
Yeah, I mean all enterprise software is replacing a spreadsheet. So let's you know start from that think about and SLO the concept is kind of as old as the dawn of you know Services, right? If you walk into a restaurant there's a certain amount of time to acceptable that they come up and offer you a table or else you leave the restaurant and this same concept applies to software Services.
If you're using, you know, a mobile tech upload, you're skimming a video you're trying to maybe buy, you know, a ticket to a concert. There's some amount of frustration that it's okay very small amount usually and then beyond that point people start to get pretty upset. Maybe they move off your service Etc.
They don't expect Perfection. I think this is the really idea that slos modern slos or harnessing is that people will accept a small amount of error and they will hit refresh. And also if you have services that are built to be resilient they can also retry to connect to an API or database Etc.
So with this kind of insight you're right the key is to build automated as well as in fact That's really the better way to think about what Noble 9 is it's an automated SLO platform. Usually I think the traditional way people look at this is they'll build an SLA. They'll put it in the contract all you know agree.
Okay. 9% of the time the details of how that will be measured is somewhat fuzzy with some special exclusions and things like that. It's looked at manually using spreadsheets and reports in contrast to the modern way that exposure done at companies like Google and others that have adopted it and the way works the slos are clearly defining code.
We have a standard called open SLO. We dealt with industry leaders like red hat Dinah train Sumo logic get lab that are all contributors to open up slow. It's checked in as part of your get-getops process.
There's no arguing over how the SLO is defined. And then the second piece is we Define different conditions. Some people might be familiar with error budgets, but basically the idea that you set thresholds that can trigger automation.
So, you know, imagine that instead of having to look at the report in real time. You can anticipate and that's what we'll be violated if the conditions continue and you could trigger an automated run book some sort of automation to maybe scale up the service or roll back or release. You could slack the team you could send an email you could kick off a ticket in a service now or in atlasting jira, or you could you know, fire off a page or to the team and so that choice based on these different thresholds and rules, but it's all grounded in trying to understand user expectations.
How much failure will they tolerate and how much are you willing to accept from the business in order to run efficiently while also offering excellent customer service and that's the name of the game not driving for Perfection, but have food to find acceptable failure in the system. Isn't me or is it becoming hard to do achieve and maintain slos in this age of microservices and dependencies and there is this nasty little thing called latency, right? Well, you know we are yeah, you're always up against some sort of speed of light for sure and there is a lot of risk management.
I would say one of the the challenges I see is that teams are companies under a lot of pressure. You know, we're under a lot of economic pressure. And people have to keep up with the competition.
There's been a lot of DIY Solutions where companies who tried to build infrastructure that maybe they weren't, you know, their core competency. They've built up technical debt over time all of those things contribute difficulty and I like to say that reliability is invisible until it isn't and so now, you know convincing organizations that uptime is critical to their reputation to their bottom line, you know, it's becoming more and more obvious. This is important.
However, you still need to make a significant investment to make that it happen. So I think one of the big shifts happening now with smaller teams and more focus more more scrutiny over spending less DIY project more people adopting, you know solutions that are proven and that is helping people kind of Direction. Now, the whole service is concept to me I think is a really powerful concept but you're right you you are what you know, the advantage is that I can do things.
I want take other things and service for example Cloud, right? I mean, it's It's not have to run data centers. However, you are taking independency on something that we're down microservices.
If done correctly. The idea is to accept failure the resiliency because you have small amounts of failure all over the place. But that aggregate into a more reliable system as a whole not everybody is able to achieve that result because they haven't necessarily designed the system for failure right expected failure or they haven't clearly Quantified the amount of expected failure that's normal versus abnormal and that lead them to noisy alerts or not designing resilient systems.
This is where slo's commit it's a really about clarifying not only you know my service. What's the expectation of my service but also my dependencies, you know, I use database acts I use cloud. Why what do I expect from those Services?
How do I get an early warning system on those Services? How do I hold my vendors accountable or if I have a cross team dependency with maybe another department our shared services or a platform, you know building that clear expectation negotiating that and defining it in code is really really powerful. It gives these teams the ability to much more clearly communicate how the service is working quantify the risk.
That's I think is one of the big changes that will allow people to over. This challenge brute forcing your way to reliability simply doesn't work. The systems are too complicated and the amount of effort you have to put in each systems up and running perfectly difficult.
You know, your CEO might want your service to be perfect, but they're not really willing to pay for it. That's the Dirty Little Secret. So you got to come up with a compromise and the compromise is, you know finding this perfect balance between Excellence customer service excellence and efficient delivery.
And that's what the solos ideally should be. If the moving Target which means you need to have a process and a system for keeping it up to date as expectations change as customers change as your competition changes, but yeah, that's a fundamental is where we're going to get resilient systems how we get to Reliable systems. So back in the day that joke was that the SLO wasn't worth the paper was written on obviously.
We're no longer writing slos in papers. So how is the accountability equation changed? And and what do I say to somebody when the SLO is not met?
well, you know accountability is an interesting word here because you know If we try to set people up for failure right where they have expectations or SLA that they couldn't possibly reach, you know might be great for the sales team to sign and that's delay, um, you know, five nines or six nine but not, you know, not necessarily great for the engineering team to have to deliver that that you know, that feature to their their customers and carry the pager and are gonna be one penalized for it. So accountability to me really start with understanding what's possible with the fact that what's you know, what's financially responsible and having that clearly defined I agree with you like if it's not written down it doesn't matter. And in this case what we really are advocating for is not putting in a contract but putting it into code and code is so much clearer because you know, you have to get it right it does it does what it's told right?
No, there's no room for sort of interpretation if you know what I mean, so that is a very good point about, you know, the importance of clarifying it having it in writing and back having it in a structured format, but you might have heard a lot about sort of blameless blameless. Word arms and you know not pointing blame one of the terms. I've heard recently as being blame aware and thinking about the consequences of blame without removing blame if we try to tiptoe around, you know blame.
It's another way of saying we're not really gonna we're not really interested in finding the truth. Right what the real cause was and oftentimes, you know things that we could say are human error are really not human error. It's usually systemic issues that we need to address.
But as soon as we remove blame from the story it slows down our investigation, I think There's always consequences to these systems. We have to accept you know that. It many many complicated factors that lead the outages in downtown most the time.
However, sometimes it's just bad actors. Sometimes people who are you know, being irresponsible checking in code. They shouldn't ignoring Pat's failures.
I've heard about mountages recently that we're just people pushing code without following the rules, you know, you have to deal with that and I think having the the paper trail having clear expectations defined having, you know, release processes defined at least you're all kind of operating from the same the same rule book and then you can say okay what happened here? Look, you know people use SLA is to hold their vendors accountable. And in fact, we even use our slos monitor our third party to get conversations with them where we can show them the data.
Let's say it's a very different conversation where I can say look here was the SLO. Here's the error budget you depleted. Here's the outage.
Here's the impact from our customers. I've got the whole record of it and we generate that very easily out of noble mind. We use it for our own, you know dog food and all of our own services and all of our vendors the very different conversation with the vendor that I've ever had before and we we've been able to Associated much better Faith because we're operating from the same data.
This is a game changer and I wish more people could experience this because I've seen too many times where a lot of finger pointing Blame Game, you know, and and people kind of, you know, pulling out the contract which never a good sign for building a strong technical partnership and we dig out the contract to see what does this lay actually say, yeah. So accountability I think is part of the equation but most Engineers I've worked with really care more about being productive involving problems, then, you know, pointing fingers and yelling, but that Often seems to be the mode we operate in. How automated can all this get because if I can see something is going awry and SLO.
Can I make adjustments on some back-end service automatically that would bring the application back in line with the SLO whether that's additional compute resources or more Network bandwidth or something. Can I you know really close the loop. Yeah, it's it's a it's a great point.
And you know, there is I would say there is always a limit to Automation and this is the general principle that you know, there's the Paradox of automation that the more you automate something the less you're involved but the criticality increases, you know think of it like you have a self-driving car. You can take a nap while the car self-driving, you know until it's about to have an active. You need to grab the wheel and automation for our operations systems.
You know, it kind of work the same way. So there is a limit how much we can automate these things. However, in some cases there is obvious automation you can do or you can do it as a as a first response to a human response.
So for example, for example, you gave a example adding computers. Let's say we have an SLO on a service that can Auto scale and we see that the latency is creeping up. Well before I wake somebody up to investigate that if I can easily trigger a proactive automation to increase the path to be on the server, you know, maybe increase the replicas in the kubernetes culture and see if that affects the latency.
If it continues then I can trigger, you know an on call to go and investigate it by hand and you just having those kinds of rules in place reduces the burden and the toil on on the humans in those situations. Another common pattern is that nuts below for Relief. So if the software rollout is happening and the SLO starts to degrade Go ahead and roll back the software maybe use feature Flags or some other method for that.
So this is the idea of trying to deescalate right gee the escalate the situation that we can deal with problems in due course of business as opposed to an on call emergencies and I think that philosophy, you know, using the Automation and that context and you can still you know, file a bugs to get investigated tomorrow. That's another great option. Right?
We have something that looks like a problem. It's just throwing itself. But listen, I've heard more than anything.
I hear people who say I got alerted I woke up. I started to investigate and by the time I started investigating the system with back to normal, that's actually the you know, the status for most people is that these big blips it trigger traditional alerting. They create a lot of noise and and turmoil for the teams and they correct themselves by the time they get to it anyway, so we say well, let's take a more sophisticated approach then that's a low over it and those minor blips won't trigger alert, but true Trends will rise to the top those will lead to more accurate alerting and this completely changes on the approach or for teams because the reacting to real impact not responding to things of the corrected themselves, and they're also dealing with stuff like in a more non-emergency way which leads to the team happiness productivity Health, you know, all the other things that we that we want.
We don't want our teams to To be alerted except when it's a real emergency, right and developers. I've worked they're all happy to be on call. They just get frustrated when they start getting page long time over nonsense, or if they don't have the tool to preventative measures, you know, they see the post more and they can't fix it because management doesn't listen to them and has all the feature pressures.
So now at the lows no, you get the regulate that you have something that can really clearly Express the situation that management team product management engineering operations can all agree on that's I think really the the trick to this whole thing. So automation is a key part of it, but it's it's about using that in concert with the human interactions and human decision making on top of the automation. So you have a more balanced approach to managing the environment.
And what you described is what I would call it danger Will Robinson feature, but do you think artificial intelligence will play a larger role in this as we go for it? I would be surprised if People don't say that. It does we'll start with that.
You know, it's interesting because personally I would never hand the keys of my critical operations systems to artificial intelligence at least not at this stage. I think that where I've seen it be effective is in things that are kind of offline from the production system. So anomaly detection.
We have looked at and are investing in some AI around SLO settings. So basically defining what the slocity from the historical data, you know, that's an area that we think that there's an opportunity which is different than saying. Oh we're gonna you know replace, you know, the human operators with AI.
So I think is Analytics tool. I think it's really powerful. I think as you know for specialized cases that are well defined, I think you can work quite well, I think you know AI for like generating tests and you know generating test traffic and synthetics and things like that.
I think it'd be useful as well. But at this point I don't see I personally don't see how it's gonna like take over operation and soon you know, so that's that one thing that you see people doing over and over again is they kind of approach slo's and you wish that they knew and we could just cut to the chase earlier. Yeah, this what I would say is quite funny to me actually excited this so many times they go.
I got the cuts or potential customers or practitioners are trying to do as well as they go. Well, we're not sophisticated enough to do that book when we're more reliable. We'll start doing at the low and I can't tell you how many times I hear this but I just think it's completely backwards because you know, the result is that you get to be more reliable because you've got a baseline you started measuring your internal improving it's it's kind of like saying well, you know, I'll start you know, I'll start going to the gym once I get stronger.
I'll start dieting once I lose weight it kind of backwards. So I want the one thing I encourage people is even if your service immature, even if you have you know, let's say, you know nine five instead of five nine. It doesn't matter you can start with slos you can start the process.
Now. It's really easy to get started easier than ever been to get started and there's tons of resources. com.
We're about to announce the speakers the lineup it's gonna be huge, you know the education around this is the men. It's easy to get started and it's simple to get started. I wish people would look at it that way as beginning of the journey not something they will do, you know, once they get around to it down the line.
All right. Well as the saying goes things measured or things done and slos are the start to the path. Okay.
Thanks again. Great to see my picture. Yeah back to you guys in the studio.