Observability in Modern DevOps Environments – Erez Barak, Sumo Logic
Erez Barak, general manager and vice president of engineering for Sumo Logic, explains why the key to reliability is observability in modern DevOps environments.
Transcript
This is Textron TV. Hey guys. Thanks for the throw.
We're here with the res Barack who is vice president observability for Sumo logic, we're gonna be talking about reliability and observability and how they all go hand in hand a res. Welcome to the show. Thank you.
My good great to be here. In some ways we've been talking about reliability for some time now ever since the rise of sres, but it's not clear to me that everybody kind of gets that as a discipline and what exactly we want them to do or think that they should be focused on so what from your perspective is the real challenge with reliability. I think what's interesting is about reliability to your point.
We've been talking about it for a while. I think we're reliability really takes on another dimension is when scale kicks in or to be precise. Cloud scale and what we're seeing is a multitude of microservices multitude of components multitude of applications all coming together multi clouds hybrid Solutions, and that kind of complexity is not just about monitoring.
It's not just about setup. It's about how you manage reliability. So reliability management, which is really where we're Focus now is how you manage reliability same sort of building blocks.
We talked about in the past around availability around performance around even security. How do you manage that at the very large scale? And what ultimately are the challenges of that because we've had tools measuring everything in the planet and we are trying to make this shift to observability and not everybody understands the difference between monitoring and observability But ultimately from your perspective.
What do we need to measure what needs to change that? We're not currently measured? And maybe I could share a few things.
We're hearing from customers and we're hearing things like and We need or we cannot measure slos in real time. And we need that real-time. Measurement in order to effectively manage the end user experience.
So if I run an end user experience and that end user experience, you know, it's in an app. Someone uses it right now when something goes wrong we want to take that into account right now. So the ability to do that and in real time is something that again as a derivative of the complexity we talked about continues to be a challenge.
and customers also saying as you pointed out Michael There's alerts for everything or just measurement for everything. But what we want to alert on is the error budget what we want to learn on alert on is how fast is that error budget consumed? Why is that interesting it allows?
me as a user it allows me as and manager of reliability to know where I should focus do I push on more and more innovation no capabilities and breakthroughs in the next month or quarter or do I take a step back? Play it in a little more safe mode, make sure we're reliable available on a regular basis. How do I answer that in a data-driven manner there's many opinions.
But once you have the error budget you get the data support for your decision making Do you think we've been kind of focused in the wrong direction as it were but we've been thinking about it systems out and maybe what we're really talking about is a more user-centric approach to managing it from the user into the IT environment because otherwise we wind up managing a bunch of components and services, but we don't understand what the actual impact is on the experience. Now to be fair, I think that's been. Recognized and true for a while.
I think you know with Solutions like real user monitoring like tracing to an extent that pull in that data they show Up to us as vendors and to us as users and to our customers we very much believe in that end user measurement. So I think that's been around for a while. I think the marriage of that with reliability management the ability to create an SLO around the customer experience.
That's a new way of thinking about my system. And you know, it's a bit of a how should I say frightening thought right? Don't look at every CPU measurement don't jump on every disk storage alert.
Let's focus and believe we have the right user View and really help. Our teams react to the right things and make sure our customers get the right priority of dealing with the issues. They're seeing we do that by raising the priority raising the urgency of collecting real user data and then raising the urgency and raising the priority of calculating the slos that are based on that first of us.
Are we not organized correctly within it organizations to achieve that goal because I still see a lot of places where yeah, there's somebody they manage the cloud resources and they're in charge of infrastructure and a couple of middleware guys over here and some Dev guys over here and do we need to kind of realign the it teams around the actual user experience slash SLO that we're trying to manage. know and in my shoes and I don't get to say our it teams aligned correctly set up correctly or not. Oh my God.
I definitely have an opinion in the area but we don't get to do that. We assume there's gonna be an organization. We assume that organization evolves maybe becomes more about real user monitoring in the future Etc.
But there's always going to be an organization and and that's aligned with it. So reliability management is really taking that as an assumption is saying this is a technology that allows us to transcend those Niche areas. To transcend the buckets and say regardless of how you're set up.
We'd love if you're optimized but if you're not we're able to look across the teams across the Technologies in a way normalize the data you have into one SLI or multiple SLI service level indicators. for those Define objectives And for those Define error budgets, so what have we done here? We decoupled the physical setup of the team.
From your ability to do the optimal definition or the optimal tracking of your slos the more those are aligned. Frankly the easier to exercise becomes but the way we've built the Sumo technology does not require a change to how you set up does not require. Hey, you got to have Team a and Team B.
It's all about normalizing the data together and then transcending that to provide the solution. Do I need to hire something that looks like a site reliability engineer to achieve this or can I achieve this with my own Squad as is per se I just have to kind of you know reorientate the team. My strong belief here and I say believe because it's not backed up with as much data as I would like it to just yet, but we are definitely seeing that brand.
is that this is one of those things that's more of a in place verse additional M types of approaches or directions and I'll explain right when you want to build something that you know, like in a mobile app. That's not in place. You gotta have experts you gotta have people who know and with technology you're gonna have people who know mobile around them in a global setting that requires the additional and capabilities the way We've built reliability management does not require additional subject matter expertise.
So that's one. two once people adopt and customers adopt this idea of error budget monitoring. That comes with less effort on overall General monitor that comes with less manual effort of calculating slos.
that comes with less manual effort of alerting based on those calculations So what we're seeing again as a trend is that this is an in place accelerator. for you to support reliability at scale with the same staff you have and with the same expertise that you already have in house. Do you think that we're reaching a level of complexity that we're gonna need some help from Ai and algorithms, but the question is always been to what degree can you rely on them or what is the state of our Collective machine intelligence that we can employ and rely on today versus where you think it might be tomorrow?
I think we passed the line of not relying on AI, you know, the collection the pattern recognition the anomaly detection. The log duplication, you know, we were talking about terabytes of logs coming in every hour sometime and there's no way to keep up without AI so we've crossed the line of being able to do these things without AI a while back. In terms of accuracy in terms of what AI brings to the table.
I think AI today is implemented in devops and sort of what I would. Categorize as safe areas. Well tested well known.
and I think once we start and we will trying to push the bleeding edge and say you know what? Will Define your slos for you? So we'll Define your slis for you.
We'll create a solos on top of them. Will start predicting where error budget is going. I think that's where we're going to have to have a lot more checks and balances because the AI were now implementing is specific.
To reliability management or SL Management in this case and not other aspects already true and tested with and and you know and other areas or true interested in Technologies or models that they have been put in place. I think the more we do an AI. The more reliable will become I think the more we do with AI the more we're able to scale and the more we do it AI we end up having humans taking really a decision making role taking really a role of hey, what are the most important things to prioritize a go after?
And really Rising above the noise of multiple alerts and billions of data points that decision-making place where we want to have them those subject matter experts spend the most of the time. So do you think we'll get to the point where I might walk into the office and they'll be a speech assisted message from a machine telling me danger Will Robinson. You're about to be fired for these three things unless you fix this media.
I hope I hope the machine speaks nicer than that and I I do hope sorry, I'm pretty sure we're gonna get there. Yes. I'm pretty sure on track to get there and I'm pretty sure that a lot of manual tasks done today.
Patron oriented then repeated of tasks done today with the help of AI get more and more automated. Note, I'm not saying that takes the human out of the loop on the contrary. It puts the human and the key parts of the loop.
It helps us free the time from the repetitive. the pattern the automated it helps reduce the amount of errors that happen when people do things manually and with the right checks and balances helps. And create a much better solution.
I would say, you know in the area of AI in general. this whole idea of compliance of fairness of explainability. Those are key general purpose AI capabilities that allows us to pressure test ai-powered solutions to make sure they are fair.
To make sure they can be explained the machine chose a why did it choose that? A machine just speak in a certain manner to me in a different manner to you. Is that a fair statement or the machine's not being fair?
Those generally applicable capabilities in AI would also apply here. So I think we need them. I think the interaction with humans is definitely an area of friction, but I think we're and we're seeing that in the industry.
There's no resistance to it. It's just a matter of what's the right way to bring that in? All right.
So let's pull it on together. We're focused on reliability. We have more observability.
We're going to have more AI augmented help. Do you think it is a job will get fun again? Because frankly in the last few years to your point.
There's been a lot of drudgery. I think I see. has very different future than where we are today.
I think it's about more and more about critical business decision making I think you know, these ideas were always talk about. Hey, let's connect the line between our containers. Okay.
It's somewhere on the cloud and how the business can do is doing I think since we're able to do more and more than that. then You know, I'm having a problem with the fun statement because some people find what I just said like a ton of fun that's still gonna be around but I think sort of the high level business view of things the control points that ID will become is a whole different ball game than it is today. I think it with riding and these technologies will rise to the occasion and have a place in the decision making and process one last note about that and the context of AI had a great discussion with a colleague the other day around and blockchain and the ability to use blockchain for really automating trusted them and decision making and processes.
And when you think about it, there are a lot of processes. There's a lot of decision making and there's a lot of trust that needs to come into the system. So I think we're just seeing the tip of the iceberg in terms of Technologies like blockchain coming into our space and you know on that front I could say it is going to become a A lot more fun when we get to bring that kind of technology to help with these critical issues for the business.
All right, folks regardless of your definition of fun. It is going to be a lot more rewarding. I think we can all agree on that.
Yes, Ezra's. Thanks for being on the show. Thanks for having me.
It's been great to speak with you. All right back to you guys in the studio.