Optimizing Observability Spend: Data Management Metrics | DataOps Day
Are you collecting just about every metric under the sun and the kitchen sink, too? Understanding the cost of collecting metrics and the usefulness of those metrics is the only way to scale in a cloud-native world. You can’t get away with just collecting everything as you grow; the logistics of storing, managing and analyzing all that data-not to mention the cost-would be astronomical. Instead, your observability teams need to make decisions about what data to collect, what to drop and what to aggregate and still be able to alert, triage, remediate and do their root cause analysis on a daily basis. In this session, Eric Schabell will explore how you can address these challenges and gain immediate insights into high-cost data (DPPS), when to drop time series data, and how to determine when the value of that data is at its lowest.
Transcript
Hello, uh, my name is, uh, Eric Schabell. Uh, I work for Kronos Feeder. And, uh, today we're gonna talk about optimizing observability spend with a slight focus on metrics.
Um, this is a, a data data ops day, uh, 2023 talk. And, uh, let's get to it. Um, so, uh, in the beginning when we talked about observability and, and what was going on with, uh, monitoring applications, and before the cloud got super, uh, uh, expanded as it is now, um, a lot of this was in, uh, uh, monolithic applications, VM environments.
Um, you saw something like this where you had a pretty good idea of your environment. You had the sizing scoped out, you exactly the data you were gonna get relatively speaking, and you had this sort of oasis of monitoring and observability that you're able to get in control of. And, uh, the experience was not so awful to have to carry a pager around those days.
Is the, the, the general consensus. There must be some something different, but the general consensus is we had this stuff under control. Um, and then you get to the, uh, cloud native, uh, environments and, and moving your company into the cloud and scaling into the cloud, where, where things happen in a dynamic fashion, we see that the amount of data that is, is generated, uh, tends to feel a lot like this picture.
And while this seems awful extreme, um, it's not far off, unfortunately, I definitely, uh, share some stories of, of, uh, pandemic, uh, uh, enterprises that, uh, exploded during the pandemic, right? So the, the, the business model jack right up into, you know, unforeseen amounts of data and, and production environments. Um, so, um, one of the things that, that, that is happening and, and you see a lot of right now is, is it's quite common to run into environments where people are struggling massively, uh, with the amount of observability data that's being generated.
Uh, uh, and they, they find out that the, the price point is being, uh, uh, surpassing their, their, their production environment. So if you're in the production, we're doing infrastructure and stuff to make our customers happy and to make our business model work, and then all of a sudden, all the observability, uh, uh, data and, and input and, and organizational stuff is just surpassing that. And, uh, that's something that's making people scratch their heads, like, full stop.
Like, wait a minute. What's, what's going on with this stuff? And just to give you a real quick idea of how this is possible and why this kind of stuff is happening, um, if you go and look around on the internet a little bit, you'll find lots of examples of some kind of baseline, you know, cloud native application.
So we're talking about a Kubernetes cluster, something that autoscale, something's been set up on. Something as simple as this is with a little hello world, uh, web app, uh, some front end, uh, uh, code to, to pinging it and act like there's actual users hitting it. And they did, uh, tracing end user metrics, logs, stuff you see right there, tracing all this stuff and, and collecting all this and gathering this metrics observability data for 30 days, sold almost a half.
That's just one application. When you look at this at large scale, and imagine this going across multiple applications, and you come to think that the average organization stores this data for, uh, 13 months, you can all of a sudden see that the price points are gonna be just crazy compared to what they were for. There's a link in the, the speaker notes, and these slides will be on later.
So, trace down how this exactly works. I love this slide because basically the comment, if you, if you have to actually can't afford it, um, we wouldn't care what it costs. Organizations wouldn't care if they had better outcomes, if their customers were happy, if they're able to triage incidents quicker.
Um, but that is not the case. It's spiraling outta control that they're being flooded with this data. And I mean, pay more for it because you're paying for the storage aspect, uh, with most of the, the vendors out there, uh, more so than the, uh, uh, the ingestion side.
And it's just a lot of useless data that you don't, it's getting in the way. It's, it's more is not better. It just slows down the loading of your dashboards.
It slows down the, the tooling you're using and really hard for your, your people to triage and dig through this. And then what you see here is that, uh, we have an, uh, uh, observability report from 2023, uh, that we, uh, data from 500 different engineers out there working in environments. And one of the most shocking things that people don't think about is you have also the data costs and the aspects of business that you notice climbing in bills.
So what you may not be thinking about is what it's doing to your resources engineers on call staff. So if you see that half, 10 hours a week or average are being spent triaging and understanding incidents, that's a quarter of a 40 hour work. It's 25% of your time is spent doing stuff you don't really wanna be doing, right?
And if you're a, a dev and a DevOps team doing this kind of stuff, you're gonna start scratching your head and wondering why you're actually here. A a lot of this leads to, to statistics around people who just don't, don't wanna work there anymore. And if you can't retain your resources, this is gonna become a massive problem too.
Uh, on top of that, you're gonna end up with burnout. It's just that simple. So, um, knowing the cost of your, of your observability and metrics data, so it's, it's a, it's a combination of the, the flood of data that's coming in, not being able to tackle what's, what's happening with the, the massive amounts that are coming in, what you can do with that, but also not not being able to take care of your resources 'cause they're being overworked and spending too much time stuff.
So who's responsible for that? Is what we're looking at here, is that developer making decisions that you see here saying, Hey, I'm gonna use this tool, or, Hey, I'm gonna do this, or I'm gonna put this new metric in my, uh, tracing code. Um, how, how am I gonna gonna, uh, instrument my applications?
Who's making these decisions? Who owns those? Right?
That's becoming the problem. When you see that, that, that that cost point is going above your production date, you don't know who made those decisions. You don't know what happened.
So what you saw was in 2023, uh, like 80% of these organizations will be having a dedicated finops. We're, we're, we're pretty much, uh, at that point now where you see the finops foundation, you see a lot of this stuff being discussed, a lot of references to these kind of things, looking for a way to put processes and phases and various tooling in place to, to have somebody that can track this down with their cloud usage, be able to keep an eye on manage quota and budgets and resources that they have for their files. This is a big deal.
And that brings me to, uh, something that we want to talk about for the rest of this, uh, session. Um, we've come up together with, uh, uh, a bunch of our customers. Something that, uh, is, is a little bit, uh, uh, recognizable from the finops Foundation and how the, how the, how the phases and stuff work.
But you see that, uh, um, observability data optimization cycle is something that you can use to, to put in place in an organization, kind of understand the data value and usage. You can use it to shape and transform your data and have a centralized government around this. And then you can continuously adjust for efficiency, and it gives you a life cycle around what you're doing with your cloud data.
And in this case, we're gonna zoom in this, uh, uh, specific, uh, uh, storyline and go through the cycle with an eye on just metrics. Um, also be aware that, uh, we're gonna look at the governance, analyze, the refine and the operate, uh, phases. Um, but it's, uh, uh, uh, a rather short session we have today.
So I will only dive deep into the analyze, uh, with a few, uh, concrete examples of what that, uh, if you have interest in, in more stuff, there's links at the end and you can get in touch and you can show, uh, we can show you some more stuff around the governance, define and the opera. So, uh, the first thing here is the governments where it starts. You see that, uh, centralized governments is, is something that you need to have in place to give your teams ownership and control of their metrics and understand, uh, what, what is happening with the cardinality.
The amount of growth that they're doing when they're, when they're deploying new applications or, or putting new metrics in place for observability. You see that there's a couple of, uh, uh, things in place here. Uh, one of 'em is quotas and the other's priorities.
Um, quotas gives you the ability to basically allocate your usage across different teams. It's basically giving 'em a check. It's not blank anymore.
It's, uh, been signed and it gives 'em a certain amount each month or each week or whatever you're doing. And, and you also see that you can, uh, prioritize which data is impact, giving you the ability to govern how you're gonna be using cloud data to be able to take, uh, some sort of guardrails in place for if something happens where like a cardinal explosion through a development cycle releases something in an application that just spikes it outta control. There is tooling behind this.
We're not gonna cover that in. Great. We're gonna go to the next one here with the analyzed data.
Um, this is probably, uh, uh, one of the most important phases, which is why I wanna spend a little bit of time on MasTec dive down to it. Being able to understand the value of observability data to identify what's useful and what is based, um, straight up. This just means, uh, have a lot of data being ingested into your organization, into your observability, uh, pipeline.
Um, that's what everybody has to pay for. But coming out the backend into your storage, that's another aspect you have to pay for. What you want to have back here and what you store there is what you're also querying and using in your dashboards, your ad hoc queries setting alerts off.
Um, so by not overpopulating that with useless data that has no impact on any of these alerts, dashboards or, or ad hoc queries or users touching it, uh, is just greatly to your advantage, not only in the price point, but also in performance of your dashboards ability of your, uh, on-call crews to be able to handle the data manner. You see here we have a, a metrics traffic analyzer, um, that's watching the ingested data live. So let you explore that real time and see what's going on with that.
You have something called a metrics usage analyzer, which, uh, filters up everything that's coming in and shows you what has no, by default, is sort of by what's not, uh, uh, important to you, uh, uh, what's the least used, I should say. And you have a trace analyzer, which will not dive. It's the same idea covering traces, spans, that kind of thing.
So let's take a quick look at the traffic analyzer you see here that, um, shows the incoming stuff live, get you a chance to break down to, uh, the biggest and the smallest contributors metric names and labels applications. And it is the direct way to troubleshoot cardinality, but immediately seems something goes a little bit through. Uh, go a little bit closer here.
See here that you have a live view. You can actually pause it if you want to view the traffic before it's stored, make decisions about the traffic, uh, how you wanna shape that, which means aggregate some stuff away or drop it before you actually pay for really important. Um, you can also break it down by labels or by metrics.
And, uh, you see here that, uh, this instance label, uh, a hundred percent on, on this metrics and, and has 62 unique values. You can select that instance label and you look down into the values of what's going on there. See that some of them not so interested.
They actually, if we go to the, uh, usage analyzer, this is after you're gonna go to, uh, the live data and now you're looking at exactly, please filter out the stuff from least important to most important. And what it's looking at is what kind of value does this metric, uh, deliver? By default, it's gonna show you the ones that are delivering.
Absolutely. Um, see here, in this case, you're looking at one, um, you're looking at one that has, uh, uh, zero, uh, references. It hasn't been directly queried, it hasn't been used in any kind of, uh, alert or dashboard as a utility score of zero.
And it has, uh, 24, uh, 24,000 plus, uh, data points per second. It's quite a lot of data to be storing for no usage whatsoever. You see here that you can, uh, sort it, uh, by default it's like at least utilized, find something that's interesting.
You wanna look at click on usage details down into more. You get down in there, you can see who's using it or what is using it. So it's either in a dashboard monitor, recording, rule drop aggregator, is it in a dashboard of what users have looked at this thing.
You also have the ability to select a time range. You see last 13, four days, whatever. Um, you're able to put this a little bit into context.
Um, so they have a utility score and a data points per second. You can see at utility scores based on utilization, how much score based on, let's take a scenario here where we dig down into one of these metrics here and we see that it has a really low utility score. So it's not being utilized a lot, but it has a lot of data points, which is kind of off, right?
So how is this being used? We click on the usage details up on the right there. Take a closer look at it.
You see here, yeah, it's not really being used anywhere. It's not in any dashboards, monitors, rules. Uh, it's not being, oh wait, here, it's being ad hoc, explored by two unique use.
That's weird. Look at what those users are doing. Oh, we know these two guys.
So here are two people that are, uh, two of our top SREs. Um, these guys must be using it for a reason 'cause they've looked at it more than once. Obviously.
Uh, we see 800 executions on the, the one by Joshua or Joseph, sorry. Um, so maybe others might wanna be using that as well. So we found an underutilized metric that two guys think are super important.
Maybe we should give them a call and find out what's going on with that. See if maybe we don't, or okay, if we move on to, uh, uh, the refining section here. Um, you can see that the refined phase has to do with being able to take action, uh, uh, without touching the source code or redeploying.
So being able to analyze and find these various problems, it's fantastic, but we don't want have to go out and have to redeploy or reach out to a, a development take care of the problem. Uh, that'll happen eventually, maybe, but you initially wanted to have your on-call team people to take action right there, right now because your, your, your, uh, budgets are going through the roof, also going through the roof and, you know, large cardinal alley spikes happen that, but the ability to take action on the, the, the ingested metrics right there before they're being stored back at the backend, uh, without touching the source code and redeploying. So we give you the options here to aggregate downsample your data, remove them, uh, by removing high cardinality labels or dropping non-value data.
This is at the ingest again, so, so all help you reduce costs and improve performance without any kind of alert query. And then we get to the, to the operate section here. Um, operate section.
Uh, it has a BA built-in capability to ensure that your queries are performing optimized. That, uh, requires no user intervention. So the ability to use to, to write queries and to query your data is fantastic, but that's not always an easy thing to do.
So there's something called a query accelerator that optimizes that automatically to ensure that while you're querying to fill dashboards, uh, to show, uh, data and, and time series samples, to investigate problems, um, you don't have to manually optimize that stuff. It'll keep your stuff performing automatic pretty easy. There's also a, a query scheduler that ensures that, uh, sharing the resources while you're doing these query.
So it's not just everybody other, no one group or user can crowd out the rest. And then there's a shaping rules ui. Um, this is helping you understand, uh, some of the refining and shaping that you've been doing with your, uh, metrics as they came in and you decided to aggregate or whatever.
It's really nice to be able to preview these policies and see what the effects can, what actually happens. And that brings us to the, to the why are we optimizing the observability spend? What are we looking at here?
So in a recent study by E S G, 69% of these companies are concerned with the rate of their ability, data growth, discuss that at. So when you're able to control and optimize your data, expanding the visibility and coverage that you have across your infrastructure, what's happening customer, um, it's also increasing, increasing the instrumentation, uh, of customer experience. So the focus goes from technology to instrumenting your customer experience, which is what we're all here for anyway, right?
Trying to get the business outcomes that we want. And then freeing up your observability time, the, uh, team's time to, to tackle real projects and strategic stuff instead of doing those 10 hours a week on, on troubleshooting issues all the time. Talk about, that's just a good example of what you're trying to get.
So the need is real, uh, need to understand and have the ability to tackle some of this stuff or that big flood of data. Just so much of this just business. Um, again, I said the, the, the, this whole uh, uh, framework and stuff we, we developed with our customers and give you an example of how it's impacted them directly.
You could see that at Snap they, uh, were able to reduce their data volume, volume by 50%. Um, reduce, uh, on-call pages by 90%. I dunno, who doesn't want 90 best pages on-call system?
Abnormal security, uh, company that, uh, has that 80 or 98% data volume reduction eight times as fast as tech problem. You see Robinhood, uh, financial, uh, trading application, they have, uh, data volume reduction, 80% free related the improvements eight times as fixed. You see that a lot of this has to do with it.
There's so much data and so much systems flooding in. It's hard to imagine that, you know, 50, 80, 90 8%. But these are, these are the kind of things that are, are, are being automatically generated that that mean nothing to your business.
This case in a Robinhood, that's pretty extreme, not normal. That's really extreme, but it's, it's how it goes. And, and here's what, uh, Robin had said to one of the senior staff engineers.
Sphe, we're able to not only significantly improve reliability performance, but we're also saving millions of dollars a year. I could imagine we reduce your observability data by more than 80%, save quite a bit of money storage. So with that, um, I'd like to share a few, uh, links here.
org. They went live before this started. So if you go over there, you should be able to find them resources, case studies link there that you can see some of the other aspects of that lifecycle go into the tooling I discussed, but was unable to show.
And with that, uh, I'm ready to take questions and I'll leave up here a little, uh, running, uh, demo. It's gonna show you, uh, a little bit more about the sticking through and being used on the live system. Ready Take.





