Scott Kelly & Peter Simkins – Four Signs it’s Time to Level Up Prometheus
In this session, attendees will leave with the knowledge needed to answer critical questions around their existing Prometheus setups and whether it’s time to consider a Prometheus-compatible solution built for massive scale.
Transcript
Hello everybody and welcome to today's talk. We're going to talk about four signs. It's time to level up Prometheus.
My name is Scott Kelly. And I am a product marketing manager at kronosphere. Peter wanted yourself Yeah, my name is Peter Simkins.
I'm an engineer cronosphere as well that I've been in the industry for just say 10 plus years, so Hey Scott. Hey, all right, so to start off I quickly just wanted to walk through what the bill background about chromosphere and who we are just to give everyone make everyone familiar with us. So we're an observability platform built from the ground up for cloud native scale and complexity.
Our origination of this solution actually started at Uber. Our Founders were the they started the observability team there and as Uber microservices and container-based infrastructure grew It became obvious to that team that all the solutions that were available both, you know in the open source world and in a from the vendors was not enough for the scale and the reliability they need or for the cost efficiency. So they decided to embark on building a new open source metric platform that would ultimately become to know known as M3 or m3db.
So by 2018 M3 at Uber was the second largest production monitoring system. And around this time Integrations with Prometheus were introduced to allow a broader Market to take benefit of the project. And so as more companies started to adopt M3 m3db.
Our Founders realized that there were many organizations that needed more than the oversource product could offer. So in 2019, they found a coronosphere to build upon the tech and expertise gained so they could help more companies that were embarking on their Cloud native Journey. And in that short period chronosphere has become a trusted observability solution for some of the largest cloud native organizations around the world.
And with that let's get into what we're going to talk about today. So our agenda here, we're gonna give you an overview of Prometheus just to make sure everybody's familiar. We'll talk about you know, why it's great for getting started.
We'll talk about you know with some of the signs that you you know, as you start to scale this that you may need to look at up leveling or making improvements to the solution and then we'll do a quick recap just to give you an overview of you know, you know this commercial and there's open source Solutions out there at the, you know, the pros and cons and how to things you should consider when moving forward so You know, basically, let's look at Prometheus. What is Prometheus? So Prometheus is an open source monitoring and alerting system.
It was started by SoundCloud in 2012 and donates the cloud native Computing foundation in 2016. It was actually the second graduate behind kubernetes a project out of the cloud native Computing foundation and the growth has been pretty dramatic, you know, since his adoption, you know, it's growing quickly, you know, a recent Cloud native Computing Foundation survey actually found that 86% of their respondents and their organizations were using for me theist for monitoring a learning and it makes sense. Right because it's really the de facto solution for monitoring kubernetes.
And so why it's very easy to get started, you know uses a single binary to use for ingestion storage inquiry and then as a separate by any for learning, but it's very simple. It's out of the box stuff the cncf recommends the tool so that means that most of the software In the cncfe ecosystem also exposes metrics in the Prometheus format. So, you know, that's why is really become the de facto tool.
There's a really active Community around it for support. So it's really easy to you know, get updates and advice, you know, it has some really cool features like, you know, dynamic endpoint discovery of on many platforms. So it's really easy for to discover any platform you're running for example, like kubernetes obviously the main one and it really makes it easy to integrate start discovering particular metrics and it also has a wide ecosystem of exporters.
So most major software projects, especially the open source ones have existing Integrations from most pieces of software that you'd want to monitor as you scale. You will run into challenges. And today we're going to walk through some of the common challenges and some of the methods steps and other open source Solutions you can take and you're used to alleviate them.
So let's jump into it. So the four signs it's time to level up Prometheus. So slide sign number one is really about you know, you're seeing your engineering overhead increase and it's getting difficult to identify and locate monitoring data when needed so, you know.
We're talking about scalability really right. So it's it's a well-known fact that Prometheus wasn't designed to scale horizontally horizontally, right Peter. Yes.
Got a typical scenario is you set up one Prometheus instance to scrape service a as an example service a gets really popular. That is the service that everyone seems to be using. So as you admit a ton of more metrics you start to overwhelm that Prometheus instance, so to ensure stability He's been up another Prometheus instance to collect from service a as you grow.
And yeah, that's that's perfect. So, you know as you get into I guess different instances of Prometheus the team has to know which Prometheus instance to go query from to know where that particular service detail is that and your dashboards and everything else start to get a little bit scattered as you as you scale up. So the next obvious thing that you have to do is start to Federate the the different data sources, so bringing it all together.
So you have one common place to to query from it'll make your dashboards a lot more consistent. It makes alerting infinitely easier. So yeah, Prometheus scales like crazy, right and Confederation to do this is the is the real trick.
Okay. So a Federated to a third instance. Does this solve all your problems?
No, it really doesn't give you a global view as you continue to Federate and grow. So yeah, there's there's a gap there. For sure.
Yeah, so one of the challenges here, right? Is that as you grow and you've got you know you Federate to third instance and where do you query right? So you you've got a subset of data in that third instance, right?
So if you point your dashboard at that, you don't have that full view. So it requires you, right, you know don't give you a subset of data you need to know which instances to point at dashboards and alerts and this will continue to happen. So, you know, it creates problems from a diagnosis perspective because if you don't have the right data you have gaps so, you know the real I think the solution here right Peter is to have to keep hiring people and expertise and staff to maintain the scale.
Yeah everywhere. I worked that's slow diagnosis is a real problem too. Right?
So the the managers start to complain that queries are taking a little bit long you don't get that real time data as you grow and Federate continue to Federate. So yeah, hiring expertise is the way to go. and I think the big problem here right is that it just keeps getting worse or scale, correct?
Correct. Yeah more complex, you know Prometheus as you continue to scale for sure. All right.
So we talked about you know, some of us some steps it's list. Let's give some the people, you know, listening a little couple more ideas and what they can do so we can recap what we talked about. But also some other things that you yeah leveling up Prometheus first scalability.
There are a ton of options and these are just some easy takeaways. Although there are easy takeaways. They do require a little bit of work.
Right? A lot of this is is something that you're gonna have to manage yourself but for example different data stores that are a little more efficient than just running different Prometheus instances such as Thanos and 3db me mer Victoria metrics, they all require a little bit of care and feeding but it's a single data store. So you might get some improve performance as well as scalability the different Cloud providers you can also use to drive down costs and improve Know fault tolerance, they all have their different products that you have to manage but they're available to help you improve scalability and then you know easy, I guess in terms of it's available.
You can continue to Federate although each as we discussed each one of those, you know instances that you create. It requires some Administration care and feeding you always want to increase your expertise with Prometheus. So having those people on the team that know how to exercise that muscle always good.
So just some some key takeaways there Scott, right? Thanks. All right, let's move on to sign number two.
So, you know, you're losing monitoring data. I needed to keep Mission critical Services running reliably. So let's talk about reliability.
Prometheus isn't highly available by default, right? No, all your data by default goes into a single prom instance. This is amazing to get started.
But if it goes down it's not ideal right to lose that you know, real-time monitoring of all your services, especially historical data. This is usually important for for a lot of services that you have to support. Okay, so, you know now we have two instances receiving the same data and different zones, obviously your storage costs and maintenance increase when you do this, right but when it comes to dashboard queries, how do you address that?
You can put a load balancer in between them. So you would Point grafana basically the load balancer instead of the the single or multiple Prometheus instances. So the read requests get balanced between the two prom instances.
So if one goes down you're still able to fulfill the requests this works great for reliability since that you get one copy of the data. but when a lot of times people will do rolling restarts of that, right? So it doesn't that create the chance for data discrepancies.
Yes, definitely seen this you can get gaps in your data. As you know, one instance is down, you know and restarting that is the problem for sure. So let's look at an example of how that would kind of look if you were to look at the different, you know data sets from the different instances, right?
So so I think you can see here in this example. If you did a rolling restart of both prom Prometheus instances, you can see that each instance is missing the metrics from where the instance was down. So neither actually has a full image of you know, the data over that given period And I think one of the challenges from you know.
With Prometheus is that there's no way to merge these easily right out of the box. So now you've got missing data which creates you know, discrepancies and and can be Troublesome when troubleshooting right true. Yeah.
I mean as you refresh your graphs possibly you can you know start to see the backfill of that data, but you're right you're not losing anything historically and you can compare between the two but yes, you do run into the concern about having some gaps in your data. All right. So let's look at some things that people can do, you know recap for people can do to improve reliability.
Yeah. Yeah, this is exactly just the takeaway of what you mentioned earlier implementing a load balancer between your different data sets really easy way to improve that, you know historical backend and as as we mentioned it's like you don't want to have a big gaps in your data definitely cloud storage between different Cloud providers. It's a great way to improve your your fall tolerance and reliability as well.
And then of course the improved data stores, which does require some care and feeding that I think is a little bit better if you want to single day to store as well such as Thanos and three Victoria metrics Etc. Cool. Thanks.
Let's move on to sign number three. so when teams need to retain more granular data and for longer period of time So this third pain point is really about efficiency. You know, Prometheus is not very efficient for long-term data, right?
I think it's going to maximum. Is it 15 days? Yeah tension.
Yeah, it's it's a big problem with no built-in down say sampling capabilities, you know long-term, you know volumes can get overwhelmed pretty quickly and totally true. I see most people storing like two weeks typically maybe a week of monitoring data. It's a it can get expensive.
Let's walk through you know, so if we talk about solution here would be to Federate again, right so to store and scrape it different intervals. So why don't we walk through? This is a put this quick example together just to give people an idea of what you know, federating and doing different interval scrapes can actually do from you know from us from a storage perspective so quickly when we talk through this, Yeah, yeah, so, you know, if you have one instance very very easy.
If you want to store that six months, you know to 30 second interval very very easy, right and even six months isn't that expensive in terms of what you're storing? You know, but if we change that to like a one hour interval all of a sudden that dramatically decreases and as you continue to grow that hundred instances gets a little bit scary, right? Like that's that's where you're you're storing quite a bit.
So definitely creating a different time series to store your metrics in the long term is a great way to I guess increase the the time without you know burning through all of your your storage budget. All right. So if we Federate the data as you suggested, what else has to be done, you know if anything like because I understand it, you know, there are some issues with this in terms of you know, having to run separate queries things like that is like, you know, so federating isn't a Panacea, right?
Correct. Yeah, you still need to as you create these different time series, you're creating a different, you know metric name if you will, so you have different, you know dashboards to look at, you know, separate views data, you know, it gets a little bit cumbersome with all the overhead, you know as you Federate out and you know, create all these different time series. Okay.
All right. Well, let's look at you know again just kind of a recap here of you know from an efficiency perspective. You know, I think we've talked about the storage you can do that federating at the different time series.
We talked about that and you know, improving the data stores. Why don't you talk a little bit about that? Yeah, I mean great Point we've we've kind of be you know, the the first two pretty well but you know having all of your data in a flexible time series data store that you have a little more control over definitely improves, you know query performance, you know, the efficiencies across the board but each one of these once again, you have to you have to manage and we'll talk towards the end.
This is an improvement. There's a little bit of overhead but there's there's a missing element just as a cliffhanger towards the end. Cool.
All right, let's move on to our last sign and it's your modern costs are growing faster than the business is actually growing. Right? So this is starting to become a problem.
Right and and when you walk through how we got here, right? We've got this timeline kind of describe what happened along these different points. Yeah.
So I've, you know been in the industry for a few years worked at large companies like Disney, for example Disney, we have our own data Center's, you know, bare metal, you know introduce, you know virtual machines fast forward to you know, the cloud providers started to come out and as you grow, it's like technology. I don't know what the rule is but, you know doubling every year every other year it basically always outpaces the growth of the company so developers love to develop and make improvements. So, you know introduce containerization.
If you are increasing the amount of you know, functional improvements to your business to your applications as you're growth grows. Let's say it's 80% year over year with the introduction of containers. If you're back end is, you know, 200% 300% growth.
It's hard to go back to the business and say, you know, hey, I need you know to support, you know, all of this and it's you know, 300% more than what the business is growing at. This is a common pattern that I see when talking to companies. I think a lot of Companies find this be surprising right they were used to this kind of steady trajectory as they moved from, you know Legacy to VMS, but when you got to Cloud native and containers and microservices, that's like an exponential explosion and I think you know, I think a lot of people don't expect that.
They don't understand that that's gonna happen at that rate and it can overwhelm, you know a system and a budget pretty quickly. Yeah, and I think the industry hasn't done a good job of you know, explaining this to to everybody. I mean the message is always been, you know, create all the metrics you need sore all the metrics that you need forever and the business, you know gets this huge cost in terms of, you know, human time just support it as well as the cost, you know to pay for it and it's unfortunate right because you you definitely need this but at the same time like technology be explicitly gross.
Yeah. And it's not just data costs and we're kind of focused on data costs here. But if we look at, you know, there's lots of other costs associated with rolling your own solution or you mining solution.
So for example, you know people engineering Talent is really expensive and in very high demand, even you know, even now and you know, and when Engineers have to deal with problems, you know, if they're manually running a system and it's got It's falling over or there's mistakes, you know that can lead to burn out right you're on call and you're getting all these pages and things aren't working and neural monitoring system isn't working sometimes right so that can be a challenge that can be a cost of having to get new people. I mean we talked about you know time is a clear one right and do it yourself morning system. You're going to be doing it all and that can really suck up.
A lot of time you have to patch test troubleshoot, you know, even the mind you probably shooting the monitoring solution itself and that's not including the systems. You're actually monitoring, right? And so and then it's really a question about productivity and You know, do you want to have your staff focused on being Prometheus experts and being a great Prometheus shop or do you want to focused on the value that you can drive for the business and The Core Business needs.
I think those are some things that people, you know, they think about their monitoring costs. They don't take those into consideration along like open source is free, but it has these additional costs that are you know, they're hard costs or trade-offs. You need to consider.
I love open source. And right this next slide is like I'm just some key takeaways where I can I can extend your permethius instance for quite a while right modifying the scrape interval really cheap inexpensive way, you know, do you really need, you know, five seconds scrape interval, you know, maybe we can change that up to every minute and now I can extend the the runway for storing these metrics for quite a while and then as we talked about, you know, improving your down sampling. Yeah right other change to a different time series.
Yeah impact data stores for sure. Yeah. All right.
So let's do a quick recap of what we kind of talked about today Peter why you talk about this, you know hierarchy of observability needs. It's kind of an interesting concept. Yeah.
This is very opinionated in terms of this is my opinion. I see in the It there are a ton of different data stores out there people have solved it with different solutions a hundred times over there's great stuff out there alerting. We've we've pretty much nailed Prometheus illiterate manager Works fantastically.
Well visualizing data everybody. Does it this is this this table Stakes right being able to to use grafana to you know, connect to your data store in a very timely and available fashion. We talked about it, you know previously and then the ability to scale, you know, that's really key to your business because I'm hoping you're growing but I think what's missing in the industry has been around this control piece, you know being able to In just data and transform it in real time and you know write out to a different time series and putting that control in in the hands of of people that have to maintain and support.
This is the real missing piece or the and there are open source solutions that that do this, but I think a lot of the the SAS, you know vendors are really coming online to start to tackle this this piece the top of the pyramid if you will. Yes, so I think you know, we've talked a lot about the open source kind of options for you know, addressing these needs and so, you know as you're you know, as you mature the belief is that most will want to move to a SAS solution right? Because it's just the overhead becomes way too much.
So I think let's just leave the audience with a couple, you know key considerations as they kind of you know, look at those sass Solutions because it can be overwhelming. You know, they have great marketing budgets and things like that. So a couple things that you really need to do when looking at solution is, you know, ask about features that help you control your data growth right without sacrificing performance.
That's key. Like what are they doing to help you control that data? And then the other idea to keep in mind as you look at SAS Solutions is to the more open source focused and compatible.
They are the less likely you have vendor lock in because that's one of the key things that people are afraid of and I think that's why they try to stay with an open source solution 100% is because they you know, they want to have that. Up, right? They don't want to be locked into a vendor and the more compatible.
Someone is the less likely that can happen because you can export and go back to something if you actually get into a bind, correct. Absolutely. Yeah.
All right. And with that we have some next steps in terms of some documentation that you can download some resources. Please go visit the kronosphere booth you can get a demo and talk to a different sees there and to help you out with any questions.
And also there's some customer case studies on this link and you can find them on our website. And also if you aren't able to get to the booth feel free to reach out to us and we'll connect you with an expert to start a conversation about you know, what your needs are. And with that I'm gonna take Q&A.





