Managing FinOps in Complex IT Environments – Martin Mao, Chronosphere
Chronosphere CEO Martin Mao explains why managing FinOps in complex IT environments, made up of emerging cloud-native and legacy monolithic applications, requires higher levels of observability.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Martin Mayo, who is c e o for Chronosphere, and we're talking about finops and observability and how there's not enough of these two things going hand in hand.
Martin, welcome to the show. Thank you, Mike, for having me here. We're hearing a lot of folks talk about finops these days, and there's clearly a lot of interest, especially as the, uh, economic headwinds get a little tougher for certain vertical industries.
The question I would have though, is it seems to me there's a lot of folks who just simply forgot how to do finops. I mean, we used to do capacity planning all the time, and now we're kind of like struggling. And every survey I see suggests that there's a lot of interest, but not a lot of know-how, what's going on.
Yeah, I, my personal view on things is, I think over the last couple of years in a, in a zero interest environment, uh, perhaps the focus on cost efficiency took a backseat to improving the top line, to building new features. And I don't know if people necessarily forgot how to do it. Perhaps it was just lower priority on the list of things that every team, every company, every engineering team, perhaps, uh, had to focus on.
Uh, I think that changed in the current macro at macroeconomic headwinds. I think that has put the, the finops function, um, and, and the focus on cost back front and center, uh, and on, on everyone's minds. Now, It also seems to me is once you go down this path, you quickly discover that these, it environments got pretty complex in the last few years, and it's getting pretty hard to figure out what's going on.
There's not a lot of transparency. There's a lot of dependencies, and I'm not sure you even know where your costs are. So we talked about observability.
How do I get in there and correlate what I'm seeing to something that relates to what it's costing me? Uh, a hundred percent. You, you're spot on there.
The, the other trend we've been seeing the last three to four years is this move to perhaps a more cloud native architecture, right? A lot of companies are containerizing their infrastructure, they're moving more towards a microservices oriented environment. And you think in that type of, of architecture, it's much harder to attribute costs.
You know, you are running and using a tiny piece of compute for, for perhaps a fraction of a second or perhaps, you know, a couple of hours, as opposed to in the older world where, you know, teams just had a bunch of VMs and that was their cost. So I think the problem of even, uh, identifying and attributing cost has become a lot tougher over the last few years just as our architectures have evolved. And observability does play a pretty key role there in the sense that imagine these observability systems, like the ones that we build predominantly are used to tell you when something is wrong in your infrastructure or in your application.
However, it also has all of the utilization data. It actually tells you how much compute is being used for a particular workload there. And that utilization data you can imagine, is key to figuring out how much resources is actually consumed for the workload.
And that's a key ingredient into figuring out your efficiency and looking for areas of, of, of improvement. I'd say We collect all kinds of fascinating metrics in the world of DevOps, but I don't think we have one that says this is how much it costs alongside the performance ones. So do we need another tab or window that says, you know, this performance came at this cost and correlate the two?
Yeah, I, I, you know, I, I think so, and I think that is, you know, one of the main functions, a lot of these cloud cost optimization, uh, platforms out there. Um, again, that may have taken a backseat of the last few years, but we are starting to see that become increasingly more popular for sure. And I think to your point, having that view side by side is, is pretty critical.
Um, having that view side by side for your infrastructure is pretty critical. So you can imagine what is my utilization of my compute? Uh, and, and versus, you know, what is the capacity I provision?
How much am I paying for it versus the value out of it is definitely an interesting concept, uh, in, in our cloud providers in general and in the compute infrastructure. However, even when you shift over to observability, and when you look at the observability data, same thing applies. You know, increasingly over the last few years, a lot of observability data is now produced.
It's very expensive. You need a similar concept of, well, how much is this observability data costing me? What is the value in utilization I get out of it and trying to look for optimizations there as well.
So similar types of concepts now need to apply well beyond just the infrastructure that a company runs today. Can I play what if scenarios? One of the things you wanna do is getting in front of this before the workload is deployed.
So can I kind of slide things on a scale and come up with a cost factor, and can we just be generally smarter? That's, that's a really great question. Um, for us in the observability space where we're trying view place, we're definitely thinking about things like that because you can imagine as soon as you've done the changes and, and paid the cost, it's almost too late.
You can only sort of, uh, uh, perhaps fix things moving forward as to prevent these things from becoming outta control. So, um, for our platform, one of the things that we are looking at is as new observability data, uh, does get produced, uh, we do want to give companies an ability to sort of, uh, assess how much that's gonna cost, uh, ahead of time, uh, before they have to end up paying for it, right? So you can, you can imagine you can produce all the data, but before you actually pay for the data, because you may not be using it, you want to give them an opportunity to say, this is what the cost is going to be.
Is this really worth the value you are getting out of it? And have a company make that decision of a yes no call there before they end up having the cost. So the pattern becomes more of a preventative approach as opposed to after the fact, here's your overage bill, uh, and maybe you can do better, you know, for, for, for the next month there, because that pattern, um, generally doesn't work out, uh, as well.
All right, here's this little thing happening called AI out there. So can we apply that here and, you know, maybe AI will save us from ourselves? Yeah, I think, you know, if you look at the, the latest, uh, innovation in, in the AI space, there's a lot around large language models, a lot around generative ai.
And I would say that innovation in those spaces are probably less, um, uh, directly applicable to, to the problems that that, that we have been talking about a lot more recently. Perhaps it could be a better interface, uh, into, into how you interact with some of these insights. Um, but I think the, the innovations there in, in, in the past year, I don't know if can be directly applied, um, to the type of problems that we're talking about.
However, I would say, um, if you look at, you know, the cloud cost optimization industry or even the observability industry, there is a lot of automation that can be done these days. So, uh, one example from the chronosphere side is, you know, we are, um, automatically detecting now inefficient usage of particular observability data and suggesting and automating the optimization of that just like cloud optimization platforms have done for a long time, that can now automate a lot of the optimizations there. Um, now it's automation.
I, I don't know if a company wants to market that as, as AI per se, uh, but, but there is definitely an automation component, uh, to things. I dunno if it's, you know, quite in the same, uh, uh, ballpark or, or or square as, you know, a lot of the, the, the large language model innovation in the generated AI innovation we've had, um, in, in the industry this past year. Who's driving this spin ops conversation?
Is it the finance people showing up and saying, you know, thou shout and are they really trying to drive down costs or are they just trying to get a more consistent cost level going 'cause they're tired of being surprised at the end of the month? Yeah, I, I would say it's a little bit of both. Um, you know, if you look at the finops Foundation, it was actually created in 2019, so it's been around for quite a few years.
If you look at the Fortune, 50, 90% of them, uh, do have a finops function, uh, there. So it's existed for a while, and as we talked about at the beginning, I think the, perhaps the focus on it, um, uh, or, or, or the sort of priority to it, uh, dropped in the last couple of years. But this year it seems to be front and center.
I would say heavily driven, um, by the finance organizations who are trying to make every company more efficient this year. Right? That seems to be the headline.
Efficiency do more with less seems to be the headline for most companies, uh, in this particular year. And I would say that's why there's a renewed interest, um, in a lot of these. And when we talk to companies, it's a little bit of both of what you suggested.
It's both, can I reduce my current bill right now, uh, because, you know, that would help me, uh, reduce a company's burn, reduce the bottom line there, um, and get a lot of companies towards profitability. But more importantly than that, which you pointed out, companies really want control over this. They more than just having a lower bill today, they want predictability and they want control over the fact that, okay, if my business does grow 20% next year, I wanna predictably know how is my cloud infrastructure spend gonna grow in correlation with that?
How's my observability spend gonna go in in terms of that? And we are talking to a lot of finance organizations where they have a calculation, you know, for each increased top line dollar I am, I'm happy to pay x, uh, in, in infrastructure costs, or I'm happy to pay y in observability costs, but as long as I have visibility and control over that, over the long run, I think that's even more important than saving dollars today. Of course, everybody wants to save some dollars today, uh, at the same time.
So what's your best advice to folks who are trying to implement observability? 'cause a lot of the folks that I talk to, they love the idea, they're just a little overwhelmed by everything, and they're not quite sure how to get it to actually function as the way that they had hoped. And even once they get it, they're not sure what questions to ask.
So how do I kind of get this into some realm that it's just more accessible to everyone? That is a fantastic question. And I'll say, um, you know, even explaining why that is happening, you know, as we talked about earlier, the architectures are getting a lot more complex.
So the tool sets that worked for us years before don't quite work anymore. They, they weren't really optimized, um, for, for cloud native workloads, for, for containers and microservices. So I'll say the first thing is probably looking at, uh, a, a, a tool set and applying the right tool set for the right environment and for the right architecture.
So if you are moving more towards a cloud native workload for those workloads in particular, perhaps picking a a tool that is optimized for that type of workload, um, would be one thing that would make it more effective. One of the other blockers we've been seeing, and this is related to cost, is the traditional tools as a company, does this shift over to containers and microservices? The, the traditional tools get a lot more expensive.
'cause a lot more data gets produced in these new environments. So cost is often a blocker. So these tools may be working great, but cost-wise, I just cannot afford to have coverage or all of my workloads or all of my hosts.
And that is really bad place for, for companies to be in. So I'll say from that perspective, um, there, there's something as, as Chronosphere joined the finops Foundation, uh, we, uh, released a, a vendor neutral framework called the observability Data optimization cycle. It's a bit of a mouthful, but essentially it's a framework to, to apply, uh, to get visibility over the cost of observability and apply particular techniques, uh, into how to control, uh, that, that growth and data irrespective of what, uh, a tool you have.
Um, that there are ways and techniques in which, um, uh, through this framework you can go get control over that growth of, of observer bleeding data. So that could be one thing that could be useful for companies out there to solve that cost problem. And then perhaps picking the right tool for the right environment would be my other piece of, of advice there.
Are we still culturally too wrapped up in chasing predefined metrics and monitoring tools and not quite going to the next level? I, I feel like a lot of people are still struggling with, well, I do monitoring, what do I need observability? 'cause doesn't monitoring mean observability?
That's a great question. Uh, I think there are a lot of thoughts in this particular area. Um, my personal view on, on this is to your point, you know, the buzzword is now observability.
The, the end result of what you're trying to achieve is the same, right? We are all trying to reduce M T T R like we have for the last 20 years, or M T T D, right? We're, we're trying to reduce the time we can detect and resolve issues.
That's the name of the game. And that hasn't changed. Uh, the thing that has changed is, again, these architectures have changed that, that the, the, the, the infrastructure you're running on is so much more complex now than it was before.
And perhaps you need a new tool set to go and, and approach that. Now, that new tool set could be, to your point, um, may maybe not predefined, may, maybe you don't want anything predefined and you want everybody just to go and access the raw data and go debug your problems that way. That could be one approach to, to the problem.
What we have found working with a lot of companies out there is that approach is, is only effective for certain individuals in an organization, the power users, the folks that know how the infrastructure runs. However, what we find is most of the, uh, the operators of, of services in production these days is the average developer. And the average developer doesn't know how the rest of the infrastructure works because it's, it's fairly complex.
Uh, so, so we do find a need to still have some, I don't know if they're predefined, but have some easier concepts, um, that anybody who's not an expert in observability can pick up and use those concepts, uh, to, to go and debug their systems and to go get the job done and not just cater to the power users. Uh, so, you know, my belief is both exist in the world and the tooling probably needs to serve both audiences. Uh, there.
All right folks. Well, you heard it here. It's an angel of truism.
You can't manage what you can't see. And if you can't manage it, you certainly can't control the cost. Hey, Martin, thanks for being on the show.
Of course. Thank you so much, Mike. All right.
Back to you guys in the.