FinOps for IT Cost Management with Eric Ethridge
Eric Ethridge, senior technical account manager for DoiT, provides insights into how organizations should best implement FinOps initiatives to rein in the cost of IT
Transcript
This is Techron tv. Hey guys, thanks for throwing. We're here with Eric Ethridge, who's senior technical account manager for doing, and we're talking about finops and how to get started with this whole thing.
Eric, welcome to the show. Mike, thank you so much for having me. Uh, it's a pleasure to be here today.
I'm happy to be talking about finops something that all of my customers are focused on. I guess, why are we all talking about finops all of a sudden? It's not like we haven't been managing the cost of it as long as I can remember, but what's different about finops?
So I think this comes in waves a little bit. Um, first we had the joy of, uh, enterprises deciding it's time to move to the cloud. Um, we want to change the paradigm by which we do our workloads, our infrastructure.
We want to move it to a more malleable platform, which would be in the cloud providers that we have now. Um, and there was a lot of settling that came with that. Um, and there's a lot of getting used to what is a cost, what's not a cost or what's a regular cost, what's not a regular cost.
Um, and now there's, now that we've discovered like what we can expect from the cloud, I think enterprises are starting to realize that, hey, since this cost is so flexible, since we have control over it in such a greater degree, we can actually, uh, with some focus areas with turning this in from a top down to a bottom up kind of, uh, engagement. Um, we can actually have our enterprises be focused on something to make it part of our day-to-day. Um, let me back up one step.
It's become part, it's become a, uh, focus because it's the, it's the model by which we can actually affect the cloud spend for the, the, these customers can affect their infrastructure spend. Um, it's kind of a wrapper on the same thing we've always been doing, but with more of a focus on how the cloud operates. So one of the things we have seen is that the costs are highly variable.
And so the question then becomes, um, are organizations getting in front of that or they more frequently still kind of waiting for the bill at the end of the month and then they get hit with something that suddenly is two x what they thought it was gonna be. I'm only, I'm seeing one happen before the other organizations, uh, if they're new to the cloud, sometimes they, they have that end of the month surprise. But then it only happens once because something else the cloud gives us is the ability to see that spend much, much faster, um, because you are racking up usage on a daily or hourly level.
So I'm seeing organizations see those costs, um, at the end of the month the first time, but then start enacting finops controls, labeling, cost alerts, budgets, um, uh, and then they start identifying where that spend is coming from. 'cause that's always the next question is not is Wow, this is really high. How did it get here?
And that always leads back to what part of our business has caused us to do this, which then drives the overall, um, the overall finops discussion of understanding workloads, understanding data flow, um, understanding where costs come from, How much of this is, you know, just shall we say carelessness versus how much of this is, maybe we just underestimated the size of the workflows. I think that's really organizationally dependent. Um, part of that goes back to are you upskilling your people when you upskill your workloads, um, as these workloads mature and turn into like cloud-based workloads or if you are even born in the cloud, um, cloud native as you will, um, are you properly like building a foundation?
Um, that prevents that carelessness because, uh, folks understand how what they do drives costs for the organization. Um, uh, and sometimes it's just because we're, uh, the same way we're, we're used to doing the business the old way. We just spin it up, let it go.
We're not paying anything extra for it because we're used to that on-prem workload infrastructure where we just have it available. Um, and again, this is usually a learning process. It doesn't take organizations long when they start getting, um, some of those cost alerts or some of those cost overruns because folks aren't utilizing things correctly to start implementing, um, a more intelligent way of doing business in the cloud by training, by upskilling their entire workforce to understand like the fit ops paradigm.
Very few developers or IT people that I know would deliberately go out of their way to spend a lot of extra money. It's just that they don't have any visibility into it. So how much of this is an exercise in just making sure that the right, uh, metrics in terms of costs are being tracked alongside all the other metrics that we normally track?
Well, I think it's a twofold question. You're right though. We, we track a lot of metrics.
We track things like usage metrics for CPU and ram, um, but sometimes organizations aren't moving down and also tracking will this belongs to customer X or we this, um, especially in a multi-tenant architecture, when you start dealing with a multi-tenant architecture, especially like in Kubernetes or any other kind of, um, containerized, uh, almost hidden workload inside of that infrastructure, you have to start using cost controls and take advantage of all of the things that have been built into these platforms such as like pod and node labeling, um, pardon me. Um, some of the major cloud providers have, uh, especially uh, some of the major cloud providers have provided, um, some very good cost and usage tools that you can enable inside of these workloads to help you more, uh, have more granular visibility there. That is the capability that's in there, but there's also, uh, do we have the engineering time to take care of these?
Um, that would be the other part of this 'cause we'll move our workloads or we'll start our workloads on the cloud. 1, once we've taken that first step, like you, once you've gotten past day one if you will, or day zero, um, there's much more of a focus on, there's much more of a focus on iteration versus, um, iteration on core products or iteration on the services we're currently providing to customers versus going back and taking those engineering hours and that engineering time to, um, label or put in tight of some sort of hierarchy. All of our workloads and all of our, the things, they're spending money as much as we can.
Um, and no organization needs to be down to like the hundredth percent on this. Um, uh, but they should have an idea of where their money streams are going inside of their, uh, um, their cloud infrastructure. Are people moving workloads once they kind of figure out their costs?
And are they actually shopping across these different cloud service providers? Or are they more likely just to be trying to optimize the existing infrastructure because well, the cost of moving something is not, uh, insubstantial, shall we say? Absolutely correct.
The cost is not as substantial. So I think when customers think, can I save more money and I'll using cloud A versus cloud B, um, one of the things they have to look at is how fast could we move this? Like if we're, are we gonna step over dollars to pick up pennies?
Is really the way that it comes down to it. Additionally, do you have specialized workloads that cannot move off this cloud provider? Um, as an example, um, Google has a data platform called BigQuery, that's great.
Um, however, it is only available on Google. Um, so if you're starting, if you're doing things like, well I could save if I moved over to cloud X from here, are you always gonna have a workload that's still on that cloud? And then are you moving data back and forth?
There are some long-term thoughts about, um, how far are you willing to go just to save that little bit of money? Um, and before you even do that, the big question is, have you optimized where you are currently? Um, and sometimes that optimization means you have to go back and again, take some of those basic engineering steps of do we know where the money is going?
Do we know how do we even have the right operating or pardon the ops agents or monitoring set up to understand how much we're using these resources. Um, is our bill our failure or the cloud being too pricey? And that's a hard question for organiz organizations to ask because they have to turn around and then say, are we doing things right?
And not a lot of folks wanna admit that they might not be. It's a hard look to take it yourself. How good are the tools provided by the cloud service providers in this regard?
And people ask this question 'cause some of 'em are a little dubious. They say, you know, well, how can the cloud service provider be an honest assessment of how much I'm spending on their platform when they have a vested interest in me spending more? So, um, do we need more of an independent lens?
So, uh, That's a complex question. Um, and I appreciate it because I think that cloud providers have a definitive vested interest in showing you exactly how much you're spending on their platform. There is no use obfuscating it because then in one aspect that would position any other cloud that was just like, by the way we provide a hundred percent visibility into your spend, that would make them a market leader instantly.
If it was, if it turned out that other folks were obfuscating it. So I really feel that the, that the cloud providers have, um, almost a responsibility to provide as much granular, um, understanding as they can without negatively impacting like additional costs to give that data to customers or their internal infrastructure to give, to make sure the customers have those data points. Um, there are, uh, as for the cloud providers themselves, their tooling right now is adequate.
Um, it might not be their internal tooling such as if you go to like a a w like AWS is cost explorer is incredibly deep. Um, it is easy, it is, it's not the easiest to use by some folks measure, but it is, it does give a lot of granularity. Um, uh, doit actually has built our own dashboard, our own tooling that ingests all of this granulated spend from the cloud providers using customers billing data.
So we're not digging around in your data and lets you provide and provide you like a multimodal view of it. Um, especially, uh, the big part is, um, and this goes back to your previous question, what if you moved clouds? Or what if you're a hybrid cloud provider?
What if you're a hybrid customer? What if you have both? Um, you know what, if you're an all three of the top three clouds, what if you have an Azure and AWS and GCP workload?
Well, inside of each one of those, excuse me, that data is siloed. Um, so what do it has done is we built a dashboard where we ingest all of that is, uh, for our customers and let you provide like a a and provide you a holistic view of it, um, while still maintaining the separation itself inside of the spend streams. Um, now the cloud providers themselves provide some great granular data.
You can also export your data from all of them. Um, AWS lets you export a cost report. Google does and so does Azure.
So you can see that same granular data. Um, so it's available to customers, um, it's available to you, me, anybody who's using them. The problem is building the tooling around it to make sense of it.
Given all that, am I trying to create something that feels like a finops center of excellence or am I just trying to drive this all the way through the organization and it becomes just part of our standard muscle memory? I think for any leader in this space, it is. So you can do it in either way.
You can use top-down, uh, you will to push this, um, without major consensus or buy-in or wiggle room. And that's all I've seen that done. And it works.
You've told people to do something and now it happens and you know, and I mean you from a leadership perspective, not, uh, you as pointing an individual. Um, but I found much more effective when leadership admits there's a problem when the folks that are calling the shops admit that, Hey, we have a problem and we need everyone to come to the table to solve it, because this is really, if you're using the cloud for your workloads and your infrastructure, everybody in the organization from hiring, well, everybody from the organization has a vested interest in making sure you're using that effectively because that's a major cost driver. Um, that's usually one of your largest cost drivers besides personnel.
Um, and much harder to get rid of that spend than it is to get rid of spend in other areas. Um, so, uh, so it behooves organizations, um, and the people within them, uh, to want a finops center of excellence to want something of that nature because then it's visible for everybody and everyone has buy-in. So the engineering team doesn't feel of doubt.
Marketing understands how what they're doing drives like revenue. Um, the executive suite then sees a team of people, and I ref I will not use the term tiger team, uh, to come together and do the thing, which is, and then it's also, in a way it's more autonomous, then the organization doesn't have to monitor it, it's monitoring itself. So leadership has created not just a, the leadership hasn't created an edict.
What leadership has done is created a tool that works inside the organization by using the capabilities already inherent inside of it. Um, also increases organizational communication and understanding everyone becomes a champion at that point. So what is your best advice to folks and what do you see organizations routinely do that's probably counterproductive and maybe should be avoided?
Um, I really appreciate this question because I've been, I I I've, I've had a couple examples of this this week and I'm gonna use one. So I always consider something to either, like to the left, it's the right of an event, either before or after. Um, and usually when there's a cost overrun, the first knee-jerk reaction that organizations have is how much closer can we get to when that cost to overrun occurred?
Um, such as like, can we implement real time monitoring of our spend that does all of this stuff and gives us a flag when we spent more than 10% or whatever. Um, and using these tools that actually might increase cost and complexity of your workloads more when you could take that same time you've spent developing all of these, all of these different alert paths and everything else about your spend, um, to actually improve your infrastructure itself. So that way these things don't happen.
As an example, um, GI ops, GI ops is incredibly powerful. It is incredibly powerful, especially for things like infrastructure changes, um, for things like data movements, anything where there's a, a, a gateway put in place that you have to adhere to certain things to helps prevent errors. Um, I recently had an experience where a customer didn't use a GI ops type pipeline, and it, it resulted in a, in a six x cost overrun for a single job inside of a data processing platform.
Because what they failed to do was make sure the tables were, were rebuilt correctly so that none of the, like the clustering or anything else was there, and that next job run cost them six figures. Um, a proper GI ops pipeline would've ensured that this didn't happen. Um, so, and there the reaction to that was a little bit, well, how will we have known this happened sooner?
I'm like, well, the money's already spent once you've ha once that money has been spent, it has been spent and that cost has happened, whether you get alerted about it, and you were talking about this earlier, do we wait till the end of the month to come in? No, you shouldn't. 'cause you shouldn't know these things have happened.
But once that money is spent, whether, you know, an hour later or the do it platform because of the way the billing is ingested, ours takes about 24 hours to show cost overruns like that. But whether it's an hour or 24 hours, it's not a break fix issue. It's not a pipeline down issue.
It is the money has already been spent. So while I can appreciate trying to get to the, get as close to that mark as possible, spending all that engineering time to make sure that you are like running in the best type of, uh, the best, uh, workload pattern and deployment patterns that you can is much, much, it has much better dividends in the long run, I believe. All right, folks.
You heard it here. Hey, it's always been true that an ounce of prevention is worth a pound of cure. And it sounds like in the age of ops things aren't all that much different.
Hey Eric, thanks for being on the show. Hey, Mike, thank you so much for having me this morning. This was great.
I really appreciate your time and, uh, happy to come on. Any other, any, uh, any other time you need someone to chat with? Alright, and back to you guys in the scene.