IBM Kubecost’s Kai Wombacher on Shifting FinOps Left to Control Kubernetes Costs
Kai Wombacher, product manager for IBM Kubecost, explains why responsibility for FinOps needs to be shifted further left toward application developers and software engineers to have a meaningful impact on containing Kubernetes costs.
Transcript
Hey guys. Thanks for the throw. We're here with Kai Wambach, who's a product manager for IBM Cube Cost, and we're talking about well shifting finops left.
Kai, welcome the show. Awesome. Thanks for having me.
It's a pleasure to be here. All right. I think part of the issue that we have seen over the years is the cost is, well, something we think about maybe too late and after the fact, and, um, yet the people who are actually spending all that money, or engineers and developers at the very front end of the process.
So, should we be rethinking all of this? And I don't know, in the case of Kubernetes, do we just need to put some sort of, um, dashboard with some metrics in front of people and let 'em know how much something costs? Yeah, absolutely.
Um, you know, thi thanks to that question, I think, you know, when you look at Kubernetes, right? Uh, it, it delivers a ton of value to users, right? It makes it really easy to deploy applications, right?
You don't, your, your application developers don't have to get too into the weeds. They have like a nice layer of abstraction, making it really easy to deploy applications. It's also really easy to ensure, you know, high availability of applications, do things like replica sets and all these things.
So it's, it's, it's really transformed the way that, you know, developers and organizations deploy and manage their applications. Um, but those layers of abstractions have, you know, come at a cost, and that typically is like increased cost and increased spend. You know, I can, uh, tell a story from my background when I, I was a machine learning engineer prior to getting into product management and, um, you know, deployed a machine learning model, a big, you know, recurrent neural network, um, let it run for a week or so, and then came back only to find out that, oh, I had spent like more than 40 k just training that model.
And, and to your point, it was super reactive, right? We didn't, we didn't know that I had spent that much money, uh, training the model until the end of the month. Um, so, you know, this, uh, problem is something all organizations face and is really, you know, at the heart of, of, you know, where we see a lot of organizations going to, you know, like you mentioned, we wanna be more proactive about costs.
We want to shift that left and give our, our developers a better understanding of, you know, what they're spending and, and how do we get that, um, you know, uh, cost visibility, you know, kind of before they deploy an application. So that's definitely something, you know, we've seen a lot of, I think, you know, to your point, dashboards go a a long way, right? I think that just having that culture of cost visibility in an organization is so, so important.
And I think, you know, regardless of what your end goal is, you really need to start with some of those dashboards to kind of provide your application teams, like, Hey, this is how much, you know, this application cost, or, Hey, we see this spike in spend, or, you know, Hey, we've set this budget for you. I think that's kind of the, the foundation for, you know, effectively shifting your cost left. I'm often wondering how heavy handed do we need to be?
'cause I think most engineers are kind of reversed to waste anyway, but if somebody just tells them what the level of waste is, they might do the right thing, uh, of their own accord. That's absolutely true. You know, I think that, uh, it really depends on organization to organization, right?
How they're structured. Um, but you know, we, we definitely see sometimes some organizations go very heavy handed, right? Um, you can, you know, integrate Q costs, let's say with a policy engine, you could use open source verno and actually start to enforce budgets on your application teams, like, you know, prevent overages on that budget and prevent changes from going through that would, would, you know, take you over budget.
So, you know that that's maybe like one end of the extreme, which is, you know, Hey, we're gonna be super heavy handed and we're gonna enforce budgets. I think you see a lot of the other end of the spectrum where it's like, Hey, you know, we're just gonna give our developers visibility into this change and, and the cost impact. And actually one of the things we've also seen a lot of is, hey, if I can show you not the, just the dollar cost impacts, but if I can show you the carbon impact of, you know, this workload, this change that you're trying to make, um, oftentimes you see developers even more, uh, you know, interested in making that change and, and re you know, reducing their resource requests because not only do they see the dollars and cents, but they see the kilograms of carbon impact as well.
And, and that can really, you know, have a big impact. And, and, you know, the developer wanting to be, you know, good stewards of, of their organization and of their, you know, broader, you know, world. One of the things that has always perplexed me about this in relation to Kubernetes was, I thought part of the purpose of Kubernetes was that it kind of lets you dynamically scale up and down, and yet it seems like we scale up, but we never scale down, and we have these over-provisioned environments.
So how does that all come about? Absolutely. Yeah.
So, you know, we, we work with a lot of organizations who are using dynamic auto scalers, like a carpenter or a GCP autopilot, right? Um, in fact, I think gcps autopilot is now on by default, um, for, for GKE clusters. So, you know, a lot of teams do take advantage of the scalability.
One of the areas we see tremendous waste is in the resource requests, right? So a, uh, a dynamic autoscaler, like a carpenter autopilot, those assume your requests to some extent are efficiently set, right? And then they come in over the top and try to, you know, uh, provision the appropriate amount of nodes, CPU, ram, et cetera.
Um, where we see tremendous waste is, you know, when a developer goes to deploy an application or a given container, they don't really have a great sense of, oh, it's gonna need this much CPU and this much round, right? So what ends up happening is they, um, you know, request way more than the application needs to ensure that application's going to run, right? Because that's their primary goal, get this application live.
And it's critical for organizations to have some sort of software or technology or process to come in over the top and then say, well, how efficiently were those requests set? Um, how much are we wasting on, you know, over-provisioned requests? And we see, you know, we, we've helped customers save, you know, millions of dollars and, and a matter of like a month or two just by, you know, drilling into that kind of, uh, dimension of waste between request and usage.
And of course, you don't wanna make that zero. You wanna have some buffer, but, you know, just fine tuning and optimizing that you can, you know, we, we've seen has been a gold mine for a lot of organizations. Mm-hmm.
And we just brought forward some of our bad habits. We used to overprovision VMs all the time because we were afraid that the app would crash. So we just assumed we would get, you know, the maximum amount of memory and CPU and storage or, you know, proverbial to turn it up to 11.
Right? But, um, in the, in the world of Kubernetes, can things be a little more finessed? They definitely can be.
Um, they definitely can be. Um, I think, you know, it, it definitely takes the right, you know, again, people process and technology in place, right? It takes having, you know, the visibility to understand, you know, um, or maybe as automation to automatically resize those requests and the people to, you know, put that in place.
Um, but, you know, I, I, I think, you know, e even to some extent, this problem still existed in the VM world and still exists in Kubernetes. And, and there's an extent of, like, it might even be, you know, more prevalent in Kubernetes because you have this layer of abstraction, right? You're not thinking like, oh, I'm putting it on a VM with these specs, right?
You're just like, ah, you know, I'm just gonna take a shot in the dark and request this much random CPU and, and, you know, pray that my application runs and I'm just gonna bump it up, beef it up if it doesn't, right? So, um, to some extent this probably didn't be more extreme in Kubernetes because, you know, you're, you're, you're, you have this layer of abstraction, you know, for your end users or your developers, they're not really thinking about, oh, is it gonna run on an M two XL note or what kind of note it's gonna run on? They're just saying, ah, you know, it might need this much RAM and CPU.
I'm gonna go, you know, make sure I, I request a bunch of resources that my application runs. Mm-hmm. Um, we've also seen folks create these finops centers of excellence, but that seems to be on the far right hand side of this equation.
Does that really work? 'cause it seems like they're pretty far removed from where the actual decisions are being made. Uh, yeah, great question.
And I think, you know, effect, in our view, effective finops takes a, a culture shift and a whole organization's buy-in. It's not enough to your point, to just set up a finops practice, you know, that gets you, you know, maybe 50% of the way there, but, you know, it doesn't, you know, you, but then you're just constantly in a reactive state. Um, I think where we've seen, you know, teams really push is, you know, can we get things like cost visibility in the pr right, that I'm about to submit, so you know that there's some cool new functionality there.
Um, you know, can I get, you know, integrations for cost visibility and cost optimization even in my CICD pipelines, right? So I think that's where a lot of organizations, they're starting with that finops kind of foundation and then starting to push it further and further left with, for like developer tools and, um, you know, integrations to meet the developers where they are. Because that's ultimately, you don't wanna, you don't want your developers, you know, having to start a new process or use a new tool.
I think the, the most effective kind of shift left, um, you know, techniques we've seen have been meeting the developers where they are and, and providing integrations into things like, you know, terraform, GitHub actions and things of that nature. We Have seen also the rise of AI workloads, especially on Kubernetes. And there seems to be a lot more sensitivity about the cost of ai, so we'll that ultimately pull through where maybe we're getting more insights into Kubernetes consumption.
'cause more people are trying to figure out what the cost of AI really is. Uh, that, that is absolutely something we've seen in the last, call it six months to a year, which is, you know, and, and you know, the broader trend of, of maybe, you know, Kubernetes and finops has been like, yeah, it used to be just developers, you know, do whatever you gotta do to get the technology up and running. And now there's a real focus on cost.
And I think you saw a similar trend with GPUs, right? Where, you know, maybe a year or so ago, there was just a big push to, you know, get all the GPUs we can, you know, you saw all the news articles about CEOs buying, you know, NVIDIA chips, um, and now there's a real thought to, well, what are these costing me? Right?
And, and am I using them efficiently? Right? So that's again, some, an area we've invested a lot of our time and, and technology into, which is, you know, giving very granular GPU usage insights and then optimization on top of that.
You know, I don't know if everyone knows this, but you know, if you, you know, go to set up a container request, you know, you can request very fractional amounts of CPU and ram, right? You can request a milli core of CPU, right? But in GPU world, you know, typically, you know, you're, you're really only to request able to request zero or one, right?
It's like, how many GPU chips do you need? And it has to be typically an integer. And that is until you start to get into some more of the advanced sharing techniques, which are now, um, you're just now starting to see the industry coalesce around things like MEG and time splicing, different sharing techniques for these GPUs so that, you know, multiple workloads can use them if you have a workload and it's not using the whole chip, right?
2 or 30% or something like that. Mm-hmm. So what's your best advice for folks to kinda address all this?
'cause I think the most powerful issue anybody encounters is just inertia. We've done it a certain way this whole time, and maybe there's another way to think about it, but, um, getting everybody to move off the dime pun Yeah. Is hard.
Absolutely. I that's absolutely true. You know, I think that, uh, you know, the best way to get started really is to start with the development environment, right?
And get, get, like, get set up on one or two clusters, get some cost visibility there. And then as you kind of, you know, I, I think it really kind of follows the finops foundation's kind of framework, which is inform, optimize, operate, right? And so you gotta just start with some, you know, inform and understanding of where my spend is going, right?
I can't tell you how many teams come to me and say, you know, my Kubernetes spend is a black box, right? Or maybe I get some visibility, but, you know, it's very limited and it's very hard to get, you know, unit costs or an understanding of what my various business units are costing me. So I think starting with, you know, can I get some visibility into my business unit spend, um, is is a great place to start.
And, you know, there's actually a lot of, um, now free and open source tools that can provide that, right? You don't necessarily have to go to an enterprise tool. You can start small, um, and then crawl, walk, run, and, and, you know, work with more advanced tools as your needs grow and evolve To that point about visibility.
Um, we'll be able to leverage AI someday to maybe get more visibility into what's happening in those environments. So that, I don't know, maybe I get some sort of, you know, danger will Robinson, you're about to break the bank alert. That's absolutely true.
Yeah. I mean, uh, there, there's a lot of different ways I think that AI and finops, uh, intersect, right? I think, you know, if you look at kind of finops tools today, what you see is a lot of dashboards, right?
It's, it's this visualization and that visualization. I, I really think the future of finops is going to be, um, you know, interacting with probably some combination of dashboard and interacting with a kind of chat bot, like chat chief pt or something like that to help, you know, you better operate that your finops tooling, right? So, you know, a good example might be, can I, can I use a chat bot to, you know, set up an alert, right?
Versus having to go in and, and click and turn the knobs that I want the alert set up for. Can I just like tell a, you know, AI agent to set it up for me? I definitely think that, you know, that's where we're going and we, we have some stuff actively in the works there.
Um, you know, I also think that AI can be powerful things for, you know, hey, look, we've detected maybe anomalous spend, right? Or, Hey, you know, this is what your forecasted to do and, uh, you've made some change and it's like dramatically, um, you know, increased your forecast, right? So I think that there's a lot of, um, space to grow for AI in, in the finops industry, and I think, um, it'll just make it easier to use.
And it, I think overall it will help, uh, finops meet organizations where they are, it will help finops be more integrated into their workflows. Is there some smart way to describe the value proposition here, other than the fact that, you know, I don't wanna sound like my father yelling at me to turn off the lights in my room, right? But the, it does matter, right?
'cause money saved on hardware goes into salaries, it goes into other development projects, but it's not clear to me that everybody always connects those dots. I to totally agree. I, I think, you know, a little bit, that's why we try and surface the carbon impact as well, because we do see developers, you know, wanting to take, uh, you know, action based on that.
But, you know, I think that, um, you know, again, it, it, it does take a cultural shift. You know, you have to get the developers to buy into, um, you know, believing that the cost impact is important. And, you know, sometimes we do see, you know, maybe this is a bit heavy handed, but sometimes putting a budget there helps, right?
So you can say developers then have the perspective of like, look, if I, you know, cut the resources required for this application, right? I can use those same resources or use those same dollars to put it towards, you know, some new innovation, right? Some new technology.
So I think, you know, sometimes that's where you can, you know, really, you know, have those developers understand the impacts that that spend is having on the broader organization, Right? I know it's hard to believe folks, but IT infrastructure is not free. So be smart about how you use it.
Hey, Kai, thanks for being on the show. Thank you so much. It was great being here.
All right. And back to you guys in the studio.