FinOps in Software Engineering with Bill Lobig
Bill Lobig, vice president of Apptio and IT automation at IBM, talks about why FinOps best practices need to be embedded into software engineering workflows to rein in IT infrastructure costs in the cloud era.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Bill Loic, who's vice president of aptio and IBM Automation.
And we're talking about, well, the impact that finops is gonna have on modern application development. 'cause well, maybe we're trying to figure out how things actually get priced and how much they cost. Bill, welcome to show.
Thanks, Mike. Pleasure to have, uh, pleasure to be here. Thanks for having me.
I feel like there's a lot more sensitivity these days to, it costs, particularly in the cloud, and people are also trying to figure out if that's a, uh, getting their value, and b, if it's more expensive than on-premise. And, and we've always come full circle on this conversation. I, I think we had an era there where no one was too concerned about cost and developers did whatever they wanted.
And now we're coming back to, I don't know, something that kind of feels like old fashioned capacity planning. So where are we on this adventure? Yeah.
You know, it's funny. I think if you go back in time, uh, it it, it's true almost all organizations assume cloud would save them money. And I think what we're largely finding is it's delivered many benefits, elasticity, flexibility, faster, time to value, developers get up and running.
We circumvent lengthy CapEx budgeting cycles. We can get the resources we need more quickly. All of those things have accelerated IT and application development.
However, from the cost and savings perspective, I think most organizations we're expecting a hundred percent, uh, you know, ROI or so they're seeing something more like 10. And, and actually a lot of the clients we work with are actually spending more overall despite those benefits. A lot of the benefits comes down to elasticity theoretically, and the fact that developers are more productive and they're not waiting on a centralized IT team for some sort of ticket to be completed.
But, um, I think we lost track of what it actually does cost to run what type of software on what type of platform. 'cause not everything is the same and certain things are more spiky than others, and other applications are more long running, it might cost more in the cloud. Um, so how do we get our arms around these different types of software applications and what runs best?
Where I think the primary focus for this type of challenge in today's modern app dev world is finops. And I dunno if everyone out there has heard about what finops is, but put simply, it's a play on words like DevOps was to bring developers and operations people closer together. Now we still have developers and we still have operations, but they're speaking the same language.
The job function is more integrated, and now developers have a more astute awareness and sense of responsibility for operating their software. And spin ops just takes that to the next level. It makes developers and operations people more astute and aware of the financial aspect of their software.
And, uh, that, that's really the movement that I'm seeing that is really gonna drive, uh, not just a, a set of technology and tooling, but a new way of working and how we collaborate between finance and operations and developers to ensure we're getting that optimal ROI from our cloud investments. How actively involved are the finance people? I mean, I can understand that they might be a little peeved about not getting the return on investment that they were promised, but are they hands, hands-on now involved in making decisions about where software runs?
Or are they more like setting goalposts? I think it's a matter of degrees. I think we're early in this movement, despite some of the popularity around it.
If you dial into the community, there's, there's quite a, I think there's over 22,000, uh, registered practitioners now in the, in the finops Foundation, which is the primary industry organization out there. Encourage you to check that out. Uh, finance is still setting the goalposts, majoratively.
I mean, I can even speak from personal experience. I'm not an IT person. Uh, I'm not a developer.
I used to be, but I run a software product management business. And our, our finance payer people are very mindful of what we're spending and why. But the real focus is on how do the IT people build a sense of trust and collaboration with the engineering folks so that they can feel that the money that's being allocated to these projects is being used as responsibly and as efficiency as possible.
And that's where this idea of finops really comes into play because it, it brings visibility. I I like to say, you can't optimize or you can't manage what you don't understand. So it really begins with clarity in where is the iot spend going?
Who has what budget for what purpose and how are they doing relative to it? And once you start to make those linkages and people see how their projects and efforts are affecting the top and bottom line, uh, it makes them more accountable and responsible and it starts to bring awareness that will improve those optimizations over time. So are we providing developers and DevOps teams with the visibility they need and then they'll make the right decisions?
Or are we going to be a little more prescriptive and enforce certain policies and controls based on, uh, budget dollars available? And, um, will all that manifest itself? Where in a DevOps workflow, how am I gonna see all this stuff?
Yeah, there's a couple. You know, it's a matter of how, how you wanna approach it. Every organization is gonna have, uh, their perspective on it.
I, I think it begins with transparency and visibility. So lemme first start with the challenge. Maybe this is obvious, and I didn't address it up front, but let me address it here.
The difference between traditional IT data center CapEx on-prem and cloud from a financial and optimization and ROI perspective is predominantly that your organization, more likely than not, has a committed spend agreement with a cloud service provider. And within that agreement, they have over, I'll take Amazon for example, over 200 platform services you could use. And it grows by the day.
Azure, Google, same thing. And your developers are empowered. They have access to these accounts, they are empowered to use those services at their leisure, and most of them have no idea how much it's costing.
They might leave systems idle overnight. You talked about the nature of workload, batch, uh, uh, uh, uh, long running shortlived. And, and all of these things have an implication to the types of resources you provision, how long you provision them, and really coming down to the rate you pay and the amount you use.
So that's the, that's kind of the problem statement. And it's very decentralized. So where we need to begin is giving these engineers and operations people visibility into how their actions are affecting a cost profile as it pertains to the budgets that are, uh, allocated to them.
And so beyond that, once you get into that kind of awareness, yes, you mentioned policies, there's, um, in infrastructure as code or GI ops architectures, we've got OPA open policy, uh, technology, you know, HashiCorp has their sentinel offering. Not to get too technical, but these are different ways to enforce security governance, cost considerations. There's a number of different things that go into this, and you could, for example, here's something that we're working on at IBM.
Uh, your, uh, uh, agreement with, uh, your cloud provider says that you have a discount on mediums size instances, but not large. So you might think that you need to downsize from a large to a medium based on the utilization, but it actually will cost you more because of the nature of your discount program. And these things are very complicated when you get into reserved instances and spot instances and on demand and all these things.
And so this is where policies can help provide those guardrails and those controls. I would consider the policy to be sort of on the more mature side, but I think starting with transparency, visibility and awareness is, is where you want to begin. Mm-Hmm.
To your point about the complexity, it does seem like there's a lot of options and there's spot instances or reserve instances, and other times people just go with, you know, your standard instances. But, um, there, how do I make that less complex for folks? Because if I'm a developer, um, my tendency is I'm just going to get as much as I can and whatever I'm allowed to get, I'm gonna run because I'm more concerned about availability than I am about cost optimization.
And, um, you know, I'm trying to avoid that 2:00 AM phone call because something went down. Yeah, I, I love that point. You know, I, i, I like to joke sometimes when I meet with clients and others that, you know, if its prime objective was saving money, they would just turn the data center off and go home.
Right? But that's obviously ridiculous. It's prime objective is delivering applications and business value for their line of business constituents.
And like you said, you know, it's, it's okay to overspend a little and have that performance assurance, that capacity, uh, buffer. Uh, but if you underspend and the thing goes down during your open enrollment period, or Black Friday, whatever your, you know, your primary in your, in your industry is that that's obviously a terrible outcome. And so I like to think of it's job is to protect the business at the lowest possible cost.
So this is where finops is more than just a, a way of working. It is a, there are technologies, there are tools. So these tools can understand all of the intricacies, discounts, agreements of your ag, uh, your, your spend profile with these organizations and calculate and process all of that for you.
You know, a lot of organizations start off on this journey using Excel and spreadsheets and quickly realize that once you get past allocation, chargeback, showback, then you get into better budgeting, forecasting, planning, and you become less reactive and more proactive so that you can put the right implementation in place for the future. So these engineers, they don't need to figure this out on their own. There's plenty of finops tools out there on the market, uh, you know, at IBMI certainly have a few favorites, but leverage these tools gi let them give you the insights.
And then from there, so the first phase in inform I talked about you can manage what you can't understand, that's inform in the finops kind of, uh, uh, way of working. The next is around optimize. Okay, now how do I leverage technology and AI to understand what my high water mark is, what my low mortar mark is, and where I really need to place my capacity so I'm not over provisioning and wasting money on idle services.
And that, that is where you wanna be and what finops technologies combined with the way of working between IT and operations and finance can, can deliver IBM, of course is, you know, involved in this little thing called ai. So, um, can we use AI at some point to manage all this or help us manage all this and kind of make the whole thing a lot less complex than it is today? Yeah.
One of the prime use cases for AI that I see, well, there's two actually. One is, and it may seem obvious, but it's very valuable, is an assistant, uh, you know, instead of most finops practitioners and tools and software are oriented around dashboarding and information, so tons of different dimensions, slicing and dicing analytics. So how can you leverage conversational AI and generative AI to ask the question, what are my top spend contributors this month?
Where is my biggest variance to budget in the last 20 days? Like, instead of sifting through to find those insights, how can they come to you? So that's, that's one.
The second, uh, is around understanding workload demands and elasticity in a cloud native architecture. Now, this applies to traditional architectures as well. You know, they are extremely complicated and decentralized.
We have serverless, we have containers, we have VMs, we have networks, databases. All of these things are part of what I like to think of as a supply chain for your, your, your application. And we have multiple clouds, things transcending on-prem.
We have batch processes. In fact, there was a statistic I read recently that said, uh, uh, 72% of containers in Kubernetes live less than five minutes. These things are coming and going all the time.
Mm-hmm. So how do you understand the patterns and what I, you know, refer to sometimes as the seasonality? And that's where machine learning and AI can draw inferences and do pattern detection, which is basically what it does around, uh, compute and capacity to help you understand where you really need to pace those workloads.
One other dimension of this is ingesting service level objectives and, and, and, and SLAs in into that. So if you think about, this is an area where I find most organizations aren't as mature as they aspire to be. Businesses will set KPIs.
99. Okay, well, but underneath of that you've got databases and networks and Kubernetes and all these things. What are the service level objectives of all of those?
I know they all line up and add up and like casket. So machine learning can have a tremendous impact here because like most things that AI is applied to, it's just not possible for humans to comprehend and discern all that data that make intelligent conclusions about it. Machines can do a much better job there.
To what degree will I need to be able to ask some intelligent questions through say a prompt that will surface something that will be actionable versus, um, can I just kind of walk in in the morning and maybe something will just tell me, here are the three things that are likely to get you fired if you don't fix the spending pattern on them. Now, I, I think we all wanna get to the latter. I I think we're mostly at the former.
Um, I, I would love that. I would love to walk in every day and my, you know, my machine tells me what all my emails, uh, were and what I need to know. But, you know, uh, fortunately I don't think we're we're quite there yet.
And, and part of the reason for that is not for lack of technology. I mean, this is not a new concept, but it, you know, people process technology, right? That's kind of the famous three.
And I think there's a cultural, you know, uh, acceptance for a lot of this stuff. You know, I, I as an IT operations person that has been maintaining and operating systems and is responsible and ultimately the, the, the person, the named individual accountable for a certain, uh, uh, uh, function in that supply chain that I described, that that technology supply chain that supports the application. Like you, you need to entrust that the systems are making recommendations that you agree to.
So human in the loop, I think it's a crawl, walk, run. I think that's why it begins with, you know, trust, but verify, ask questions, what are the things I need to worry about? I go look at those things.
Yes, you're right, I need to worry about those. And then ultimately over time, we can let the, the machine take more accountability and responsibility. So I, I think it's more about cultural and, and, and acceptance and kind of comfort is the word I was looking for around, you know, offloading these things to technology.
All right, folks, I think you heard of your, it will get better. It's challenging right now and we'll be able to sort through this, but I don't think you need to rush out and get a finance degree just yet to add to your computer science degree. But, uh, hopefully the conversation will become less tense as we go forward.
Hey, bill, thanks for being on the show. Thank you so much. Thanks, Mike.
All right, and back to you guys in the studio.