Reducing Cloud Computing Costs with FinOps – Harish Grama, Kyndryl
Harish Grama, global cloud practice lead for Kyndryl, explains why FinOps is gaining traction as a methodology for reining in cloud computing costs.
Transcript
This is texturing TV. Hey guys. Thanks for the throw.
We're here with Harish grama. Who's Global Cloud practiced lead for Kindle and we're talking about bitops in the cloud and rainy and all those costs that are starting to spiral out of control hurries. Welcome the show.
Thank you. Mike. Get to be here.
What happened on the way to the cloud Forum? We were supposed to be all saving a ton of money and now it seems like suddenly we're getting Cloud bills and they're unexpected and they're higher than people anticipated. So, what are you hearing from customers these days?
Yeah, you know that's a great question Mike, you know, as you know, I've been on several different sides of this if you will, I was a cloud service provider and my prior life. I was a CIO Cloud transformation in a major bank now in the services business helping clients migrate to cloud and of course helping them with their corporate priorities, so to speak with the cost being one of them, but you know also Innovation, right? So what that does context Mike I will tell you that, you know, since I've been on the customer side on the technology provider side, and now on the services side, I think historically people have always approached this a little differently, you know, the the people being counter so to speak and the people further up, you know in the business side.
I've always looked at this as A a cost saving opportunity right? And then if you look at the other side of it the developers and you know, the technology people in the lines of business have really looked at it as a way to get rapid Innovation a way to start to produce new applications rapidly without waiting for you know machines to show up and then being loaded and the right the middleware being loaded and so on and so forth, but it took months up to six months before they could write a single line of code. So they were looking at it from a speed and Innovation perspective the bean counter so to speak and the senior people and Enterprises were looking at it as a cost savings.
And so that's really the Dilemma that you see here, right and even back in 2016 when I joined the bank. And we started to go down this route. We learned very quickly that you have to be careful on what you move to the cloud and more importantly how you move it.
Right lift and shift for the most part is going to send your bills the wrong way. Um, you know, it might work in in cases like just moving VMS over or what have you but if you're gonna take honking big monolithic software that was written for on-prem and you try to run that on the cloud. It's going to send your bills the wrong way, right?
So I think it really is a combination of doing it. Right. So in some cases you're saving some money but more importantly you're getting rapid Innovation rapid time to transforming your business and so on as opposed to just a pure cost saving statement and we're still you know your costs getting worse, right?
So I'd say that really is the balance over here. We've had capacity planning since time began. So what makes spin Ops different.
Is it a really a fundamentally different approach or is it just capacity planning for the cloud Yeah, so I mean it's it's a whole lot more than capacity planning. Right? I mean you could look at finopsis something that just simply watches your spin.
Uh, but I think it's a lot more than that. Now it certainly will do that and it will flag it and if you've got the right technology platform like we do as a company it will even give you you know, predictive a Trends on you know, something is going the wrong way that's going to land up with you getting a bill that's larger than you expect. We do all of those but I really think of Phenom, you know getting a discipline into how you operate in a company because when you start your modernization your it modernization and application modernization What you you want to do is you want to take a very measure approach to it, right?
Because you've started out on Prem you perhaps are in a virtual environment in the past few years. Maybe even have a private Cloud. Now, you want to move to a public Cloud you may want to go to another public Cloud as well because you don't want to vendor lock-in or your regulator's may say that you have to be on multiple clouds.
And if you're a buyer of SAS applications, which most large Enterprises are that could be on yet another Cloud because you know the application runs wherever you're SAS providers provide them, right? So I think getting the discipline of really figuring out what your priorities are and not just quote unquote moving to the cloud. But you know what you're trying to do.
What are you leveraging a particular Cloud for and then right-sizing your use of that cloud phenob starts to give you a view of all of that, right? So think of it more as a disc of how to consume these various Technologies for the best outcome for your business. Will we be able to predict the types of costs we might incur by application types.
So you mentioned lift and shift earlier. But if I turn my app into a set of microservices will I get a better handle of what my cost will be as a dynamically scale things up and down using kubernetes or whatever. It might be but it seems like different applications are gonna have different kinds of usage patterns and cost patterns.
Yeah, and that's a great observation completely agree with you, right? So again back to my time when I was a client and now being a Services guy helping clients move, you know clients typically look at this. The more sophisticated ones the larger ones they look at it as Look, I've got this application estate that's thousands of AppSec.
Right? So the first thing I got to do is take a look at them and say what do I want to do with that application estate? Some of them you're not going to invest in anymore.
And so therefore, you know, you don't want to go and do a transformation on it. But some of them perhaps are better off running in VMS. Some of them are better running on VMS on Prem or maybe you want to run on VMS on a public cloud or maybe you want to go further up the stack and start containerizing them and therefore in a sense modernizing them or you may say hey look, you know any augmentation that I do on this particular app.
I wanted to be completely Cloud native or this new class of applications, even though I'm writing it brand new up. I'm gonna write it Cloud native and not go one of the other routes right? So Figuring out the landscape of your application set and figuring out the rule around.
You know, what you're gonna actually do with each one of them. Starts to give you a handle on a you know, if I'm going to do this kind of transformation. It's roughly going to cost me this much.
And if I'm going to do this kind of transformation, it's going to cost me, you know something else right? So that's one level of sophistication the next level of sophistication and you'll see a lot of the larger companies especially, you know Banks and stuff where I grew up or at least where I spend time. Is they take it one level further and say look if an application looks like this so, you know think about what is the most prevalent Apple there?
It's essentially a web Java transaction app, right? I mean arguably 50 to 60% of the application out there end up being that so you now know you need an app server. You need a web server.
You need some kind of data store and you're gonna write your business logic in some microservices fashion that runs on the app server, right? So that's the kind of application pattern that you come up with and you say now not only do I know that I'm going to write this in a certain way. I now know exactly what I'm going to consume on the cloud because it's gonna be this app server this web server and it's going to be this database so I can even start to model what my cost will be right.
So getting it to what we call a cloud application design pattern gives you the next level of Receipt and you could do that with big data as well. I mean, you know people doing big data AppSec are typically tend to can Shard their data, right? That's a horizontal slicing of their data and they have an application set that they containerize that they put on these charge and then they have an Uber app that kind of brings all these the results of these container sets operating on the shards back together on a bigger business picture.
So that is also an application design pattern. So that starts to give you you know, what does it cost to write an app that looks like that and the same thing can be said of batching or it can be set off, you know, just going to the cloud when you run out of compute, you know, if you're doing things like Risk modeling Etc running a Monte Carlo simulations, that's a design pattern in itself. So that's really the next level of sophistication that starts to give you all handle on not only what you're going to move.
But how are you going to move? And then what are you gonna start consuming in the cloud? So each one of them is not a one-off but having said all of that right, I mean they'll be variations in all of this and sometimes if you've got developers you get a little you know, over enthusiastic or perhaps, you know, they do it for a very good reason you want to make sure that you're proactively looking at what you're using where you're using it and what the cost profile today is and as importantly where it's heading.
And that's an important point right because applications change over time and that which was less expensive on one platform this month next year might be entirely different as the platforms themselves change. So do we need to make sure that the workloads are portable enough to take advantage of that because it seems like a lot of times workloads wind up getting wrapped up and proprietary apis and even if you wanted to move them you couldn't yeah, and that's a great point. Right?
So the the vendor lock and aspect is always front and center for people and when you write an application design pattern as well, if you're doing a job right most times nine out of 10 times or maybe you know, 95 out of a hundred times you want to ensure that that basic Cloud application design pattern will run on multiple clouds, right? So if you're starting to go up into the past layer and you're starting to use, you know, more important the secret sauce of a particular cloud service provider even there you can get away with it to a An extent because you know, if you're using a nosql database or you're using a SQL database you tend to have them everywhere but if you're going to go out and use like a specialized AI engine that is prevalent only on one cloud and they start to raise the costs sometime in the future. Then you can't necessarily to your point move that one other Cloud because the equivalent function isn't available there.
So your observation of you know, what you use in a cloud is extremely extremely important. The other thing also is how do you behave when you write application on up on a particular Cloud right? I mean you could be well-factor Cloud native microservice microservices.
You could be all of those things. But when you write an application on Prem for example, right and if you're a bank, you're extremely risk-averse needfully, so you're dealing with trillions of people trillions of dollars of That's not yours. And so your risk-averse and your your logging every single thing that you do right.
Now, even if you followed all the right Cloud development principles and you start logging every single thing that you're doing in a cloud, you're most likely funneling that back to some on-prem, you know monitoring system or a sock or what have you because you've got this Cloud you've got on-prem you've got another cloud and you can be looking at a picture and multiple different places. So people tend to have their socks on Prem and they're not on Prem right and if you start to funnel all that logging back, you're gonna pay that data Ingress charges, the clouds are notorious for and then you're gonna find out that hey, even though I architected my app perfectly. I coated it perfectly and I use the cloud perfectly.
Well one of my behaviors here a funneling everything back whether I need it or not. Was not the best idea because now I'm paying like a million dollars a year for logging for one application. Right?
So it's that as well that you would want to really take a look at. And by the way, phenobs will help you with all of that, right because as you start logging more and more it'll take a look at it and say are you sure you really want to be doing this? Because you're headed towards a path of you know, incurring or recurring bill of a million dollars for this particular application in any given given year and it'll tell you that well in advance right because you've got the right triggers and so on.
Is it your sense? The customers are looking to save actual money's from what they've spent year over year or is it more that the finance team just wants some sort of predictable sense of what the costs are going to be so they can plan accordingly because it seems like a lot of times with the developers. You're not quite sure what's going to happen from one month to the next.
Yeah, and that is certainly true. Right which is why you want to forecast these cost things now, you know having worked in multiple different Industries and and companies. I will tell you there is no Finance person that says I'm happy with predictability.
I don't care what the actual absolute cost is. Nobody says that they careful care about both of those things. They want you to keep to a cost envelope and they want predictability and how you reach that cost envelope, right, you know if you spend everything in one quarter and you know the next quarter you say hey, look I got to do something different here.
They're not going to be happy with you. So I think both of them are really important but you know, that's the discipline that I'm talking about right you get the rules right you get the disposition of your applications, right? You get the methodology and how you're gonna move these things right?
You have this phenops platform and our processes and our people, you know plugged in in place to kind of give you that ongoing picture and allow you to build that discipline into your day-to-day development and deployment activities, then that's when you land up in a position where you're not surprised at the end and you're within a certain cost on and I think they're both important. Do you think also as we go along? At least in devops processes.
We have lots of metrics but not many of them are related to costs and you think we'll see cost metrics kind of start to be embedded within all the other metrics that development teams use to figure out whether or not something is performing well, Yeah, you can start to do that. Right? Because when you're in your devops model and your SecOps model, you've got a CI/CD pipeline, you know, what starting to flow through from a cloud from a bits and bytes perspective, right?
And you got your tool chain locked in and everything else. Now there's two things right one is how many resources are you using? And where are you using it and phenops will certainly give you a very good view of that.
But even in your deaf psychops model or your compliance secops model, you know, you want to know what code you're deploying out there because if that code runs away from you or you know, it's got a vulnerability where it's just multiplying itself and you know, using more and more resources and so on and so forth. Not only do you want to catch it from a cost angle. You want to catch it from an angle of hey, I don't want this to happen again, right?
And so you you can trace it back to your ci/cd pipeline your development time operate or your development time. You part of the cycle and you can start to see where am I going wrong? Maybe I didn't have the right test coverage on the code that I'm deploying.
Maybe I didn't check it against the right vulnerability databases that are out there to know that hey, you know, it had this security issue and this security issue. We can tell proliferated costs and made it unsafe. So those are the kinds of things that you can start to pick up.
Our SecOps cycle that will feed into the phenob cycle, but will also give you the benefit of fixing how you actually write and deploy code and the same thing is to be said with compliance secops as well. Right if you're now deployed on a particular infrastructure stack and you've got a set of compliance met parameters that you've set you want to know when that starts to drift away because drift can add to cost it can add to vulnerabilities it can add to security issues. Right so it can add up to all of those things.
So you have to take these things as a whole because they're all actually interconnected. They're not discreet things. You think are a sustainability goals will be bolded into finops as well because we're going to be tracking carbon emissions and all kinds of fun stuff and it seems like it's a similar process.
Absolutely is right. I mean it starts off by observability. How much are you actually using resources?
Because resources equals your carbon footprint at some level right? So you certainly want to track that and again that starts to become part of your ongoing operations and your ongoing optimizations. Because when you talk to Enterprises and they all have a carbon footprint goal, you know, do do this business transformation for me and this particular project and then come back and do a sustainability project transformation for me in this project as a different project said no one ever, right?
They're all like look. I've only got so much money and I need to transform my business and oh, by the way, when you're doing the transformation project, please make sure that I'm meeting my sustainability goals as well. Right?
So it has to become a part of that and finops helps you in that as well. Right? Because the more resources you're using on a cloud the worst your carbon footprint is going to be at some level right?
So it's it's all connected Mike. So that's a great observation as well. Last question will AI save us from ourselves when we just have a bunch of algorithms bigger this out for us or how much cognitive load do I have to have to make all this work?
Well, I mean you've seen what's happening out in the press out there and several notable figures saying they should you know, we should pause as an industry on AI until there's better governance Etc. Now look I mean those guys are all multiple times as smart as I am at least and so, you know, they say there's there's got to be something in there, but I will tell you from my experience. He not everyone's gonna get this right right.
You may have the most powerful technology, but if you don't wield it correctly it's gonna have disaster results. We've seen that in Cloud as well right in the same thing is gonna happen with AI whether it's Chad GPT chat gp4 or something else and and that's what makes the people that do this on a day-to-day basis, you know have pause when it's suddenly proliferates the way it is. I mean even the last six months The state of the art of AI has progressed like tremendously and it's only going to accelerate right?
So is it going to be used and wielded wrong? Absolutely is is that going to cause you problems absolutely is so I think the the call for better governance and education and Skilling people is is a good one. All right.
Great. Well folks you heard in here finopsis here the question of how you wield it to get the most benefits out of it. But if you're not engaged you're probably paying more than you should I reach thanks for being on the show.
Thank you. Mike pleasure to speak with you as always and look forward to our next one. All right guys back to you in the studio.