gMaestro – Asaf Ezra, Granulate
Granulate, recently acquired by Intel, is launching its new and free Kubernetes cost optimization solution: gMaestro. gMaestro joins Granulate’s suite of optimization solutions and provides DevOps, SREs, and FinOps teams full visibility into their K8s clusters allowing them to eliminate over-provisioning and reduce costs by up to 60%.
Transcript
This is texturing TV. Hey everyone, welcome to a text drunk TV interview my interview or my guest today is a soft Ezra. He's CEO of granulated esophage is been on our show and in person with me before so it's a pleasure to have them back on.
Hello. It's off. Well, it's afternoon for you.
Good afternoon. How's it going? Good afternoon, Alan or good morning.
It's always good. It's nice to have you on a soft for folks who may not be familiar with you. Or granulate.
Why don't you give him a little quick background on yourself in the company? So they um, so my co-founder and I have founded run late about four years ago. The companies main goal was to create a what we call the self-managing workloads.
We started from the optimizations within the VM itself. So anywhere from the colonel stack all the way up to the runtime. And from then on out we moved into two different spaces one was through a profiler we aim to give customers the ability to use a free open source, continuously monitoring profiler to optimize their own code something that granulates without the prior knowledge of what the business application is would never do and on the other hand how to utilize performance optimizations or and in this case a lot of it comes from the just innate overhead of over provisioning and so you have a full stack from your own code all the way to the infrastructure level starting with kubernetes and expanding to VMS and data centers as well as Cloud installations.
Actually, very very good. Actually before we even go further, let's tell people right now if they want to get information on granulate the website. We're running late.
Don't IO an Intel company now for the past four months. That's going to congratulations on that as well. All right.
Now that we've kind of laid that Foundation let's talk about today's news is off. I'm gonna let you share it with our audience. I'm sorry and go ahead have at it.
Two things. So when we started talking to more and more customers running kubernetes, I'm sure you're familiar with the fact that it has become D-Day Factor standard for orchestration in the cloud and we've seen in more and more customers the case of it's nice to see an optimization, but I have a name over provisioning where I'm requesting from kubernetes and x amount of CPUs or an x amount of memory for my container where I'm actually utilizing about 30 40 percent and targeting about 60% for the auto scalar. So for a lot of customers is that innate over provisioning and the inability to dynamically change resources resource requests and limits.
That alone ended up sometimes with over-provisioning of about 40 to 60% We've seen actually even higher cases. So right now for those customers that are inherently over provisioning we want to give them the free version which is take the insights across all of your clusters. Tell us.
What the buffer that you want to keep we start by the default at 30% because you know what sometimes your birthday sometimes things can change. So the default would be 30% and the age of the workload that you want to look at. So this could be a whole week.
If you're saying you know what I have this tremendous change over a whole week or you know what you're going to say over the course of five minutes. This is the Traffic characteristics that I'm seeing not necessarily the max traffic that I'll see at night or on the weekend. But the traffic characteristics will remain the same so I can take that as a good representation of what would happen at Max workflow.
And then we'll give you the new recommendations. You can apply them directly to a running cluster. It will initiate a rollout a gradual rollout of the new request or you can take the yaml code put it back to your GitHub and you'll have your OneSource of Truth the next version which is coming in about two weeks would allow you to let granulate automatically change that through a right permissions right now.
All you need are read permissions and the next version you'll be able to add write permissions to allow the automatically apply from the UI itself and we're also going to add the ability to create a pull request directly to your GitHub based on the ammo. actually, so this is I mean this is really important stuff for folks who who know right for people who know about kubernetes deployment and provisioning. This is you know, a very very matter of fact.
Wow, what a great thing for folks who don't know though. Let's let's kind of explain it to them, you know, one of the One of the advantages of cloud of course is the burstability the elasticity. And the idea is you really should only pay for what you use.
Right, but if you need more it's available to you. well practice that that sounds great as an ideal, but in practice what often happens is. You need to provision for elasticity you need to provision, you know, if you know, you're going to be having a burst your you know, some sort of peak usage.
You need to provision to be able to handle that and you know in a perfect world, we would be constantly moving that provisioning based upon what we we are anticipating in usage. And as you said you want to leave 30% 20% whatever right overhead in case you exceed your your estimates, but what happens in real life is a lot of people said it and forget it. Right, they they set their provisioning and we have so many different instances running and clusters and everything else.
No one remembers to go back. And and so it's kind of like, you know, the marks on the tide at high water. You have your high water marks, but it no one that very few people ever go back down and move off that high Watermark, and it's a problem.
so that's good. That is that is absolutely a part of it. We see the cloud like what you say.
The cloud should be here for the speed for the velocity of the developers. We want the engineers to be as fast as they can. But we also added some complexities.
First of all, like you said the it's not a perfect world. You cannot provision a machine it'll be instantly ready. You sometimes have to load data.
You have to load your application. You have to load your container those things next take time. That's what you set the target at about 60.
Let's say 80 70% So you have to be ready when I happens and who is the one who said in all of these percentages you have your developers. Your engine is your devops your SRE team and all these functions expect to to have a different goal. Sometimes your SRE team will be the one who's managing the production, but the one who's provisioning the resource request for containers.
the engineer who wrote the application So the engineer has an interest that the container would just run. So we he would or she would sometimes over provision just in case you know, why not and yes already team wants to make sure that everything is fully up a hundred percent of the time. Let's say five nights.
And so they have a very similar interests that it's better to let's say pay a little extra but don't have that pager Duty call them a 3 AM because you have out of memory for example, or you don't have enough containers. And so and so all of a sudden you have your Cloud costs that are way way over what you imagined and not necessarily the right people are responsible for those requests and limits because if I'm a necessary team that is responsible only for the production. How do I know that the change that I'm making for the resource requests is actually legitimate and I'm going to create a page or Duty call to the engineering team all of a sudden or even to my own teammates.
So all of a sudden you have this Miss miscommunication or because the responsibility is not shared between both themes. So what you have is the this? Great, like you said great idea for a perfect world where you could just elastically grow and come back down.
But you also need to create those governing policies to allow you to make sure that you don't do it you're responsibly and to allow you to find those gaps and automatically remediate them when you can because it wouldn't hurt the SLA indeed. and this is basically what we're trying to give the organization and we believe it should be free because you're not going to change the resource requests and limits every day at most if you deploy a new version or there was a change in traffic characteristics, you might need to change it a little bit but We believe that this would not be every let's say day for every deployment. Sometimes you have a daily deployment.
Usually it's not for the same service. So the kubernetes so you'll have deployment said that is updated once today and different one tomorrow and maybe then you'll need new recommendations. So that's why we believe we should be free and that's why we believe that people should be able to take it back to their GitHub for their own single source of Truth when the engineers could also see the changes that have been made and it's not drift that you haven't production.
That's excellent. Great great great description. So it's off.
Let's just talk business with people say hey, this sounds great. I want to try it out. How do they go about to you know, let people know how did they get started here?
So you can either put Gmail in Google or go to granulated IO to Gmail. So you have a free sign up page. It just requests you to you know, like every other website and choose your username and then you'll be able to deploy as a single single pod deployment set in a in your cluster.
And then from there on out you choose the buffer you choose the age to default defaults are five minutes and 30% and you'll get those recommendations on a per cluster basis. We also do that according to the highest impact, obviously. So we've already have really good really good examples with customers some some of them are seeing in the in a single cluster sometimes six and seven digits and on an annual basis.
Really? It's a huge. Yeah hugees, right and you know, it sounds pretty easy.
I mean look this is you think there's a 30 minutes set up. See it's even less and the walkthrough is pretty much immediate because the first string that you see is setting those Headroom and age and after that you have all the recommendations in front of you. Excellent.
And now this is for kubernete clusters, wherever they run makes no difference AWS Google Microsoft up even even bare metal maybe in your own data center. Yeah. So we have some customers who are running their own managed kubernetes locally in the data center.
We have some Federated clusters across clouds the Let's say neat thing is it's all based on the kubernetes API so you don't have to rely on certain infrastructure. That is that's a real that's a beautiful thing, isn't it? Yeah, right.
The obstruction layer is amazing. Actually so soft we mentioned before it's galvanized that IO and this is Gavin. I said IO and just look on there for G Maestro.
We are already gearing up for coupons nativecon in Detroit in October. We'll be there live, you know texture on TV will granulate be there. I would hope so or not.
I would like to see a scene every cubecon that we can make live or not. But yeah, definitely. All right.
So if you want to see it maybe in person in October in Detroit as well. We could go so make sure this is live right now on the website people can go. Sign up for free right now.
io or Gmail on Google. Fantastic a soft sounds like there's a great offering. But thank you.
Thank you to the whole granulate team. This is the kind of stuff. I think people will really flock to because look the whole the whole finops thing right with people trying to control Cloud cars and everything else is is mushrooming people.
Yeah, everyone's recognizing and with so much as you said with so much of the new development new workloads going on cloud being kube-based. You know, it's almost kind of falls into that fin I'm saying where you can really I mean get a handle on what you got here and you can try to be smart about what you spent. Fantastic.
Congratulations. Congratulations on the Intel. Movie which I we've covered and hope to see you soon.
See you soon. Thank you very much Ellen. All right, sofas are of granulate here on Textron TV.
We're gonna take a break. We'll be right back.