Optimizing Cloud Computing Environments – Asaf Ezra, Intel Granulate
Intel Granulate co-founder Asaf Ezra explains how to continuously optimize cloud computing environments in a way that improves application performance while reducing IT costs.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with ASF Ezra, who's co-founder for Intel Granulate, and we're gonna be talking about how to continuously optimize cloud computing environments.
There's a lot of concern about costs, and turns out there's a way to do that in real time. Hey, Asif, welcome to the show. Thank you, Mike.
Uh, pleasure to be here. People, I think basically lost sight of cloud computing costs because, well, I mean, the economy was booming and it was kind of viewed as almost being nearly free, but today I think a lot more people are sensitive to the cost. They want 'em to at least be stable, and we see a lot more interest in things like finops and a lot of technological approaches to controlling costs.
But walk us through some of the other approaches that you guys have been working on and how does that work? And maybe put some context around what we can do in terms of actually taking more control of these environments. Absolutely.
So, um, I think everything stems from what you just mentioned. You know, in a, in an economy with zero interest rate and money is being cheap, everyone's talking about growth. That's the only metric that they care about.
So a lot of things go out of hand. Uh, one of these being, let's spend as much as we can to, to earn as much as we can. And at some point, it's gonna catch up.
And when interest rates started to bounce back, money became a lot more expensive. And right now we're seeing a lot of investors and a lot of boards just talking about profitability, talking about, uh, um, cogs and talking about margins. So, um, when, when you look at cost reduction, we see it as sort of a layered approach.
You can start from anything like the best practices and, uh, utilizing the tools that are cloud native. Uh, for example, if we're talking about cloud specifically, you have spawn instances, you have reserved instances. You have EDP or CUD depending on the, uh, cloud of your choice.
Um, but the, the important part is you have to leverage those, uh, um, sort of in a way that complement each other. And there, there are layered approaches. You start with the best practices.
You start by auto-scaling. You start by, uh, um, maybe even, you know, segregating to the right, uh, uh, instance types or depend on, depends on the, the cloud provider. You might be able to vertically scale your machines.
Uh, so a lot of what we see is customers that are bottlenecked or CPU or on memory, uh, but they are locked into SKUs that are available to them in the cloud. There are cloud providers that allow you to do vertically scale and, and sort of sidestep that issue. Next layer, uh, is, is pretty much what you said around finops.
This is where we see a lot of the customers right now in the journey. Uh, they're trying to, first of all, at least have a visibility into where their spend is from multiple clouds as well as on-prem. Um, a lot of the larger, uh, customers, enterprises are still at least partially on-prem, and they need sort of some sort of a unification unified view.
So we see a lot of providers around that. Very good ones. Uh, I think that, uh, when you talked about the, the position of a finops, uh, I think it started maybe a few years ago, as, as some sort of, uh, a nice to have on the DevOps side right now, it's a must have for any organization.
And we're, we're seeing finops return their investment tenfold, uh, each year. So I think, I think it's a, it's definitely a, a, a right approach. And then on top of that, you have everything that has to do with optimization because you did all the hygiene in terms of the financial operations.
You did the right discount negotiation, you chose the right SKUs, you utilizing the right tools for the right approach, right? I'm not gonna use, um, Lambdas for everything because lambdas are more costly. But if I have something that is reactive and it's, let's say even once in a while, I may not want it always on.
And I may want the flexibility of having it trigger at any point by something. So I may not want everything always running, so I have to use the right things. But once I'm using the right things and I have my application running, and most organizations are happy with their SLAs, this is where we have to fight for the optimization.
And the organizations, uh, have a decisions that, that they have to make at that point. Am I going to invest more into my business logic, more in the features that I'm going to develop for my customer, or I'm going to invest in margins? Most organizations, when they face that sort of trade off, would opt in most cases to develop more features to market, to bring more, uh, uh, things to the top line revenue for obvious reasons.
Uh, you can infinitely increase your revenue line. You can, you, you have a cap on the amount of costs you have, you can reduce. We want to help them finally sort of remove that trade-off with performance as a service, which is a pillar of Intel software.
Intel is offering not just Intel ate, but uh, a full, a full stack optimization solution for the application layer. Anywhere from the infrastructure itself, all the way to the runtime. Uh, where we last spoke, I think, uh, about granite specifically.
So this is where I expect organizations to come in and say, you know what? I want to maintain my SLA, but I also don't want to have to spend all my resources on increasing the margins because I still have to deliver to my customers. So we're trying to find a way to, um, have our cake and eat it too, as it were.
Um, to that end though, walk us through, uh, I don't think a lot of people are familiar with granulate. So what exactly are you guys doing and why does Intel kind of get involved in this whole equation? Uh, that's, that's a, that's a fair question.
So I think what you see, or what people historically expected from Intel was, Intel was a silicon company. It just delivers processors and maybe delivers, uh, um, some sort of improvements on, on the, uh, compilers to allow us as users to leverage those instruction sets and the processors to the max. But in reality, that's never gonna be the case, right?
Intel has about 20,000 software engineers. It's, but by definition one of the largest software companies in the world. And the, the idea of we will push everything to open source and the user will find it doesn't work.
We need to help customers achieve the maximum from our hardware as well as our software. Um, at the end of the day, if on the cloud you have five different generations of just Intel processors, not even talking about everything else, you are not going to compile against the latest and greatest every time. And when you create, let's say, a container image, that container image needs to be able to run on every machine that you allow the cloud provider to provision for you.
So if you are, if you want to leverage all the latest and greatest of Sapphire Rapids, but your container might also run on a machine that runs Ice Lake or Broad Lake, which is still available, you are by definition losing a lot of the performance benefits. And so we, and Intel, not only we know best, our Silicon, we also can help you leverage that because we have multiple layers in the stack that allow us to leverage the best out of every silicon. Um, and I'm, I'm sure you're familiar with, uh, one API as well, which also leverages that, Well, I may be familiar with it, but not everybody watching this is, so walk us through where you guys integrate with one API from Intel, and then, um, what exactly am I deploying per se, and how am I gaining control of these runtimes?
That's, that's a great question. So, uh, thank you. Um, so first of all, when Intel looks at the software, uh, um, software suite, uh, it's creating, it's divided into three groups that form sort of a unified solution.
Uh, later on. Uh, one is, uh, the trust authority. One is the performance as a service, and one is ai.
And when it comes to performance as a service, our mission is to not only give you the right tools for you yourself to debug and improve the performance of your own software, whether it's from, uh, our profilers, uh, to our PMU counters, to, uh, v tune to, um, and and so on. It's also to, to bring you, uh, uh, the performance, what's called as a service. And the way that it works is, is as follows, granulate, uh, built its agents to run in a production environment with almost no overhead.
1% CPO, and we're talking about 20 megs of ram, give or take. Um, and so on an ongoing basis, the overhead is almost unfelt and by loading Granulates runtime modules into the application runtime itself. So when we're talking about Java application, for example, that would be the JVM.
When we talk about Python, that would be the Python runtime. When we talk about, go by the way, we can also load our runtime module into go that loading mechanism was the basis for granulates on optimizations that help with the resource allocations on the machine itself. So the memory accesses the, uh, disc access, the CPU and thread allocations and so on.
But it also allows, uh, uh, for delivery of basically any optimization that Intel has, because Intel has created a huge amount of optimizations across the years. But pushing them just to mainline means that for customers, we use the mainline, it will take a probably a few years until they upgrade everything. Uh, I don't need to tell you probably how many people are still using Java eight, um, or Java 11.
Um, and the same goes for even even the Linux kernel, right? 6, but let's say four, right? Um, so we have to help our customers gain not only, uh, the performance when we can, but also the security.
So we wanna load those, uh, as well at runtime. And that ability means that we can also do this and differentiate between the different processors. So if historically Granulate only cared about the runtime, now that we're part of Intel, we can also leverage the Silicon enhancements and do everything in real time.
And so that, uh, uh, sort of trade off that you had to say, oh, you know, I might run on Ice Lake, so I can't leverage Sapphire Rapids accelerators, for example. Granulate alleviates that need. And so when we talk about performance as a service, drag lake doesn't become just a performance or the application performance optimization module that it used to be, it also the delivery mechanism for anything Intel can give you.
So it would seem to me like one of the benefits of being with Intel is you can make this technology embedded with other ISVs that are out there. So is that the next great thing? 'cause it's, can I make this capability pervasively accessible through those folks?
So that's one of the, uh, uh, um, one of the avenues that we're taking with Intel. Um, and I will also note that since joining Intel, uh, grand Lake has been tasked, tasked, sorry, with, uh, enhancing the capability out of just the application itself. So, uh, we encountered a lot of customers.
Were talking about less of a performance optimization, but more of an overhead and a complexity from running, you know, thousands of machines, uh, in the cloud with dozens or hundreds of services. So how do you even manage that? So ground light was actually tasked to just also take those performance optimizations and translate them into cost reduction by leveraging the orchestration capabilities.
So we started with Kubernetes, and we also do that with Spark and Yarn. And this is not just those, uh, uh, providers, but also, uh, um, this is an avenue that Intel is pushing rally towards with its own partners, right? Intel sells most of its, uh, processes through partners, and there's a lot of avenues to go about there, And ultimately will be the role of a DevOps team in all of this.
'cause it seems like if we're gonna give them more control programmatically, they have to work this into some sort of, uh, workflow or process that they manage. So how do you envision that all coming together? So I Think that when you look at the, the role of the DevOps, uh, and how it developed, it was originally a lot more about the infrastructure as code and started as, as we could take it even further, back when Kubernetes was super hard to deploy and super hard to maintain.
And over time, that became a lot more, a lot easier, I would say. And, and Kubernetes took care of the, uh, what we call the availability layer. So when you look at the, what you expect from your application, it's first of all to be available then to be reliable and then to be performant, I would say, and, and so on.
Um, I think that a lot of people originally expected, uh, um, some sort of no ops solution to, to emerge in terms of we'll have this incredible AI that would make all the decisions for us. But I think realistically, when you think about the future, you think about how how we translate the, the conversation from resource-based to business metrics. Because me, as the organization, I don't, I don't care if I'm running at 60% utilization or I'm running at 70% utilization, I care about what the customer experience is.
And right now all of our conversation revolves around the, the resource metrics. And so I think over time, the DevOps, uh, um, DevOps position would change into translating the business metrics into machine language that would allow it to make the, the decisions, because I don't expect somebody from the DevOps, even no DevOps team to manually maintain 15,000 applications, right? It makes no sense.
But there's also no chance that an AI would be the one to dictate what business metrics I want to satisfy. And that translation, I think, like you said, some sort of a workflow management and some, and, and also like policy enforcement, if behind the scene, uh, an ai, uh, uh, based solution would make better decisions across the entire, you know, hyper parameter space. Yeah, of course.
Right. It's almost infinite. So on that point, it seems like AI is gonna enable us to manage these environments at a higher level of distraction, maybe at the business metric level.
Um, does that mean that the DevOps team needs to evolve? I mean, are there skills in there that they don't have? Because it seems like today we're very focused on optimizing bells and whistles for tweaks in performance and whatever else there is, but is there a different way of thinking about this?
So we, we already see a lot of it changing today in terms of, uh, um, flow creation, like you said, API based, um, flows not just, you know, manually sending policies. And we think that realistically, uh, um, you know, choosing out of the insane amount of knobs and SKUs that you can, that will probably be done by, by something that is, uh, let's call it an AI that can go through all of those, but discussing how to translate my business metrics. So my actual revenue to my, let's say P 99 latency, that's where I expected DevOps to start to come in.
And I expect them to start saying, uh, um, you know, be more of a, um, involved in the AB testing of infrastructure and, uh, um, workload performance to revenue, and not just, uh, let's say be less reactive to what, uh, uh, engineers expect from their application, but be very proactive on how they can impact the business metrics, not just, uh, um, expect the, let's say availability requests or SLA requests, but push them proactively through the stack. Uh, back to the, back to the revenue line, actually. Do you think the shift towards building applications that are infused with AI is gonna force this issue?
Because we are already seeing that, um, processors like GPUs are scarce, but there's gonna be a lot of different classes of processors that I have out there, and I'm gonna try to mix and match those things. So at the end of the day, ai, lots of data costly. Does this kind of drive everybody down to something that's more efficient?
Yeah, it has to, uh, otherwise it's not sustainable, right? It's, it's, it's taking basically what we experienced in the past five years to the max, because right now money is super expensive, but you can't let the, the, uh, AI wave just go over you because you don't have the funds and resources to, to work on it. So you have to do it, but you have to do it sustainably.
And when you look at the, the era of, let's say, uh, data is the new oil, there was an era like that. I think we sort of shifted from data being the new oil to data driving decisions to data driving business, uh, um, decisions. And now it's going to be AI is, is extracting better, uh, um, let's say better anomalies, better decisions out of our own data.
So we cannot stand aside, but if we simply, let's say, lift and shift our data science from our own data scientists to ai, that is going to be costly to the point where it's going to, to hurt the business. And when you look at what happened with the cloud, when everybody lifted and shifted, that was about four to five times more expensive than what they experienced on-Prem. And I think it's gonna be the same here.
So we're, we're, we're having this, uh, interesting situation where they have to, everybody has to implement this somehow. It's new. So there's not a lot of, let's say, easy tools to implement everything, and they Have have to do it sustainably.
'cause we don't have a lot of, uh, uh, resources that we can just throw at the problem. All right, folks, you heard it here. We're gonna have to raise the bar of the DevOps game.
We're gonna have to start with the business metrics and then programmatically work in the technical metrics and align them accordingly. It's been a long time in common, but cross your fingers, it's about to happen. Hey, asf, thanks for being on the show.
Thank you so much, Mike. All right. And back to you guys in the studio.