Huamin Chen, Red Hat | OSS North America 2023
Mike Vizard spoke to Huamin Chen, R&D for the Office of the CTO for Red Hat, about the need to study the tradeoff between the performance and energy cost of running huge foundation models on Container Platforms, specifically focusing on Kubernetes.
Transcript
This is Techstrong tv. Welcome back to the Open Source Summit. We're in beautiful Vancouver as always.
Remember Woman Chin who is heading up a project Kepler effort at Red Hat that's talking about energy consumption and Kubernetes clusters and the infrastructure and everything that goes with that. Warm and welcome to show. Nice to see you Mike.
Nice. You. What is the issue that we're trying to get at here?
Because in theory I have a Kubernetes cluster and things are supposed to dynamically scale up, right? Scale down. Exactly.
So, so why am I having an energy consumption issue? What do I need to measure? Yes, That's magic.
When scale up something, you actually get more power. Power in terms of performance become powerful. Power also means the energy.
Using more energy is speed of the power. So what we want to share with the community, with the end user, with the develop that uh, how much energy do you use by scale up and down and whether we can help you to make a better decision to scale up your cluster without using too much power. So that's the end goal.
How we can make that happen. Because use the same equip nett is principal, observe, decide and display the metrics so you can make a decision. Is this gonna become another metric that I'm tracking in my DevOps environment?
Is that how that will play out? Yes, That's yes. Or another metrics you're gonna see, well many of the cook nett is using the same ous metrics and that's the of the infrastructure we're using as well.
So we created a project called Cap Ka using a lots of these to under the food, collect lots of information and fast information to manage consumption by the workload. Whether that is a container apart any other entities. And then this metrics is exported.
The permits will sit and end users can use it and decide what they want to do with the metric. Whether they can be using as thematic schedule or scale. And that's where we very used to see what kind of innovation they can use to conserve energy.
And is that also gonna get connected to something that's counting carbon as well or how does that come together? Yeah, Spot. So carbon is our, in our spot rights, we wanna see how we can correlate the energy with carbon.
So not energy is created by the same uh, source. So somes are created by renewable energy, which has a lower carbon food cream versus some created by fossil fuel, which by carbon food cream. So if you are seen amounts of energy used by your workloads and you schedule this workload as a carbon friendly area, you are gonna use the less carbon.
So yes, your question, we want to correlate with energy as well as carbon to give the end customers. End users the best option to choose. Rather they can run [unknown] Historically in it we've been, shall we say a little sloppy, we tend to overprovision things because you know we're afraid something might crash.
Exactly. Um, is that become a crutch in a bad habit? Yeah.
So it's time to make that happen. It's not corrupt. No.
Oftentimes the microservices can help you to reduce I power. So meaning that's if you are not using your server, not use your infrastructure that much, they can be scaled down the same matter. We can use it for power.
Even if you are using your microservice and during only time of the day, we can tune microservice using less power by using the lower frequency or GP frequencies. So you can still write a workloads, get things done, but using less power, how that can be reported. This can be reported by capital.
How can that be done? That can be done by overstates. Are there other projects that are gonna address other aspects of this?
Because a lot of times the code that somebody wrote is also a little on the sloppy side and consuming too much energy. Can we me, can we monitor and measure that? Yes, That's exactly we can have.
So when the developers use the SEE UST pipelines to release their code, they can plug capital aggression inside of the city pipeline so they can see what kind of power they're using. It does any regression or is that like green lights of red lights kind of scenario? That's too much regression happens in a power consumption.
Maybe they can roll back the changes and make better decisions. So that's where help for the developers. This sounds like a nice to have, but is it gonna become a requirement?
We're seeing a lot of legislation floating around. So is this gonna move past the, you know, this is cool and interesting to the you can't live without it. Yeah, So the, as you said exactly in the perspective, this could come from OI or habits if use side of the story to the production system, that's a supply demand as vaccines.
So that's as long as as right now we are seeing more and more demand from legislative landscape requiring the company to report their carbon footprints from it, from their supply chain, from their own operation. This requirements will be translated to supply of technology software stack improvements, measurements stack like capital. This technology will for foresee that's will happen in the next few years to help the customers build up the better technology stacks.
Aren't we gonna also look at this in terms of our deployment models? Because a lot of folks, um, they don't like to share, right? They have one application running on a cluster and you know that's probably inefficient.
So can we figure out which apps might lend themselves to running on multi, on a single cluster in a shared model rather than giving every application its own cluster? Right. I think that is very interesting use case as well.
I see that is has a lot of impact. Many of the clusters as we see are single purpose with classic you just one cluster for single round. So how efficient is you want run in that mold?
You want know run a shared mold, the concern can be validated by using a loss of metrics. The metrics we are providing for the coupler can help you figure out from the energy perspective, once if you share the cluster to multiple applications, would that be saving you more energy? Then you run the application cy own dedicated environment.
So this kind of things can be faked out as long as we are supplying these perspectives and metrics for them to analyze. This is also good for the decision making. If they have concerns on different suspects, cost, energy and security, we provide them the fact so they can inspect each one of them make a bad decision.
How automated will the decision get going forward? Because if I have the metrics, can I just plug them into some automation framework and it will execute whatever policy I decide to set. And you know, I don't have to think so hard about this.
Yeah, All policies will be driven by the metrics. But data, so as on, so for example, my policy is to run a green cluster. So how we define green, so we can say performance per was wise.
If you are able to, for example, a database, a database application. If I can run my 1000 query per one was that is called green. How can I guess that numbers capital can help to get there?
So these are all policies driven, policies driven deployment, how we are going to get there. It's about a lot of aggregation of ships. So capital can plug in into that picture.
But providing this information, You can't walk down the street these days without somebody leaping out and telling you about their great new AI thing. So Uhhuh, will AI get applied to this in some shape manner reform and how might that look? Yeah, Absolutely.
Spot on. So we just have this uh, open source talk about AI and energy conservation and we show them very, some very interesting discoveries. We show them that uh, your AI environment is running in idle power regardless whether you're serving anything or not.
But if you are using as idle power to do some active service, we actually can give you the response without losing a loss of power. So this is very mind blowing discovery. Similarly, we can help you to optimize your AI configuration.
So if you are able to, if you are okay with your performance objective, we can tune down the GPU frequency so you can use less power while you just satisfy your workloads very much. These are the new discoveries we can share with the community, share with the operators so they can be more greener in the operations. And as far as I know, the bulk of those AI workloads are running on what Kubernetes clusters, right?
I agree. So that's the beauty. So AI runs thousands of Kubernetes cluster thousand knows of Kubernetes clusters models and they use the simple architecture to solve the models.
So in that environment, Kubernetes auto scalings and hopefully capital to have a long way to optimize the something that environment A lot of workloads are running in the cloud. So will this be fed through some sort of report so I can analyze what's going on with my cloud service provider? That's my best hope.
I hope the cloud providers, they can open up their chest but help us to understand what is going on with the GPUs with their workloads so we can use more accuracy to try to my workloads and fine my infrastructure configurations and I can save money, energy for the best part of the world. Well if they don't share that, well we just move our workloads to somebody who does, right? Yes.
That's if you have compilers who are ready to do similar things, many of data centers actually can provide you some eristics, some APIs you can get some sense of how much energy you use. Mm-hmm. So by end user, I will choose the one that's provide us this information.
We can fine tune api, fine tune the workers, fine tune the infrastructure to save more energy. We've been talking about this in the context of data centers, but Kubernetes is moving out to the network edge and we're seeing more deployments out there and those environments are compute sensitive. So is this becoming become something that they're gonna want to use out of the edge as well?
Yeah, I agree. So, uh, for just for one reference, uh, we are actually working with people who are providing the 5G flex kind of services. So this is a really power hungry applications.
You see this a 5g, well, they're running busy all the time. Can we do something to measure the workloads, uh, energy consumption? So people have some ideas.
What is the best configuration they can provide? Can we supply them some of the eristics so they can fine tune the application or configuration? So that's all the big questions we want to address.
Plus, without master sitting behind the policy makers can hardly make any decisions. So that's what I see is a community driven efforts. We should work together to have the world have the ICE operators to drive up to new innovations.
We talked about regulations and that's the stick, but there are carrots in the world. So will organizations, um, give it people bonuses for being more energy efficient? I mean, you know, efficient.
How, how would this play out? Yeah, That's huge. I mean these are, people do hard work, not, they're not only just saving the energy card, they actually help us in the long run.
If you look at how much energy is consumed by data centers, the carbon emission is equivalent to the airline industry. So that's huge. If you are aware of due increas, the carbon emission reduce the energy consumption.
That's not only have the customers modernize, it's also have environmental sustainability for humanity in the long run. So I should, I think the IT department in organizations should work towards that direction, that more metrics optimize their workload, you know, making decision making more intently. This all more automations to help drive down the energy consumption, reduce the carbon footprint.
What are we gonna do to get developers and IT people more conscious of energy and carbon and their role in the larger green movement? We should work together. So not everybody knows how things can be done and not every IT operators know how to cast the information from the infrastructure.
So the messaging, we are coming up to the community through the end users data. Uh, we provide you some mechanisms you can use very easily to use and you can cast very rich information also. It, and there is up to you to make how things could happen was objective we want to achieve.
So that is the collaborative way we want work with end users operator. All right. What do you see as the most common mistake that people are making there for?
I mean, what's that one thing you see them doing that you just shake your head and go, guys, we don't need to do that. Yeah, So that, yeah, lots of people using different, uh, framing of frameworks, uh, different version because the energy information, what I would say, make sure this methodologies sitting behind the tool chain makes sense. They have scientific foundations and scientific validations behind it.
Not everything work the same way. So that's why we chose to be the methodologies that's provided by scientific research community over the past decades. And that's the models we use in capital.
So capital is not something invented from this own, it's actually built on top of all this scientific research and we extend the research, the research to then use community so they can use this as a tool. Environ. All right folks, you heard it here.
We can save the planet using math. That's right. It's a very unusual and new idea.
Yes, wain, thanks for being on the show you. So My pleasure, Mike. All right, Guys, that's it for today.
We've been doing this all day long. We've been talking about all kinds of exciting new innovations in the open source community. We hope you'll come back and join us again tomorrow because we'll be here and we're looking forward to it.
Till then, see you later.





