Monitoring Energy Consumption in Clusters with Kepler at Cloud Native Now 2024
In this session, Sergio Méndez will talk about how to monitor energy consumption metrics with Kepler. In recent years governments are trying to achieve environmental sustainability to reduce the impact on the planet by reducing carbon emission by reducing the amount of energy that clusters use. Sergio will explain to us how to use Kepler to get metrics of energy consumption in clusters, and how Kepler works to get these metrics by using eBPF. Sergio will also explain how Kepler is fundamental in bare-metal Kubernetes implementations and the use of virtual machines, and how initiatives like TAG Environmental Sustainability from CNCF is looking to bring tools to achieve environmental sustainability for cloud native environments.
Transcript
Hi, Joel. Uh, this is my session monitoring, energy consumption in cluster with Kepler. This is cloud native now, now 2024.
Uh, about me, well, I am an operating citizen, professors, also a cloud native enthusiast. I have a cloud native group in Guatemala, uh, coming from the CNCF Cloud native groups, and also ambassador from CNCF and out a book about HCPD. So I want to get started, uh, with a phrase coming from the software foundation.
Um, uh, they are like two ways of looking clean the software. So the software can cause climate problem because of the carbon emission, but also the software can be part of climate solution to resolve the, or reduce the carbon emission for this. Um, because in the context of this, um, conferences, um, now Kepler has the mission to measure the carbon emission for cloud native applications running on a cluster.
You will need how a Kepler does this job. It's like recollecting data, uh, recollecting data from the processes running inside the containers, and these containers are running on the nose on, uh, Kubernetes clusters. So basically it's recollecting all these informations around the cluster, right?
And it uses, uh, E-B-P-F-A pretty popular, uh, technology right now and EBPF because it's like, um, can run some sandbox programs inside, can recollect all this information about CPU and using the Linux kernel, right? Also, uh, recollect some information about cgroups CS FS, um, using this sandbox program, running, using EVPF and thus some prediction using ML models of offline ML models to finally, uh, estimate the energy conceptual for the bots running on each node in a Kubernetes cluster. So this session is, has like some basic theory around just to understand what has happened with Kepler and how it works in order to motivate people, also to create breaks, to measure, um, this energy consumption.
So EVPF, because it's like right now pretty popular. Um, but, um, it's something that maybe, um, a few years ago, uh, this technology were used like to run the different stuff, but become pretty popular to run and to perform some networking stats and this kind of stuff. Um, so EBPF, uh, basically what it does is running a sandbox programs in a privileged context in the operating system garner, right, or Linux.
Uh, so because you can run, uh, this kind of sandbox programs, you can create some part of programs that can get some capabilities to get information. So Kepler, what it does is run a kind of program to recollect this information in order to calculate, like energy consumption, carbon emission, um, without regarding to change the current code, code and record by other things or load models, et cetera. So Kepler, um, technically let's say that is listening into the finished task switch function from the ker and with the kernel or DBBF catches that happens.
It's like a triggering when, when a finished task switch function from the kernel is like triggering, is looking for when it's triggered and getting information from from, from this place, right? So when this function is called from the Linux carna, um, well, it's called when a new task is going to switch and to use the cpu. So it's going to, let's say, release the previous, um, the previous processes and in that moment, because in the UX kernel, uh, programming with seal language has like a task structure that contains this information about the task, which is leaving the cpu, right?
So contains that information when it is like switching between like the one process is like getting out of the CPU and the other is coming in that moment. There is information in the task structure, so EBPF with that kind of problem, uh, programs that can run, uh, in a privilege of context in the current can catch this information, right? So with that information using, uh, running in the CARNA is recollecting information and starting is started this information in a series, in a time series database, call it promeus.
Um, so also has a PROMEUS exported that is listening to this information and insert inserting information to promeus. Finally, um, because you are diss you are creating like kind of, um, values that you can access as a kind of variables across the time. And finally you are going to, uh, show this information in a graph andana dashboard.
And also this information is calculated using these, uh, machine learning models and the data that they recollected in that market. So this is how Kepler does and is using EBPF. So of course you have to have like EBPF installed, or c or let's say for example, in Google, like the, um, ceiling version from Google, et cetera, in order to run Kepler taking a look about the machine learning part.
Um, Kepler runs as a diamond po. That what it means is that a pod is running on each node, so it's like a replica on each node, recollecting the information from the containers from each bot running on that node and recollecting all that information, that sends an information to the service, uh, that has the machine learning model in order to do the prediction about the energy consumption. Then this data is moved to, uh, Promeus and it started in taking par, uh, or using the, the time about when that carbon information about the emission was like calculated which time, right?
So Kepler works in, in that way. So it's like a little bit like a kind of machine that is, uh, doing estimations and using a machine learning model and capturing this information using the kernel by using this part of EBBF, right? So it's turn technology, some stuff about machine learning and that things, this is like the kind of variables that you can find when you are storing, when, when you can find about ke learning certain information on app promeus.
So this is like for this year and this, uh, version of Kepler, this is like the several, um, variables that you can find in, in the Promeus, uh, to calculate or do some calculations by yourself instead of using the default dashboard. Like, like the Jules Jus is a measure of energy. Um, so you can find the total of JUULs for container, the, for the course that are using the container, the RAM packages, other information, GPUs, energy statics and statistics, and, and that can stop, right?
Um, so how you can install Kepler in a Kubernetes cluster, you can use the regular, um, Kubernetes manifest how operate the operator of Kepler, and you also can install, uh, Kepler using an RPM package, a Linux package based of Red Hat in order to start cap player like in a regular machine, virtual machine or, uh, on-premise machine, right to install because it's like the, maybe the easier way or referred way, uh, you can install, uh, Kepler using helm. Uh, this is the link. Um, the other thing talking about, well, let's say that this presentation includes information coming from two groups of sites from the Linux Foundation.
Uh, one is like the green, so water foundation, and the other one is like the technical adversary group for environmental sustainability. So in the tag environmental sustainability that I'm going to talk later, uh, they have something in a blog post about mentioning that there is a, a place that you, where you can find information about a kind of parameters about how many energy is used by a data center, uh, about, or maybe data transmission networking and this kind of information. So there is a place about the energy agency, um, from United States, I think, um, where they have a table of information about, uh, the energy consumption in a different ways, and it's like the base basis of how to compare how the carbon emission is going in your data center, like going bad or, or going, going good like, right?
Um, so this information is available in, in the internet, and they have like the last measures from, from the 2022, and I think that is an update for the last year. And they are constantly updating this information, doing calculations, recollecting information about different data centers across the, that country. And yeah, so that's the place to have like a starting power point to compare your carbon emission to a, let's say, to, uh, an average carbon emission from different data or data centers or important, uh, data centers, right?
Uh, continue with this, the Green Software Foundation, uh, because it's a non-profit organization, uh, under the Lenux Foundation, it's like trying to build like a trust ecosystem about green solar building, right? So they defining a kind of a standard to calculating the rate of carbon emission that is called it the salt carbon emission intensity. Because let's say that Kepler is not like, um, something like I have to do a lot and it's more like it's showing this information and I have to take some adjustments to my system in order to reduce my carbon emission.
So the important thing of keslar is not, not like complicating styling the software and that kind of thing. So how to interpret that information. That's the goal I think when you are using Kepler because it's like a graph and what that graphic means.
So the Green Silver Foundation define a kind, uh, standard, call it the silver carbon emission and a formula to calculate, uh, this carbon emission. This is the formula and how what it looks, uh, remember silver carbon, carbon emission intensity and it's like mature by this, the e energy consumed by a software system, the I region specific carbon intensity dm, the embodied E emissions and hardware needed to operate the software in the system and the r the reference unit for comparison. So the reference unit for comparison could be like maybe, uh, the information from the energy agency and other is called like coefficient of gas emission and that kind of stuff, right?
So if you visit the website that contains this formation, you can see like the, in the last, um, column, you can see the pumps per kilowatts or, uh, using measuring the coal, the natural gas and the petroleum. So this co efficiencies, uh, or values, well in the dashboard is going to appear as a coefficient are the points of comparation about how green is your software running, and you are going to see similar values, um, in the graph dashboard of Kepler. So continue with this, uh, and taking this information from the drug environmental sustainability block.
Um, this standard, uh, is for the use of granular data and it's going to give you information from virtual machines and individual services, et cetera, uh, about the energy consumption and carbon emission, right? Uh, so in the context, so of this, uh, is there is a formula too to calculate is more, uh, in the focus of, of, of er let's say about how to calculate the energy consumption that represents the accumulate power reduces over the time, right? So there is a formula, uh, you can see the PU um, size P based on et cetera, right?
Uh, how this measure, so let's look a little bit about this formula and what that it means. Well, the final result is energy consumption, right? But what each leather means, the p is a power consumption of the virtual machine, the geo, the authorization.
And it's pretty interesting because you are, when you are studying virtualization, you know that virtualization re uh, helps you to reduce the amount of physical machines both also one machine that is using virtualization cons consume more energy. So you are impacting more or producing more energy or energy consumption per machine using virtualization, but it's better that using more physical machine, so you are reducing in some way the carbon emission and, uh, physical space and maintenance and cost. So you are saving cost, cost like using virtualization.
So the p was the poor consumption, you utilization how many times the virtual machine is running the size of the Bert Thomas machine p uh, for P nine dynamic power range of the device and the p base, the idle time of the time of the machine that when the machine is sleeping, right? And the end is the number of VR 12 machines running in the, in the, in the server. So let's look again the formula, right?
Um, then this is how Kepler, the dashboard Kepler looks is like going to show you like the general carbon food footprint, the carbon emissions for all the name spaces, and you can take a look at the upper side. You can also select, uh, different namespace, uh, the bot that you want to measure or maybe all the bots, all the name spaces, uh, and value of comparison about, for example, the coal coefficient, natural gas coefficient, petroleum, that are the values, uh, found in the previous PA page for the energy a agency, right? And you are going to find all this information across the, the Grafana dashboard and, uh, let see, yeah.
So right now after that, let's talk about the tech environmental sustainability because they are like, it's a technical adversary group that supports projects and initiatives related about getting an environmental sustainability in cloud native applications that are running maybe in a Kubernetes environment, right? They are looking for building, packing, deploy and manage it, operate them, giving like best practices and this kind of stuff, support different projects, initiatives, uh, do some kind of small investigation recollecting data like doing, uh, recognizing like the state of art of disinformation and in some way collaborating with this, with, uh, green Software Foundation. Uh, so this time is going to identify, uh, values initiatives that can help to reduce the carbon emission, a carbon fronting, uh, footprint, energy consumption when you are using cloud native softwares, applications, tooling, et cetera.
And of course, we are calculating this using cloud native software. So Kepler in the other size, let's say that it's like the one of the popular projects that runs on a cloud native environment in, you know, Kubernetes and calculate this information. And let's say that maybe is one of the most matter, uh, projects in this moment to calculate this information, but also the sustainability as we call it, like that, uh, can proposal ideas, a majority model to to run a project or to implement something or give ideas, et cetera.
Uh, look, they are looking, uh, and doing blog posts about how to create a QEs cluster that could be like sustainable, um, how to measure that car carbon emission energy consumption. And there is like, they are creating right now a kind of learning plan of, uh, about cloud native sustainability in the ecosystem and investigating which other breaks. Um, you can just for that.
So this technical adversary group have like different way to move the things. So, and also you can collaborate to improve Kepler. That's the idea, to invite people to do block posts, participate, eh, collaborate, contribute to code, et cetera.
Or I can put you in contact of different projects, uh, from the attack environmental, well, from, from this mission of reduced carbon emission, for example, Kepler. And, but they are another programs. So because our group, I am part of that, of this technical group, uh, as I'm looking for for the website and that, that kind of thing.
So automate stuff, but I am pretty getting in touch with this and participating a lot. So you can participate giving ideas hosting events like, like cloud native now, um, no, sorry, like cloud native, uh, sustainability week, right? In that week we are, I think that is in October, we are going to talk and spread the word about environmental sustainability on a Kubernetes environment.
So you can also participate in that way. And there are a lot of ways that you can participate of this group in order to reduce the carbon emission and have a green water, right? And in this website you can see the different programs that can help to environmental sustainability.
This is a landscape that or tech is like looking for and constantly updating about the different changes of programs, new programs, documentations, et cetera. So finally, um, I can say that just green software to achieve environmental sustainability or that list measure how your Gloster is running on the carbon emission and do that adjustments to reduce carbon emission, right? Uh, let's go to a small cap demonstration.
It's basically, uh, a dashboard to take a look at the dashboard you have to pull forward. Once you have Kepler install, you have to pull forward the Grafana dashboard and you can like, eh, going around there. So here you can see like the different information about the, the bots and this kind of stuff.
So you are going to take like different information here, um, from different things and for example, change about each bot or a specific well information about it and the different, for example, GPU, et cetera, and all the information about the bots and, and this kind of stuff. You can manipulate also the, the coefficients in order to take a look if it's like safe or it's like impacting too much in the environment. And so this is like how Kepler goes there.
Also, let me see if I can move there or something like, yeah, for example, in, in this dashboard, well you can see the different variables that I show it. You can create like your own like query to permit use in order to show the information and you can use all that variables that I mentioned it, et cetera, and create a regular dashboard because there are a lot of variables that maybe you want to create something more specific for your needs, right? So this is how Kepler goes or moves and let's finish the presentation.
Uh, of course the links important things. The sustainable computing IO is like the website where Kepler lives there. Or you can follow information about Kepler, the TAG environmental sustainability from CNCF, the block from the tag environmental sustainability.
I want to invite you to blog post there and the grid software implementation and the official page where you can find, uh, this starting point to compare how sustainable is your system in this, um, energy agency, right? And the slides, you can find the slides on my website. I think that cloud area now maybe is going to post something about that and my email contact and my Twitter, I can, you can also find me, uh, LinkedIn, uh, Sergio Menez and see you there.
And thank you to Clouder now to give me the opportunity to per for participate here. So thank you very much.