Exploring Cloud-Native AI and Container Runtimes at KubeCon SLC 2024 – Techstrong Unplugged EP33
Ricardo Aravena, with more than 20 years in technology, leads the tag runtime group and cloud native AI working group. His group emphasizes the machine learning lifecycle and open source solutions, focusing on environmental sustainability. Ricardo invites participation in the working group and shares insights on container runtimes, enjoying the positive discussions around AI at the event.
Transcript
Welcome back to Text On Unplugged. My name is Cassandra Chin, and today we're here with Ricardo Ravina. And we're actually, today we're at CubeCon North America at Salt Lake City.
So can you do a short introduction of yourself? Yeah. I'm Ricardo aa.
Um, I currently lead, uh, the tag run time group, but in the CNCF and I also lead the cloud native AI working group that falls within the tag run time also with, uh, in the CNCF. I've also worked in many different companies in the industry, in tech, primarily for over 20 years. You know, small companies, large companies, um, large companies like Cisco, VMware, and small startups.
Like, uh, recently one cultural era where we had a machine learning monitoring and observability solution. Can you tell me a little bit about how you got into technology? Like what really inspired you?
I got into technology very early in life. Uh, I can't remember how old I was, but must have been around maybe nine or 10 years old, uh, when I started going to some bookstores that had computer magazines. Uh, so they, they talked about this old computers, like they eight bit kaba door 64 and Atari a bit computers.
And for some reason, I, I, I got excited about them. It, it also tied in with like gaming. Uh, so I got interested in gaming initially.
I don't do a lot of gaming now, but I, but I, I got super interested in gaming and together with computers. Then that led to me being super interested in, in technology, right? And, and computers.
So that led me to have a, a college degree in computer engineering. Right? So, and, and I like the part about gaming is still true today.
Like how I got into technology is I used to run a Minecraft server and like, whenever it would break down, I would have to fix it. Yeah, I think that's, yeah. When you get started, it's, it's something that, you know, grabs your attention and then you, you just wanna keep on learning and, and, and getting, trying to understand how things work underneath too, right?
Like, uh, like how do they build these games or something, right? And, and, and then, then you start going through a different path. So what kind of games was it for you, which really inspired you?
I mean, you talk about early games such as like s space invaders, right? Like the, the little ones where you had like a, like a tower at the bottom and there were this, these invaders coming at the top of the screen and you were shooting them and trying not to actually invade you, right? So, uh, really basic games, right?
Um, and, and then later, um, maybe some of the more popular games like World Warcraft, you know? So I got involved in that. But, uh, to be honest, I haven't gained in the last maybe 20 plus years, or maybe 10, 15 years, I can't remember.
But, but you know, I, it sort of got, kind of became sort of like in the notion in the back burner. It, I still like it, but, but I was doing much, but it, it, it did actually help me get interested in, in, in, in the technology field. Yeah.
It's still important to you 'cause it's really what got you going. Yeah, yeah, yeah. So today we're at CubeCon, like you run the ar working group.
Yeah. Can you tell me a little bit about what they do? Yeah, so we have, uh, several initiatives going on.
We got started with, uh, cloud native ai, white paper published at CubeCon Paris or right before CubeCon Paris. And that talks about the machine learning lifecycle and how you can use open source to enable that machine learning lifecycle, meaning, uh, how you prep, prepare the data, how you train your machine learning model, and how you serve your machine learning model. So all, all the different aspects and, and additionally observability.
Uh, so that was what the cube con and that kicked off a bunch of initiatives in the, in the, in the group. So today we have things going on, like, uh, scheduling challenges, white paper. So how, uh, organizations can use open source to, to schedule, uh, uh, AI type of workloads on top of Kubernetes, on top of cloud native projects.
Uh, we also have a security, uh, white paper, uh, some challenges around safety around, uh, hacking attacks or, or, or, or vulnerabilities. Uh, and then we also have, uh, environmental sustainability, AI white paper in conjunction with the tag environmental sustainability. So how to, uh, be more mindful about running this type of AI workloads because they're very expensive and they use a lot of compute, uh, and yeah, a lot of, a lot they consume a lot of, a lot of power, right?
Oh, You mentioned the sustainability white pair paper. Yeah. Is that just like a single paper or a section of a different paper?
It is like, like a single paper. Uh, but it's, there, there is an existing tag, environmental sustainability, uh, in the CNCF that is, is collaborating with this, with our working group, right? Um, uh, yeah, so that's it.
It, I mean, those are some of the examples. There are some super exciting things around how you can use AI to help cloud native, right? So not, not just how you can run AI on top of Cloud native, but like how you can use the output of an LLM to improve the, uh, Kubernetes pots running in your cluster, right?
Like, uh, like how, like, how to utilize them more efficiently, how to find gaps or how to find like, uh, problems or issues, uh, or how to maybe, uh, make it easy for engineers to, uh, issue like, um, or write like a prom that says, tell me something, tell me something is wrong with my Kubernetes cluster. And coming up with like this, you know, very, uh, descriptive, uh, natural language response that can help people troubleshoot some of these, these problems in, in, in infrastructure, in, in cloud native, uh, uh, yeah, uh, environments, right? If someone watching this interview today wanted to get involved and join a working group and help with these white papers, how can they do it?
I, I say like, uh, start joining our meetings. Uh, you can get on our Slack channels and look at our agendas, look at, uh, what we have on topic. Anyone is free, free to add an agenda item, you know, related to whatever the interest area is.
Uh, and they can just bring it up and, and if someone has a specific idea that is not there, feel free to bring it up. And, uh, there, there, there's still a lot of gaps. Uh, and if there is something existing, I mean, we're welcome anybody to just join and join in and, and, and, and, and start contributing and start working with the other folks that are interested in, in the same space.
Uh, the idea that all we also have is that we can join efforts to avoid duplication, right? So if somebody else has the same idea, then why not actually work together to to, to join forces and, and, and make something better. And I would like to add to that, like, I joined the AI working group about a year ago, and it's a very welcoming environment.
'cause like, even in the meetings, like it may look intimidating, but actually when I speak up, like people will stop and listen too. Yeah. So we, we tried our best and, and, and that's how we believe it, it, it's going to be going forward.
And, and also we think it's a path to getting people excited about the technology and continuing to grow the, the working group and the space. So, moving on, do you wanna talk about your involvement with TAG Runtime? Sure.
Yeah. So with Tag Runtime, I became a co-chair maybe about four years ago. And the CSCF has a structure where they have a governing war that typically manage, manages the finances of the CNCF.
And then under the governing board, there's this entity called the TOC, and then the TOC. Initially, I think it was like maybe seven members, five or seven members, and then continue to grow. But the cloud native ecosystem kept on growing, um, behind Kubernetes.
So Kubernetes started all this, but it kept on growing with different projects around the Kubernetes technology. And some of them later may not necessarily be related to or directly related to Kubernetes, but it's just the ecosystem started to grow. And so the need, uh, uh, started to, uh, arise in terms of like, uh, how do you break down this different type of streams of, of, of different areas, right?
So the TOC worked on creating this tag structure around different technology areas like runtime, observability, application delivery, storage network. Uh, and those are, were the five initial, uh, tags, but they actually created some more afterwards. And then, so when that opportunity came for tag run time, they were looking for co-chairs and, and I jumped in and, and I, and I put myself up there to, to be nominated and, and, and basically became a, a, a co-chair.
Right? Uh, uh, similarly, there were other folks in the other tags, uh, that get involved in, in as co-chairs for the other different tags. What does the runtime part of it do specifically?
Specifically the container runtimes. There's, there's several aspects, but initially I think what people think about the most is the container runtimes, like docker itself or container D or cryo, just projects that help you run containers in, in the s or help you run, or any, any bare metal machine you that you need to run a container. But the scope is a little bit wider and the CCF F is actually doing some restructure, uh, around the tag.
And maybe there will be something happening in the next few months, uh, because there's a lot of within tag run time right now. So you have, you know, talk about the basic container run times, but there's also like the storage of these containers, like the image registries. There's also the AI type of, uh, workloads, like ML lops type of things.
Also the GPU capacity, uh, management or the orchestration of batch jobs. So there, which are very related to, to AI and, and processing large amounts of data. So, uh, yeah, so the scope is pretty wide and, and, and there are many different things, but then, then it, what you think about initially is, you know, just like basic runtimes, but, but there's, there's a lot more than that.
That's pretty in depth. How are you enjoying CubeCon today? It's great.
Uh, I actually flew in last night, uh, 'cause I've been so busy with work that I haven't been able to attend the full conference, but, uh, uh, it was fun. I, I had the chance to attend the keynotes this morning. There are a lot of conversations around ai and I think that'll continue to happen maybe in the next few years.
Right? And, and, and ex exciting. There's the, there are are a lot of challenges.
I mean, they, uh, it, I think Nikita, Ragnar, you know, mentioned in the keynote this morning that, uh, you know, you can think about challenges at least in the next 10 years. And, and I don't, I don't think it's this is ever gonna change, but, uh, there's, there's always gonna be stuff in exciting things to work on and, and, and, and new things in this space. And with the, all the AI hype, like this is the place to be.
Yeah. Yeah. Exactly.
Yeah. So thank you for talking to me today, Ricardo. Thank you.
I'm glad to be here. Glad thank you for the opportunity. And stay tuned.
We'll have an interview after this with Melissa McKay.
