AI for Kubernetes – Cloud Native Now Podcast EP21
Mike Vizard and Paul Nashawaty, practice lead for application development for The Futurum Group, dive into the degree to which generative artificial intelligence (AI) might accelerate the adoption of cloud-native computing by making it easier to manage Kubernetes. Then, they turn their attention to platform engineering, usage of virtual Kubernetes clusters in the building of AI applications and an Operator from Slack that makes it simpler to deploy stateful applications.
Transcript
Hello, and welcome to the latest edition of the Cloud Native Now podcast. I'm Mike Ard, and once again, we're here with Paul nti, who's, uh, what is your formal title? I don't think I ever used that before, but let's, let's, let's establish that.
My great to be here, you know, as the practice lead and, uh, uh, principal, lead principal analyst for the AP dev practice at the FU and Group. I cover, uh, day zero, day one, day two, build, release and operations. Yeah.
All right, cool. So we've seen Red Hat start talking about Lightspeed, and there's a couple of flavors of Lightspeed, but Lightspeed is in the context of their Red Hat. OpenShift platform is a gen AI tool that helps you essentially better understand how do you build stuff on top of and deploy stuff on top of OpenShift.
It doesn't actually manage the clusters yet. That may be something handled by Lightspeed for Ansible or something else in the future. We'll see how that goes, but clearly we're starting to see Gen AI come to Kubernetes and arguably, uh, this is a platform that needs it more than many others.
'cause it was complexity. And so are we seeing a new way of thinking about how to manage Kubernetes clusters at a higher level of abstraction? Yeah, Mike, you know, I think this is an interesting perspective.
I, when I, uh, was at Red Hat Summit earlier this year, uh, you know, we, we talked a lot about, um, uh, OpenShift and, uh, gen ai, OpenShift, ai, and, and adding in the, the, um, uh, functionality of automation. So Ansible Lightspeed was announced a year ago, right? That was, that was something that is part of the workflows for the automation platform on Ansible.
It really, it started approach wisdom that, that turned into Lightspeed, which was a, which was a great approach to say, basically saying, okay, if you're gonna build workflows, we wanna automate the building of workflows, work well for automation and work well for Ansible. Taking that same kind of, uh, approach of saying, okay, well, with OpenShift we want to bring in, uh, and gen AI capabilities, so OpenShift ai, and they also had rail AI as well. So there's a number of different ways of kind a sudden expand into, uh, gene driven AI capabilities for, um, OpenShift, for Kubernetes, but also for the operating system as well as automation.
And I think that's a great move for, for Red Hat. I've, I've talked to Stu, been many, many times about this. Uh, we have, uh, we just did recently a DevOps dialogue on this, and it's on our bond, uh, the group's website as well.
But like, basically it's the ability to speed up the, the, uh, delivery of the, and the orchestration for, uh, for these Kubernetes clusters. When you start building out the adoption of these Kubernetes clusters, you really need to have a way to do this in a, in a more automated fashion because of the sheer of scale and growth. And I think that's where OpenShift Lightspeed or ai, uh, was OpenShift kind of comes into play.
ai, um, talking about the rise of software intelligence and the ability to use AI to kind of understand, want the flows are between different components and what the services are. 'cause one of the biggest challenges that you see with Cloud Native is it's just complex, and it's beyond the ability or the cognitive ability of people to keep track all the relationships between all these things. So, will AI play a much bigger role as we go forward in helping us, you know, sort out the chaos?
Yeah, Mike, I mean, I think when we look at, uh, you know, adding AI into the workflow of the SDLC and acro across, uh, across the, uh, not just the software development life cycle, but across the CI I CT pipelines and everything that kind of goes into play, understanding the shared delivery. What we find in our research that we find that, uh, just two, basically two thirds of respondents indicate that they're working about, um, almost a hundred percent faster in delivery of code and, and applications now compared to just three years ago with the same or fewer resources. And if you start extrapolating out where that's going in the market, there's no way that you can, you can, uh, uh, throw more hands at these projects to get these, these, uh, you know, code delivered out the door faster without using some type of tools like AI to do that.
So the interconnectivity between the, the, the workflow and the, and the information that ties it all together, uh, that will be defined and developed by the, uh, by the innovation of the, say, the DevOps teams and the innovation of the developer. But once it's created using AI in order to do the execution, it makes logical status and it, and that gives you the economy of scale as well. Mm-Hmm.
So who's gonna lead the charge on this? Is it gonna come from the vendor community, or will the open source community kind of lead The charge is a couple of open source Kubernetes AI projects that are floating around out there. Uh, it's still early days, but sometimes I feel like, you know, the vendor community plows ahead and then the open source community is shows up a couple of months later and says, uh, you know, we'll just take it from here.
Well, I mean, this usually, uh, means to an end, right? I mean, if, if, uh, if the vendor or community sets something up and it gets you faster time to value, but it doesn't give you full functionality, that's might be an opportunity for the open source community to expand on it. The question I have is how much, well, two things.
One, how much advantage do you have, or what's the delta between what the open source community provides versus what vendors provide? Is it 10%? Is it, is it, is it 30%?
Is it 50%? What's the delta between the functionality? The second thing is, is if organizations are actually deploying and how they're deploying based on these, uh, automated platforms, are they fully taking advantage of what the vendors are providing versus adding that additional functionality that you can get for open, open source?
The reason why that's a big factor is because if the organization takes on the open source initiatives themselves, they own that initiative all the way through, and they have to own the support, the delivery, the creation, all of it that goes along with it, that means that that 2:00 AM call to the SRE because the, the pipeline's down, uh, is what will happen. And they have to own it all the way through. Now, the difference is if you get that 2:00 AM call and you are working with a vendor, you have a vendor of enterprise level support, potentially, if you use an open source, you're basically stuck to the community.
So with all that said, I think that, uh, it makes a lot of sense for, um, understanding the gaps between what the faster time to value is versus what deltas it needed for the additional functionality. Mm-Hmm. How do you think this will impact the skills required to ultimately implement all this stuff?
Because I can see a day when, you know, that mirror mortal IT administrator doesn't necessarily need to be a DevOps expert to manage Kubernetes clusters. Yeah, I agree. Uh, actually I think that, but we're a long way from there.
Take it as you went out of the loop right now is just not, is not, uh, it's not really an option because I don't think it's mature enough to do so. However, with that said, I do think that you take your skilled, uh, talent that can develop in these workflows, create those templates, and then use tools like we were just talking about lightship, right? Like Speedway and take those, uh, uh, workflows and put them in play.
If you have those workflows in play, you don't need to have that, um, uh, that, that, that high level of talent to keep pushing the, the big green button, so to speak. Um, so I think that the workflows have to be developed. That is, in my opinion, working in collaboration between the organization's bench strength as well as working with vendors that can deliver the solution, Right?
Um, so that, and are we getting to a point where the DevOps teams will go solve more complex problems? And I'm asking this question because one of the things I hear from DevOps teams all the time is they go, yeah, we could create a template for that, but we're so busy all that, everything else together, we don't have time to go create the template. Yeah.
I mean, we see that in the research, our most, our latest research that we see, and it's, it, it, it's, it never ceases us to amaze me. I've been doing these studies, uh, for the last, uh, four or five years now. So I have trending data, I chose this.
But we see that 33% or a third of the time that organizations are DevOps teams are spending, uh, within organizations, uh, are spending a third of their time on innovation and creating new applications and 66% of their time on maintenance and in maintenance mode, that's a problem, right? Because that means that they're not, they don't have enough cycles in the day to innovate, to create these new applications. And frankly, that solves two, it, it, it causes two issues.
One, um, job satisfaction becomes into play. Uh, DevOps seems they don't want to go to work every day, and developers don't wanna go to work every day working on maintenance. It's just not something that's interesting.
It's not fun, right? The other thing is, you don't have skill gap development or skill development. You, you end up having, uh, you just have to continuously work on doing the, uh, you know, tedious tasks, not fun.
That's not what organizations want. That's not what, um, the, the staff wants. The other side of it is, if you, uh, take those pieces in play and, um, provide more innovation, then you can actually be, be, uh, produce higher levels of output with those workflows.
And by producing those higher levels of output, you can put templates in place that produce better results. The problem is, like, I haven't seen this in the last four or five years. I see the same results every year.
I run this trending study that a third of the time is spent on innovation and two thirds of the time spent on maintenance. And I don't know if that's gonna break anytime soon. Well, let's hope that AI will break the vicious maintenance cycle, because that would be like the coolest thing going that we can immediately do, I think, or sooner than later.
'cause some of the stuff that we're thinking about doing is a little on the far end of fanciful, and I will get there one day, but there's a lot of scope work that just gets in the way. Definitely does, definitely does. All right.
io on Textron tv, and he's just pointing out that I think maybe the rise of cloud native is what's driving the move to platform engineering, because, you know, the, the complexity begets a need for a simpler approach. Are these two things kind of joy at the hip in your mind? I mean, yeah, I think that there's a lot of, uh, synergies between the two, uh, you know, uh, methodologies and relationships between delivering SA cloud native code and such.
But here's, here's where I think the rise of platform engineering really kind of comes into play. Uh, I, I mentioned, uh, just a minute ago above the trending data that I provide, right? And then I was talking about in this one particular cloud native study I do, uh, we, we have this, um, uh, information gathering information around the CICD pipeline, and we're looking at, uh, the, um, continuous integration testing and in 2022, so yeah, I'm going back a couple of years, but in 2022, we saw only 29% of the respondents did continuous integration testing.
In 2023, that number jumped to 66%. And part of the reason behind that is when you start looking at DevOps, SREs and platform engineerings and development, there's the con uh, the, the concept of moving or shifting left, moving things closer to the engineers, to the developers, um, but by having the ability to shift left the responsibilities and have the platform engineering teams develop those, uh, closer to the source of application creation solves a lot of problems that, uh, that organizations run into. Now, what I mean by that is, in the 2022 study, what we were fighting was DevOps teams, their business, the business KPIs was to push the code out the door faster, push the big green button to get the code out the door.
So QA was going, Hey, we got a problem, they're going, don't care. Push the code out the door faster right now, which ended up potentially causing problems. So I, my my analysis there, talking to those companies, the, the response was basically, look, we have sprint reviews every two weeks, we have a SaaS based application, we're gonna continuously update, and if we find, uh, problems by the, with the end users, we'll make those updates every two weeks.
They won't see a problem. The problem with that is the attention span for, uh, end users is instant gratification. If it doesn't work the first time, you're at risk of losing that client forever.
So that's the problem that we see. If with the rise of platform engineering, uh, it's solving the problem of harmonizing the business KPIs, you may hear a lot of organizations talking about, uh, uh, SLOs are, uh, service level objectives versus service level agreements, right? So they're moving towards these objectives where everyone's rowing in the same directions toward that north star.
That's where the platform engineering team kind of comes in and shines. That's where there's an alignment between platform engineering and cloud native application development. As part of that, are we seeing, Um, A different mindset here from, uh, the folks in charge of the platform engineering team?
Because back in the day, a lot of folks embraced DevOps to get away from centralized it, and now we're back with platform engineering. And to a lot of developers, it smells like, you know, the new boss is the same as the old boss. Yeah.
Well, I think that that's a good point. It's a good question, but I also think that the evolution of the tech stack changed over the years, right? Platform engineering historically, uh, and I would say in a heritage kind of view of the word platform, whether freeze platform engineering was more about the infrastructure and less about delivery of the applications.
Now we see platform engineering moving up stack into the business logic and worrying about the end to end. So going, uh, going, uh, looking from the business logic all the way down to the endpoint devices, right? Looking through doing tracing.
When I look at our observability study, we see cloud log monitoring and tracing being one of the top things that, uh, organizations are, are considering historically the tracing and the monitoring down at the, uh, the infrastructure layer was platform engineering, right? And now we're seeing that happening at the application level all the way down through the stack. So I think that the, there's an elevation of skill moving up stack into the business logic where historically it was focused predominantly on infrastructure Mm-Hmm.
Um, as you kinda like, think this all the way through, then, um, will AI be embedded into the platform engineering strategy? And these three things are now joint that they have platform engineering, the shift to AI and the rise of cloud native. They all kinda are, you know, one is the ankle, one's the knee, and one is the shoe goal, right?
I like the body analogy there. Yeah. I mean, look, I think that, uh, you know, the right tech stack, the right process and methodologies that use to achieve your business goals, um, AI definitely has a place in that, right?
And ai, whether you are looking at, um, you know, uh, using natural language to, to, to create things. So you don't need a highly skilled developer to create those applications, and you can just put it a natural prop in there, a language prop in there, or if you want to use AI to do actionable insights based on logging, alerting, and monitoring, if you see something happen, then the AI would do some, um, some action based on workflows and work, uh, and templates that are created that we talked about earlier. So I think that that's, um, that's an area that, um, it's an enabler for the success of the business and it's an, it's certainly an enabler for Cloud Native because as we, we talked about, you know, I was just doing a study this morning, I was looking at this and, um, the research, I was finding that over the next three years organizations responded that there is gonna be 500 to a thousand production applications running at the edge locations.
If there's that much, uh, production applications running at the edge locations, it's, it's very difficult for a human to be in the loop to make sure that everything is running properly. You need to have some type of lens into that with automation, and I think that's where AI plays into it. Yeah.
Um, so we also have this interesting shift happening that's related to ai. We're seeing, um, more uses of containers and Kubernetes to go build the AI models and more consumptions of GPUs that are hard to find. And there's a lot of obsession about utilization rates of those GPUs.
One of the articles we have in cloud native now is kind of suggesting that virtual clusters for Kubernetes are the way to go to optimize those GPUs. And, um, I think we've talked about virtual clusters in the past, but, um, is this gonna be like the killer app for virtual clusters? 'cause I kinda, it's, we talk about this tech minute and I know it's widely used, but not pervasively used.
Yeah, I mean, look, I think virtual clusters, um, is a way to take advantage of the steps of utilizing GPUs, um, in a multi-tenancy environment. And it also allows for, um, you know, the, the architecture to, to drive more horizontally and scale across the cost, cost effective way of delivering. Now, the one thing that I like to kind of bring up, and when we talk about these kind of conversations around Kubernetes clusters and cloud native and, and, and GPUs versus CPU versus TPU versus whatever you going to use for some type of, uh, uh, processing.
The, it's important to understand what level of processing power you need for your AI as sensation. Um, if you're using a, uh, A-A-G-P-U and you know, it's super expensive to get it and, um, you know, it's, it's, it's you're tying up resources that you may need. But does your application actually need a GPU?
Can you get away with using A CPU, um, or do you need a CPU? Do you need something that's that's higher capacity than what A GPU can do? The question really comes down to is balancing the time, resources to your, to the time that's needed for their, the, to, to build out your learning models.
If you can do that appropriately, that's where you're gonna to utilize the best technology for the, for the solution you're trying to achieve. If you're not doing that, you are really tying up resources that's gonna be, you're over to some extent, you're over provisioning for something you don't need to overprovision for. Mm-Hmm.
And it's always been surprising to me how much over provisioning there is in the land of Kubernetes. 'cause theoretically you're supposed to be able to scale up and down dynamically, and that's a core feature. And yet, um, we don't And is that just a bad habit?
Well, I think it is a bad habit. I think that organizations leave it up to, um, uh, maybe DevOps teams and not to, not to talk bad about DevOps teams, but you know, they have a goal. They have a goal to produce the result of, to have these applications out the door or work through the CICD pipeline.
And if it requires provisioning resources to do so, DevOps seems don't necessarily are, they're not necessarily associated with budget. So they basically say, okay, well I'm just going to grab the resources I need. And they may overprovision because they're getting the job done, but they're overprovision to get the job done and it has to be some belts and suspenders in place in order to check those things.
Right? They're rewarded on making sure the application is available and the last thing they want do is be bothered by some incident that occurred because there wasn't enough memory. Um, so they are gonna just kind of default to the most expensive item, or they might not even be aware of what it actually cost 'cause nobody told them.
Exactly. Exactly. And then if there's no pressures on 'em, then they gotta con then that behaviors will continue.
Alright, there we go. Now, last topic. Um, returning to the subject of stateful applications, but, uh, we saw the folks over at Slack have come up with a operator to help to deploy some of these.
Um, where are we on this whole adventure with stateful applications on Kubernetes? Is, is it now a standard thing? 'cause earlier on it was kind of like this whole philosophical debate about don't ever do it there, it's only for stateless apps and, you know, store your data somewhere else.
And how much data are we seeing on Kubernetes clusters? Well, state pole, uh, environments within Kubernetes clusters is not, certainly not new. I mean, this is something that's been coming out for, this has been, this was the challenge when going clusters were built, Kubernetes clusters were built because, you know, you, you can run, certainly run stateless environments, uh, and, and have, you know, applications that, that don't have to worry about the data integrity and such, and that's all, that's all well within you immediately just processing.
But the minute you wanna run applications that move away from, um, your heritage environment, your past kind of architectures, and you wanna move those, uh, data sensitive or, um, check the data integrity of those, of those, uh, applications, you need to have a state stateful environment, right? You need to have a way to have the, uh, cons, you know, consistency in the data that you have, snapshots that you can take, you have, uh, uh, you know, ways to roll logs and such. And that all requires the, uh, the ability to have, uh, access into those APIs.
So among many storage vendors have using the CSI driver or CIS driver, uh, to, to make sure that the, the, uh, Kubernetes clusters working properly. But you know, what we're seeing is the more and more functionality that you see in traditional architectures that's moving to Kubernetes environments, even though Kubernetes is, what, 10 years old now, or a 10-year-old anniversary. And, um, even though Kubernetes is out there, it's been out there for a long time, there's still some situations where Kubernetes just doesn't fit the need for heritage applications until that functionality is brought into those Kubernetes clusters, that's when we'll see those, uh, the full functionality of Kubernetes and, and, and the workloads.
Yeah, I feel like to your point, there's still work to be done in this whole area in terms of optimizing Kubernetes for these types of workloads. What is the overall sense of the level of maturity here? Well, maturity is all in the eye of the beholder on what you're trying to achieve.
I mean, like, when I look at, when I look at the way I define a maturity model within the applications, I mean, I look at it in the five phased approach and applications and organizations can fit in any one of those five phases. Um, and, and organizations can have applications that sit in any one of those five phases. So you, you know, an organization itself may have, uh, phase one to phase five and have, you know, hundreds of applications that fit in any one of those stages.
But I think when you look at maturity and you think of, of, uh, uh, moving from, you know, a heritage monolithic architecture, a three tier architecture, and you wanna move that into a, a cloud native state, uh, with a stateful, uh, application that's running there, the storage and all the associated pieces that go along with it need to be tracked appropriately. All the, all the, the lines need to be, uh, in imperfect alignment in order for, for, uh, the connectivity of the data, for the connectivity of the application. And then the access points.
The, see, the big question, Mike, in my opinion here, is when you start breaking this apart and you start taking, you start refactoring when you have those heritage applications and move them into a stateful environment in Kubernetes, and you build this entire architecture to support this new staple environment within Kubernetes, are you build, like what are you gaining by going from your heritage environment to this new environment? Are you gaining cloud native functionality? Absolutely.
Are you gaining microservices? Absolutely. But are you gaining any performance?
Are you gaining any, um, you know, BCDR scenarios or do you have, uh, um, uh, governance and compliance for data sovereignty issues? I mean, all these things gotta come into play when you're looking at moving these, uh, large dataset applications into Kubernetes clusters, but you also have the underlying storage that has to align to it as well. So maturity depends on the application.
That's how I, I guess I'll leave it at that. It really does depend on, um, uh, uh, what's going on in that application. But it also depends on how the organization, uh, may have multiple applications that are running the space.
All right, well, data has gravity and there will always be greenfield apps. So we're definitely gonna have state fill apps on Kubernetes clusters. This is just a question of how many and how often.
Hey, Paul, as always, enjoyed the chat. Thanks, Mike. It's been great.
All right. And thank you all for listening to the latest edition of the Cloud Native Now Podcast. And we hope that you're having great adventures with cloud native computing in general.
And until, uh, we talk to you, we'll see you next time.
