IBM, NVIDIA, FinOps and Observability – Cloud Native Now Podcast EP11
Mike and Paul, a practice lead for application development at the Futurum Group, discuss the cloud-native implications of the IBM move to acquire HashiCorp before diving into why NVIDIA acquired Run:AI. Then the conversation shifts to applying best FinOps to rein in Kubernetes costs. Additionally, Paul previews some of the results of an observability survey conducted by Futurum.
Transcript
Hey folks, welcome to the latest edition of the Cloud Native Now podcast. I'm your host, Mike Zure. Today we're with my Good bunny once again, Paul Nadi from the Futureum Group, and we're talking about all the things that happen in the cloud native world this week.
Paul, how you doing? Good. Good.
How you doing? Having a good day. I'm doing great.
I'm doing great. I am though, trying to figure out what the heck is going on with this IBM HashiCorp thing. It seems to me like on the one side there's this Ansible framework, I'm the other side, there's Terraform, the Terraform community is kind of at war with itself.
How do you see this all playing? Yeah, it's a great, great, uh, kind of thing that happened this week. I, you know, in my opinion, I think it's a, it's a good move.
Certainly a great move for, for IBM to kind of get a, a good handle with the market. You know, back on March 20th, I wrote, uh, a, a research note that anticipated a lot of this activity happening. And, um, you know, I I, I'm putting, I have a new research note that's going out now that's referencing that.
So, so I've been kind of tracking this for some time and watching what's happening in this space and, and seeing what's going on. You know, really when we look at, um, what's, what's going on with this acquisition, uh, it, it, it's, it is a, it is a way to help simplify, uh, the complexity within, um, organizations that are trying to, uh, grow in their maturity model of their, their tech adoption. So automation, security, service mesh, whatever it may be.
Uh, you, you know, you do have the IBM per, uh, take, but you also have the Red Hat piece, right? Which you mentioned Ansible, you mentioned Terraform for the Hashi side. Um, all of that plays into the picture.
But what, what I think is kind of brilliant about this is IBM is giving customer choice, right? They're giving customer choice on the direction they want to go on the platform tech stack that they want to use and, and, and move forward with it. Now, there's a lot that's happening here, um, but really this is showing that it's, it's strengthening the strategic move for IBM to kind of, uh, focus on their cloud computing software services and align to the overall, you know, efforts and company values that they're trying to evolve to that competitive landscape.
They have to, right? They have to drive in that space. And, you know, when I, when I look at this acquisition and I look at, you know, our research and what we're, what we're talking about in our research, and you can read more about this in the research note as well, I'll just kind of touch on it, but, you know, we see that there, uh, out of this, uh, 848 person respondent survey that we did, we see that 75% of the respondents are using six to 10, six to 15 tools to organize the data in the organ in their, uh, organizations.
And one of the things that this, um, move does is it does provide those tool sets and the tech stack, but hopefully IBM will be able to harmonize those tool sets to make it less complex for organizations to, to achieve their goals. How do you think this will play out with the CNCF, which sponsored a fork of the Terra foreign, uh, source code project when it was determined that they had changed the licensing policies and that created an uproar? And I wonder, do you think, um, IBM is gonna turn around and contribute to reform back to the CNCF to kind of unify those forks?
And, and I'm asking the question because on the one hand, red Hat is a supporter of A-C-N-C-F, but you know, I look at what they're doing with Ansible and I don't quite see that one going any consortium. So it seems like they're got a mixed message here. Yeah, it's, it's a great point.
Uh, you know, I'm assuming you kind of are alluding to the BSL kinda licensing and the commercialization, the licensing. Yeah, it's, it's a challenge here. So, so the way I view it is this, and again, this, this is based on our research, so I'm kind of basing most of my comments on data that we have.
Um, you know, we see that 68% of respondents, uh, in our survey want to use, um, they wanna work with vendors that sponsor open source projects, but those organizations don't want to go out the alone, right? They want to make sure that those, uh, the vendors have enterprise level support to go at it with them. So the organization is looking at these open source projects, and there's a, um, you know, uh, an opportunity to use, uh, enterprise level support.
That's what organizations, that's what organizations really want to focus in on, is using that support level. 'cause they're not in the business to build those applications. They're in the business of doing banking, finance, whatever they're doing, retail, et cetera.
Um, but with regards to the c ncf F and the open source community there, in my mind, it's not an, and it's not, it's not a, or it's an and or, right? So like you can have a commercial version and you can have a, uh, supporting commercial version that maybe a completely hands off, maybe, um, you know, organizations are gonna look at it and go, we just want a fast time to value. We, we, we work with IBM Hashi and, and Red Hat as a trusted advisor, and we are gonna deploy this solution in the, in the, in the whole stack.
Cool, faster time to value, don't worry about what the underlying code, or maybe they wanna have completely pivot the other way, work with open source, have it completely flexible work within the ecosystem. And we'll talk some other things on ecosystem later in this chat. But like, but basically the ecosystem gives that customer choice.
Which tools do you wanna put in your tech stack, right? And what I like about this solution here, oh, there's this opportunity that's coming for forward for IBM is, it gives that flexibility to do the, and or right, it gives the flexibility to say you can use, um, open source projects, you can use commercial projects, and you can kind of use the best of both worlds. Mm-Hmm.
I think it's gonna be a mixed play, but I gotta say there's a part of my soul that's looking over at the AI community a little bit, and this whole conversation may be rendered moot by those folks. They are making a case that says, Hey, instead of using Terraform and giving that to a bunch of developers who are probably gonna misconfigure something anyway, it's easier now to just kinda look at what the application requirements are, expose that to an LLM and have the LLM configure the application and, you know, let's just follow today and all go home happily. Yeah.
I mean, look at the recent announcements with Nvidia and run ai, right? I mean, we see that the, the, the, uh, ecosystem that we were just kind of chatting about there, you know, run AI is providing that API approach, right? That, that control plane and the cluster engine to put everything in place to, but put up those guard rails, so to speak, to make sure that, you know, exactly, to your point, run it through the ll run the application development through the LOM optimize your development processes, right?
Um, are we going to get to the point today where it's completely fully automated where applications are created? Probably not. I mean, there's going to be a human in the loop, right?
There has to be at this point, and most organizations feel like they, there must be, right? You know, we don't wanna just like, let automation just take over the world, right? That's just not, not, uh, not real.
Uh, but I do think that this, this is a good example to what you're talking about with run AI and Nvidia that backs up your claim, right? I mean, NVIDIA's driving that message. ai.
ai, that's basically an orchestration tool that allows you to maximize the u utilization of your GPUs. And that's been an issue for folks because, well, they're expensive and hard to find. And I guess my first question to you is, um, you know, it seems counterintuitive that Nvidia would go buy something that requires people or helps people use something they're providing more efficiently when theoretically they're supposed to be in the business of selling more processors.
Yeah, no, that's a good point. I, you know, I think that it's, uh, it's less about, Hey, let me, let me sell you more, right? It's more about, let me make sure you're getting what you need to be successful in your solution into what you're trying to do for your business.
You know, for example, all the, I I, I would say from an Nvidia perspective, all the hyperscalers that are offering, uh, you know, uh, AI kind of solutions have the ability to use, uh, you know, tpu, GPUs, CPU in order to kind of optimize depending on what you're trying to develop against. And what I see with the Run AI acquisition is it really gives that, um, proper allocation of which tech stack, which, which underlying, um, process or which kind of flow that you're running. Um, how do you optimize your development process and what, what resources do you need to do that?
And that's what I was talking about, that control plane and the cluster engine really kind of allows for the, so the proper usage of each one of these, these areas. Mm-Hmm. It also seems to me that we're kind of starting to move towards a world where maybe we're trying to figure out how to be less dependent upon GPUs because, um, increasingly folks are saying, well, I'm gonna run the inference engine on anything but A GPU, so I'm only gonna use the GPU for training purposes, and how often am I training?
So, um, how much of this GPU scarcity is really holding things up? Well, again, I think it's right sizing the, the tech stack underneath the application to use, uh, the right tools, the right processors, the right tech stack for your application. If you can use a CPU that's adequate for the, the, the, um, development of your ai, uh, you know, uh, workflow workload, then that is going to be optimal for your solution, right?
So that's how, that's how it works. I, I, I just think that selling GPUs just continue to sell GPUs. It, it's, it's basically putting a, a very large engine possibly in a solution that you don't need such a large engine on, Right?
And even if you could find the engine in the first place, so That's true. That's true. Um, all right, well, let's shift another gear here, because there was other news last week and, um, that included Storm forges aligned with Cloud Bolt, and it's another example of this, uh, finops move, but essentially, storm Forge, if you're not familiar, makes a, uh, an engine for helping you optimize Kubernetes to increase again utilization.
And they are gonna work with Cloud Bolt now to pass along that data to the Cloud Bolt automation framework so that you can actually see what things are costing you as you're going along and hopefully maybe put some controls in place. And my question to you is, maybe I'm getting too old for this, but I seem to remember in, in the era of on-premise, I would know how much I was spending, and then we went to the cloud, and now it's like a monthly bill, and I'm surprised, and the developers are, for lack of a better analogy, or kind of like drunken sailors, just like consuming resources as they go. Do we need some adult supervision?
And what does that look like? Yeah, this is, this is a good one. Um, you know, I've been working with Yasmin and the team over at Storage Storm Forge for, for quite some time.
Um, you know, the, the thing that's interesting is I, I I, I think developers and DevOps get a bad rap when it comes to, uh, you know, you know, the, the resource allocation and, and utilization. You know, I, I I think that a lot of, uh, what I see in our research is the business KPI is to release code very rapidly, right? We see that organizations wanna release code, uh, 8% of respondents wanna release code in an hourly basis, right?
So the, so DevOps just really are about getting the code out the door, right? And, and they have to get it out the door and the, the things that block them from getting the code out the door may be resources. So do they over-provision?
Sure. Do they know what they need to do exactly when they're starting to build the application? Not necessarily, right?
So they build and they get the, the overprovision, they have the resources needed, they meet their business KPIs. The challenge is they don't always, well, they don't have to worry about the bill, right? So they just keep bringing in the resources and not worry about the bill.
The thing that I find fascinating is the, uh, CNCF ran a sur a survey that found that 69% of respondents in their survey, um, basically they don't, they don't monitor Kubernetes spending, um, at all. Right? And that is fascinating to me because with the sprawl of, of, of applications and Kubernetes pods and clusters, that it seems like it needs some bumpers and guardrails.
And that's where we look at, you know, the, the, you know, collaboration here between Cloud Bolt and, and Storm Forge really is helping, um, not just, um, helping the organization, but helping DevOps properly size what they need, right? So these tool sets are putting the, you know, the right finops actions in place to say, here's, here's what we need to develop and push our code out the door, and here's the resources that are needed to do it. So I, I like this approach.
I like this, this, uh, unified front. I definitely think it makes it frictionless and seamless. It's not a heavy lift to use to forge and cloud build to kind of get you to where you need to go.
So I think it's the right move in the industry. How much of this is a, a political issue within an IT organization? We know that there's not always a lot of love loss between the, uh, idle, the ITSM crowd and the DevOps crowd.
And a lot of those folks will say, you know, the DevOps stuff is great for going quick, but you can't manage at scale. And to your point, they get a bad rap for maybe not managing consumption as efficiently as possible. So how much of this is like a tension that's manifesting itself within the IT organization?
'cause the CIO is under more pressure to reduce costs. Yeah, I mean, look, CIOs are their number one concern is to modernize their applications, right? They're moving forward to keep the business operational and run things forward.
Right? Underneath that concern is the challenge of skill gap issues and complexity and, and, and, and resource allocation, right? And it, and you can either do this with, you know, your bench or you could do this with work with service delivery partners.
But you know, in, in my opinion here, this just gives data to work with. So this gives the information to say, look, we're, we're gonna provide you what the, the spend is going to be. You can do proper budgeting, but not just go out at like, you know, what can your finger in the air and waiting for, which way the wind's going to blow?
The CIO can go, okay, we can actually budget an appropriate amount knowing that the, you know, we're gonna extrapolate this, this growth and this monetization effort, um, to move forward. Now, we all know that not everything needs to be refactored. Not everything needs to move, you know, like, it, it really kind of goes to what is going on within that specific business and how are you going to, you know, build net new applications, et cetera.
But the tools and the underlying, um, tech stack that's required really needs to have some type of understanding of what the budget is to do it, right? Because if you just say, I need unlimited budget, then that's where it kind of goes outta control, right? Uh, and I think that's where this, this relationship here is really tying it back to giving those, the visibility and the optics on what the spend would be for extrapolating those projects.
Right. We only got a couple of minutes left, but I know you guys got an observability report coming out, and I suspect that cost is tied to observability. So, um, you know, give us a couple of highlights from that report and what should people look for and I mean, work.
Yeah, Love this report. It's great. It just came back.
The, the data's coming back. We have a bunch of research to share with anyone that wants to see it. It's on our future intelligence portal.
Um, but it's also, we also have some research notes and briefs that are out there. You know, absolutely. Cost is a factor when we look at, uh, you know, monitoring and logging and tracing and, and, and, you know, those actionable insights, understanding what is happening within the organization is imperative for the success of the organization, right?
One of the things around cost, we found that early adoption of, of, uh, um, observability practices was about collecting a bunch of logs, collecting a bunch of data, which was bringing up the storage costs, bringing up all the resource costs because you're collecting all this stuff, but not knowing what to do with it. And what we're finding is most organizations are moving from that monitoring and alerting stage to really, to the tracing and logging, and then taking that information and, and using those actionable insights. I know we only have a little bit of time here, and I can speak for hours on this topic 'cause it's just, it's, it's near and dear to my heart.
But I will tell you that there's a lot of information to be found here at the, in the Futurum, uh, the Futurum Group website. There's a lot of research notes and there's a lot on our intelligence portal to kinda dissect a lot of this data, right? And there's a full interview on this subject on Textron tv.
So by all means, go check that as out as well. Paul, as always, good catching up and we will chat next week. Likewise, thank you for your time today.
All right, thanks everybody for listening to In Good Luck. And of course, you know, if it wasn't for Cloud Native, we probably wouldn't need observability in the first place. Take care.
