IBM on the Upswing, Intel Struggles, and the DeepSeek Saga Continues – Infrastructure Matters – EP69
IBM had a strong quarter, and Seagate sets the hard disk drive storage record. Intel cancels its new XPU and pushes out its 1.8 nanometer process chips, while DeepSeek continues recalibrating the generative AI inference world.
Transcript
Hello everyone, and welcome back to Infrastructure Matters, episode 69. Uh, and, um, wow. It's, uh, been a pretty interesting year so far.
Uh, it's been, um, lots of revelations on AI and new administration, lots of executive orders. Um, but let's, let's get right into the news. So I, I think we, we we're talking about IBM, uh, had a very interesting, sorry, um, IBM or IBB, what is it?
IBM? That was my mistype, That's what I thought. 0 version.
Um, so I've watched that closely, but I haven't been tracking, um, their earnings. So, Kimberly, can you catch us up? Yeah.
So they, they did very, very well. 7 billion of free cash flow. Um, you talked about, Arvin talked a lot about, you know, where they were in 21, where they are now of setting the plan that they had, um, for mid single digit, um, growth.
And they delivered that, which is a 6% com, uh, CAGR over that period. Um, what really, uh, really grew was their software, um, their bus, that business has grown, um, significantly driven, not just by Red Hat, because over the last few year, you know, quarters, we've been talking a lot about the Red Hat stuff, but also about their, their AI initiatives with Watts and et cetera. Those have all done very well as has transaction processing.
And, um, as I, you know, people know I keep talking about the mainframe, it's not dead, and boy, it's not dead. Um, this year they'll be releasing the next generation of the z um, and one couple of things he said, um, this is the 11th quarter of the Z 16, um, after three years, I'm quoting exactly three years in this product, has outpaced prior cycles and program to date with over 30% of the client of, um, increased over 30% of his client's capacities continue to grow. And that, I think the number that they quoted was 60% or 70% of their clients have increased their capacity.
That's also driven into some of their, um, some of the things that are going on there is, you know, the, the AI work that's going on with, uh, transformation of the COBAL technology. Um, so expect this next year as they release the next release is going to be downing. Well, the other thing that I lo loved what they said, and um, they have had a huge momentum with Red Hat.
Um, fourth quarter was 17% growth. Um, so they continue to do double bookings, um, double digit bookings. But one of the things that they alluded to is that some of this is getting driven by VMware.
They didn't say VMware, but they said they continue to see the increase of the OpenShift, um, virtualization engagements. Um, and they didn't quote a number there, but, you know, so you and I, we've all been talking about, you know, how much is this happening? Are they really getting off of it?
Is this really changing? And, um, of course, IBM is in the core, core enterprises that is, uh, you that Well, and, and their I offering is quite, is is quite popular. And, you know, at their, at their recent analyst day, I think, uh, some of us attended, they were real clear that they had a migration offering to make that easier.
Mm-hmm. That some sim CIOs wanted to take them up on that. Yep.
So that's done really well. So expect, um, you know, next year they are expecting a banner year. I mean, they've got a huge book of business on, on the AI side of the house, um, and, uh, is both on the software and, and the consulting business that's going on.
The only thing that, yeah, so, so I'll stop there, but that's really, really good. Well, no Mean, I just finished a major piece of research looking at all the different, uh, uh, a agentic ai, um, uh, frameworks and solutions include a Watson X was included in that, and they came out, they scored, I think, uh, number two, number three, uh, across the whole field. Uh, and this is a stark, uh, departure from, you know, kind of their, their very flad results with, uh, with Watson way back in the day.
Uh, Keith has, I mean, looking at all these results and, you know, people wanting to put their, you know, run their cloud on IBM and use their ai is, is IBM back? I mean, this is, yeah, they seem like they were heading towards the, the long good night, you know? Yeah, yeah, Yeah.
I think IBM has always had a seat at the table, right? The, they bought PWCs, uh, legacy consulting business, and so they've always been strategic. When you have a RFP for a ERP solution, IBM is going to be at the top of, if not providing the a solution, implementing, implementing and process change.
I've long said that the value of AI is in, in the ability of solution integrators to be able to help with the change. And IBM is just, you know, they have it within their fabric to understand both the technology and the, uh, process change needed. So it's not a surprise.
I think a lot of us who've watched the industry for a long time have said that IBM is probably one of the best convents position to capitalize off of ai, and we're starting to see them get out of their own way, quite frankly, and, and telling their their cloud story and their end-to-end value prop. One of the line items that, um, Arvid talked about, I think it was him that responded to the q and a from the investors, was that they, they talked about the smaller models, and I think this is a little bit of, you know, we, we were gonna talk a slightly about deep seek, so you may wanna go to that sec that now, but talked about the value of the smaller models and that they have seen a 60%, um, cost, um, or efficiency number and using the smaller models for inference engine and inferencing. And, you know, that's, you know, all the noise on the deep six things is what the cost about it was.
And we talked about the speed of it, but what the cost around it was that we talked last week and, and this is, you know, again, that innovation that we're gonna see that's gonna drive down the cost of delivering of what we're doing with ai. And I think IBM has, you know, got some, got a, you know, definitely is in the game. 0 has got, uh, very close to the performance of the top models, but one can cost, right?
So they're, and, and you know, we, we see that it budgets right now are, are, are ballooning because of the expectations around what it takes to run ai. And so IBM is saying, Hey, don't, you don't have to take that hit if you, if you work with us, which I think is, is, uh, a really smart move. Although I think, uh, you know, trying to try to get to the, you know, top spot, uh, uh, based on pricing is not necessarily always the smartest move.
Yeah, Yeah, yeah. And, and, and again, I I wanna reiterate that Iny is not going to take the GPU and power requirements of training. Uh, I think we, uh, constantly, we just did AI field day six, I should have put that in the rundown, but, uh, as part of AI field day six, we talked to VMware member and a small company, comma, that helps us internally with some of our signal 65 testing.
And the consistent theme is coming out that customers are still asking the question, how many GPUs do I need? We'll get toe, but I think IBM granted is doubling down on this idea that, you know what the, on training, yes, you're going to need the multiple megawatt or the one megawatt rack at some point, but for infra small models, uh, more traditional air queue systems are what is probably gonna win a day in the enterprise. Yeah.
Well, let's, uh, let's do Seagate real quick, and then let's dive into, uh, what's happening with deep Seeq. Well, this, there was two, um, companies that announced earnings in the last week on hard drives, and I bring that up, um, also because I talked about mainframe, which is not dead yet. HDDs are not dead yet either.
And that was Seagate and, um, Western Digital Seagate stopped, jumped, had a big jump, um, from it, you know, and it's kind of like ours is, who cares about hds? Um, and, uh, WDC was kind of like, uh, not that great because they have two parts of their business. One is the flash side, which is not doing well, and the other is the HGD side.
So why are HDS still in play? Because all the enterprise, if we're talking the enterprise, they're migrating to, you know, solid state. Well, it's a big portion of that is what's going on with the cloud.
And so much of the data is on HDDs. They are not moving off those devices. Um, and they are shipping them.
So the talking about, you know, as soon as they start talking about opening more data centers, well that's going to, you know, they're just rolling them into big, huge racks of hds. I've got a friend of mine that's a sales, you know, sales rep for them. Do we talk about, you know, you're just gonna roll these big huge racks of rise in there and they're gonna, you know, deploy them themselves, et cetera.
Um, so that's what's happening there. Um, it's not dead yet. Yeah.
Well, I think the yet a little bit of life left and them given their huge patent portfolio in storage and things like that, so Yeah. But we're also saying the enterprise is still buying hybrid systems. I mean, I just had a call with NetApp and they were talking about, you know, their hybrid systems, the solid state and hard drive for secondary storage is still very alive and well, it's, well, I think it's well over 50% of their business still in terms of that, you know, the, the, um, on tap arrays that they're shipping.
So, well, I, I Think we Really forget, I mean, I, I would just talking to, uh, some CEOs in the Midwest, uh, so CIOs in the Midwest, and it was surprising how, how very much, uh, in the two thousands, they, they still are, right? You know? Mm-hmm.
Yeah. So if, if we dig into it, like computers becoming dense, uh, we're seeing it in GPUs, et cetera. And the, the, one of the things that we're battling from a architecture perspective, how much, how do I, how much do I sacrifice in space?
Obviously SSDs are denser than HDDs, but from a practical perspective, if the rest of my computer's shrinking, if, uh, what do I care about performance, uh, uh, energy costs, et cetera, there is a argument that if I'm, if I have SS, Rackspace and HDD costs are lower per terabyte than SSD, okay? So this is something that I think we'll have to ping our signal 65 peers on, you know, what, where, where's the value, where's the real value point of SSDs versus HDDs, and help customers make better purchasing these decisions. 'cause I think, I don't think this is a obvious, you know, price per, uh, terabyte question.
It is a question of, you know, how much is it costing me and what's my operational overhead to run one stack over the other? Well, you could get me Louis on here from Signal six five, because he's done that Ana some of that analysis and, um, it's, you know, we, it's, it's an interesting analysis in terms of what you can pull, you know, what you can do in the data center. So, yeah.
Yeah. Well, the, so the, um, you know, we're, we're talking about data and, um, you know, the, the big cyber from last week, uh, is, is still ongoing. And that is, um, the arrival of, of, you know, highly capable frontier model that apparently, uh, doesn't take a lot of, uh, uh, compute, uh, to, uh, for, to use the whole question now, still centers around, uh, how did deep seeq, uh, create its model.
And that's the, the huge debate. Uh, and, you know, some data came to light that, um, 20% of all of NVIDIA's, uh, GPUs go to Singapore, which is probably highly unlikely. Um, and, and that's probably the vector.
Uh, you know, that's the, the, the, the, the forward shipping location for GPUs into China. Uh, and so, um, that's, that, that's been top of mind, um, for a lot of folks, um, because it's really greatly impacted, um, uh, the stock market, um, VC investment, um, a a bunch of things on, uh, in terms of, you know, where AI is really gonna go, and it's coming clear as we really analyzed, um, deep Seek didn't use Cuda. That was, I think, one of the most interesting revelations.
They, they went bare metal, uh, for their AI stack. Uh, so they could do all kinds of, of, uh, of super optimizations to make inference really, really cheap. The, the list, I'm tracking the list of the optimizations that Deep C has on, on ai, um, and I keep saying longer and longer, and a lot of it's because they said, well, we're not gonna use that.
We're, we're, we're gonna do re really extreme optimizations, uh, in, um, our use of these GPUs to squeeze out every last, uh, bit of capability. Uh, what are you guys seeing in terms of, uh, how, how deep seek is kind of recalibrating the AI landscape? Yeah, one of the things that I'm paying attention to is, again, the interesting from a technical level, really interesting tweet or ex post whatever we call tweets anymore on Twitter, is, uh, uh, one of the engineers, uh, at large engineers posted the configuration for running deepsea on dual processor epi, uh, a MD Epic CPUs.
So deep seek is memory bound. So we talk about the importance of, you know, GPU versus CPU architecture. And it seems that the biggest factor is in its recursive way that it does reasoning, is the amount of ramp.
So this, you know, 700 billion parameter model acts like a 45, uh, billion parameter model. So you can run it on systems with less oomph. This gets you about, you know, this dual CPUs, uh, air cooled, uh, system with no, uh, GPUs 768 gig of RAM looks very much like a typical virtualization server.
I have 'em in my lab to run deep seat at about 10 tokens perspective, which is not pretty usable from a production perspective, but from a lab perspective, very achievable. You put, uh, uh, Luke Norris from, uh, Kaza noted that if you ran this on one of the digits, uh, developer platforms that Nvidia announcement covered a few weeks ago, you can get a few thousand tokens per second with deep seat. So that is, you know, oh, one level A GPT oh one level, uh, That, that, that, that's thing you run top of the line ai, right?
A single machine and right in front of you, right? That, that is on your desktop, that on essentially a desktop that is game changing, which goes to the argument around what IBM is saying around, granted, you don't need for a trusted model that does the job for your them to use case. You don't need big GPUs to do this.
And that's, and the thing, what I, I think I said this last week, um, scarcity is a mother of invention. And when you're scared when things are scarce, we now all, you know, we start optimizing, you know, which we were talking about. What are they, you know, you're gonna look at all the areas that you need to optimize in order to deliver on.
And so, you know, this is, uh, we're gonna see that in many, many places, especially as we saw, you know, the dollar signs on this. My concern around that is what does that mean to these forecasts that we're having about how much it's gonna drive? Well, then I go back to saying, you know, I can remember when we would start shipping PCs or whatever with this massive capacity, how you would ever use all that capacity.
Well, you know, we always do, you know, if, so, one scarcity is a mother invention, and we will use whatever capacity we have, you know, darn it. I used all my closets in my, in my very, very large house to do it. I think that's a great proof point.
I mean, the, the real issue is gonna be is how long is Deepsea gonna be with us? Just because yeah. You know, Italy has now blocked deep seek, uh, the Department of Defense is blocked deep seek, and it may be the hitting the way of TikTok, you know, um, just because, yeah, go ahead.
I Think the big difference is that you can run it locally. So unlike, you know, unlike TikTok, unlike see some of these other Chinese platforms, it's open source. I can just download it and run it and, you know, on my desktop, literally, you, it's sure it doesn't have a call Home button.
You, Well, That's, you sure it does Have a call home button you don't know about? Well, you know, we, we have some pretty, uh, good computer scientists here. We'll figure that out.
But I, I, I don't think users are really worried. Uh, we, there's what, a hundred million TikTok users in the us? I don't think we really worry much about our data being extricated to the C ct.
We probably should, but we don't. Yeah. Well, is gonna be, is the big story for 2025, and AI for sure is, uh, is this the commod, uh, the long expected commoditization of, uh, generative ai.
Uh, and this will, of course, a real rethink. Uh, you know, we talked about Project Stargate, uh, $500 billion to train these models. And so if, if, if they've got the shortcut to, to creating these very lean, you know, uh, do lean training, lean inference, it's gonna be, it's gonna change a lot of, a lot of things.
And it could put China in the leader, uh, position, uh, with AI by the end of the year. That's, that's the real concern, I guess, in a lot of quarters. Alright, let's, let's move on to, um, uh, uh, other news.
Uh, uh, so, uh, Intel, um, you know, man, um, you know, speaking of chip companies that, um, have had some challenges, uh, recently, um, they had some, uh, delays, um, and some product cancellations, uh, announced, uh, yesterday. So Falcon Shores, which is Intel's much touted, um, XPU, which is a combination of CPU and GPU and a, a single unified package, uh, they decided not to run with it as a product. They're gonna only use it as, uh, a test chip inside the company.
And this, this has a lot of raises, a lot of questions about, uh, um, Intel's, um, capability in terms of really competing in the, in the AI era. So, uh, I don't know, Keith, you, do you have a, any reaction to that news in particular? Yeah, the, I think the architecture has promised the big GPU farms that, again, we see in meta, we see in Z and we see in across all of the big producers or models, we're not gonna replicate that in the enterprise.
So we need some something in between, uh, a, a, a small, a service provider actually in Australia noted that after deep seek, they saw a doubling in request for a MD, uh, accelerators based on the demand for deep seeq, I think we are going to see these market shifts on how we use accelerators in a data center, whether these are on chip, uh, I think CPU on chip, uh, accelerators route, while stuff like, uh, running deeps seeq on two Epic or Intel Xon CPUs is interesting. It's not what's going to happen. We need some combination of, of Accelerator and CPU, and I think Intel is deciding that Falcon, Falcon Shores, isn't it?
I, I, I hadn't even heard of it until it was killed, but I do believe, uh, Intel is going to return back to its days of the, the, the specialty asic, the programmable ASIC combination of that and CPU. It was the right model, probably a little early, but, we'll, I think we'll see the return of it in some, in some way. Yeah, I, I've been bullish for on FPGAs for a long time, um, uh, just because they have such potential to, to, you know, create a hardware performance, uh, for software based applications.
8 nanometer process, which is just amazing to me. I didn't, I don't know how many more nanometers are, were to shave off the process. That's astounding.
Um, uh, and I, I don't think it was entirely, uh, unexpected there. The current Zs have 288 cores or the, the versions that'll be released this year. So I was bringing it a little bit close in, uh, you know, if they were gonna, they were gonna ship this, uh, first quarter, 2026.
Now it's pushed off to, to later in 2026. Uh, and it might be later than that, but that's, uh, that's, that is an amazing advance that there, that Intel's, uh, bringing on the, you know, the ultra advanced, uh, process, uh, ahead of much of this competition. So, um, so let's, uh, let's talk about, uh, Kimberly, you had some, some, um, news, uh, around, um, uh, edge computing infrastructure.
Uh, why tell us about That. I Mean, that's actually, uh, Keith, he's, he's running with the edge stuff. Yeah.
So he's very, the edge, uh, Elon Musk, and we don't miss Elon much, Musk, Musk on this, uh, on this podcast, but Elon announced or mentioned that, uh, unfortunately if you bought a Tesla pre 20 22, 20 23, uh, and you were expecting, uh, full self-driving FSD, you'll need, you'll need to upgrade the computer system. 0 been a lot of controversy around this promise of kind of full self-driving, being able to be just a software download. 0.
It not So infrastructure matters. So infrastructure does matter. So it, it is a question about an example of lifecycle.
While I bullish on the idea of running deep seek on CPUs, that's probably a little bit aggressive to what's going to actually happen at the edge. So as we push AI further out to the edge, we really need to consider what our, what is our lifecycle strategy around managing those physical components? Because as we depend on our business processes being enhanced by AI and more compute at the Edge, software will eventually outstrip the hardware.
Do you know why Tesla is saying you have to upgrade or to, I guess, swap out the hardware that's in there and what, Yeah, it just, they underestimated the, the hardware needs of, uh, FSDI think this is, uh, yet another lesson. This is something that as advisors, if we always talk about, don't buy hardware or software based on the potential, what it can do, but the value that it adds today, because that potential may never come throughout the depreciation life cycle of the equipment. And this is just a classic example of buying on the promise versus buying on the actual capability.
So it, I, I'm assuming what this was is the processing power was not strong enough. It wasn't strong enough or process The amount of data that they're gonna need to process, um, whether it's memory or processing power or whatever that's going on there. Okay, yeah.
Whatever CPU that's needed to, uh, process the real time data coming from off the cameras, it just wasn't enough based on, you know, their continued development of FSD software. So that gets us into, you know, we've talked for many years and the market has talked is about the software defined data center mm-hmm. The software defined world.
Everything's software, everything's software. And it's kind of like, yes, but you still need to have the m behind it to drive it. Literally, literally in this situation, And this is why I've been fairly bullish on FPGA, because what if, what if you could recompose the hardware?
That's, that's the thing. If you come up with superior designs, uh, you can just reprogram every, everything down at the right, right, on the chip, uh, to run more efficiently without having to ship new chips. So that, that's, that's, you know, if you're doing really softer defined, that would, that would be the, the, the ultimate potential.
But we're not, we're not there yet. And, um, you know, I I, I'm glad to see you guys actually run recompile C code now into FPGAs pretty easily. But, but, um, yeah, uh, yeah, it's interesting.
And we Deep seek is that example, right? They pa bypass the abstraction layers of Kuda, Kuda iss very convenient for developing, you know, platforms on top of Nvidia chips. And these researchers say, you know what, we're going to go the old assembly level language map, uh, route, and we're gonna program directly to the hardware to get out every inch of, of efficiency.
So a hundred percent efficiency versus 60 to 70%, 70% would be outstanding for GPU performance. So they're claiming a hundred percent efficiency. So from the networking, storage, and compute capabilities, you can only do that by going as low level as possible.
And to your point, Diane FPGAs allow you at least the option to do that. So this gets into, um, 'cause I think of the life cycle of what you're talking about with Tesla. We're only talking, I mean, are we getting into a point now?
We're looking at a three year life cycle on this. I mean, that's costly. Then I have to look at an ROI very, very differently because I've been typically on any kind of server deployment, I'm looking at us and I use, is it five or three, five years on a server, or is it, It typically five years?
So the five years before the server becomes free, I mean it, that basically, and we then we, we, we, we, we sweat those assets for as long as we can. If it's running and we have support on it, why get, why, uh, uh, why accumulate additional depreciation costs? And the, I think to your point, the bigger point is lifecycle management.
We've set up our environments in which we can switch out this hardware on a rolling five year schedule. If now I have to go down to three years or less, that's a good challenge for the enterprise, Right? And I start to think about, okay, so what use cases are we talking about this hitting?
Okay, let's start with retail. And you know, I, I've worked retail, I understand, you know, the pain of, of train, you know, doing a train, you know, doing an upgrade into all those stores. You're talking thousands of stores or something.
Um, I'm Hope Depot, I'm implementing, If you wanna move to, let's say, you know, the, the ai, uh, powered, you know, uh, checkout list, uh, uh, checkout, which is you can just, there you go, grab the items and go out. But that requires spatial recognition, item recognition, uh, a connection to the ERP system, the all that, all those types of things in real time. Uh, and, uh, you know, that that will consume, consume a lot of power.
That's the kind of use case we'll say. Uh, that's a step function up to a whole new level of compute for an average retail organization. Yeah, It reminds me, I did reached back out to the chief architect of Chick-fil-A a few years ago.
They rolled out Intel nooks, uh, to run Kubernetes at every store. So mm-hmm. Every store was a software platform.
Well, that's the Intel Nook is not even offered anymore. So how are they managing, uh, that platform lifecycle out at the edge is a really, it, it, it raises a really interesting question Yeah. What to do.
Well, that brings us, uh, to time for, uh, this episode. Um, and so thanks for, uh, sticking with us, uh, this long, uh, next episode's, episode 70, and, and we will, uh, we'll see you next week.