Are the Hyperscalers AI-Capacity Constrained? Gaining Efficiency with New LLM Releases – Infrastructure Matters EP70
Google, Azure and AWS all claim AI capacity constraint — is it impacting their growth? What is behind this? Plus, more on the latest LLM and other models from Google, DeepSeek and players from Stanford and UWA. Lastly, we get into file systems for AI and what we are seeing from key and new players: Weka, Hammerspace, NetApp, Dell and Vast.
Transcript
Good morning folks, and welcome to Infrastructure Matters, episode 70. We are back again with your favorite friends, Keith Townsend and Diane Hinchcliffe. We are gonna be talking all earnings.
Um, we're gonna be talking earnings as it really pertains to this Google Clouds or the Cloud. Not this Google Cloud, but the cloud stuff and ai. Um, also we're, uh, gonna get into some kind of really cool stuff that's happening with the AI in terms of where the cost is and what's being going on.
So with that kicking off, I think, um, Diane, you had some things to talk about with, in terms of this, uh, the earnings, what we're hearing about. Well, the, um, you know, we, we heard from Amazon's earnings, uh, uh, yesterday. Uh, and they, and they came in just a tiny, tiny bit, uh, under estimates and, uh, getting really punished for that.
Um, and we saw similar things with, uh, Microsoft and Google. Uh, they're, they're, they're just struggling to meet cloud demand is the story that we're hearing from folks like Andy Jessi, um, from Microsoft CFO, Amy Hood, saying that they just don't have the space, the, the capacity, um, the AI demand is, is, you know, off the charts and they simply can't build, um, enough. And Amazon's reporting yesterday said they're gonna put a hundred billion dollars in the CapEx this year.
That's the biggest number we've heard. Uh, Microsoft, uh, is, uh, saying 80, Google is saying 75 billion, that's a quarter trillion dollars in CapEx. Um, which as I go really digging into it, uh, about half of that is AI infrastructure.
Yeah. Um, and so this, this really, you know, there, there's this narrative emerging that the hyperscalers can't keep up where, you know, but we talked to a lot of practitioners and it's not clear that, um, there, there, there really is that demand there. This is maybe a lot very speculative.
Keith, I, I know that you had some, you know, a sharp point of view on this. Yeah. So, you know, I, I've used both Gemini and copilot as desktop assistants.
I'm going to say not good. The, the, the promise is the, uh, the capability isn't meeting the promise. Both of these companies are asking quite a bit for it.
When we talk in our C-I-O-C-T-O panels, uh, most of the CTOs I talk to question the value for, uh, the overall capabilities. And then there's still the debate around ag agentic ai and where is, uh, these workloads going to run. David Litum, who's a long-term cloud watcher and, uh, management consultant, uh, pretty famous in the space, was, uh, taking up victory lap this morning.
He was saying a year ago, he warned that AI isn't going to be the growth engine that these companies expect because they're not building what customers want and how he assumes they will, uh, consume ai. So there's a little bit of debate. Is the, is the demand truly there, or is the numbers disappointing because the, the services aren't meeting the expectations?
Well, if I go and quote, um, directly from their, their, uh, earnings calls, um, the Google Go Google on CFO said, new capacity, um, issues, impacts revenue, you know, and they were off 30% year to year in terms of their Google Cloud. Now, specifically on the ai, I'm not sure which ones the numbers were. Um, and then if you go to the Microsoft, um, Amy, I believe his name is, is the AI capacity constraint Q3.
And so when they're giving guidance for this next quarter, even though Azure was, you know, they're expecting a 31 to 32% growth on Azure, the AI has some impact. So that's, this is kind of, you know, where they're pulling it back. I mean, it's, and I guess what it is, is because we, the, the market may be just overvalued right now or something, um, along those lines, because you think about, you know, when Azure talks about, or Google or, uh, Microsoft talks about their AI growth, it's $13 billion is what they're equating to that.
And that is up 175% year to year. So there is growth there. It's just a question of how big is this gonna be?
I was about to say a customer. Well, and is it strong growth? So when I've had, we've done some other work for other cloud vendors, um, uh, and with CIOs, and what we're learning is that, that they're giving away large amounts of promotional credits for some of these cloud ai cloud services, uh, in hopes that you'll build something around it and then have to keep paying for it.
'cause you'll like what you build and then you'll keep it. Uh, and, and, and that is what's driving this, you know, incredible demand. You know, we, we know how big the, the cloud market is now.
They're saying they can't build fast enough. They're gonna put a quarter trillion dollars. We just, by far the highest watermark ever.
Uh, is it, is it strong growth though? Is it, are they really giving away a lot of services right now on the books, has cost of sales and hoping, hoping to, to hook enterprises on these AI services? And it's, it's the total speculative bubble.
I mean, I think that's the real risk. Very. Yeah.
And, and you know, the time will bear this out, right? We will start to see the results in the market, just like when early cloud providers, Microsoft leveraged, uh, credits early on in their cloud Azure journey. And we, uh, saw the numbers, but we didn't see the use, uh, that eventually changed as the services got better.
So we could, you know, we could see a little bit of both that they're seeding the market by giving away the credits, booking that revenue as cus and they're getting ahead of the, and they're getting ahead of the demand from a capacity perspective. We don't know. We'll, well, but what time will tell Yeah, we'll find that soon enough.
Well, and that balances back to the research work that you've done on Diane, and also that we're seeing, you know, both from the C the CIO and the CEO insights work. Um, and the conversation that has been having, you know, even coming outta Davos and the, the ability to really move, um, true impactful AI technology is dependent upon on the data and that data preparation, that data cleansing that data, you know, all the things that go around making it, um, safe and, and mm-hmm. Secure Absolutely.
That slows the process down. So even if the capacity is there, or maybe if, if I, I think which, you know, which is it, you know, is it capacity that's not there to be able to do the training that you're looking to do? Or is it the data that is having problems getting to the point that we can do something about it?
Or is it, you know, purely being able to see, as you, you've talked about as the ROI of this particular applications? I think all three of those areas have potential to slow this down to what I would think is a normal pace Yeah. Record.
Well, it is a breakneck base. Well, we did see the, a clear signal, uh, in our, our CEO survey in particular that, uh, those reporting, uh, uh, a significant number of stalled or failed AI efforts, uh, 60% of those were reporting data as the primary issue. It's the, it's the leading issue, is the holding back AI in the enterprise is the data story is poor in many organizations, have been under investing.
Um, and, and they may be trying out all these AI services and discovering that their data isn't, you know, isn't there. So yeah, it'll have to be our next line of inquiry. Yeah.
And then, you know, coming up in this month, we're gonna have, um, maybe the end of this month, the end of February, we'll see the Dell announcements. And so we'll hear more about where they're going in terms of the server. We saw the IBM's announcements on their earnings and, and what they're talking about in terms of their growth.
It's been fabulous. Um, I'm more or less expecting the same thing coming outta doubt, um, in terms of how they've been performing in those areas. So I, it's when you're saying the capacity is not there, and what if we're looking at what's happening with the shipments coming out of the folks that are going to on-prem, and they're shipping as hard as much as they can, if that's what we're seeing now, I'd say that issues around the cloud, Keith, would be more along the lines of what they're saying versus they're not building the right things.
The, the other thing that's hidden in the numbers, because I think so much of AI is focused on training right now. Mm-hmm. We don't know how much a actual inference and agent adoption is happening.
Correct. So how much of the dedicated, uh, AI services are being consumed versus people coming to cloud and building their own and just using EC2 instances and CPUs to do the lightweight referencing that they need to do. So we, we do need a lot more clarity, a lot more data to actually make this, you know, make this, this, this Decision.
Well, and you bring up good point. How much of this, uh, demand is actually artificially driven by venture investment, and, and these are not necessarily your traditional enterprises. These are actually AI companies building, trying to build quickly in the cloud, um, and creating artificial, uh, you know, uh, demand that may not, may not hold, you know, depending on how all it's got.
So it's all, it's all very interesting. You know, what's gonna happen here. Well, that kind of gets me to this next topic, and we'll come back to a little bit on, well, OpenAI and that kind of things.
It's, um, on their benchmarks. 0, um, one of their models, um, for 50 bucks. I don't think the Stanford, Stanford research were being paid.
So you have to add in the cost of labor, which, you know, add a lot. Well, I'm still gonna go with my promotional credit story. I'm, I'm hearing this a lot, so, we'll, yeah.
So, um, it was a 50 bucks and somebody else said somebody else did it for, you know, 450 or whatever. So maybe one of you guys could take, take, do a little rift on what is distillation, what does that mean? Because I bet you 90% of the people that are listening to this don't know what we're talking about.
Yeah. Well, I'll be honest. I'm one of the 90, I'm one of the 90%.
What, what is distillation? Well, the distillation is, is taking some of the, the data that you have there and, um, and using some of the techniques to do the analysis on it. Um, so it's not using as much processing power as, and this is where AI o Open AI came from, um, not open ai.
I'm rather rather the, uh, deep seek came from in terms of what they were doing is, is a distillation of somebody else's model. And that's the model that they, they're assuming that Well's, right? That's the accusation that some people would say is that they, they, they distilled open AI's model down into a more efficient model.
Yeah. So my understanding is distillation is taking a large model and saying, all right, let's, let's create a smaller, more efficient model that may doesn't have everything, everything in it, but maybe aimed at a specific use case or domain or something like that. And that way you run, Uh, distillation.
I feel like Homer system, when it was says Simpson, when he looked at the word Jim and pronounced it guy, and then when, uh, he figured out what it was, he was like, oh, guy distillation. So the, so I've read this open AI paper, and it's really interesting from what they did. They mentioned that they've taken larger models and they, so they're not, um, they're not saying that they did not take, uh, uh, a larger model in distiller.
What I, what I don't think that they do in the paper is call out OpenAI directly. And I, I think this is, uh, an the amazing feat, whether they did or not. I have no, I personally, they take my data and they, and OpenAI uses my data to inform, uh, their models.
I don't have a problem with it. And I think generally speaking, none of us should have a problem with the ethics behind what this process does. But what's really interesting, the, the, the innovation is the reasoning.
And there was a lot of news, and we'll get to that around the reasoning capabilities that they then come out of with, and you come out with better models. So with this Stanford, uh, research and this Stanford Pro project, we're starting to see the fruit of the, just the advancement of AI reasoning. And is that, um, because how does that have, where, where do you see that impact then?
Yeah. So the one, one of the biggest challenges with AI is getting it. And I, I talked to, I, I think I talked to this about a year ago during, uh, my a hundred days of AI is getting AI to think like your organization.
Like, if that makes sense. The, the, a manufacturing organization thinks differently than a healthcare organization. A matter of fact, two organizations within, um, within the same business unit might think differently, or as, as they're calling it reasoning.
So how do you get a model to reason, to your, to, to the needs of your organization when you send through a prompt or you have a agent working, how do you get it to make decisions that are consistent with your, uh, with your organization's strategy? And on top of that, do it at a reasonable cost. And this, we're getting to that point where, uh, techniques just, just as distillation and, uh, advanced reasoning, the low, lower cost of retraining these models is going to have, uh, I, uh, my opinion, uh, really great advancements around getting smaller models that are able to be positioned in a way that make better decisions with, uh, within your organization.
It was also Arvin. Um, the CEO of IBM had a post on LinkedIn, which I thought was kind of an interesting thing. It was fairly long, um, talking about how there are and, and, and, and reflecting on deep seek, um, and how these models were finding new ways.
And I talked about this last week, new ways to lower the cost of development. And he reflected on how we have done this over and over and over again in our industry. You know, it's, call it Moore's Law, call it whatever you wanna call it.
But, you know, he, he cites all the other technologies that we have been faced with, and initially were extremely expensive. And then innovation brought us to this place. This seems to be moving faster there to, to that, that, that place than we expected.
The, uh, perhaps, um, but I'm gonna go back to the same thing. We still are running headlong, even though we get more efficient and less maybe possibly need to use less energy, et cetera, to get to this space, possibly being able to do all the things on somebody's A IPC as opposed to, you know, an entire data center of those things will probably end up with the big, huge, massive language models. Those language models, and then some unique other models that are unique to the industry, which we've been talking about.
And then the distillation of what this is talking about, which is very cost, cost efficient, um, and reasoning. So then the question's gonna be, how is the quality of the reasoning coming out of these guys? So we, we've seen some of the reasoning coming out a little bit interesting.
So, Well, this is what everyone's right now talking about something called, uh, uh, Devon's Paradox. Um, we're talking about the Big H hyper, uh, scalers. Um, and saying that, uh, that all of this, uh, will leads to more compute and storage, not less.
You never use less compute and storage, no matter. You get much more efficient and it's cheaper than you find all new applications and all new things that you can do with it. And that's, you know, this is the problem also with energy.
When you create cheap energy, people create, well, I got all these new applications I can do now that was not affordable before. I can now run off and do it. And so the demand never goes down, no matter how much you try and conserve, no matter how much you try and get, uh, more efficient, it just creates more demand.
'cause you can two more things. Same closets are the same problem. Yeah.
Right? Yes, exactly. Don't forget the closets.
You'll always fill them up, whatever you have. Okay. Let's, uh, have we hammered on this one enough?
If not, we can go to a couple of other topics. Yeah. Well, I have one that was very interesting, uh, is I, one of my predictions is, is that we are on the cusp of what's called super intelligence.
Um, and I, I've been, you know, promoting this idea of, uh, uh, you know, since it seems clear, if you look at the benchmarks, uh, if there is, you know, sweet, that measure the iq, the human super Intelligence is, is is cloning you, is that what that is? No, no, no. By no means.
Um, uh, and, and this is when AI is smarter, uh, uh, and than any human that that is ever lived. Mm-hmm. Uh, and that's where we are.
We are on the cusp of that. Uh, and, uh, proof of this is open AI just released a new service called Deep Research. Uh, this is a model that does extended in-depth inquiry, uh, formal in extended in-depth inquiry.
It doesn't just consult it, um, the models. It goes out there and, and looks at research. Uh, it queries the internet, um, and it does it, it can take, you know, dozens of minutes to hundreds of minutes to run a signal, uh, query.
And it's only available in the, in open AI's, uh, highest, um, uh, tier of service. And what's interesting is that, uh, it's, it, it, it, it blows the doors off of the hardest AI benchmark, which is called, um, human humanities, um, final exam, or sorry, human's Last Exam. And this is a, uh, AI benchmark created, uh, uh, uh, and all the questions on it are extremely biblical.
Like take this very hard to read, very obscure obs uh, archeological, uh, uh, uh, inscription of an, of a dead language that almost no one in the world knows and transcribe it in different, in, in these different languages. Um, and that's the kind of test it has on there. It has, uh, you know, math and geometry questions that most humans can't solve.
3%. Um, but, you know, things like oh three, uh, open AI's model passes at 10%. Well, um, deep research can, uh, pass, uh, it at it 25% by far the har the highest scoring, um, uh, model on that benchmark.
This is something that probably none of us can answer any of the questions. We probably couldn't answer a single question on that benchmark no matter how much time we were given. So, yeah, so when I saw this announcement, I'm like, oh, well, it looks like my career pivot to doing something other than research has started.
Exactly. And, uh, the, we're not quite there yet. Rera, uh, SAS Main Manian, who is a, uh, fellow independent analyst, a pretty sharp guy, uh, spent some time with it trying to get the, uh, the model to write a report type of output, not just answer questions.
And we, we, we see a little bit of a capability in the lower end model. The mini at the, uh, the o uh, O three Mini does a little bit of this on the light side, and I played around with it, and it is absolutely great at answering questions and giving you a start, but, uh, he claims Shera reclaims that he wasn't able to get it to do anything more than like a thousand word blog post. And the, and the quality wasn't that great.
It was still, uh, hallucination, it went off on tangents, didn't keep a consistent methodology. So I think it is a great indicator of what it will go to be one day. But, uh, right now it seems like much more of, again, a human in the loop type tool that gives you a good starting point of where to start your research, but not, you know, conduct deep research as if well Also raises the question of, you know, how much do these AI benchmarks really tell us about what it will do when the AI model will do for us in the real world?
Right. You know, that's the thing. Yep.
Because what you were citing, or equations and getting answers to a, an equation that you said very few humans can do, but those are known equations as well. The question is, is can it solve the equations that we can't solve? Well, and that's my theory around super intelligence is that, uh, well, yeah, it can solve, I mean, what they're arguing is that, uh, those equations are, are known, but most people can't actually do the math.
Uh, it, it's too involved with too many variables, um, mm-hmm. And, and come up with the answer, whereas supposedly this can. And so, um, the, the question is again, how useful are these benchmarks and telling us what these models are gonna do for us in the real world?
And, you know, uh, it's gonna be interesting to see, but nevertheless, uh, we see Marvels like oh three are coming at about consistently about 135 iq. Um, we're gonna see models that can do in excess of one 60, which, which now puts them well outside of, you know, uh, you know, six Sigma in terms of how, you know, how many people can, can perform at that level, or this is something that anyone can use. We're all gonna have, you know, certifiable geniuses in our pockets here, Sam, On your iPhone.
That's right. Here we go. Um, okay.
So, uh, let me switch to a couple of other things. Of course, on the data side, we have to talk about that because it's me. Yeah, of course.
Um, um, wca, um, we, we saw some of the companies that are on the data side that is particularly are in, we're still in the AI piece of it, um, that had some layoff on some people. Um, at the same time, we have a flip, you know, flip side of Hammer space, which is another, um, software, um, data, data offering, um, that announced 10 times growth. Granted, they're a fairly small company, you know, but still, that's a very, very significant number, um, to grow on and could really straighten the company.
So, um, Keith wca, you were, um, you've been plow into that one. I didn't get plowed into that because of my travel. So Yeah, so w is an interesting story, right?
Uh, we, if you follow HPC at all, you know, wca, they, uh, are, they do storage fast. So let's just that they are a Fed storage, uh, p uh, uh, system, you would think that would naturally translate to being a good system for ai. But according to our, according to reports, the overall challenge is that they didn't pivot quick enough into ai.
They have strong revenue, a hundred million dollars, uh, a hundred million dollar run rate. 6 billion, uh, based on their last, uh, round of funding. So fundamentally seems well, but they're part of this larger d debate, how much to rotate to ai and can you over rotate to AI and Wiki seems to be the case where they didn't rotate to AI enough.
Well, there's a couple things, um, that were saying that the traditional files, traditional file systems like the, the NetApp, the Dell, vast leading, leading the charge as well have gone to, which is announcing some work that they're either integrating with a Starburst or some other companies that are doing vector databases, et cetera. They're integrating with all the pieces that need to come together to do all the data management that we're talking about in terms of privacy and, and those kind of things. They're developing methodologies to import data more easily into there.
So they would be the holders of all the data. So very varying strategies, you know, vast and NetApp are building it in the systems. Um, Dell right now is partnering.
Now take that with where WCA has been, WCA has been partnering, but they haven't been partnering on that kind of level, which is probably where they need to go. And that does take some engineering and design work. And so that's what I suspect is going on with them, although I haven't had a briefing with them.
So that's gonna be my next thing here, um, to get some time to look at what, where the roadmap is going. So that could be where they're going because WCA is a parallel file system and competes very well with BGFS and GPFS, and which is IBM scale. Um, and, and some of the others in Luster.
So that's where they traditionally have played out. Now they're shifting to some other, to playing out and saying, am I gonna compete? How do I compete with Vast?
Which is, I suspect is what they're feeling. So the second part of that is where, as, as I mentioned, hammer Space has taken off, um, and is doing very well. And the reason why they're doing well is this chair tier zero strategy that they have, which is not dram, um, dram, it's tier zero for them, is what they're doing is they're bringing their capabilities, um, and integrating their metadata management into gr bringing this very, very fast tier on top of a regular file system.
So they're acting like a parallel file system, and they're bringing the speed of what WCA brings to the table, along with some of the capabilities of a regular file system and an ease of use of those environments. And so that seems to be really taking off, um, also with our partnerships that are out there. So, we'll, we'll watch this space, as I said, you know, my predictions that this year is going to be the year that we're gonna talk about how do we do scale out and what are the differences between these kind of technologies and why WCA is doing this, why Hammer is succeeding, why, you know, we're, we're seeing all these, these investment areas that are going on.
So Interesting times, uh, in infrastructure for sure. Yes. Very interesting.
So with that guys, do we have any, uh, anything else? I think I've got my list here. We've been through it.
We're kind of down at the bottom of the hour. Um, We, uh, sorry, ke let's go real quick. 0, their most advanced model include, oh, Forgot about that one.
Yeah. Yeah. And inclusive also, uh, uh, Gemini Flash, um, which, uh, smaller model, um, you know, able to do things, uh, cheaply because we are seeing, this is where IBM for example, competes is with, or granite is being able to provide high quality answers at much lower price points.
0 has been at the very high end of the benchmarks, but it hasn't been out yet. This puts Google now with a, with a shipping AI that is at the, uh, at the very, you know, in the top three slot in, uh, on AI benchmarks. So Did they.
And then just from a tactical perspective, I, I've played around with Google's runtime engine, and so GKE and using that to call the services, it is a really interesting dichotomy of approaches. You can use Gemini to, you know, run curies like you would chat GTP, but less functional from a capability perspective. But the platform to develop applications top-notch, it's, it's really interesting how easily you can call Gemini and other models from Google's runtime.
So again, the competition of how to provide these platforms and how to enable 'em, the, uh, uh, that I think in some of, uh, my internal discussions folks have called LLMs Commodities. I don't know if it'll go that far yet. No, not Yet, but heading That.
But, but the, uh, platforms definitely are not in this ability to consume these models. I think that is a, a, a battle that we should probably talk about one day. Yes, indeed.
Maybe next week while I'm on vacation. I like how you do. It's fine.
We'll have to duke it out. Yeah, you'll duke it out. All right guys, thank you very much for joining us on this Infrastructure Matters episode number 70.
Um, we will be seeing you. Well, I won't see you next week. The two gentlemen will see you next week.
I'm on vacation. So with that, thank you very much and have a great week.