AI Infrastructure Wars: AWS, IBM, Nvidia, and the Race for Scale | Utilizing AI Ep. 7
AI infrastructure is turning into the primary battleground for the next phase of enterprise AI, and the biggest players are moving fast. In this episode of Utilizing AI, Stephen Foskett, Brad Shimmin, and Nick Patience break down the latest strategic developments from AWS, IBM, and Nvidia and what they signal about where the market is heading.
The conversation opens with AWS’s push to expand both models and infrastructure, including Nova models and the broader concept of “AI factories.” The panel discusses how cloud providers are packaging compute, orchestration, and model access into repeatable platforms designed to shorten deployment cycles and lock in enterprise workloads.
The episode then turns to IBM’s strategy, including its acquisition of Confluent and what that means for enterprise AI readiness. The panel highlights how the ability to move, manage, and govern data in real time is becoming a competitive advantage, not a back-office detail. As AI applications move from pilots to production, data pipelines and operational integrity matter as much as model quality.
Nvidia’s direction rounds out the episode, with a focus on optimizing mixture-of-experts architectures to improve performance and efficiency. The panel discusses why these efficiency gains are increasingly important as organizations try to scale inference without runaway infrastructure costs.
Beyond company announcements, the group explores the evolving economics of AI, including the idea that models are becoming loss leaders used to drive adoption of platforms, tooling, and ecosystems. The key takeaway is simple: winners won’t just ship models, they’ll deliver complete, integrated environments that enterprises can actually run at scale.
Transcript
AWS made major announcements at Reinvent, including the Nova two family of models, inclu and Sonic for speech, as well as emerging AI factories featuring Tanium custom Silicon and S3 vector storage as they try to stake a claim in the AI infrastructure market. Companies like AWS are building deep infrastructure capabilities and platforms like SageMaker and Bedrock are delivering AI powered applications as our companies like Google and Microsoft. IBM also made waves this week with their announced acquisition of Confluent, which is yet another modern AI application arrow in the quiver of Big Blue.
And Nvidia told us a little bit more about how to use a model of mixture of experts designs to increase performance, all that and more on this episode of utilizing ai. Welcome to utilizing ai, the podcast focused on practical applications of artificial intelligence from the Futurum group. Each episode brings together diverse perspectives to explore news and use cases in the ways in which AI is transforming enterprise IT and the industries it serves.
I'm your host, Stephen Foskett, president of the Tech Field, a business unit here at the Futurum Group. And before we get started with today's discussion, let's meet who's joining me on the panel today. Steven, thanks for having me.
Hi, everyone. Brad Shimmin. I am an analyst with the Futurum Group looking at data intelligence, analytics, and infrastructure.
And I'm Nick Patience. I'm the AI platform's practice lead at Futurum focused, uh, solely on AI really. And as mentioned, I'm Steven FoST from the Tech Field Day Business Unit.
I, uh, host the AI Field Day events, and of course, uh, this podcast, I, I have to actually say I am here in at the New York Horological Society, which is hosting me. And if you're wondering what that is, Google it. It's amazing.
So let's dive right in. Um, I'm surrounded by antiquities here, but we've got a lot of cool, uh, modern ultra cool stuff happening here in the future. Um, we just got done with AWS Reinvent.
Now we're not a news podcast and we don't wanna make it sound like we're, um, you know, kind of reporting the latest. But Nick and Brad, you all have, uh, had some time now to digest the announcements from Reinvent. Um, perhaps, uh, you can regurgitate a little bit of that knowledge.
Uh, I'm gonna not stretch this metaphor anymore. Um, Nick, thank you. Uh, what was announced at Reinvent and what was interesting to you?
So, yeah, we, we, um, Brad and I were there last week and, um, we wrote a, a note for a few from clients. I actually wrote the headline. And the headline was, um, wrestling Back AI Leadership.
And I think that kind of sums up where we are the end of 2025. Um, it's been, I guess, fairly well known that, uh, Google, at least from a narrative point of view, has kind of, you know, taken a bit of a lead in outta the three major hyperscalers. Uh, and, and AWS is, you know, uses reinvent every year to do, to make these major announcements, obviously.
Um, and, you know, this was, this was more about that. So I guess, you know, one of the things were the NOVA two family of models, um, including the, um, yeah, there was Nova, uh, Nova, Nova two Pro for Advanced Reasoning, um, our Omni for, you know, long context, um, workloads. And then Sonic, which was the speech to speech model, which was pretty impressive, that last one.
Um, and because Amazon, um, AWS rather has its contact center applications, which embeds, you know, which some well-known content center apps out there are built on. And I think that that's a, you know, is obviously a real proving ground, um, for that kind of stuff. So I think it's, I mean, I, I certainly believe, you know, it's fair to say that a Amazon HA wasn't at the cutting edge of models.
Um, and Nova was released, announced that last year's invented, this is NOVA two. Um, you know, and these weren't, you know, these are, these are strong models, but I think they, yeah, each one of these kind of, um, vendors are gonna find a, a niche. Um, there, I think the, the other interesting thing I thought was the, um, AI factories, which is essentially the, um, Amazon's infrastructure on the, in, in the client's data center.
Um, and as part of that, they announced, um, you know, theran Ultra Servers, um, based train Train two, um, they also announced Tanium three, the, the, the chip and then roadmap for four. Um, so it was very much, I think from my point of view anyway, it looked, um, you know, like Amazon on a AWS trying to own the kind of AI infrastructure narrative at least. Um, there's lots of other things and there's lots of other things there.
Announced to maybe Brad, you can, uh, talk about a couple of the others. Yeah, for me, um, it, you know, just building on what you're talking about with Nova, we all know that, you know, AWS is, is not going to outdo anthropic and even really Google, uh, with Gem, the Gemini family in terms of outright performance. And I, I I to admit that when we were, when we first got to the show and they were talking about everything being Frontier Scale, you know, with a Capital F and air quotes around it, I, I was a little, you know, uh, taken aback and, and thought, nah, come on guys.
You're, you're not really doing Frontier scale models, you're just ruining the name and it's really not true. Uh, they're actually building frontier scale models in terms of being multimodal as one checkbox and having super long context windows at a million tokens, and another checkbox. And, and I think that when you look at that, um, coupled with another announcement that came out, uh, and that is their, uh, what they call Amazon Nova Forge, which is this basically a facility instead of tools of what Amazon likes to call recipes, uh, that you can use to fine tune the Nova family of models, and not just in your basic, you know, here's the open weights models, good luck, you know, because we, we've been doing that for years now in terms of fine tuning and instruct tuning and, um, aligning models in the training during the training process and after that.
And what they're doing is, is kind of unique in that they're productizing operationalizing those mechanisms, uh, those data science techniques to make 'em a lot more accessible to the enterprise marketplace. And they're opening up the Nova models to, to actually let you, um, sort of work with those as though they were your own. So they're letting companies basically grab different checkpoints.
Checkpoints are steps along the way to creating this final frontier model so that companies can more readily bring their own data to those models. And that's, that's pretty cool. And very quickly, uh, I wanna say one other thing that I really caught my attention, and, uh, this is, again, me being dragged forward because, uh, you know, I, I've been, I'm a database guy and I, and I think database management systems are kind of important, and yet, uh, we are seeing a very sizable trend in the industry towards pushing down database functionality to the storage layer itself.
And on Amazon, it's the AWS S3 layer, so their object storage, and they released it to ga their, um, AWS S3 vectors capability, which is, as you might imagine, a vector database. One that can do, you know, super huge indexing tasks and do so at scale with a comparatively very small cost footprint, which is quite impressive. And one of the reasons why you might want to push that functionality down to the S3 layer.
And if you couple that with an announcement they made earlier in the year around S3 tables, which is basically to bring structured data to the, you know, non-structured na nature of, of object storage, you've got yourself kind of a database, you know, disguised as an, as a file system. So it's, it's really interesting times. So just picking up Brad on the Nova Forge thing, I think you're right.
I think that is at the moment, unique. You could do that, couldn't you? And like Google Model Garden and, and things like that, you could tie it all together and Microsoft, you could do the same thing and maybe IBM as well, but I think, um, it's, do you think, I mean, I, I kind of think this will be copied by the others by its direct rivals pretty within, you know, I would be surprised if it takes them a couple more than a couple of quarters, but they'll do it.
But for now, it's the only, it's the, it's the only sort of productized way of doing what they're trying to do, isn't it? Yeah. Everything else has been data science, you know, open up a note Jupyter Notebook and grab your favorite PyTorch version and have at it.
And that's not something that's accessible to every company. So if I can Ask your opinion Experience, uh, yeah, Lemme, lemme ask your opinion on the, um, the direction that A AWS is going here and, and what we can learn from that. I see, um, something of a similarity as well.
You, you know, you mentioned Google, of course, uh, they're have a similar approach to this. Am I reading this wrong, or are they seeing, you know, companies like AWS and Google that have a huge infrastructure CapEx investment? Are they seeing, uh, models as almost a loss leader to attract people to their environments or to enable, uh, the build out of this, uh, new industry, which I guess that could be a second point?
Or are they seeing these as products that they will make revenue from? I personally think it's the former, I, I've long thought voice thought it as the former. Really, the, the models are, are not models are not, the application models are not the product.
Um, they are very much a a means to an end. They get the attention as, as Brad knows, I mean, every single time any major company releases a model, journalists will always be asking us about that. And sometimes I'm kind of wishing they'd ask us something about something else.
Um, but, you know, I think they are very much, um, a loss leader and then you build the value around that, um, either below it with silicon or on top of it with the various tools and then applications. I, I dunno what you think, Brett. Yeah, I feel the same way.
You know, we were just talking about before we came on air that, you know, whether you can believe this or not, but, uh, it is, has been studied and argued that, uh, perhaps, you know, the further we get with the transformer architecture, the less differentiation we're gonna see between, you know, frontier scale models and actually all, all scale models, simply because they all converge around the same patterns that they're trained on. So Yeah. Is Nick's totally right?
It's it's really just a part of the tool chain and it's, it's an important part, but it is just one. And we're seeing, we have seen a lot of effort from all of the vendors that we've been talking about in terms of trying to at least standardize on how you interact with those models from an API perspective, so that whatever tooling I'm using in that tool chain, you know, I I can bring in the model that's most cost performant, accurate, efficient for me. I mean, that's not to say at the moment, you know, Google, for instance, with Gemini, and that's the only place you can get hold of Gemini, uh, does, yeah.
The model that is, um, does have, yeah, that, that kind of, that kind of appeal. Um, and I think it, but I think it waxes and wanes. I mean, everybody's got a model garden, a model, um, switch switchboard.
I like to think, you know, you know, how am I, how my, how my, I direct your prompt. Um, do you want this model, that model, um, model. So that that ability to, to offer, you know, first and third party models usually isn't much of a differentiator, uh, unless your model is something, you know, quite spectacular.
Um, and I think those, the, the kind of window for being spectacular could last days, maybe weeks, um, if you're, if you're lucky, and then it, then it shot. So what does this mean for the rest of the industry then, if, if, if models are loss leaders, what does this mean for companies like philanthropic and OpenAI? Um, and, and I have an idea, um, I think that Anthropic and OpenAI are trying, they, they don't, so let's, let's back up again.
So AWS and Google understand that infrastructure is, has always been where the money is, the revenue, the bulk of the revenue. And I think that, um, a product like S3, for example, is a very sticky revenue magnet in a way that very few things in our industry are. And so it makes wonderful sense that Amazon would want companies to have more data in S3 to make S3 more AI friendly and attract people to using Amazon's infrastructure to build their applications.
I think that OpenAI and Anthropic, it looks like what they're trying to do is instead build a platform. You mentioned sort of that app store kind of approach. I think that's kind of where those guys are going.
Or am I missing that, uh, point as well? And, and our Amazon, is Amazon playing there too? Yeah, I, I think, yes, I think they are.
I mean, OpenAI I think is trying to build, um, a completely vertically integrated technology company from chips through, you know, all the software layers we've talked about, um, to devices, um, and everything. You know, you're working with Johnny, Ivan, all that kind of stuff and everything in between. And then anthropic, you know, purpose to a slightly less, um, lesser extent, but still, you know, they're building out tools.
Um, so yeah, I think they're trying to build new software companies essentially. Um, whether they, whether they will succeed or not, we will see and we'll be trying to track it very, uh, very closely. But I think that's, um, I think that's, that's what they're trying to do anyway.
That's my opinion, Brad. Yeah, I feel, I feel the same way in terms of, you know, these companies are, have been for quite some time trying to build beyond the model itself. And you look at how Anthropic Claude family has evolved in that regard, and you can see that what they've been focusing on is how do you build the attendant tooling around, you know, using their models.
You know, how do you get from just a chat bot to a full fledged agentic process running in your line of business that's, you know, where they're building for. And that means, as Nick just said, building out a platform. Everybody's got a platform.
So, so that sounds more like a, um, like a, a a, a Salesforce or a, a ServiceNow kind of play as opposed to more of a traditional enterprise tech kind of play. And you asked about aws, are they kind of trying to do that? Yeah, I mean, SageMaker has what hundreds of thousands of, of, of customers and, um, has been around for a very long time.
And Bedrock seems to be doing pretty well, um, as well. So I think they, yeah, they all are doing it. It's, it's obviously leads to, it's, it's sticky, isn't it?
It like it leads to, you know, getting out of those platforms once you're, once you're in them, um, to that level of depth. And S3 is the best example, isn't it? Well, the first kind of cloud, well, the first cloud storage product around.
And, um, you know, when you talk to, when you talk privately to some, you know, people from some of these companies, it's like anything we can do to get people to, you know, put more stuff in S3 or in Google's case use BigQuery, which is a very successful product, um, is, is an absolute, um, is a, is a kind of mother load for them. So I think, yeah, I think they're all, they're all trying it Microsoft's slightly interesting. Um, and yeah, maybe, maybe the canst one in some ways.
'cause obviously outsourced the development of, its of the models to OpenAI and it's now sitting there as an investor in, on a, you know, potentially a, um, eventually one day some sort of great exit. Um, meanwhile it's got obviously its own software, you know, stacks and franchises, if you will, which, which dominate, um, you know, c you know, the, uh, the corporate world. So it's, uh, yeah, it all takes slightly different approaches, but I think, you know, largely they're all trying to do that as well.
It's, it's funny, isn't it, that mi Microsoft has a number of times in its history been that sort of weight and then pounce, uh, kind of approach instead of trying to be the one on point taking the, taking the fire Done by Microsoft product or version three type thing. Sometimes Yeah, definitely not 11. Yeah.
I mean, but, but you look what, how far they've come with Azure. I mean, I, I don't, I'm old enough to remember, and I think you guys are as well, that in the cloud wars Azure was seen as sort of an afterthought. A Oh really?
Microsoft, you know, yeah, you're gonna bring Windows to, you know, you're gonna bring a knife to a gun fight here, gun. Well, nobody's laughing now because Azure was so incredibly enterprise and developer focused that they were able to build it into essentially, um, I dunno if it's a bigger business than their traditional Windows business, but it is a monster. And I think they're trying to do the same in ai, right?
Yeah, I'd say so. And I think it's, um, and you can say the same thing about Google, even at the beginning of 2025, people were saying, yeah, that's not a serious enterprise play. And, um, and now, and now look at it and, you know, and realness in the space of 12, 20, 18 months.
So yeah, I, I think it is, um, yeah, I think, yeah, there's, it waxes and wanes. That's what makes it interesting. That's what keeps us in the job, right, basically, isn't it as well, Indeed.
And the waxing and waning, by the way, I just wanna pause for a second there, to, to honor the fact that if you do invest in a given model maker, um, that that can be the most, you know, beneficial and frustrating aspect of that at the same time, because it's every time there's a new model that comes out. 1, lets OpenAI for instance, um, it's, it, the models are a, are a different employee, they become a different, you know, beast that you have to contend with. And you may have perfectly running, you know, uh, workflows that you've spent, you know, months painstakingly building towards something, you know, that they do exactly what you want and next day they don't.
It is, it is crazy. Sorry, Nick, was that Yes, that you do have something to say about reinvent or ne Yes. That you want me to No, I think, I think we're done.
I think We're done with reinvent. I sense a transition point that will work. So here we go.
So, you know, you talk about companies like Microsoft who are seen as sort of, I don't wanna say stodgy, but sort of, you know, they're coming later, they're more enterprise focused, they're more, you know, more of a traditional, uh, IT vendor. Um, but there's of course, a, a big daddy traditional IT vendor that made some waves this week as well. And frankly, there again, I feel like the market, um, the zeitgeist isn't right when it comes to IBM because you say those three letters to a lot of people and they immediately go, oh, mainframe.
Well, yes, mainframe still exist and they're still relevant, but that is not IBM and IBM this week, uh, made huge waves if it's possible to make waves, uh, outside of an AI announcement. They made huge waves, um, with the announcement that they, uh, are going to be acquiring Confluent, which is a company that maybe not everybody's familiar with, um, but they're acquiring Confluent. And those of us who are kind of insiders in the industry, we're like, oh yeah, like, I feel like the Kool-Aid man here.
I'm like bursting through the walls saying, this is the, this is such a great move for IBM, but I think most normal people would be listening to this saying What? So, um, um, first off, uh, what's your reaction to confluent IBM and second? Um, this isn't about ai, or is it, It's totally about AI and yeah, $11 billion of any sort is a big splash, is it not?
And the fact that they, they paid this much for a company that, that basically built its business on top of a, uh, very well known, uh, but nonetheless, a piece of open source software in Kafka, uh, which is a, a streaming, um, you know, service, um, says a lot. And what it says is that IBM gets infrastructure and that, that's like my, like top level give, you know, takeaway from that. And you can see this acquisition building on the HashiCorp acquisition they made earlier this year, and all of that says, infrastructure as code, code is infrastructure.
And if you're going to build ai, um, you're gonna need to, to be able to, you know, architect something that can be, you know, posted, hosted and scaled anywhere, any cloud provider, any premises, et cetera. And you're gonna want something that you can orchestrate in the most effective manner. And by orchestrate, I mean bringing data to ai.
And that's what this Confluent acquisition is all about. You know, we've been building, um, AI systems with, you know, context windows and, and supplementing those, you know, and supplementing the training of the model with context window data, like through rag pipelines and such. And that's great, and it can bring, you know, you're not gonna be indexing that stuff instantaneously, but if instead you could bring data in real time to the models, um, as they are going about the business, and also, you know, bring data out of those models, especially in an agentic workflow, that's gonna mean, you know, whether or not you can actually build and succeed with an idea that you have as a company.
If you're gonna build with ai, you need to be thinking about not just static data that gets fed in through the context window. You need to be thinking about streaming data in real time. Yeah, excellent, excellent input.
This is Brad's area very much. He's the expert here, but, um, I, I, I think it'd be, it will be interesting to see, um, how it kind of get ties into the watsonx, um, AI story. Um, and I think they're kind of, you know, that, that the messaging around there is sort of getting, um, you know, is, you know, getting stronger, but I think yeah, may, may change a little bit, um, and the, you know, and the orchestrate product as well.
So, uh, yeah, it's very much, it did remind me HashiCorp different, different use case and different technology. Um, but, uh, but similar, you know, IBM has always had a formidable AI m and a machine, I remember as an analyst event, it must is well over a decade ago, we got to sit with their m and a team, I think I remember and sit there and they were basically telling us like, we round this table. Let's, let's imagine this, if we were gonna buy a company and made up company, and this is how they go through all due due diligence.
And it's, uh, it's quite a process and it's quite something to see. And, uh, yeah, good to see. It's still, it's still executing.
Yeah, they're definitely looking ahead once you say Steven, they're, they're definitely looking for the long play here in terms of not being a hyperscaler in terms of, you know, having data gravity, but instead being a hyperscaler in terms of, you know, having an infrastructure that can run anywhere and enable anything. The the thing as well that IBM does so well is, at least in modern times under this current administration, not so much in the past, but in this, this current, uh, ad administration at the company is, uh, run these things well in a way that doesn't run off customers, um, that is actually accumulating customers and bringing people into the fold instead of e excluding them from it. Um, and I think that we've seen that certainly with Red Hat, uh, we've seen that with, uh, HashiCorp.
Uh, another acquisition I want to bring in here that I think is, uh, parallel here. You're talking Nick, about basically the, the great big acquisition machine. Uh, all of these acquisitions rhyme, essentially IBM is looking for companies that have just incredible recurring revenue that have, um, you know, customers that all, you know, can, can come into IBM Fresh and, and expand within the IBM portfolio of products.
But, uh, they recently purchased Data Stacks as well, which is another, um, incredible open source. I mean, maybe not so high profile, but a open source purveyor, um, for enterprise customers from an Apache product, um, Brad, um, data stacks plus, uh, uh, Apache Kafka. I mean, how, how does this work?
Yeah, like, like Nick said, it's, it's all about enabling wa the Watson X portfolio. So watsonx, ai, Watson, Watson, X Data, Watson X Governance, all of that is benefiting from every one of these acquisitions. And to my mind, it isn't so much about what these products do, what these technologies do, because you can get Kafka from anybody, you know, I could, I can r up on GCP and, and Azure, you get it for free, It's open source baby, right?
But, but I mean, like manage, host it. You know, this, this is what Confluent makes its money on Yeah. Is managed host, right?
And, and yet what what is really interesting to me about these acquisitions that we're talking about is that each of them have a very well regarded and global ecosystem, um, that is mature. And I would say, like if I was to look back at IBM over the last five to seven years and say, what's their biggest weakness? It would be lack of ecosystem.
I mean, think about what drives the value of Azure that we've been talking about. It's, it's not the greatness of the software, it's the breadth and commitment of the ecosystem that builds on it. And for it, you know, not just the the get and, but it, it, it's right.
It's huge. So if I can, um, break in with a timely quote to those of us old enough to have seen the film, it reminds me of hand solo Luke Skywalker says, you know, that hunk of junk and Han Solo says, who's gonna fly it? Kid you?
Well, that's IBM, right? That's Red Hat. That's, you know, what's going on here?
You know, yeah, you can do this, but who's gonna run it? Well, we are, we are gonna run it in a way that works. And frankly, that has been a very compelling argument for these, uh, for these companies.
Now, another thing that I wanna bring in here, um, hopefully without any more Star Wars quotes, is, um, Nvidia, of course, now I have been really, uh, excited about some of the models that Nvidia has developed. Um, parakeet, absolutely. Rocks, I think I mentioned that recently.
Um, but Nvidia is out now with a, uh, you know, mixture of experts, uh, models. Tell us a little bit more about that, Nick. Yeah, That it's, it's not supposed to, the mo we were talking about the models and the efficiency of models earlier, but they're talking about techniques in which to optimize mixture of experts models.
And they were, uh, they put out a blog post, um, late last week, I think. Um, I thought it was, it was interesting. They were talking about how, um, 60% of ai, um, of open source AI models released this year are, are mo let's call it MOE, so we don't have to spell out every time, but mixture of experts, models.
Um, and this is kind of taking over from, from kind of dense transformers, but there's problems, um, with those in terms of, um, memory, bandwidth pressure. And this is on their own, on their H 200 systems they're talking about. So on their own, uh, GPUs, um, there are limitations around memory bandwidth, um, from, uh, constantly loading all the expert parameters and then communication latency, uh, when experts are distributed, um, across more than eight, uh, gpu.
So they were looking for a way to, to solve this problem. And is there MV link technology? So they call this the GB 200 MVL 72, which connects 72 Blackwell GPUs, um, and delivers, you know, I'm not gonna go through all the, all the numbers, but a lot of it's very quick.
Um, and it enables, um, you know, to, it enable, gets rid of the bottleneck by distributing the experts, uh, across up to 72 GPUs, reducing the number of experts on each, on each GP, and that's that that relieves the pressure on, on the, on the memory. And so, you know, given the, essentially the, you know, Moes are, are the kind of the way that everything's working now, um, or at least, yeah, this is, the models are gonna be, um, you know, rolled out. Um, they're looking for obviously, you know, ways to optimize the infrastructure on their infrastructure.
And of course, you know, with their MV link technology, um, actually interesting one on the, um, was it train three or four is gonna have, um, of the AWS ones is gonna have, um, uh, Nvidia MV link in it as well. And so I thought, I thought it was just quite interesting I how we can, 'cause you know, in the early days of the transformer models, you know, we're obviously talking about, um, how these models obviously cost an enormous amount of money to, you know, to train to train. And that's partly because in that the parameters, the way the parameters work, the way the data constraints work, um, and you know, the, the MO the MOE is sort of becoming kind of the standard way of building, um, frontier models.
Um, and there's obviously, there's very specific reasons why you would use other niche models, and there's like, obviously diffusion and things like this. Um, but I think it's, it's interesting. Um, it's obviously, you know, it's self-serving for Nvidia to say, you know, use our, um, infrastructure use I hb two hundreds and use SEM link 72.
Um, but it's interest, I think how they, you know, and they claimed, I should have mentioned, they claimed it was 10, a 10 times performance increase. I should have probably put that higher up in the, uh, in my little, uh, ramble. But I think, I think I thought it was worth mentioning anyway, because, uh, anything that can optimize these things, um, is, is gonna be of interest to, to a lot of, uh, model trainers.
And as we were talking about earlier, there's still a lot of those around. Yeah, we're, we're in the scale out era, are we not? And, and that optimization is, is gonna make or break investments in data center as well as individual projects.
And the skeptic in me, by the way, Nick want, wants to see, um, them do this comparison on a chip that isn't three years old. Um, you know, but, uh, so it is, it is, like you said, it is a bit self-serving, but I, I think they're bringing up some really important points that, um, you know, the architecture of these models, you know, really dictates what you can do in terms of scale out and up with these things. You know, if you have the greatest model in the world, but it sucks so much VRA just to set it up and you can't scale it across clusters, how much concurrency are you gonna get out of it?
You know? And by the way, it's not just big models. I, I know, um, IBM with their, um, granite models, they, it's like a month or so ago, remember that they released a mixture of experts that's like, uh, under 4 billion parameters that is meant to do the same, to bring the same benefits of an MOE to, you know, on device, you know, inferencing.
That's awesome. Yeah. They're doing a lot of work with small language models with IBM, aren't they?
That's very, it's very, yeah, yeah. Differentiator for them as they say it. Well, I've just, uh, accidentally demonstrated the power of expert mixture of experts models by revealing that I didn't know what this story was all about and calling in an expert who did, which is exactly what mixture vaccine.
See, I meant to do that. You routed, I routed it properly. Um, you know, I wonder, Nick, does this have anything to do with the, uh, uh, rising price and, um, lowered availability of ram?
Uh, there's kind of a RAM crisis right now. Um, I'm not a, I'm not a RAM expert. Um, but I think yeah, it probably does.
It has just, and in, and also just, uh, in terms of, you know, squeezing the most out of the, uh, of the assets that are out there, obviously there's, there's shortages of all sorts of things, including GPUs themselves. And so I think it's, uh, you know, they, they were talking, um, specifically about, you know, performance per WA and then being able to, like in Nvidia language, them being able to generate more tokens from your AI factory. These are all the terminology they like to use, but if you can squeeze, if you can get a 10 times, um, increase in performance per wat, that's, that's extremely important in these kind of power constrained and GPU constrained, uh, times in which we, uh, live.
Indeed, the data center is limited by the number of watts that it can consume. And, and by the way, I, I understand that that RAM shortage is, is down, so it's a one individual named Altman, uh, Mr as in, Are you trying to say that he doesn't have a good brain? Or are, are you trying to say that, that No, That he, he bought up.
He bought up. He bought all around. I know, I know.
Sorry, His memory's not what it was. No, he's got the best memory. Um, alright, well, on that note, um, thank you very much.
I'm gonna have to route that query to another expert here at Futurum, uh, to find out more about that, uh, hardware crisis. Uh, thank you both for joining us for this week's episode of utilizing ai. I have to say, uh, our producer just ran the numbers and we've got some, some great viewership already with this new podcast.
Thank you everyone for listening. Um, please do, uh, drop us a line if you're watching. Uh, I know you are because we can see the metrics.
So, uh, we would love to hear from you. Uh, you can find me on LinkedIn, uh, as s FoST. Um, Brad, Nick, where can we find you?
Brad? Yeah, Brad Shimmin on LinkedIn and, uh, on the future website itself. Uh, yeah, Nick, patience on LinkedIn and um, Twitter X and I'm on Blue Sky, uh, and all of our stuff is published on futurum group com.
Excellent. And, um, and again, uh, thank you to the Horological Society of New York for hosting me here today in their beautiful, uh, library in Manhattan. And thank you for listening to this episode of utilizing ai.
If you enjoyed this discussion, again, please subscribe. You'll find us on YouTube as well as in your favorite podcast application and do consider giving us a rating and a review since that helps visibility for all podcasts. Uh, this podcast is brought to you by the analysts and experts from the Futurum Group where Insight meets ai.
ai, which is our AI news site, the utilizing AI YouTube channel or the text TV app on your TV or smart device. Thanks for listening and we catch you next week.