The Evolution of AI Infrastructure: Trends and Insights with Nadav Eiron
In this Techstrong.ai Leadership Insights video, Nadav Eiron, senior vice president for cloud engineering at Crusoe, following the formation of an alliance with AMD and the picking up of a $750M line of credit, dives into how the consumption of IT infrastructure is evolving in the age of artificial intelligence (AI).
Transcript
ai Leadership Insights series. I'm your host, Mike Bazar. Today we're with NAAB Iran, who is head of cloud engineering for cruso, and we're talking about AI infrastructure, and a whole lot of things are going on with these folks.
They got a $750 million line of credit partnership with a MD, and I think they're expanding into Europe if I read all the announcements correctly. But nab, welcome to the show. Thank you.
Thank you, Mike. Happy to be here. Um, it's a funny thing, but I think we've talked in the past and it felt like this was gonna be some sort of interesting little niche, and now I've turned around and it's kind of this billion dollar industry and everybody and his brother's talking about more efficient ways to consume AI processors.
Is there some sort of rapid maturation process going on here? I I, I think there is. I think, I think we actually talked about it last time we talked, uh, a little bit about kinda like the pace of how it happens and, you know, this hockey stick we're in.
I'm, I'm not horribly surprised. Um, maybe just a little bit. Uh, definitely the, the pace is picked up and, and I think the, the market also looks a lot more interesting in terms of like how you see, uh, both the consumption and the supply of, uh, compute infrastructure.
Um, uh, that's, that's focused on AI for sure. I feel like it's also changing in terms of who's responsible for what. I think early on it was a data science team, and there was a couple infrastructure people attached to that, and they built one application.
I think we're at some point now where people are trying to figure out how to operationalize AI infrastructure across multiple applications, and that's changing the nature of the conversation. Is that a fair assessment? I, I, I think to some extent it is.
I, I think there is, um, also the, the focus is shifting to some extent from AI as a technology to the applications that are enabled by ai. Um, it's, it's, you know, it's less interesting what the AI is and more interesting what it does for you. And I think we're seeing that more and more as people start to use AI in more and more areas.
We, we do see that, we do hear from our customers more and more, like, I just want this thing to work. I have a day job to worry about not building AI infrastructure, not figuring out how to optimize the performance, not figuring out how to, uh, you know, find the right device driver for the hardware I'm using. Um, I have a day job and, and that means building a business or building an application or something like that.
AI is a tool and, and not the, uh, not the objective of the work. I also feel like there's more separation between the process of training an AI model and the actual running of the AI model. In terms of the inference engine and where that might be, and are, are we seeing some shifts there or, you know, how is that evolving in your mind?
There, there, there's definitely a lot of, um, uh, focus, you know, like, like I mentioned as, as people try to use AI more, that means they do more inference. Um, inference is becoming more and more important. I think last time, um, we spoke, we, we talked about manage inference, you know, that's a, a product that we, uh, we have out there that is seeing more use even within our, uh, straight up infrastructure, uh, users.
We are seeing more and more focus on inference. And, and to me that's just natural. Again, because people are using AI more than they're focused on making AI better, bigger, et cetera.
That's still an important thing. But as a percentage of the slice of a pie, uh, you know, that slice that is using AI is the one that's driving the growth. There's also a point where, um, you know, certain things are at a good enough place.
Um, and again, there, there's always the, the, the leading edge, the bleeding edge, the place where you are finding new capabilities in, in the technology. But there's also, you know, to your point about maturation, there are parts of ai, you know, large language malls have been with us, um, roughly with the same general direction. They're, they, they have today for, for a couple years now.
And so at, at that point you do see like the, the, the core LLM becoming less of a focus and how you use it becoming more important. Mm-hmm. And to your point, I also feel like l LMS now coming t-shirt sizes, they're small, medium and large, and, um, and people are getting smarter about which to use when, and then what the infrastructure resources are required to support those.
Uh, for sure, 100% that, that's actually something that, that has happened. Um, I, I think pretty since, pretty early on, um, you know, the, the full scale, the large LMS are super expensive to, to use. Uh, they, they're slow.
Uh, you need, you need a lot of hardware to make them work, and for a lot of stuff you don't need the power that they provide. I think more and more we're gonna see actually systems that, you know, dynamically use, uh, the, the right size model for, um, the, the situation you're in, depending on, you know, your context window length, depending on the prompt you have, et cetera. All these things to me are, are natural optimizations that are happening as, um, as a technology matures.
Um, you know, there, there's, there's a rule that I've learned in, in, in building infrastructure over the last couple decades, which is the, the closer you are to the end user, the bigger the leverage you have over optimizing the problem. So, you know, we started by optimizing these things by looking very close to the hardware. Um, that's important early on to understand how the hardware works, et cetera.
But you really get the leverage where you look at a problem from a user or you look at an application in general and say like, you know what? I don't need a 400 billion parameter to serve that need. 4 billion is enough.
That's a huge optimization you can make. And so I I, I do think we see as the technology matures people more and more looking at that place in the value chain saying like, Hey, you know, I have a 10 x improvement I can do here by just using something else. Mm-hmm.
Now you also partnered with a MD recently, so are people looking at GPUs from different vendors now and mixing and matching, or are they saying I'm either an all Nvidia or all a MD? Or is it becoming a little more nuanced That that, that, that's a great question. I think, um, in general, as a, as a cloud service provider, um, and I think it's good for, for the customers as well, the users, the end users of the technology, uh, competition is good.
Um, I think right now we're still at a place where people have a preference for one or the other. Uh, there is like the, the level which that we work with our customers, they would be aware, Hey, you know, I'm going to get a MD hardware, I'm gonna get Nvidia hardware. To me, the really interesting next phase, again, going back to that point of like, the further up the stack, the more leverage you have is actually building services that can mix and match.
Um, we're not quite there yet. Um, but I think that is forthcoming and I'm, I'm, I'm super happy, you know, to have our hands on on those a MD machines. I think just by having two things from two different vendors that are different from each other, we're gonna, we're gonna find the niches, we're gonna find the place where one is really useful, where the other is really useful.
We do see, um, healthy demand for, for both, you know, we also have, uh, B two hundreds that, that we started, um, uh, offering to our customers a few months back. We have the A MD that, that is still a couple months out. Um, both of these we see healthy demand for, um, and I, I think this is recognition in the market that, um, competition is good as well as specialization.
You know, uh, this, this workload is with this model on this type of hardware, this workload acts differently, needs different type of hardware, different model. It's all good. Um, speaking of different classes of processors, are we also gonna see more usage of things that aren't GPUs to run inference engines?
And what might that look like Perhaps? I, I, I think, um, I think in general, um, as the market grows, it would make more and more sense to try and hyper optimize for a fraction of the market, right? So like, if you have a market that's a hundred billion dollars, whatever, just to, just to pick a random number, uh, an easy round number, um, you know, 20% of that market is a 20 billion market.
It's worth an investment to, to create hardware that's hyper optimized for that. I do think that hyper optimizing hardware makes the problem simpler and therefore allows you to get, um, uh, more benefit out of it. And, and as I mentioned before, also from the usage perspective, as your software stack becomes smart and is able to direct the traffic to the harder that's most useful for it, that, again, will lower the bar for people to come up with with new types of hardware.
I, I would say today we are seeing, um, more and more interest in the asics that the, uh, big hyperscalers have. So like, you know, tra and, and infinium with, with, um, AWS TPUs with Google, you see them taking more and more, uh, perhaps of the market share. Um, we've been, uh, exploring things with, with several startups that, that work on interesting hardware solutions.
Again, all these solutions are trying to limit the scope of their problem to be able to do better than the big GPU vendors. You know, if you wanna build something that will accelerate any AI workload under the sun, it's a very tall order. But if you're just focusing on, let's say, inference for small to medium models and, and you make that simplifying assumption, I think there's a lot to gain there.
And as the market in general grows, even optimizing for just such a slice, we'll become lucrative enough that we will see innovation happening in the hardware in that space. Mm-hmm. Is it me?
But I feel like, and maybe it's just part of that whole maturation process, but people are more sensitive to the cost to AI as they try to figure out what to operationalize. 'cause it's one thing to have a thousand experiments, it's another thing to figure out what can I actually afford to run in production. Yeah.
I, I think that's fair. I think, um, the, the way we sometimes refer to it here is there's, there's sewers and there's harvesters and, uh, I, I think we are leaning more towards the harvesters, and the harvesters do behave somewhat differently. They're not just, um, more, uh, cost sensitive.
They're also, um, less tolerant of complexities and, you know, the need to gain expertise and to learn and to experiment. They want something that just works and hopefully just works and doesn't break the bank. I, I think for ai, this is actually something super important.
You know, when, when you think about using a new technology, um, there is, there is, you know, the 5%, 10% kind of improvement in, in cost performance, which, which is good. That just helps the current players, you know, uh, improve their products by a little bit or make a little bit more money. We, we are still seeing, and there's a bunch of data about this, about, you know, the cost of inference over the last two years, basically going down by orders and orders of magnitude.
When something becomes order of magnitude cheaper, that opens up a whole new, you know, universe of possibilities in terms of the kind of products that you can use it in, the kind of business model you can use to monetize it, et cetera, et cetera. And so as we're seeing those cost improvements, um, you know, compound, um, and, and continue to improve, uh, I'm, I'm super excited because I think those harvesters will have all, all kinds of different things they could harvest all kinds of different areas where AI will all of a sudden be a great solution just because it's now cheap enough to justify its use. Mm-hmm.
Can we simplify the stack? And I'm asking the question because one of the great things about the cloud was it made it relatively easy for developers to stand up infrastructure and build and deploy software. I look at AI today and I still feel like there's, you know, ML ops people, data engineers, security people throwing a bunch of developers and shake and bacon.
Maybe by the time I get the village together, something good happens. Can we streamline that? I, I, I think we can.
I think, again, this is partly the, the transition from the sewers to the harvesters where, like I said, I think people are looking for more simplicity, so there's more and more demand for that. Um, I, I, I will say a lot of the complexity even starts before you get to ai, for example, you know, um, a as you know, we, we have Chris who are building the world's favorite AI cloud, but we're not pretending to the, to be building a general purpose cloud. So many of our users have a footprint on a general purpose cloud as well, to run the known AI parts of their workloads.
And even just that, you know, the, the ability to work together across two different cloud providers, sometimes three or four, um, and, and hide the complexities there. That's, that's a demand that we see coming from our users and that we're working on helping them with. So the, the complexities start even before you get to the, you know, the science of ai, if you will.
Um, so we're trying to fix that. We're definitely seeing that in the, in the AI space itself. Like I said, we're seeing more and more demand for managed services, for managed inference, for, you know, a a a point and click kind of interface of like, here's my data, um, do something with it.
Um, you definitely see much more demand for that. I think, again, that that part of the journey is still in, in its infancy. We see some first steps in the right direction, but I think it's still a long journey ahead.
So what's the plan for the funding? And as part of that, what's coming next for you guys? Um, so again, we're, we're excited, uh, as, as everybody knows, um, this is, uh, a pretty CapEx intensive, um, uh, industry and, and so like having that funding secure is, is great.
Um, I, I'm not here to talk about future financials for sure, financials in general. So I, I, I, I'm not gonna say what's, what's ahead in terms of the funding, but I can say in terms of our expansion, um, as you mentioned, we just announced, um, a new data center in Norway, um, that uses 100% renewable energy. Um, we, we announced a big deal with, with a MD.
Um, there's more things coming, you know, the business is growing for sure. Um, it does require a lot, a lot of capital to feed it, but, but we do see a lot of growth. Um, I would also point out, we had, um, an announcement, um, a couple weeks back with Redwoods materials, which is a company that, um, uh, recycles EV batteries where we have, uh, a pilot deployment with them, um, that is 100%, uh, powered by solar energy and recycled EV batteries off the grid.
Um, so all sorts of interesting things are happening in terms of our, our growth. The, the business is definitely, um, uh, on the up and up. So what's the thing you see people doing today that kind of just makes you shake your head a little bit and go, folks, we could be a little bit smarter than that.
I, I think, um, this is a great question. I, I, I think one of the things, uh, is, is something that we've touched on in terms of like the, the hyper focus on the model versus the application. Um, you know, like I said, there's, there's a lot of optimization that can be done at the top layers of the application in terms of like, what actually you're gonna use for when, uh, we don't have good infrastructure for that, and therefore a lot of the people in the space, I think, overlook the ability to optimize at that level.
Um, I, I also think, um, that there is still, um, maybe I would say focus on expanding the capabilities where the capabilities are good enough. There is a little bit of like chasing like, oh, let's, let's, let's, let's work on transitioning to this new model, uh, where actually the system that you have is good enough. Um, I, I would say it's in general, in my experience, something that you need to watch out for with software engineers.
Software engineers don't like good enough. I like to be the best. And so a lot of times, you know, there's a point of diminishing returns, those last 2% of squeezing something, um, is oftentimes really, really hard.
And, and 98% may be good enough. So I, again, this is part of moving from like expanding the envelope of the technology to just using it. Um, and I think shifting that focus will allow us to iterate a lot faster, because again, those last 2% take forever.
Very hammer folks. There's a thing called AI optimization. It's coming to a store near you soon.
Nadav, thanks for being on the show. My pleasure, Mike, my pleasure. Thank you.
All right. And thank you all for watching the latest episode of the Techron AI Leadership series. You can find this episode and others on our website.
We invite you, check them all out. So then we'll see you next time.