How Neoclouds Are Driving More Sustainable and Cost-Efficient AI GPU Consumption
In this Techstrong.ai Leadership Insights interview, GMI Cloud CEO Alex Yeh explains how the rise of neocloud providers is reshaping access to GPUs required to train and run AI models. He discusses how alternative cloud infrastructure models can improve cost efficiency, optimize resource utilization, and help address sustainability challenges associated with large-scale AI workloads.
Transcript
Hello and welcome to the latest edition of the Techstrong AI Leadership Insight series. I'm your host Mike Biz. Today we're with Alex Yeh, who's the CEO for GMI, and they are a provider of a cloud service that hosts a lot of GPUs.
And we're gonna talk about, well, how complex things are getting from an infrastructure perspective. 'cause we have multiple types of AI models. Many of them are, shall we say, multimodal.
And it all needs to be run on some type of infrastructure that well is increasingly hard to get access to. Alex, welcome to the show. Thank you, Mike.
Thank you. Having me. I feel like there's two things going on that are kind of constraining our ability to operationalize ai.
And so let's just jump in. But the first is complexity. This seems like there's a lot of different piece parts and things that I gotta kind of organize.
People are using multiple models, they're using text and video and all these other things at the same time. And it's, it's hard to stitch it all together with a bunch of APIs. So how do we simplify this if we can?
Yeah, absolutely. So I'll, I'll give you some, some, uh, I guess some background, right? So we, uh, we are a GPU Cloud company, but having GPU infrastructure is not enough, right?
We need to have a scalable infrastructure and easy to use tools. And that's how we come up with the API platform, where api, everyone can have access to full modalities, full access to every single models out there, from open source to closed source, from all the open apps, the world and traffic of the world, the Google of the world. And here's one thing we realize is having these, uh, snapshot of these models are not enough because people are building workflows, right?
That's what agents are. It's a workflow, a string of different models, and having a strings of building these workflows is the true application, what people are actually building, not the model itself. And so this is where we come up with a solution called GMI Studio, where it's basically a workflow builder where you can bring up multiple models APIs together so that you can string them together and build a actual application so that you can be, it can be used for not just professional ML experts, but, uh, not just, or even coders.
You can have, you can be a creator, you can be a influencer, YouTube influencer, and you'd be able to build your own agents or tools. Mm-hmm. The second thing is, there just seems to be like the entire AI supply chain seems to be constrained.
I can't find GPUs. There's, we we're waiting on data centers to be built. Heck, I can't even find enough memory.
Um, so, um, is this gonna be a year where we kind of just deal with some of those constraints while we wait for, you know, them to kind of play themselves out? And maybe 2027 is a, a much bigger year for, um, deployments in production environments? Or how do you see this all playing out?
This is definitely a year of constraints. I thought this would be, uh, less constrained. Uh, you know, and when I was, uh, in 25, I thought 26 was gonna be okay, but it turns out everything was rising.
You know, gold price, silver price, now potentially copper price, memory price, GPU, price, everything and data center, uh, everything is constrained. And if you are a, I would say ML company or agent company or in general, you have to plan ahead. This is not where, hey, I just raised and I'm gonna find a couple thousand cards.
You have to plan three to six months in advance even before you start fundraising or before you're, uh, you're planning to write a check. I think that's the, the best advice that I can give to people. Uh, and here's something that we're doing on our, our end, which we are locking in memory pricing forward 26 and 27, as well as data center capacity all the way to 27 and 28.
So we're planning three years in advance in order to secure these precious supply in either power or, uh, or, uh, memory and GPUs. Of course. The other thing you, we hear a lot about these days is various AI accelerators that are being positioned as alternatives to GPUs, especially for inference.
What's your perception of those alternatives? When do I use them? Are they really mature enough yet?
What's your sense of what's going on here in terms of our processor options? Yeah, so training Nvidia still dominates the world. Uh, I would say 95 even more percent in terms of market share, uh, to training workloads.
But inference, they're starting to see quite a lot of different usage. But those are for much larger business. I would not recommend any smaller startups to use anything that is, that is outside of, uh, NVIDIA's system because if you have any bugs, you can just find any CUDA code, not any, but find cuda code to support you.
But if you're using other ASIC systems, then you would need actually a team of people to debug for you. So this is only for the end profit of the world. If you raise less than 10, less than 1 billion, then you shouldn't, uh, think about using anything that's not, uh, Nvidia because that would just waste too much time.
AI companies should just focus on pushing products out and building amazing product instead of thinking about their infrastructure scaling. Let that hard work to infrastructure providers like us. You mentioned Cuda, it's a pretty awesome framework when you look at it.
A lot of thought when into it, but a lot of people are also concerned that maybe it just locks us in a little bit too much. So what is the balance to be struck between taking advantage of something like cuda or alternative frameworks or I don't know, might be one day see Cuda running on things other than GPUs just for Yeah, so, so again, right? I think startups should be thinking about scaling, right?
Building products and killing your competitions in other companies that is building, building the same same product in the same sort of same category. I think that's the number one focus, right? When you think about diversification, that's usually for much larger businesses, right?
For example, like the, well, like Anthropic or like SpaceX, well not SpaceX, uh, X ai, right? These type of larger giants, then they can start thinking about, okay, maybe I should use TPU where you should shoot uh, a MD or, or uh, Nvidia, right? Leave that infrastructure decision on ops, right?
We are highly incentivized, we're aligned to build you the most cost efficient, call it token per dollar, right? For, for you or else you're just gonna choose another provider. So we want to take care of that entire infrastructure and scaling, stability, reliability, and obviously cost.
That is something that we are extremely good at. Um, and so I think that decision shouldn't be the main focus for startups. That's my opinion.
Um, are we getting to the point now where maybe I need to narrow my number of use cases that I'm gonna really push this year because, well, it's going back to that constraint question, but um, do organizations need to maybe pick two or three things that they're gonna focus on? 'cause right now, I feel like last year everybody was in this kinda, let's let a thousand flowers bloom and everybody had a pilot. But if, if I'm gonna really make an ROI case for something, do I need to narrow my choices and kind of put more wood behind a couple of fewer errors?
If you're a tradition, if you're a CIO from a traditional enterprise that is non-digital native, then I do not suggest you to buy GPUs. I would only suggest you to start building with APIs, start building application and do your test case roll out internally, roll out to your, call it sandbox customers and see and iterate. Once you're done with that, then you can start thinking about scaling and putting things in a dedicated endpoint with GPUs or even thinking about, uh, fine tuning.
I think that would be the step 1, 2, 3, uh, to, uh, to roll out your AI application. So in that capacity, I think you should just think about how to field your workloads will be the central message to, to DCIO. Don't think about the these GPUs constraint, you can just pull APIs from us or from other API providers, uh, or even the, the hyperscalers, even though they're very, very expensive, um, on building your application first.
Alright, we'll come to that. Um, the hyperscalers, of course are having GPUs and any other thing you might imagine that you want to access. So what is the difference between somebody like GMI who's very focused on GPUs and the cloud and the service and what I might get from the hyperscalers?
Um, first of all, I like to, uh, you know, I've interviewed a ton of ccio and CTOs and not one has said that their cloud bills is too cheap. Uh, so I think that's number that, that's number one. So I think cost is a significant, uh, hurdle for, uh, for companies to use.
So, um, I tested on a, um, particular hyperscaler and just spinning up an instance, it costs $150 a day, just one server, one task, one person. Imagine you have 10, 10 agents building and you're rolling out to thousands of your employees or, or, or, or customers that cost would be inhibited. So, um, I think that's number one, right?
Number two is we highly optimized for ai, and that's what I was mentioning about token cost as well as token performance. So what we do, two things really great is we will drive that cost down. Second is we will increase that throughput.
So A GPU is a GPU, right? Everyone has the same H 100, H 200, but we're able to increase that throughput by three x to six x. Why?
Because we're building on the new AI native way where the hyperscale is built on the old CPU codes that VMs everything that they lose its control over the GPU power where it we extract and amplify the same system with much better, uh, underlying infrastructure that makes the GP performance triple to six x. Mm-hmm. Um, as you kind of think this through for a minute, is there, um, some notion of, uh, am I, do I need to be smarter about what GPUs I'm using when, and I ask the question?
Because a lot of times I go talk to these data science teams and they seem to be very obsessed with the latest and greatest GPU that comes from Nvidia. But if I look at the workloads, particularly if I'm starting to build things that are maybe smaller models, aren't there older generations of GPUs that might just work just fine, that are less expensive and more available? That's absolutely right.
So, uh, it, the gps are application specific, right? If you're going to racetrack, you should use maybe a race car, right? But if you're going, going on, you know, uh, a mountains, you should use a, you know, a four by four maybe, right?
So like a four wheel drive. So it really depends on application you're building. So if you're focusing on low latency, then you should use smaller models, right?
If you're building a large reasoning, deep research model, then you probably use, you have to use the, the latest and the greatest GPUs. So it really depends on application. And then I think the great thing about using, for example, our studio is you don't have to think about any of those constraints.
Literally just pull API and think about what you wanna build. That's what we want people to think about. It's like, stop thinking about all the infrastructure.
Just think about like what exactly you wanna build. Okay? You wanna build a voice agent, what you wanna build some, a customer support agent you wanna build internal GPD?
Just think about that, right? Leave that complexity like what GP to use, we can optimize behind. Like customers won't care about like what's, what's actually powering the GPUs just as long as it's, it's, it's uh, you know, it's it in, in the right, you know, regulatory frameworks like the GPUs in, in America or in a, in a have the right license, right?
We, we handle all of that, right? And we optimize this. We may mix and match different gps, which we are by the way, that lowers that cost together.
We do with a crazy science, like we, we, we do a cluster inferencing. You're not running GPUs, like you're not running inferencing on a single GPU. We use a full cluster of different GPUs.
But again, we, you don't have to think about any of that. Just think about like what application wanna I build? How do I scale my customer?
How do I scale my application? That's all you have to think about. And it's like, oh, how do I market it?
How do I sell it? That's it, right? Le leave that things to us.
Do you think there's also, uh, the market and the way we manage buyer's perspective and the customer's perspective is evolving? And I asked the question because early on it seemed like there were a lot of these tiger teams for ai and they included the data scientist and somebody who was the infrastructure specialist and off they went. But as we kind of operationalize this more, I wonder if we're starting to see the IT team play a larger role on the inference selection side because they're gonna look at things like Kubernetes and things that they can scale.
And a lot of folks are saying that the infrastructure conversation for AI is moving over to the IT department while the training stays with the data science folks. That's absolutely right. So you're talking about, okay, so the company has built their application out there, think about scaling and for enterprise, uh, they have a much stringent, uh, requirements than startups do.
And I yes, that typically what we have as a is your observation is correct, is typically with IT team. And the first I would say mistake that they will do is, Hey, I'm just gonna buy like four servers and put it, put it in my data center or, or like in, in and in their office. Typically I have like a server room.
And then when people are using it and then it crash, because first of all, it's very difficult to manage the servers, the infrastructure below OS level and then above OS level on how to separate those into different containers, which is meaning that different users will come in and each user shouldn't see other users, uh, workloads, right? So you're talking about like security solutions and softwares, uh, and, and then you have to build the actual application system sits on top and which crashes, which conflicts with I think maybe the Kubernetes instance and or even directly on, on bare metal. The long story short, the traditional, so CPU is very easy to manage and IT team is pretty, pretty good for that.
But when you're talking about GP scaling, it becomes a difficult problem. And so what I've seen that they're now moving on to is basically, okay, I built my application, I wanna scale, and they'll talk to infrastructure provider like us and they'll discuss, Hey, here's my base load, here's my base usage, maybe like a hundred GPUs. And then they'll say, Hey, at peak we will hit 200 GPUs.
So I would like to flex up to that and we can design those solutions for you and even cage up these servers, obviously with the base loads and flex up, it's all difficult. Um, so, and we can build direct line. So all the regular regulational hurdles we can, uh, and data security, they are privacy issues we can handle for you.
And, uh, I would say having a localized private cloud or hybrid cloud would be the best for any enterprise start. Once they have their MVP and they are thinking about scaling and use, they start getting user attractions. You would, you need to start talking to your, uh, GP infrastructure providers and think about, uh, how to best accommodate kind of the well best tailored towards what the, the, the user actually needs.
So we can give you that solutions, uh, while not losing kind of, uh, uh, I'll say, um, reliability and cost in mind. So as you look into the coming year and there's a lot of possibilities and what's your crystal ball telling you? The supply chain constraint will continue plan ahead.
That's I would say supply side. On the application side, we are seeing the explosion of inferencing starting early last year. Now towards 26, there are multiple companies building really amazing applications in agentic workflow that optimizes your kind of your day-to-day work, completing tasks to optimizing kind of customer service, right?
Uh, and so this is really the year of true application as all the infrastructures are ready. I still remember, I think last year, uh, I was on on the show and talking about GPU scaling and people renting GPUs, and now just on our platform, we have 146 models available and workflow. I wouldn't had that last year and just so many, so many models available, so many toolings and so many solutions.
Um, and so I truly think that this will be the age of scaling. Uh, and there will be two, I would say gen AI applications, uh, that would amaze people, uh, in I would say a year ahead. Yeah, and let me ask you one last question about all that.
So, um, how automated and can all this get, because you were talking about APIs and not having to know what's going on with the infrastructure. So ultimately, can I just kinda express my intent for my application and the infrastructure will just automatically take care of it, including figuring out what models maybe to use dynamically based on what it is I'm trying to accomplish? I mean, how smart can smart get, You're right.
So in, in the past, AI is middle to middle, okay? Human is end to end, meaning that, okay, if I give you a task and you go research and you do ask people, I think about it things, read books something and you'll complete a a task, right? Or you ask me questions, get feedback, but you'll complete a task.
Human is end to end, right? But AI currently is middle to middle, meaning that it's a tool that people use for productivity gain. It only exists on the, basically on your screen.
But what I'm seeing a lot of companies are building, which is quite exciting, is they're making it end to end, right? It, it will control your mouse, control your screen, do research, ask you questions, and complete, actually execute a task. And, and for example though, and then they'll open your email, write the email and ask you, it's like, is this okay?
Can I send this out? Right? So I think I'm seeing the expansion of the AI task from middle engine towards the end, right?
And so I human can just be now approver, basically. It's like they'll give you a bunch of options. He said yes or no, yes or no, yes or no improve 'cause and all you do is making high level decisions and let all the execution be done by ai, which is, which is sync and amazing, right?
And so people who can use AI tools really well, you're gonna excel at your work. If you're a company, if I'm talking to a company, you're gonna crush other enterprise if you're able to use these tools. Well, All right folks, well, I think you heard it here.
We're pretty much, I'm not even sure we're at the end of the beginning yet, but we got a long way to go before we operationalize this AI stuff to the N degree, but it's gonna happen. Hey Alex, thanks for being on the show. Thank you, Mike, for having me.
All right, and thank you all for watching the latest episode of The Techstrong, that AI Leadership Insight series. You can find this in episode another's on our website, and we invite you to check all those out. Until then, we'll see you next step.