Breaking the Compute Cartel: Democratizing AI Infrastructure
The AI revolution is colliding head-on with a massive infrastructure bottleneck, as a handful of hyperscalers hoard the lion’s share of global compute power and leave startups scrambling for scraps. Jack Collier, Chief Growth and Marketing Officer for io.net, argues that the solution isn’t building more data centers, but rather tapping into decentralized networks that aggregate the 85% of global compute capacity currently sitting idle. By breaking free from restrictive volume licensing traps and embracing distributed open-source models, organizations can not only slash their inference costs but also reclaim their data sovereignty before big tech monopolies swallow it whole.
Transcript
Hey guys, thanks for the throw. net. And we're having a little chat about, well, is there gonna be enough IT infrastructure to go around?
Because, well, the big tech companies seem to be, well, buying it all up. Jack, welcome to the show. Thank you, Mike.
Good to have... Good to be here. Thanks for having me.
I think we've seen the spend now on AI infrastructure is somewhere north of $650 billion or so, or will be. Um, and the question then becomes, is there gonna be enough infrastructure for everybody else? Because it seems like everybody who builds those foundational models is gonna buy up a massive amount of it, but it's not clear to me that there's enough capacity in all these data centers, and we're making some major investments in that area, but they might not come along for another three to five years.
So what's your assessment on what we're gonna see? Yeah. It's a, it's a huge problem in AI, and actually, to me, it all boils down to access and, like, true innovation with, with the people.
It's... What we're seeing at the minute is the, the three main providers, cloud providers, account for 70, 80% of all global compute. And what that means is, you know, it's a, it's a little, uh...
Like a special club at the top there. Big companies work with those hyperscalers. They, like you said, block up a load of compute, load of machines, and what we actually see is that most of global compute isn't actually utilized because of that exact reason.
Uh, 10 to 15% of, global compute power is actually utilized at any one moment in time. So what we see is that 85%, say, redundancy within the whole global compute network, and the people that lose out are the developers, the startups, the people that are trying to innovate, get access to this space. And so what it creates is, like, an inequality for, for people to get going.
And, obviously the, the, the main hyperscalers have little interest in letting the little person, innovate. They want to get the run on everyone. They want to build these big, huge mod- you know, like, huge organizations and dominate the AI, AI landscape.
net where I, where I, where I work, is we're trying to think about how we can solve that utilization problem and access problem in, in novel ways. So what we see is that, you know, we, we want to be able to utilize that 85%, you know, the, the data centers that the hyperscalers haven't dominated, like, and give access to people, in a fair and, and, in a fair and cheap and accessible way so that they can innovate. And so at io we essentially work with all of the, the data centers that aren't the top three.
We aggregate their supply in a common place, on our, on our network, and we serve that back to end consumers. And so, you know, whether... Even if someone's blocked a, a, a, you know, a, a data center for...
Or blocking a data center for 12 months and they only need it for, you know, a, a day a month or something, they can offer up the rest of that, capacity to people and monetize that. So yeah, we're trying to, instead of mining a lot of new resources, building a lot more data centers, et cetera, which I think inevitably we'll have to do, but we actually want to solve this problem in a different way and actually get that utilization up closer to 100 so that we're getting more out of the stuff we're extracting and being able to act- give it back to startups and innovators to be able to, you know, get, get access to this space. Because, because if we don't do that, then AI will be controlled by a very few, organizations, and it will be, at the end of the day, the people that lose out.
Do you think that the shortage of infrastructure resources will drive people to shop around more? I think there's a bit of a default mindset where people are like, "Well, I'm going with AWS, Google, or Microsoft," and maybe they'll start to consider other options. Yeah, absolutely.
I mean, I don't know if, if you've, if you've tried to get compute from those guys, but it is actually a bit of a nightmare. It's... You know, you're often on wait lists for the best type of devices.
I mean, just look at what happened when the H200s came out. It was, you know, all of the main, the big tech companies got the first access, got the run on everybody with the latest, and then it's like, "Yeah, we'll serve the remainder to some companies at some crazy price," you know? So you're on a wait list.
You're paying loads for the access to it. Um, you're also putting all of your data through one location, which is a, which is an issue, you know? If that...
If... We've, we've seen AWS go down. We've seen GCP go down.
Um, and so, you know, if you don't have... I think people are now looking for ways where they can protect themselves against that. net, we're, we're trying to solve this in more novel ways, you know?
Like, you don't need to book yourself a six-month block in a data center anymore. You can literally spin up a cluster in, you know, in a sec- in a, in a minute or two and provision it for as long as you need to do that job. And so what that means is you're able to be more cost effective.
It's more flexible. You can spread your risk out. You can spread your compute out over many different locations, many different devices.
If something does happen to go down, it's not like your service goes down. It moves to another, you know, location. You can move your compute to another location.
And so it's that flexibility and, and, and that access that I think really turns people to other providers because, you know, i- if... The more the hyperscalers try and control access, the more people will, will look out there for more novel solutions. And so, you know, that's...
net I think some organizations may be falling into what is... I might call the volume licensing trap, where they think they're getting a good deal from AWS or wherever, and they might be, but then that requires them to consume compute resources from them, and they force that down on everybody else in the organization. And-Then when those resources aren't available, everybody kind of like sits on their hands and winds up doing not as much as they could.
So do you think that, there's a different way to think about licensing compute resources in the cloud these days? I mean, I know we have spot prices and reserves and all these other things, but a lot of folks I talk to find it hard to manage all that stuff. So what's the smart way to go about doing all this?
Yeah, it's a good question. I think people were stuck in the way that tech used to work. You know, you used to have to like buy software for six to 12 months in, in time.
You used to have to provision the tech stack for that, because you needed reliability and you needed that security that all of your tech would work. And it's the same, I think people just rinse and repeat that when it comes to compute. It's like, right, I need to...
You know, especially when they read all these articles out there about there being a compute shortage and, you know, prices going through the roof if they don't lock them in and 'cause energy prices are going up and all this. And actually, I think what, what generally will happen if, if medium and large i- businesses fall into that trap is, especially in this, the AI space where d- you know, the inference costs cost so much for those organ- for that organization, they'll start to lose the competitive advantage because startups, like we've seen lots of startups come to, to solutions like IO, they're saving 70, 80, 90% on their compute costs because they're thinking novelly about the s- about the problem, right? It's like, why do I need to block six to 12 months in a data center when I only use it, when I, when I use my compute in more patchy, spikes as opposed to like a consistent, usage over time?
And so they're, you know, provisioning through IO Net, for example, would provision devices for a day, two days, spin it down, spin it back up again. You know, so they can actually use on demand what they require. Um, and I think that gives, that saves them money, it saves them time.
Um, it, it means that they can get the devices they require for the job at hand. You know, not everybody needs H200s all the time, you know. So maybe they need more consumer grade, or maybe they just need, you know, not the latest chip set to train their model or whatever it is.
And so those companies, I think will have the time and cash advantage over some of these larger organizations and, and that means that they'll be able to get the run on, on businesses that still think about the problem in a, in a traditional way. Do you think a lot of organizations are also defaulting to GPUs and not considering other processor classes that have emerged lately? Or are they starting to find those and use them?
Um, I mean, on IO we, we, we stick to GPUs at the minute. We do... Especially when new devices come out, they're very difficult to...
You know, there's always the cutting edge of that that are often provisioned by just the large scale organizations. You know, the, there's, it's a little club at that sort of level. You know, chip, the latest chips created by Nvidia are sold to the companies that Nvidia have a vested interest in, and it's all one big sort of web.
Um, but eventually they become available and, and you know, it, it's often just a few months and then they, they'll be p- they'll be bought up by the other, the other hi- the other data centers and then served up on, on IO. I mean, very few... I'd say the m- large majority of organizations don't require the latest tech all the time.
They don't need to be on the edge when it comes to processing. You know, what they need is good, reliable inference, often with open, open source models or models that they've trained themselves. Um, and they need it to be served up to consumers in a fast and reliable way, and that problem has been solved for today, you know?
And so unless you're sort of at the frontier and the pi- like pioneers of what a model can actually, can do or process, I, I don't think... I think, you know, saving a tiny amount on speed in the grand scheme of things is, is not really that interesting to you. Are people also starting to revisit their data strategies in the era of the cloud and AI, and are they rethinking some of the services that they might use over there?
We live in a world, or lived in a world where all, like, all of our data was ring-fenced and secured, and you know, v- we're very protective over it. And with what we, what we've been seeing is like people have become a bit more fast, not necessarily through, through our, our protocol, just as a global trend, that people have become a bit more fast and loose with a lot of their data. I mean, people are very happy to plug Copilot or Claude or OpenAI into their organization and don't fully understand the consequences of what that actually means.
You know, when you're getting Claude Code to find a bug in your code, you are essentially giving all your code to Claude. And, and you know, there may be agreements in place from, you know, that, that, that these organizations, you know, won't necessarily actively read your data. But you're training their models, you're training their infrastructure, you're training all of their systems.
And so I think it's a real risk. I think what people need to, to, and, and we're seeing definitely a trend on this, is people are starting to realize that that is a r- a risk for their, them and their organization, and they're starting to reali- starting to figure out how can we actually build this type of technology in-house? How can we bring, instead of leveraging these super models that, you know, the big three have created, why don't we use some of those open source models?
Why don't we train our own models? Why don't we self-host them? Why don't we run it in a trusted execution environment so that it's secure through the data centers?
net is really important because we have all of that, it's secure. You know that you can spread your compute out over many different devices. You can host whatever models you want on those devices.
You're not training other people's systems when you're using those models. And so I think then people are coming back to the secur- like, they're starting to realize that, you know, their IP and their cons- customer data and all of that, they need to try and protect it and ring-fence it again. net as well.
It's very much a privacy first, a privacy first infrastructure. Are you also seeing something that feels like maybe sticker shock when I start to use some of these services in the cloud where I'm consuming a boatload of tokens all of a sudden, and then I get a bill, and I'm like, "H**y c**p, I didn't think it was gonna cost that much"? Yep.
Yeah, I mean, I, I remember playing with Claude Code, back in the day and, just left, left an agent do a bit of research for a little too long. And, you know, you look at the bill, and you're just like, "Oh, d**n. " Um, yes, and I think again, that's where, you know, at IO we have, we have the cloud platform, but we're also trying to develop other solutions which, give people more protection, like, who, people who don't really know much about token usage, or, like, they're just getting into AI.
Like, we're trying to create tools and services that, bring them in. So we have our IO Intelligence, our IO Intelligence product, which is our inference product. We self-host all of the open source models, the ba- main open source models out there.
Um, you pay a subscription, and you get a certain amount of tokens each, you know, hour, day, et cetera. Um, and then people can implement those models into their product or service using our API. Um, and what that does is it just gets people, you know, used to playing with the models, exploring.
You know, if they hit their daily caps, they can see why tokens were being consumed, and, and they can then play around with, web... You know, their prompts can play around with their, the training loads, et cetera. Um, and so it gives peop- It teaches people more about how to use AI more responsibly because at the end of the day, if you're using a pure API, say, on Claude, there is no cap.
The cap is how much money you have, and so one wrong training run can cost you a lot of money. Um, and so what we need to try and do is give people safe environments to be able to test in to understand how all of this works so that they can, you know, get it right first time or get it more right first time so that they can save on cost and some can ultimately save on energy at the end of the day. So what's the one thing you see organizations doing that just makes you shake your head a little bit and go, "Folks, maybe we wanna be a little bit smarter than that"?
So I find that people implement AI in their organizations in a patchy way, right? It's kind of like a legacy organization. They're using sp- spreadsheets or whatever it is.
" Um, and what then happens is it spreads horizontally. People do that, that same thing. They just rely on a single agent to do that type of stuff, and nobody's thinking about the power of AI holistically within an organization.
Like, how do we implement this in a way that's data secure? How do we implement it in a way that isn't feeding these massive models and, you know, leaking our data to these big organizations? How do we do it in a way that's most cost effective?
How do we do it in a way that's sharing context of the business and each department and each use of AI with each other? You know, and the learnings within AI being captured. Like, that's definitely the next theme that we're gonna see in AI, and we're already seeing it sort of now, is context, context building, context management, like being able to, to be able to effectively create a memory for an...
for the organization and for it to self-evolve. Um, and so really, it kind of needs s- like, conscious AI strategy within an organization to really think about all those things, think about, you know, how we implement this and the infrastructure also to provide to it, you know? I think, like the stuff we've talked about, a lot of people are worried about rising costs.
They are worried about data security. They are worried about reliability. And so yeah, it needs, it just needs a conscious effort from an organization, I think, to put all those pieces together and almost marry that bottom-up AI discovery with, like, a top-down AI, well, you know, or emergent AI strategy within, within an organization.
So do you think that this all is conspiring to limit the pace of innovation because ultimately organizations are gonna have to maybe pick a smaller number of projects that they can either, A, afford or, B, find the infrastructure for? Um, I think ultimately the big issues in AI will continue to get bigger and surface, and surface their heads in weird and wonderful ways. I genuinely think that, there will be data breaches in the large organizations.
There will be big shocks. You know, people's data will show up in a place where they didn't expect it to. There'll be big outages.
There'll be, like, the reliability problem will become more prevalent. And so, you know, I actually think that people will become more curious, and I don't think there'll be... I think, I think the days of, like, homogenous AI where everybody's just using one particular model or one particular data, like, data pr- center provider, I, I genuinely think those days will fade out, and people will want a more distributed version of that.
People will be using lots of different types of models for lots of different use cases. Uh, people will... Like, AI g- Like, to use AI the most effectively, you have to be curious.
You have to be, creative, and I think people will start to break away from the, the big companies unless they learn to adapt and offer, offer people a fairer and more flexible and cheaper way to get involved. Um, we always see it, like I always say, like, open source always wins. You know, I, I genuinely think that in the long run, and I think that's the same with models.
And I think in terms of infrastructure, you know, these models will become more effective. They will become easier to run, and with, with novel solutions like IO, people will be able to, to get the infrastructure they need at a fraction of the cost. Um, and so yeah, I genuinely see it becoming more of a mosaic of infrastructure as opposed to, you know, you can pick from these two or three providers.
All right. Well, folks, you heard it here. The AI ride, well, it's already been bumpy, but it might get bumpier still, so buckle up.
Hey, Jack, thanks for being on the show. Oh, no, cheers, Mike. Appreciate it.
Thanks for having me. All right, and back to you guys in the studio.