The Rise of AI Inferencing with Cirrascale’s Dave Driggers
The industry is shifting from bulk training and fine-tuning towards comprehensive inferencing. Prior to ChatGPT, the focus was predominantly on training models, but now the emphasis has pivoted to deploying and using these models at scale. This shift underscores the significant growth of inferencing in the industry, marking a transformative phase.
We have seen explosive growth of AI and this trend will continue to thrive for the next 4-6 years, unaffected by economic conditions. The increasing prominence of inferencing reflects a fundamental change in how AI technologies are applied, shaping the future landscape of ML applications.
Transcript
This is Textron tv. Hi everyone. Welcome back here to techron tv.
In case you haven't had enough AI information in news today, I got, I wanna introduce you to Dave Driggers. Dave is the CEO and CTO, as well as founder of a company called Cera Scale at C-I-R-R-A-S-C-A-L-E. Just the way it's spelled in his background there, I guess on his shirt.
Dave, welcome to Text Drug tv. How are you? You're very good.
Good. Very happy to be Here. Absolutely.
Dave, we're gonna jump into S scale and we're gonna talk about inferencing, but before we do that, I thought we'd give the audience a chance to hear a little bit about you. Sure. So, uh, s scale and, um, and me, we come from a really deep, deep history of being a, actually a hardware vendor prior to being a, uh, a cloud service provider.
And we actually designed and built, uh, the largest systems, uh, at the time that were on the market. I mean, we're talking about doing 30 kilowatt racks for HPC, uh, in 2000. So really 30 kilowatts, 50 kilowatts kind of normal now.
It certainly wasn't then. Uh, we also designed and built the, the world's first eight GPU server in 2012. So, and at that three years before anybody else, and at that time we were told, you know, why would anybody ever need eight GPUs in a server?
And I don't know if you were at GTC, but about 80% of what you saw on the floor was eight GPUs and a server. Absolutely. You know, it's funny, I was up in, uh, SCON last week, Dave, it was the week before already, and they, they had a guy from a data center provider, and I forgot his name, I forgot which data center provider, but he went over the power and cooling requirements for these GPU racks that they're using now.
Yeah, it's, it's insane. It's, it's pretty incredible. No, it's, it's insane.
And it's going up. So, I mean, we were stuck at, you know, 17 to 30 kilowatts for pretty much 20 years, and now the standard almost overnight became 50 70 and now even 120, and that's in a year, you know, a good solid year and a half. And we've, we've more than doubled the amount of power that we need per rack.
And that's the standard now Yeah. For these systems. Well, so it's, it's, it's for good thing.
The, The issue is you, you can't air Cool. That, that kind of, you know, usage. No, not, certainly not air by itself.
No. I mean, um, we've gotta, in our, in our data centers, we mandate that we're bringing liquid to the rack. So we, even if it's up to the chip, we definitely bring liquid to the rack and use rear door or above rack, uh, ex, you know, heat exchangers still air inside the system, but, um, But we're, I, I think Liquid, liquid is can't do it All liquid's the future when it comes to that stuff.
Absolutely Crazy. So, uh, so that was something we recognized pretty quick when we were building these systems, uh, that half of our customers were asking us to keep the system in hosted for, and that's why we created a cloud service, you know, eight years ago was because that's the way that people wanted to consume the product. Excellent.
com. Yeah. com.
And so you will build, you know, these high performance computing rack systems for people that they could host in their own data center, or you could host it for them? No. Okay.
No, we're, so we're not a, we're, while we, while we come from a hardware background, we're used to design and build platforms. We pivoted to be a clear, pure cloud services company about eight years ago. So we, That's, I wanted to clear, Sold off the hardware business and then focus solely on cloud.
So we, uh, now we'll build and architect the entire solution, not just the server, but the networking, the, the, the connectivity, the storage, the backup, the dr, you name it. And, uh, and we are unique in this way as a cloud service provider. We do allow our customers to buy something if they need to.
Uh, we can deliver pure cloud, we own all the assets, or they can, we can deploy a hybrid where some of them buy part of the assets, uh, to have kind of a balance between their CapEx and opex. Love it. Okay.
Yeah. There, there are quite a few, Eric, quite a few industries with businesses out there where it's a mandated of their funding. You know, we, we handle research, um, AI two, which is also getting a lot of attention right now.
Alan Institute of Artificial Intelligence, their funding mandates that they own some of that they own something and then can pay for servicing it and supporting it. But it's a, it's a blend and a lot of education, a lot of research has that as a requirement. Sure.
Makes sense. So We're, we're definitely finding that a number of VCs, especially the, the bigger, bigger boys who have done this more than once and have numerous, uh, numerous, uh, uh, invest investments that are in the AI space, they're looking to have their, some of the money that's being invested, not be pure cloud, but actually own some of the assets so that if in a year they need to pivot, they don't lose everything. Yeah.
You know, because they've spent some of the, some of the money on, on an asset versus just pure cloud, they still have something after a year. Got it. Dave, if you don't mind, I'd like to pivot ourselves and start talking about inferencing a little bit.
Um, sure. You know, our, our audience is pretty AI savvy. They know, I think a lot of us have heard anyway, how, how these ais are getting trained.
LLMs used to train the, the, the model and, and then inferencing, but we an inference, we, we think we know, but how would you describe what inferencing is? Sure. So inferencing is when you actually use, when you're actually using the AI to do something.
So when you say, Hey, Siri, or hey Google, that's, that's inferencing. You're actually running a model. In this case, you're running it mostly at the edge, but also when you're, when you're driving your vehicle, that lane detection, that's inferencing.
You know, we work with Tesla early on to train that model on GPUs, and then it gets deployed and does inferencing inside your vehicle. So there's two very different, uh, types of processing going on when you're training. It's one job across a big huge machine that's got lots of GPUs, lots of servers, it's one job, one person running it.
But when you're inferencing like in your vehicles or your phone, it's billions of little jobs that are all executing at the same time across what could be millions or billions of endpoints. Yep. And, and yet, yeah.
How do we distinguish that, Dave? Um, well, more and more we're, we're getting to a situation where, uh, inferencing occurs in more than one spot. You know, you'll try to do as much as you can at the very edge, but then often it has to get support from the near edge or even back all the way to the core.
Um, you know, you don't, you don't store chat GPT on your laptop. Your laptop is being used as an interface that's going back to Microsoft to actually run it. Yep.
Well, you know, it's funny, I had a discussion this morning. We were talking about, um, uh, connected devices on the edge and, and that, so the ai, the inferencing, if you will, is not taking place on those connected devices, but some of it is starting to take place at the edge, right? As we're moving resources out there.
Oh, even, even on the devices itself. I mean, if you think about the latest Apple and the latest Samsung phones, a big chunk of that AI lives inside that phone. Now, in the United States and in other, you know, top countries, we're unique in that we are, we're walking around with a thousand dollar computers in our pocket, you know, whereas in, in India or or third world countries, they've got a very simple interface that's barely more than a screen.
So an awful lot of their inferencing is gonna have to take place in the cloud versus actually at the device. I don't know about all those starving children in India or anymore story. They got some pretty good phones out there too, I hear.
Um, but you know, if you look at the amount of iPhones and, and, uh, and the Samsung high, high level ones, it's, it's global. I mean, that's the crazy thing about this. When we, when we think about the, the, the scalable, the scalability of, of what we need for everyone to participate in this AI revolution.
A absolutely. But, and that's, and as these models get more and more complex is they do more that's gonna be done in the cloud. I mean, that just like Google glasses, when we start moving to augmented reality, you know, we're gonna have, the glasses are gonna have to be very lightweight, very simple.
Yeah. And work with us for just like our, just like we want our phone to last all day long, those glasses are gonna have to last all day long. And that means they're not gonna have much compute in them.
They're mostly gonna be communication devices that go to the near edge where the actual AI is gonna work. So Dave, so here, here's Dave Driggers, C-E-O-C-T-O, founder, sir, scale. You've been in business since the early two thousands, right?
Yes. How do you, this is a strange new world, right? What, yeah.
How do you, how do you rise to the challenge here? What, how, you know, your mission, not you personally, but sir, scale. So we at, Yeah.
So we always look at what, uh, what's impeding scale. We've always, our name sir, scale for a reason. I mean, sir, is is a sirs clouds, and then scale, obviously is scale.
So we've always been focused on trying to solve problems of scale, scale up. We built the world's first HEPU server scale out. We've been doing Linux clusters since they started.
So, and we, we definitely see the market. And, and Jensen talked about this the other day in his keynote. The, the world constantly goes back and forth between scale up and then scale out and then scale up again, and then scale out.
So you're always, you build a big box and then you scale out with a lot of them, and then you go back and build a bigger box. And the same thing's happening with, with ai, you know, we build a bigger model and then we have to scale it out, and then we've gotta go back and build a bigger model. And the other thing that's happening though, is that these models have gotten so big that, uh, we now need to make them smaller so they can scale out.
Um, and that's where we see a lot of the enterprises that they've been shown what they can do with a chat GPT, they've been shown what they can do with a copilot or Gemini, but frankly, they want their own. You know, I don't if, if I'm doing my own little, uh, copilot, I don't need 95% of the information that that copilot knows what, what to do with, uh, and I, frankly, I want my own. So I just need a small version of that for me.
So we definitely see companies looking at, um, smaller models now. As, as the models have gotten bigger and bigger, bigger, they're looking at small models, which are still massive. By the way, 7 billion, 13 billion, 32 billion parameters was unheard of just a few years ago, and now everybody's gonna have one.
So, um, but in order, the, the big challenge there becomes once you do train this type of a model and you own it, you have the problem of deploying it. Because if you're a global company, you're gonna have to deploy globally if it's, if it's a critical application to you. And that's where, that's what we're targeting to solve right now, is enterprise.
Enterprise is trying to deploy their own models. So not relying upon a hyperscaler. Got it.
Crazy. I mean, it, you know, interesting times, certainly Dave, you have to admit, back in when you first started this, you didn't, did you even comprehend we'd be here? Uh, n no, I did not comprehend it.
We'd move this quickly. I thought it was gonna be, I mean, we got to watch the, the first wave, you know, which is the, we were working with Nuance. So, uh, Hey, Google, hey Siri Auto correct on your phone, you know, uh, when, when, that was really cool and new.
We were there during the auto, uh, auto revolution, the second wave. You know, we worked with Tesla and, and Uber and Samsung and other real, really innovative guys, you know, trying to make cars autonomous. And this wave is, but this wave is bigger and massive.
It's, it's just crazy how it's coming with, and it, it kind of happened two things at, at once. The, you know, the LLMs, but also the true Gen i Gen ai, you know, where we're, where we're creating images are all of what feels like from thin air, and they're amazing. We're creating audio and music, and it's, it's, it's, it has progressed so quickly, um, that it's hard to tell the difference between AI and real when it comes to Gen AI and LLMs.
I, I agree with you a hundred percent. I mean, it's changing my business. Dave, I want to thank you for coming on here, telling us a little bit about Sir Scale.
Keep up the great work, right? 'cause Quantum's coming. Get ready.
Yeah. That, that's, well, hopefully that one's takes its time. I'm, a's I'm a little nervous About that, I'm starting to see a build, I'm starting to see a build, man, it's gonna be here.
Get ready. But, you know, one day at a time, man, keep doing what you're doing. We appreciate you.
We mentioned the website, but it's C-I-R-R-A-S-C-A-L-E, right? Sarah Scale Yes. Dot Com.
Yes, That's it. Absolutely. All righty, thank you very much.
That's it for Techstrong TV with Dave Dreger, C-T-O-C-E-O, founder AtScale. We're gonna be back with more. Stay tuned.