Cloud Architecture Trade-Offs with Qumulo’s Ryan Farris
This discussion focuses on the trade offs cloud architects were required to build around using traditional file systems when constructing AI infrastructure. Ryan Farris, VP of product at Qumulo, also discusses how Azure Native Qumulo resolves those trade-offs, decreasing GPU time and significantly lowering costs without sacrificing performance.
Transcript
This is Textron tv. Hi everyone. Welcome back here to Techstrong tv.
Um, and our next segment I want to introduce you to Ryan Ferris. Ryan is VP of Product over at Culo, and we're gonna be talking about Gen Cloud storage for AI workloads eventually. Before we get to that, I wanted to kinda give you all a little bit of Ryan's background as well as, I'm not sure if all of you are familiar with kumalo and what they do.
So, Ryan, welcome to Text Drug tv. It's great to have you on here. Thank you so much, Alan.
It's good. Be, uh, back on again. I think I was on your program ly about, uh, nine months ago when we actually launched Azure Native mlo.
So thanks for having me on back then. It's good. Be back on now, Alex.
Just know what, thanks for reminding me. And hey, if he was on here, it's on Text Drunk tv, so if you go to Text Drunk tv, search Ryan Ferris in it. Good luck.
Nice. Nice Ryan. Yeah.
Opportune plug. Alright. Um, uh, so, so a little bit about my background.
Uh, I kind of cut my teeth in the technology industry doing storage stuff. Uh, and back then I was an Isilon from 2005 all the way through, um, IPO and acquisition to 2013. Wow.
So a lot of the, a lot of the same people that I worked with at Isilon, which is now Dell Power Scale, uh, Kenneth, uh, went from that world and into Qumulo. In fact, one of the founders of Qumulo, actually, the co-founder is from Qumulo, uh, came from Isilon, came from that world. And my current boss, he's the CTO, his name is Kiran bpo.
I used to, uh, rub elbows and work with him back at Ison. So a lot of these folks, uh, particularly from engineering and product, have come into Qumulo and over the last many years have built this system, uh, which is a hybrid cloud hybrid. It's on prem, it's in the cloud.
You can do a lot with this system. Uh, for those have that, uh, have not heard of Qumulo. We're a data platform that provides file and object storage, uh, for software-defined storage, uh, on premises and in the cloud.
And so we allow our customers to scale their data anywhere. com, you'll see that phrase scale anywhere kind of peppered throughout our site. And so storing data in hybrid IT environments from the edge to core data centers, and of course in the public cloud, uh, is what we do.
Yeah. A little bit of, Uh, I hear you and it's a great, it's a great s story in there. com.
Yep, that's right. com. Like the cloud, if you're familiar with that name ullo and UL Cumulus Clouds.
com. Yeah. Excellent.
So Ryan, look, everybody's about AI today, right? And yeah, though, interestingly, and I don't know if you're seeing this, we're now starting this, you know, that Gartner hype cycle is always dead on, right? So it we're, we're starting to see may, I'm not gonna call it the trow of disillusionment.
Yeah. But people are, you know, saying, well, it's not working the way I thought it was gonna work, or it's not as useful as it was, or some of the things we're building around and aren't kind of being loved as much as we thought they'd be loved. Uh, but that's normal for these kinds of things.
Nevertheless, it is certainly transforming a lot, a lot of how we work, a lot of things we do. And, and from an IT perspective, it's really, I mean, the promise of being just game changing is, is just, it's there. So I would imagine we've got, you know, just last week, right?
Uh, uh, Nvidia became the most valuable con company video on the, in the market in the world. Yeah. More than Microsoft, more than Apple.
Yeah. Yeah. I know.
Almost Got some big plans. You just can't get away from it. Any tech article that you read, it's likely to include either AI as a side note or have that be the topic.
Uh, or just It's written by it. Yeah. Or, or that's right.
That's good point. Or it might have been written by, or mostly by chat PT That's right. Well, I wanna talk about, uh, AI in the cloud.
Uh, let's, let's talk about AI in the cloud and then maybe we could pivot to AI OnPrem from a storage point of view and from the through the storage lens, uh, since we're storage Company. Okay. Yeah.
Good. So tell me, Yeah. So I wanna talk about the role of storage, uh, for AI solutions in the cloud.
And that's kind of where our, our largest enterprise customers are pushing Qumulo AI in the cloud. So let's first talk about that. Um, if you are a customer with a bunch of GPUs that you either procured, either by reserved instances, or maybe you want to pay by the hour and use Nvidia GPUs and either AWS or in Azure, it's very likely that you're gonna be training data that's derived from cloud native object storage.
Uh, cloud native object storage are like, like nutrient rich soil, uh, in a garden. So they provide all the essential elements for AI models to develop and grow. So that's kind of where all the data is sitting.
And the challenge is that you have to get at that data, at that object data such that GPUs can be kept active. And so imagine you have a couple of petabytes in S3, or maybe it's Azure blob, and training that data from file and pix dependent applications is a big challenge. Just the getting all of that data from object or the data that you need, um, into these repositories where you can keep NVIDIA GPUs active, right?
And, and one of the things that we do in Azure, native Qumulo in particular, we solve that general problem for AI workloads. And so one of the advantages that we have as a solution is that we provide performance elasticity and we're able to perform elastically because our final system sits on top of object, right? Mm-Hmm.
So we're acting as, uh, transactional cache, if you will, for SMB and NFS clients, uh, that are pulling data from an object layer or from object storage into this transactional cache. And this architecturally effectively accelerates that performance of transferring massive amounts of data from object storage to file storage, such that you can keep those GPUs as busy as possible. Okay.
Does that make sense? Makes all the sense in the world to me. Yeah.
Yeah. So yeah, that's one of the challenges that we hear quite a bit, um, s**t, is we, we wrote a blog. It's always nice to look at a picture that kind of depicts this general problem.
We wrote a blog about, uh, two weeks ago, and the blog talks about a benchmark and your readers and some of the folks that might lean storage nerdy have probably heard of spec SFS and now that's renamed Spec storage. Have you heard of this benchmark, Ellen Becca? I, I, I remember hearing of it, but not far from, uh, an expert on that.
No worries. All right. So it's explaining, yeah, this benchmark kind of came on the market like in early two thousands, maybe mid two thousands from the NetApp founders, and it has grown and sort of, it's still a very relevant benchmark, all right?
Mm-Hmm. Um, and they have an AI specific benchmark that they shipped about, it was probably like two, maybe three years ago, right? Where what we were calling this ML back in the days when this used to be machine learning, right?
So it's about 2021. All right. So this benchmark, it does a really good job of synthesizing, uh, common file sizes and IO patterns from AI workloads.
And this includes like large ingest operations from object, for example, large ingest operations that require pulling much of data from object into, into cash. Uh, it includes model trading based on IO parameters from like the ubiquitous TensorFlow AI framework. Um, it includes model checkpointing so that, uh, you can issue smaller transactional rights such that if that model fails for any reason, then the application knows where to reference back to and restart that job.
And so this spec, SFS or spec storage benchmark includes all of those things, and synthetically gives customers an idea of how fast your storage will perform. Okay? So we ran this benchmark recently, and, uh, we were pleasantly surprised in the first run.
Like, well, let's actually make a go of this. 84 milliseconds. And that is the fastest benchmark of its kind in the public cloud.
And what that means, and kind of what that depict on is that you're scaling, like the mental model that you can develop is that I have one AI job and I'm scaling to 700. Okay? So incrementally, I'm just getting more and more stuff, uh, that I'm, that I'm training to keep those GBU busy, okay?
Mm-Hmm. And the, probably the, the speed is one thing, but the, perhaps the more important part of this is there's always a cost benefit trade off to the performance that you're paying for and the price point that you can expect or incur. And throughout the life of the benchmark from zero to 700 jobs, we recorded the cost of the benchmark, and it was about $400 to run the entire thing.
Now that, that might, is that good or is that bad? Is 400 bucks a lot? The reason why that's cool, and then why that's novel is that, uh, after that job is done, you're not paying for anything else.
So $400 is representative of the elastic nature of our system as your native Qumulo and what you pay for, for the performance when you need it, right? So let's say after that job is done and maybe wanna do one, two weeks from now, again, can tear down the entire infrastructure of GPUs, not pay for those. And Azure Native Qumulo is also no longer charging you anything.
So you get this great set of performance capabilities, an awesome set of, and predictable set of, uh, uh, performance metrics, right? Can a cost that's just totally unbeatable. You know, I, I mentioned before we were talking offline that I just finished through wrapping up a Textron gang, uh, episode.
Yeah. And one of the subjects on the episode was finops. Yeah.
Right? And cloud, cloud costs, and how AI is now being applied to, you know, make our cloud usage more efficient and elasticity and everything. Yeah.
And, and you know, you said, well, it's 400 a good or a bad number. Yeah. It depends.
But I'd like to know it's gonna be 400 every month, not just spin a roulette wheel and, and see where it lands, which unfortunately is the state of, of being for many companies today, right. Month to month. Right.
If they can figure out their cloud bill, because they don't, they don't make those easy to read. Right. You don't, you know, you, you, you expenses, they seem to always go up.
They very rarely go down. Right. And that's, that's part of that too.
Now you are saying you're calling it Azure culo, am I to take it that it runs exclusively on Azure at this point? Yeah, that's right. Azure, native Qumulo, uh, Azure, excuse me.
I call it Azure, but Azure. Azure, whatever. Potato, potato.
Uh, Well, it's my fancy Fred accent, but, but so it is, it's it's Microsoft specific at this point. Yeah, that's, it's very much like a first party service. But what we've done is we're, we're deeply partnering with Microsoft, and we have partnered with Microsoft to launch this, um, looks, feels, and acts like a first party service.
Um, if a customer has a giant a Mac or an ED an EDP like program, and then they've spent 10 million with Microsoft, uh, spending some of that money on Azure native Qumulo counts toward that Mac, um, customers want to purchase it. They can do it through the, a Azure marketplace, just like they would a first party service. And it's built on Azure native, uh, infrastructure primitives, uh, like object storage.
So the background, and the reason why we can scale up to a hundred gigabytes a second, not gigabytes, but gigabytes a second, is because we're using those primitives from Azure and Azure's infrastructure network so efficiently that we can scale up from zero to a hundred gigabytes and then back down again. And then really truly only charge for that performance that's needed. Same thing on the storage scale as well.
If you wanna scale, if you have reason to burst up and, and put a few hundred petabyte, or sorry, few, few hundred terabytes on the system can delete a lot of that data. And after you delete it, you're not paying for it anymore either. And I'm kind of continuously underscoring that point, Alan, because, uh, in, in other cloud native storage, uh, systems and vendors, really what you're getting is fixed storage in the cloud.
So it's a fairly similar kind of experience that you would get with on-premises storage. Um, it is provision a 75 terabyte volume or something like that. And that's what I'm paying for annually.
And it's kind of kind of crazy that until fairly recently, until Azure native Kumu came along, um, that cloud native kind of service and privitives were not offered for file storage specifically. Now they are. Yeah.
Fantastic. You know, and that is, you know, one, are the advantages of using cloud, right? And some people thought, oh, I used the cloud to save money, but no one ever held the cloud out as a, not necessarily cheaper than on-prem at, at certain scales and stuff like that.
Yeah. Yep. But it was the elasticity, the burst stability.
It was also the idea of let someone else maintain that infrastructure because yeah. A we may not have enough resources to do that in-house. It's much more expensive for us to do, and it's not core and critical to our business.
Yes. Right? Yes.
Microsoft, the Azure team's pretty good at maintaining Azure and, and Yes. And continuing to innovate on that. Yeah.
I would rather ride on top of that than have to reinvent the wheel down here as just an individual company. Yeah. And so that's important to remember.
And, and, and so when I could have real elasticity, not just burstability, but real two way at elasticity, let's call it Yeah, yeah. Now, but Ryan, for me, the, the, the, the real, you know, nitty gritty is how hard is it to elastic to shrink back down? Do you have some sort of ai, whether it's Jen or or ML or whatever, that's gonna remind me that says, Hey, it doesn't, you know, you're not accessing this Sure.
Or whatever that I makes it easy for me to bring it back down, Right? Yeah. I think there's kind of a two-part answer there.
Uh, we actually do have AI embedded within our architecture, um, and it's, it pertains to our prefetch algorithm. Uh, and now this is a little bit in the weeds, but basically what, while we're doing this, we're always guessing what blocks to prefetch into cache so that we can serve those files the most efficiently and with the highest performance as possible. That's insider architecture.
It's something that users need and worry about. It just works. It's one of the reasons why our cache hit ratios are well above 95% across the board.
Uh, so all workflows and kind of all industry verti verticals benefit from that algorithm. So that's one part. The other part is that, uh, you're kind of alluding to the fact that you can get into trouble in the cloud if you're not careful.
And you're absolutely right, Alan. Um, I spent a half a decade at Amazon, and when customers complain about their bill, they would either usually, uh, avoid best practices and setting up billing alarms and things like that, and they would kind of abuse some fundamentals of cloud, setting up cloud infrastructure. And I'm sure you're like your DevOps, your SecOps, uh, folks that are listing are probably very appreciative of that.
Mm-Hmm. Um, everybody's been bitten by cloud builds. So with Azure Native Qumulo, we wanted to create, uh, the, the simplest experience as possible and, and, and kind of eliminate that burden of having to worry about what am I paying for?
So it's, we're giving you the elasticity when the application needs it, and when it no longer needs it, we're backing off and we're metering by the minute. Right? So if, if you have an application that is going from a gigabyte a second to 15 gigabytes a second in some high performance compute workload, the moment you're scaling back down and contracting your compute farm, your application, uh, infrastructure, we're no longer charging for it.
And you don't have to worry about it. So the customer doesn't have to do anything particular or special. They don't have to set up alarms or traps and things like that.
They just get it when they need it. And we don't charge when they don't. Excellent.
Yeah. Good stuff. Um, Ryan, where can people find out more information on this?
Yeah. com. Mm-Hmm.
com, you'll kind of look at the, our mega dropdown, and you can see all of all of our cloud products, not just Azure Native Qumulo, but we also have an AWS offering that we're rapidly evolving as well. Uh, and that you could look at pricing as well and all the aspects of our pricing Hold, don't jump in on the AWS stuff. Sure.
Kinda what's the difference between that and the Azure? Yeah. Uh, yeah.
Uh, yeah, I kind of avoided the, uh, channel, the AWS offering, mostly because our benchmark pertained to what these results that we achieved on Azure. But, uh, a little plug on AWS for sure. Um, so our AWS offering is in public preview right now, and we're doing this case by case.
So, uh, any large enterprise customer that wants to kick the tires and even put in production, we're allowing that to happen right now. But one big difference is that it's not a managed service on AWS, so we're deploying via Terraform and we're deploying, essentially giving a lot of these keys and a lot of the capabilities that I described to the customer to manage, it's a little bit different, right? It's not a service, it's you deploy your own infrastructure, but your EC2 instances are completely disaggregated and separated from the, the persistence layer that sits in the object layer or S3, right?
So you still get the same amount of elasticity and the same capabilities and value that I described earlier. It's just in a different form, and it's not a managed service. So I think around Q3 ish, uh, is when we're gonna make a much bigger splash about that and a push into the market, uh, and be much more aggressive about what that is.
Uh, you'll start to see it kind of pop in into our website and be more prominent there as well. That's cool. I I'm glad to see it expanding like that.
And it'll be, as you said, let's be clear, the managed service that sits on Azure, and it's kind of well settled at this point. It's benchmarked, it's it's known commodity. The a not the DAWS service will be less than, yeah.
It'll just be different then. It'll be different. Yeah.
Right. Yeah, that's, yeah, That's right. Yeah.
Uh, we made a little bit of a push, uh, in NAB, uh, with AWS in fact, we were in a W S'S booth with, uh, a couple of their fantastic essays or solutions architects that were demoing this. So you can even find things on AWS's website that show Qumulo as the data store on top of which they're running render farms and, and some pretty intensive m and e workflows. So, Absolutely.
Good stuff. Hey Ryan, we're about outta time. I want to thank you for coming on here and keeping us posted on all the latest and greatest from Culo.
It's an interesting time. You know, a lot of people say that's storage, never nothing new in storage. It's an interesting time and AI's changing a lot of, of the games here.
So keep up the great work and keep us posted. Okay. Awesome.
Thanks so much for having me, Ellen. Appreciate It. My pleasure.
Hey, Ryan Farriss, VP of product at Culo. Go check out the, uh, Culo is Azure offering as Azure offering, as well as at least AWS offering coming up. We're gonna take a break on text Trunk tv.
We'll be back in a moment.