Introducing MinIO AIStor – Object Storage for AI and Analytics with MinIO
MinIO’s VP of Product Marketing, Jason Nadeau, introduces the AIStor object storage solution, designed for AI and lakehouse analytics environments, at Cloud Field Day 23. AIStor distinguishes itself from object gateway approaches by being object-native. Nadeau highlights the importance of object storage for AI, as evidenced by its use in building large language models and various data lakehouse tools. In contrast to the complex, multi-layered architecture of retrofit object gateway solutions, AIStor presents a simpler, direct-attached architecture, leading to superior performance, data consistency, and scalability.
Nadeau emphasizes MinIO’s object-native architecture, which provides strict consistency and SIMD acceleration, resulting in significant performance advantages. These architectural benefits translate into tangible storage outcomes, allowing customers to scale from petabytes to exabytes. AIStor’s architecture facilitates real-time data services even during hardware failures. This object-native approach enables optimal hardware utilization and cost-effectiveness. MinIO offers direct engineer support, bypassing traditional support queues and providing customers with direct access to experts. The company is seeing strong enterprise adoption and growth in headcount.
The presentation features examples of AIStor deployments in various use cases, including generative AI, high-performance computing, data lakehouses, and object-native applications, as well as an autonomous vehicle manufacturer, a cybersecurity company, and a fintech payments provider. These deployments are achieving desired performance and are helping to control costs. MinIO plans to offer a channel bundle skew, which will simplify the acquisition of AIStor by bundling hardware and software into a single SKU.
Presented Jason Nadeau, VP Product Marketing, MinIO. Recorded live in Millbrae, California, on June 5, 2025, as part of Cloud Field Day 23. Watch the entire presentation at https://techfieldday.com/appearance/minio-presents-at-cloud-field-day-23/ or https://techfieldday.com/event/cfd23/ for more information.
Transcript
I'm Jason Netto. I lead the product marketing function at min io and I'm excited to talk to you today about our latest offering. We call it AI store, uh, it's object storage for AI and analytics, uh, as well as object native applications.
We're gonna dive into a lot of what makes it different. Um, and, uh, we're gonna, uh, start by talking about a little bit the context behind why object native, uh, storage is so important. Frankly, AI and analytic analytic storage is object native today.
And this may or may not be, um, uh, news, but, uh, we definitely find that there are some folks that are surprised that like every single LLM maker built those models using object storage. And, you know, you can even ask the models themselves, or you want to ask Chad GPT, or you want to ask Claude, Hey, what type of storage were you built on? They'll respond and tell you that it's object storage and not just the big AI models, but all the data lakehouse analytics tools have become object native.
And this, you know, there's, there's a bunch of reasons we're gonna kinda get into that. But every query engine, the OpenTable formats like Hootie, like, um, iceberg, these are all object native formats. They want object storage as their primary storage.
And so for that as a bit of context, why, why is that? What is it about object storage and particularly object native storage? So let's come, let's contrast with what the alternative would look like for a lot of people, and certainly on premises.
And I want to, I want to stay right off the bat. This alternative doesn't exist in the cloud. Okay?
Those folks that are building those frontier models, you know that they're, if they're working in AWS and they're using S3, that's object native storage. Uh, so that stuff is purpose built. But on premises, what's the, what's the alternative?
It's a retrofit object, gateway architecture. And roughly speaking, it kind of breaks out into some layers that look like this. And as you can see from this slide, what these layers come with is a whole bunch of constraints and negative implications.
So that retrofit, uh, you know, unified architecture, it starts off at the top with a gateway layer doing that translation between S3 API, and, you know, underlying, uh, you know, kind of posix translations, right? And, uh, and whether it's file or block kind of underneath the covers, um, to, you know, handle those, uh, those requests. There's a metadata database of some sort that's managing where that metadata is and mapping it to, uh, the individual objects themselves.
There's a san or a NA backend because these gateway storage, you know, implementations, whether they're appliance based or software defined, they're, they started with file and block storage implications or, uh, architectures underneath, and they added object later, right? And then there's a, a storage network behind the scenes in there. So think about NVME over fabric, for example.
So a lot of layers and complexity, and what we're trying to show here is, is what does that actually mean in terms of those constraints? So for sure, architectural implications around performance, okay? Coming from, especially that translation layer, the object side, that, that NAS backend, the metadata layer, everything is impacting, uh, performance data inconsistency.
Once you start working with network storage on the backend, you lose the ability to actually guarantee strict or strong consistency, depending on how you wanna talk about that. Uh, and of course, just overall capacity scale, any database can only scale so big. Uh, and so this architecture fundamentally comes with these limitations.
And like, that's why these big frontier models, all of the sort of, the heavy duty, uh, AI that's happening out there isn't using technology like this, because it just doesn't, it doesn't work at that scale. And so, what, what does work? A very simple, this is what we mean, very simple architecture.
This is what we mean when we say object native, right? You can just visually see the difference here, right? There aren't a bunch of layers.
What it's, what we would call a single layer. Uh, the storage is direct attached. It has software that is, Hey, Jason, running across all of this.
Yes, sir. How do you, why do you think strong consistency doesn't exist in these other unified architectural solutions? Yeah, so let me, let me back up.
So, uh, for example, in, uh, a direct attached model, you can use a, you could take, uh, uh, software that, uh, can take advantage of IO direct calls through the Linux operating system to guarantee that, uh, strict consistency. But you can't get that, those, those direct calls to guarantee re, you know, strong read after, right? Strong list after, right?
Consistency. If you're going over a network, And most of the San Nas backend products today, they are data integrity is, is number one. I, I, I, I just don't understand, uh, the challenge.
I understand that there's networking problems with respect to pulling together the sequence of, of blocks that have been written. But, you know, once they get to the storage, they can all be sequenced. So, I, I guess I don't understand that.
If You, if you wanna jump in, this is, uh, ab iss here with me. AB is the co-founder, CEO of Min io as well. Yeah, everyone, uh, I'll let him also kind of chime in on this question.
Yeah. So it is, it is strictly consistent. Like block store built their foundation on strict consistency for DB two Oracle type databases.
But you see there, the strict consistency is only for a single lung. How large can you make a lu? Maybe a hundred terabytes, even a hundred terabytes put NTFS, 400 terabytes, XFS 400 terabytes, they start shaking up right in object store.
We talk about a bucket in tens of petabytes, the deployment we made last month, single cluster of what? Uh, 1,088 nodes, all 400 gigabit. It has like more than 60,000 drives, right?
Try to create so many, each drive is 60 terabyte drives. Try to create 60,000 lungs, and to be strictly consistent every operation, you're not writing to single drive, you're writing to multiple drives across multiple racks. And to do that, at automatically, SAN or NASA has no foundation.
The only way to do this is that object store owns that responsibility object store owning this. Inside this layer, we can actually make sure that all operations are automatically committed across multiple drives. And one lead after flashing, we do OO direct, which is DMA operation.
We flash it only then return 200. Okay? So the application is guaranteed, nothing is lost.
You're talking about pulling the plug on an entire data center in the middle of busy transactions. You won't lose a single transaction. You cannot do that with that many lines.
Sand has no context. Only the object layer has the context of a transaction. So with that, now as kind of the background, uh, it's not surprising.
A lot of enterprises are saying to themselves, Hey, if I wanna do AI and I wanna do analytics at the scale, um, and, and, and the scale that they want may not necessarily be at the Frontier model scale, but it could still be quite large. Uh, they wanna have something that looks like this, that they can run in their own environments, and that's where Min IO comes in. So we are the software defined object storage leader, and we're completely object native.
We are used by the world's enterprises. They run on min io. You can see here 52% of the f uh, fortune seven, uh, for, sorry, fortune 577% of the Fortune 100 using, um, min io, nine outta the 10 largest automotive and so on.
MIN io has incredible adoption across, uh, across the globe. And developers love it. Right?
That may be something that everyone in this room is familiar with. Uh, the open source heritage of Mano, uh, goes back, uh, you know, years and is extremely popular. Uh, and now what the company is doing is moving into the enterprise in a much more, uh, kind of directed and targeted way.
And we're working to convert a lot of these, these open source deployments is not surprisingly, into, uh, commercial paid deployments. And so the way that we're doing that is through a new offering. And that offering is called AI Store.
And I've, I've put up a slide here that shows, I'm not gonna go through all, all the lines in it, but fundamentally, the, the difference between the Open Source Community edition from Minow, which probably a lot of people, uh, have familiarity with and has all this broad adoption and Minow, a AI store, AI store is the enterprise version. And it is actually quite a different product. And that's the thing I, I think is a key takeaway for everybody, uh, here.
It's, it's not just the, uh, the Open Source community edition wrapped with support, okay? The code is now quite different. All of the feature innovation that we've been, um, putting into AI Store has only been going into AI store, not C edition for, you know, almost nine months or so.
And certainly going forward from there. So we launched AI Store last fall, actually last, uh, November. And we're super, that's why we're here, and super excited to share, share it with you today.
You can see just from some of the top line items there, like the fundamental difference, like Community Edition is meant for test and development. We want teams to continue to take that software, use it, get familiar, uh, with it, uh, build their applications. But when it comes time to go into production, that's what AI store is for.
It has all of the features that enterprises need to be successful. You know, whether it's the console, whether it's encryption, full replication, quality of service, and so on and so forth. Encrypt, you know, you can see all the security capabilities and so on.
So, AI Store is now Min iOS flagship, okay? And we're super excited about our customers, um, are loving it. And in fact, it is changing the Trajectory.
Are you gonna go into more depth for Those in some, but not, it depends on what it is. So what, what, what are you guys Curious, I'm curious about the Kubernetes integration, like, you know, the upstream Kubernetes and Upstream plus. What is Upstream Plus?
Like, is it like additional to Kubernetes? What, what we're saying is here in addition to upstream Kubernetes, okay, You do the Plus, the Ag store also has Red Hat, OpenShift, kuber, you know, version of Kubernetes Rancher and so on. And so what integration are you using?
Is it the, it's not CSI, right? How you integrate with the Kubernetes For that? I don't know.
I'm gonna let AB take that as well. Yeah, Yeah. Sorry.
Don't mean to like, you know, jump to another thing. That's okay. Yeah.
Yeah. But we, just to be clear, we won't be diving into that detail later, the presentation. So just A quick thing, right?
Yeah. You come. So O Object store is, uh, is basically a layer, right?
In the end, we still, everyone has to write to drive, right? Serverless still requires server. So Object Store still requires drive, and drive is a block store.
But you're talking about local drives and managing lots and lots of local drives. CSI, Uh, Uh, drive CSI is a generic block storage interface. It's not a SAN or NASA or anything.
It's just generic, right? To the SAN vendors, they always want to actually segregate at that layer and give a CSI driver. But that's not, that's not really useful.
So like, trying to build AWS S3 on top of E-B-F-E-E-B-S or E Fs, it'll fall apart. So what's the answer? How do people deploy mean IO ai store on Kubernetes?
The drive management, we actually wrote something called Direct pv. It's, uh, the technical term, the pro, the product name is Direct pv. We, that's called, uh, at a higher level.
It's a volume manager, AI store Volume manager. Hmm. Is, think of it like a distributed LVM, right?
It, it discovers all the drives when you have hundreds, a hundred thousand drives in a single pool, just discovering them, right? The right drives, ignoring the bud drives, formatting them. Yeah.
And then knowing which drives are down, is it safe to replace? Can you enter that into maintenance mode? We wrote a layer for Kubernetes, that's called a store volume manager.
It is also a CSA driver, but it's more of a local volume manager. Beyond provisioning and volume management in provisioning, sharing of isolation between multiple tenants, when you actually read and write, you are doing the, the tenants are doing direct, uh, io into the drive. Uh, we actually worked early on with VMware and, uh, uh, we helped them design VMware vsan direct.
So vsan Direct is only available for VMware. What about Kubernetes people? We actually brought, we implemented this for everybody else.
This is not only useful for Minio, even if you're running Cassandra or Elastic or Cockroach, tvb, anything you put, they need local drive. They take care of distribution, distribution, high availability, everything at their layer. They are already doing replication at their layer.
They need a local drive. So this is basically a distributed local volume manager as a CSA driver. So that is what actually, uh, uh, uh, customers deploy in OpenShift or anything else, or small deployments.
If you use Sand Bill, it works. Your, it's any CSI driver will, DRI driver will work, particularly the sand, like, like block store. Uh, if you have, if you're talking about serious performance, you need this direct PV volume management.
Mm-hmm. Am I remembering it right? Maybe I'm being confused.
Maybe I, you know, I remember it wrong, but haven't you guys been involved in the COI that was this like, yeah, Ozzy, yeah. Oh, yeah, yeah. That one, Yes.
So where, where is it? Like, I remember it sort of kind of like went silent a little bit. Yeah.
Since then, I've never heard of it. So It kept telling everyone, right? I think it generally, I think that it's kind of, um, cozy.
The, the first part, everybody thought, oh, it's going to be a unified object layer. Yeah, no, you cannot unify an object layer. Because even though the Azure Blob, S3 Swift, everything, there are multiple object APIs, they look alike, but they cannot, you cannot do one-to-one translation.
There is no equal in functionality. So there is no way to unify, even if you unify as a, a p proxy, just doesn't work. So cozy was supposed to be just a CSI equal, and that's provisioning attached detach.
But we always kept telling them, even like the, the gen, the, the users understood, but the general infrastructure community, they thought Cozy will solve this as a CI side driver. What they did not understand as CIS side driver is to actually pro, um, uh, take a block store, which is always protected. It was always built around the block, right?
Exactly. Right. So you need always support.
You need, uh, Kubernetes support fundamentally to make, to have access to block store. Mm. But when your application, you're talking to an object store, which is a htt, B-T-A-P-I, all you need is a cert, and Yeah.
Is The additional complexity, right? So You really don't need, so do you actually have a cozy driver for Postgres for Elastic, right? Like, you don't have, so there, you have basically Kubernetes jobs to create db uh, apply schema and stuff like that.
When you provision a tenant, you actually don't have a CSA driver to create a database, uh, create tables. Mm-hmm. And why do you need something for object store?
You don't need it. We kept saying this, and some users, customers started recognizing, they said like, you don't need it. And then we are like, thank you very much.
Then we can put this debate outside. And cozy kind of quietly disappeared. But what matters is expect every object to a vendor, just like any other, any other database vendor.
There need to be an A-P-I-C-R-D to provision tenant, sorry, the provision bucket apply policy, uh, uh, and then that's integrated with Kubernetes access controlled mechanism, uh, that everybody should do. It's the Kubernetes has clean mechanism to do this, already use that, because these are all user user, uh, like a PL layer, uh, uh, semantics. You don't need OS level Kubernetes core, uh, Support.
Yeah. So, so you can do the API call to provision the buckets security, you know, and like, you know, you can like assign the tiers to the buckets if you need to, right? Yes.
You don't need like a ali, like, you know, something like the CSC. Yeah. CC, yeah.
So what we did, like, what people used to do typically was they, they were actually scripting around mc, uh, in the, in their container startup. When they provisioned, they actually said, no, I can give you a cleaner Kubernetes native approach. You have a YAML to make a request, just like a CSI volume claim.
You have a AML in the yaml, you request, I want a bucket with this much capacity, this much, this kind of access control, I'm authorized to do it, check my Kubernetes credentials. How would you do that in a YAML file? How do you give it to Kubernetes?
At CRD, there is an admin job. Mm-hmm. So, admin job seems to be the most native Kubernetes way admin job through A CRD.
You don't need Kubernetes core, core support. Kubernetes already have the mechanism to allow this extension. So using the, you are using the CD?
Exactly. Oh, okay. Yeah.
So CRDs are extensible, and that's why it's customer resource definition. So any CRDs can do all of these without making core changes to Kubernetes. But why do you need that?
Why do we need a CSI driver for block store without Kubernetes and daring, you can't access low, low level resources. Yeah. Yeah.
Okay. Cool. Thank you.
Alright. Thank you. That's a, that's a deep conversation.
Oh, yeah. I didn't mean to hide. He'll Go as deep as you want.
Maybe we, We can take it off on afterwards. So, AI store has changed our trajectory as a company. It's really exciting.
It's one of the reasons, um, why I'm here. And so, uh, tremendous, uh, revenue growth, uh, some really big deals. And I'll, I'll talk about, uh, a couple of these deployments here in a little bit.
And headcount, right? So over the last 12 months, we've already grown, uh, headcount by about 50%. Uh, it's really, uh, things are changing every day.
It's really quite, uh, quite exciting. And, and where is Min io AI store getting used All over, frankly. And so this picture, there's a lot on here, but like, we can be used across that hybrid multi-cloud environment.
So running at the edge, you know, gathering sensor data, shipping it in the core, of course, is where we are meant to run in that private cloud, uh, on-premises data center, whatever you wanna call that. That's the AI data infrastructure where all the, the action is happening in terms of aggregating that data and then doing analytics on it. Uh, and, and AI model training as well.
Storing not only the, the, the raw data, but the vectors, embeddings models, you know, kinda you name it. Uh, but of course also using AI store for perhaps more pedestrian things like archives, right? Backups if people want.
So it's a, it's a very flexible platform. And then even pushing data out on the right hand side, you know, back out to the cloud, back out to the edge. Uh, our binary is a little less than a hundred megabytes and can run in all sorts of areas.
So not just on-prem, um, also in the cloud and at the edge. So, uh, really a foundation for all of the AI and analytics that, uh, that organizations want do or organiza organizations. Sorry, I want to, want, want to, uh, achieve, so now I wanna get into, we, we talked about this object native architectural challenge, right?
With object gateway type solutions. What are we doing, frankly, different? So the foundation of our architecture goes a different direction.
So there is no gateway, right? There is none of that translation that's happening. It's stateless.
There's no metadata database. All of the metadata is written anatomically with the actual data. And in fact, there's a, um, uh, a determinist to cache approach that avoids the need to actually have a central database at all.
Really, really fascinating stuff is direct attached. So those three things, gateway, free, stateless, and direct attached form this foundation that eliminates the core challenges of that multi-layered object gateway architecture I showed earlier. And then it enables higher level, uh, capabilities, also architecturally that are unique to us.
So, for example, the strict consistency guaranteed, as we just talked about a little bit ago with AB up here, guaranteed read after write list, after write correctness. That's only possible at this large scale that we're talking about because of this object native foundation. Similarly, uh, we take advantage of a bunch of SIMD, simm D acceleration capabilities on at the chip level, other, uh, uh, offerings that are out there.
They've got bigger performance bottlenecks to worry about and limitations. Doing this isn't gonna help, but because we've already eliminated the, the, the majority of those bottlenecks, just from the foundational architecture point of view, we see real material gains there. And that in, in improves the overall performance to allow data services like, uh, encryption like, um, erasure coating to happen in real time, even through failure.
So like gets on objects, even as hardware is failing in it, it's gonna fail significantly, you know, in a large scale environment with, you know, a thousand nodes. That stuff is failing all the time. The, the super fast performance enabled by these different layers of the stack, um, become, uh, critical to supporting, you know, those applications and, and, and keeping them up and running and keeping them performant.
So that object native architecture itself then delivers actual storage outcomes, which then of course will deliver, um, uh, AI and analytics outcomes. But at the storage level, this is what is enabling us and our customers to scale from petabytes through, up to and through exabytes. And we'll all, yeah, I'll talk about a, a a couple examples here in a, in a, in a second.
But that ability to scale is a, a huge differentiator for us. So is performance. We are gonna deliver maximum, uh, performance from whatever hardware that our customers choose to put under, uh, AI store.
So they want more iops, they can put, they'll put, you know, kinda like smaller faster drives if you will use those NVME lanes, um, uh, you know, or they can go for more capacity and if they want more throughput. So it's whatever the hardware is that, that we're, that we're given, we're gonna take maximum advantage of and get line speed throughput. Uh, you know, even on these, as AB mentioned, 400 gigabit networks, for example, even going to 800 best in class, total cost of ownership.
So this performance means we can get the performance that, uh, uh, um, uh, an efficiency out of a smaller amount of hardware than our competition can. And that drives down cost of ownership, of course. So does the simple fact that we're software defined and cus our customers can choose whatever, you know, commodity, uh, industry standard hardware they want to use behind it.
And then we're able to, because the, our architecture is so simple and so reliable, we are able to pair that up with direct engineer support that then, uh, it totally skips like things like L one, L two, uh, uh, support queues and gives our customers direct access Of the, uh, the object store, the AI store compared to something like GPU direct over files and things of that Nature. So, so AB will actually talk about GPU Direct and what we're doing with GPU direct in his session. So if it's all right, we'll hold that question for now, but e even I would say this, even without GPU direct, we're saturating the network through to those GPU servers.
Okay. So GPU Direct will help us in, in, in some different ways, um, but we're already, uh, as, as fast as the network, uh, is today. And so, so by any hardware you mean what?
Any hardware? Well, so like any, Like any servers, like, you know, what do you mean by any server? Yeah, so we have, I'm sure there's not like, you know, not any server.
That's Right. So we have recommended servers that, uh, that we, you know, we suggest to our customers. So we tend to partner with Super Micro, with Dell and with HPE.
Mm-hmm. And we will recommend particular configurations, uh, for them, depending on, again, whether they want to say, for example, optimize for guess, It will depend like, you know, HCG, like, you know, SSD and VME Yes. Or That we can support That's correct.
People can use if they want. And we have customers that say for, you know, they want a cheap and deep deployment, right? And they're gonna use hard drives right in, in there.
They can't, and, and perhaps they might use that as like a cold or colder tier. Mm-hmm. Um, but mostly these days are customers are choosing nv, ME mm-hmm.
Right? And going for performance. I Wanna add a little bit actually.
Yeah, please, please, Yeah. Qualify. So actually Jason is right actually about any server, we mean it, right?
Because there is a difference between us and others. Like we see most storage companies now, they are actually claiming they are software defined too. There is a big difference in saying software defined in our world, right?
We actually, if you look at m it's a key value store. It's a data store. Amazon understands this, Azure, Google understands this.
If we go and tell a key value store, like imagine if I as a key key value store, which is a data store, uh, it's like the simplest of all databases. Mm-hmm. If you're a, if you're a a, a cockroach DB or say elastic or some vector database, you go and tell your, your architects, Hey, this is a software defined database.
They would laugh at you, right? But mean I will being an object store, which is a key value store, we have to explain it is software defined because historically storage was sold as hardware appliance, right? In our case, when we mean any server, it is very much like a system, minimum system requirements.
Like there are same people in our, uh, in our customer base, they actually run miao on their QAPs analogy, even raspberry pie. And then you see their daytime work. These are like, uh, 64 nodes, 64, like, uh, these are like the, some of the lar larger clusters are a thousand plus nodes.
Same software. And, uh, we mean, uh, any, any server means software defined in our world is downloadable software. Just like you download any software like Microsoft Fire, like it's Firefox or browser or anything you download and run, as long as you meet minimum system requirements, Minow being super small can fit anywhere.
But if you're in a serious production environment, a typical requirement, we, we do have recommendations you today per making the most value. You get a direct attached storage, a server with a bunch of NVME rice, and if you have NVMe Rice, even 400 gigabyte is not going to do the justice. But that's the best we can have.
There is always going to be some bottleneck a hundred gigabit. Sure. If you want to put 10 gigabit, you're lose, you're not taking the value out of it.
Right? In, in today's times, actually our customers, we, we actually tell them no longer hard drives. We don't, we don't, we don't want to support hard drives anymore, even for archival, what customers are doing.
But very large scale, when you are talking about these are like exabytes of single cluster like exabyte and why can't I have a tiered approach? They find that hard drives are, don't make sense anymore. Their capacity is small.
The operational cost there, uh, the failure rate is so high. People are more expensive than hardware. They actually went all flash.
We actually, for, for us, we actually now we tell our customers do not buy hard drives anymore. And even customers only started telling us, telling this to us, as long as it's a, it's like a, it's like a industry off the shelf hardware. Every server vendor from super micro, Leno Wood, HPE, Dell, name ita, all of them have a server, a bunch of drives, and a Nick and a CPU.
That's all we want. Mm-hmm. Okay.
Yeah. Thank you. A question here regarding support, you said that you don't need L one L two, so that means that, uh, you support, uh, remotely, uh, minio, uh, infrastructure.
It, what it means is our customers have the ability to contact our engineers directly through a portal we call subnet. So, yes, okay. Yep.
So that, it doesn't mean that, uh, you manage infrastructure, Correct. We're not managing their infrastructure, but they, they're, they're not waiting in a queue to get into, you know, in touch with the right person if there's something, um, that, uh, that they need to talk to us about. Yes.
It's a very different support experience. Sorry, I didn't mean to jump in. Last Question on the hardware X 86 x 64 Arm.
Yeah. Arm. Yes, for sure.
Rb. Okay. Yeah.
Even Power. We actually are portable everywhere, not just that we, we support, we actually have sim d acceleration on all those platforms with assembly language. Okay.
Okay. Cool. Thank you.
Alright, so, uh, not only storage outcomes, but actual workload outcomes. So here's, I'll talk about just a few examples. So you see the, the use cases that were really being used for, uh, more and more are generative ai, HPC ML type stuff, data lakehouse and analytics, and then object native core applications.
So, uh, there's actual, you know, real companies behind each of these logos. These are anonymized. I, I'll for example, talk about a large autonomous vehicle manufacturer that's using AI store.
So this is one of the largest private cloud AI deployments on the planet. And, you know, they are have a custom training application that's, you know, using the data that's coming off of all these cars, right? Video telemetry, um, logs and whatnot.
35 exabytes. Okay? Now, they, before they, they used AI store, they tried pretty much every, you know, call it AI storage platform you can think of, they all fell apart at roughly 20 to 50 petabytes.
Okay? Uh, but today the largest a a B actually mentioned the largest, uh, cluster. They're actually, uh, using it as 1,088 nodes.
It's 700 petabytes in a single name space. And they also have multiple 500 node clusters, uh, each with about 350 petabytes in a single, in a single name space each. So really big scale there.
Uh, on the Lakehouse side, uh, also, uh, this is a, uh, a leading cybersecurity company, and they've got one of the largest, uh, private cloud data lakehouse deployments on the planet. And so for them, it's all about security analytics and incident response, which is, you know, that's mission critical, uh, for their business. It actually is their business.
25 petabytes on AI store. So, uh, an incredible amount of data there as well. They, um, lemme see how, I wanna make sure I got my facts right, uh, here.
Uh, they had actually repatriated initially a few hundred petabytes out of AWS and, uh, as in, in order to do that, they needed full S3 API compatibility, which is one of the key things that, uh, MIN IO has delivered for years and inclu and, and, and continues to do so with AI store. And so that plus the ability to, to to scale as big as we do was really what led them to, uh, to AI store. In fact, they saved so much money in this migration, uh, and, and subsequently that they improved their gross margin about two to 3%.
So really actually material for their, for their entire business. And of course, the, the low TCO, uh, and simplicity of AI store allows 'em to continue to scale, um, very cost effectively. And then on the right hand side here, we think about object native business application.
So a, a large leading, uh, FinTech payments provider, you know, is, uh, using AI store. They have literally almost half a billion merchants on their platform. Like that's a lot of merchants.
And about every month, every week, and on an ad hoc basis, they're generating transaction reports for each and every one of those, uh, uh, merchants. And so it's a bursty application that they have. Uh, and it's generating billions and billions of small files.
And the object gateway based appliance that, uh, they actually had appliance and some software that they were using just simply wasn't able to meet the SLAs for data processing and replication that brought them, uh, over to, uh, to MIN io. Uh, and at that point, now they're running about 30 petabytes and planning to go to about 50 petabytes over the course of the, of the next year. So these are all like significant, uh, sized deployments.
Uh, and they, um, uh, the last thing that I will say is that they're absolutely achieving full performance for their environment. They're not missing any SLAs, and they're getting lower cost of ownership than they had before. Um, with such a, um, substantial deployment, uh, you have thousands of hard drives right?
Under or Drives anyway. Yes. A lot of these MVME drives, yeah, not much hard drives anymore.
Think, Uh, uh, this proactive, uh, um, failure, uh, recognition before it happening. 'cause I assume if you have one point 25 exabyte of data must be thousand. Like really a Lot of this drive things are, yeah, there's a lot of drive.
How, how you, You picked that Something will break before it break. I don't know the answer to that. Ab is that something you want to be comment On?
Yeah. Yeah. So, uh, once you cross just even a hundred drives, right?
You have to assume that in a scale out environment, failure is normal. And then if you have to send people middle of the night, these are software people. Like if you send the people who also don't understand software, they pull the wrong drives.
Oops. It always happens, right? You have to assume failure is normal at, uh, at like a hundred thousand drives.
You, you like, what we are also seeing is the drives are getting denser and denser. They, we went from a TB drives now 30 and 60 tbs. Now the new rollouts are all 60 TB drives.
And then we are already evaluating 1 22 TB drives the clusters. We, we would, you would assume they will shrink. They're only getting bigger and bigger.
So the, how do you deal with these failures? You have to assume they're just normal. If they fail, let them fail in place.
Once in quarter, once in a, we actually once in a year, you can send somebody to just replace these tribes early on in fer quite a bit. Uh, we actually kept the default as eight parity because we always said people are more expensive when they fail. You're talking about half of the machines in your data center can die and your data is still safe.
But when you're around 35% failure, you have, it's long due, around 40%, you start decommissioning, the machines don't troubleshoot. We kept this for a while and then we started seeing a behavior in the enterprise. They already have all the RMA process set up.
They didn't have a problem if you told them these are bad drives, they already had the warranty and they would get it replaced. So what we are telling them, maybe once in, once in three months, once in six months, we can even tell you like, say four, drive failed is nothing. Right?
Even 400 drives failed is nothing. But if you drive, if, if you have four parity and four drive pay failed in a singular era set, we will tell you these are the drives you should just send somebody to go replace. We have that knowledge.
That's why end-to-end hardware, operating system, everything that supportability is baked into our product. Mm-hmm. But then when it comes to hardware, customers know how to get RMA done.
But we, they, if they expect us or if they expect the hardware vendors to un understand and troubleshoot and do these things, nobody has any idea. So we want it. Yeah, sure.
But, um, when you have this a hundred thousand drive system, uh, you give information, uh, about predicted, uh, failure of those disc to the customer before it happening, or it's just like, okay, it failed wherever. Next one. Yeah.
They, the drive vendors promise that they have this metrics and they can tell you inside, right? Never work. Never.
Yeah, you could. So we couldn't rely on it. Instead, it is okay, if I don't have the visibility, because this is hard, can they, they're, they're supposed to give us even NVME commands.
There, there is enough, enough provision for them to tell us that it's now reaching. Its, uh, the enough wear level happened and can't hold up anymore. They're supposed to tell us information.
There is no clear standard. There is standards, but are they following the standards? I think still long way to go.
Instead, it's a software problem. I think here is where the meta engineers, Google engineers, like working with them, collaborating with them the right way to deal with this. That if fix software problems, software is easy to fix, hardware problems, just, uh, uh, don't even troubleshoot those problems when they fail after that, you, it's, it's a, unfortunately, it's a reactive approach, but software is able to handle it because of erasure code.
Mm-hmm. Mm-hmm. Thank you.
Thank you ab And so last slide. Uh, we launched AI store last fall and it's been a, a, a stream of innovation since. And in fact, the, the next two sessions, we're gonna kind of dive into this so you can see some of the things we're gonna talk about.
Some, uh, the left hand, uh, side of the slide has already been, uh, ga. Uh, but we're also gonna talk about some things that we've pre announced, uh, coming, going into the future. Uh, and so a lot of innovation specifically focused on AI and on analytics.
Uh, one thing that's actually unique that's a bit different though, is if you notice there on the, what we're calling horizon one into the future, here is a channel bundle sku. So, I mean, IO AI storage is software defined, but we do have a segment of the market that would like a simpler way to, to purchase something that is kind of bundles the hardware and the software in Aku to, to, to make it easier to acquire. So that's something that, uh, that we'll be, uh, excited to roll out, uh, in the, in the coming months as well.
So, uh, with.