The Vital Role of Data Orchestration in AI and GPU Workloads | The Six Five Summit
There has been a lot of focus in the industry on how to deliver the performance needed to GPUs as infrastructure teams embark on LLM training, GenAI and Enterprise HPC projects. As the projects expand, data orchestration is crucial to provide efficient and timely movement of decentralized data to available GPUs that likely are not local to the data. By automating and managing the data flow, data orchestration minimizes latency and maximizes GPU utilization, enabling faster processing of large datasets and complex algorithms. We will discuss how existing AI infrastructure teams are leveraging data orchestration to gather larger data sets and improve their AI results.
Transcript
Welcome back to our continuing coverage of Infrastructure in the Cloud here at the Six five Media Summit. I'm delighted to have Molly Presley, the one and only Molly Presley. Join me from Hammer Space, where she serves as Chief Marketing Officer.
Molly, welcome. It's great to be here and great to be with this audience for our first time. Yeah, it's great to finally meet you.
In fact. Tell us, tell us about Hammer Space for, for those who aren't familiar with Hammer Space, what's Hammer Space all about? Yeah, absolutely.
So there's kind of a fun, memorable way to think about a hammer space. A hammer space, if you've ever watched cartoons was the place where, like if Bugs Bunny was pulling a mallet out of his pocket, the mallet was massive, and yet Bugs Bunny was quite small. The Hammer space is the place in comics where something big comes out of a small spot.
That's Mary Poppin's purse. You could think of a lot of different analogies. And with the latest Spidey Burst, they actually talked about the hammer space specifically for Spidey.
And I bring that up, that kind of origin story of our name, because it's very analogous to what we do as a company. That idea of we want to use data in a lot of locations, and we've historically felt like we were really constrained by physics on how do you move large data sets? Is my networking big enough?
And yet, with Hammer Space as a corporation and a technology company, you can bring these really big data sets out, even if you have a very constricted space that you're bringing it through. So what we offer is a global data platform that organizations are using for big data processing AI types of initiatives. Uh, one of the, one of the better origin stories for a company name, I have to admit, uh, So fitting, I wish I could claim I came up with it, but it was actually one of our engineers.
It's very, it's very, very clever and apropos, um, often when, you know, once you kind of have a feel for what Hammer Space is all about, um, you're gonna hear people start talking about orchestration, data orchestration. Uh, talk a little bit about that. What do you, what does Hammer space mean when they talk about data orchestration?
You bet. So when you think about if you're an infrastructure person, this will make a lot of sense to you. And even if you're not, I think you, the concepts are pretty straightforward, that if you think back the 30 or 40 years ago as IBM and the big mainframe companies were building storage systems, everything was built into the hardware platform.
Everything was subservient to the platform of hardware. It was brought into ensure we've evolved to scale out systems and hybrid cloud systems, but still the infrastructure has really dictated how data can be used. And you take that a step deeper and it's because the file systems are embedded in the infrastructure.
So if you want a high performance, high throughput system connected to your GPUs, historically people would've thought, okay, I need to go buy a high performance cluster of storage with certain attributes. And as you, as the world has evolved, hybrid cloud came about that led to, okay, which is my cloud instance, which is my cloud region, which is my storage cluster. And they were all isolated from each other, but served different purposes.
What Hammer space has done is brought in a global file system, which manages any infrastructure. So any cloud region, multi-cloud, any data center cluster, and provides the performance of like luster file systems in the HPC space, but provides that magic part that doesn't exist at all Today of, in that global file system, we have automated data orchestration. And what that is, is software policies that say this data needs to exist in a certain place, um, based on recent access, the project, it's associated with the job scheduler or data, our, um, GPO orchestration tool like run AI says, I need this data set.
We gather, put it local to the compute, it needs to run in and orchestrate it to that available compute. And that might sound kind of trivial, but if you think about in a machine generated data world where you're dealing with billions of files that you don't know where they're located, who owns them, which project it's associated with, we have all that really smart metadata to know which files exist and then intelligently only move the ones you need to the available compute, which may be in one location, multiple locations, depending on where your compute is available. So it's the whole idea of moving data to the available compute resources that you wanna use.
Yeah, I think anyone, um, can relate to the idea of location being important. If you wake up at four o'clock in the morning, uh, and uh, an outhouse three miles away, not not great, you know, it's, and so, so no, I I mean, it it, it should be self-evident that having data close to where the compute is is important. Uh, but are you talking about whether the data is in sort of a traditional on-premises environment or in cloud or both?
Does data orchestration In both? Yes. And so we run in any cloud, any virtual machine, you can spin up an instance of Hammer space quickly.
So what you find the most AI architectures today and you know, AI's on pretty much everyone's mind is there's a shortage of a couple of things. GPUs, power and data. Um, I think most organizations, even some of the big, like 4G 10 tech companies say they have a hard time getting access to enough of their own data to train and a large language model.
So, you know, there's GPU availability, nvidia, and companies like that needs to sort that. Power companies will sort out power. What Hammer space really helps with is getting access to the data, so knowing what it is and then placing it where it needs to be to be processed.
And so it's a really big deal to be able to figure out through intelligent metadata, which, which files exist and which you need. And then being able to intelligently and efficiently move those across potentially small network pipes, um, locally give high performance the GPUs and applications using 'em while files are being moved. There's a lot of kind of magic that's happening behind the scenes.
That's a really, um, been a difficult to solve problem. And that's why there's a big debate in the industry is data gravity such that you must move your compute to your data. And that's kind of crazy if you think about it, the thing that's digital is the data, not the infrastructure that should be the easy thing to move and it's been solved in the consumer space.
Think about how your iPhone, I'm an a Mac, an Apple user, um, but I know this is true in other technologies, deal with Samsung and whatnot, that you have pictures you take on your cell phone and you wanna see them to edit them on your iPad, and then you're on your computer, you're receiving text messages. That's all happening because iOS is a data orchestration tool for the consumer world. It, I don't know where my files sit, I don't know where my pictures sit if, and I know they're in the cloud somewhere, but I don't even know which one.
But I do know I can access and see all of 'em anytime I want. Sometimes there's a bit of latency as it's being pulled to a device, but it's largely unnoticed. It's really what we're doing in the enterprise world that makes that data viewable in accessible anywhere from any DEI device compute application versus it being dedicated to one iPhone or storage system and having to figure out how do I get backup?
How do I move my files over to something else? So it's really following a trend that we've become so dependent on in the consumer world. Yeah, it's interesting you use that as an example, the consumer device, um, on a daily basis, we don't care where that data is as long as it's accessible.
Similarly, in the, in, you know, in the era of cloud, um, we have come to a point in time where it's appropriate to say, I don't care what the infrastructure looks like on the backend. I want the service that will be delivered on top of that infrastructure. Now we're talking infrastructure here.
Camera space is, is more of an ethereal infrastructure thing as opposed to the hardware that it's orchestrating data upon. What does that hardware look like underneath? How much do you care about that?
It's specifically in the ai? Specifically in the AI space? Yeah, it's a really good question.
So we're software, we're a global file system with a bunch of policy engines that move things around. Um, we have two Linux kernel maintainers that work for us. So we try to drive as much of this as we can into the Linux kernel so that customers aren't required to use proprietary drivers, you know, figure out how they can get systems up and running with us.
So there's a lot we're doing to make it very accessible within the organization. But when you look at the hardware, there's a full spectrum. We don't care.
Um, but you know, hardware has physical performance limitations and MVME device is always gonna be faster than a tape drive. I mean, that's just reality. So when we look at it from our perspective, our customers typically have multiple different types of storage systems.
They may have a 7-year-old Isilon, a brand new vast, um, some Google Cloud and some Amazon S3 and a Spectra Tape library, as we were talking about Spectra a few minutes ago. Um, those are all, all that data and metadata is ingested into our file system, and you can view it all as a single entity as far as which hardware customer chooses, you know, we certainly help them. Um, we require some resource for metadata services and things like that.
But then the rest is just the performance and cost attributes of the storage. So Facebook, I'm meta, I still sometimes make that mistake. Meta uses us for a LAMA two, LAMA three training, and they prefer very specific hardware that's very standards based.
Um, as we all know, they're very involved with OCP, they have a big partnership with Pure, um, so that we don't care. There's pure, there's OCP, there's some other just kind of white box stuff sitting underneath us for LAMA two and LOPA three training. But then you go look at maybe an enterprise that's been up and running for a hundred years and they have a bunch of old net apps and IBMs those also can be presented and we unify the data.
And then maybe you have this old, let's just say IBM box that's been in production for 10 years. You still want the data, but you wanna have some faster access. So when the files do need to be faster, we'll orchestrate them to something fast.
Um, but you don't have to throw out everything you've already invested in. Ah, yeah. You uh, you stole my thunder on the next question here.
'cause the, you know, the, that's Good. Exactly, Exactly. It's because you're a pro.
It's your, you're a professional. So, um, yeah, because it's one thing to say, oh, fantastic performance. We're, you know, we're here in the future.
The folks charged with managing infrastructure in the backend are thinking, okay, yeah, what about, what about all this legacy stuff that I have? Do you help, uh, migrations essentially, and when I say migration, I mean the removal of old technology, you know, the, the unscrewing of the incandescent light bulb and the screwing in of the LED light bulb in a way that is completely transparent to the enterprise. You basically just said that yes, you do that, but I wanna confirm that that's what I heard you say We do.
And if you're like me, at least I haven't heard this, the light bulb analogy, but it's a good one. I just made it up. You know, I've replaced most of the light bulbs in my house.
I still wait till the old ones die, you know, 'cause I'm too cheap to just pull 'em out and throw 'em away when they're still working. Um, and that's kind of the way enterprises run too. They spend a bunch of money on infrastructure and for economic greenish initiatives and just time because migrating, doing a data migration is so time consuming.
They try to leave it in place as long as they can. And what ends up happening is they start to fall behind the organizations who are doing greenfield new deployments because they're slower, they're less nimble, that type of thing. Um, so two data migration, yes, we help a ton in that.
You bring in hammer space, and again, we're the active file system when you imp implement us, connect the applications, the users, the GPUs to our file system, just standard N-F-S-S-M-B S3 connections. So we're standard space as far as how you connect to us. And at that point we assimilate the metadata from that old infrastructure.
And you can add in new infrastructure as well, but that's all transparent to the users and applications. So let's just say you do have that old IBM and you know, it's finally just petering out and gonna die on you. That's okay.
'cause the data and the metadata are already in hammer space and the files that you wanna migrate can be migrated over days, weeks, months is however long it takes. And the users and applications will not be affected because they're working with metadata, not the files themselves. And so the files and movement, while to the new fast system or the cloud system, it doesn't interrupt operations.
So the problem data migration is a couple things. You can't do. You really want plan downtime for the entire weekend or a week as you do a migration.
That's very disruptive and it just takes a lot of time. You know, it takes a lot of it's time manually copying things. And so that's completely removed when airspace comes into the environment, it and infrastructure are decoupled from the user experience with the data because we become the data layer.
Molly, let's talk a little more about performance and how Hammer space can deliver the kind of performance that's needed by these data hungry GPUs today. Yeah, you bet. So certainly the first thing organizations think about is they're design their infrastructure is, okay, I've bought a few hundred GPUs.
If you're meta it's tens of thousands of GPUs. How am I gonna keep those busy? And having a high performance file system, typically a parallel file system that you would use in, uh, historically in the HPC world is the first thing people think of.
Um, hammer Space absolutely provides performance. Right now we're streaming, um, in large language model training at Meadow, well North 24,000 GPUs at a time. So the speed matters.
Um, there's a couple companies who are kind of bubbling to the top, us being one of them, um, of how you keep that local performance really rolling to never have your GPU slow if you have the data. But the cool part that we add to that is what if you need access to the remote data sets, to other data sets from other organizations, other locations? And our integrated orchestration is a key part to these more decentralized AI workloads.
So with, with the minute or so we have left, um, gonna hit you with something that may be more than a minute long answer. So we're gonna see just how professional you are here, Molly. Camera, space, camera space's development kind of has preceded the dawn of at least the current AI era that we are in.
Um, how is it that the architecture has been able to stay fresh and what kinds of changes have you had to make in order to make this truly appropriate for the age of GPUs and ai? Is any, is there, have there been any hurdles you've had to overcome? Well, I would say there's always a little bit of luck, um, when your technology comes out and the need for, and the need for our technology.
Na AI is massive. Um, but there, you know, as you think about what we've needed to do, the technology's been generally available for about two years and in development for 10, which is about how long it takes to build a file system. Um, our CEO was the founder of Fusion io, which was the company who actually decentralized data to start with by putting those fusion cards out in servers.
And he's had a vision towards this of great, I forced this decentralization, but now how do you use the data? And we're very closely, um, partnered with Gary Rider at Los Almos National Labs, who was kind of the grandfather of a lot of the HPC file systems. And we've been working in conjunction with them and the Linux community of how do we solve this decentralized data problem.
And it's just been exasperated by now. We have a lot more remote workers through covid now we have GPU shortages, and so there's more need for it than we potentially expected originally. Um, but it's just kind of a perfect storm of why hammer space is needed.
Um, and the technology is just showing up in the market at the right time, to be honest with you. Yeah, no, very interesting. Serendipity is also is, is always something, uh, that could be, that could be a beautiful thing.
But no, that, that, that history back through Fusion io in the sort of indirect way and direct way that the thought that the thought process is linked, uh, really has set you up in this market very, very well. It's very, very interesting. Molly Presley, thanks so much for joining us here at the six five Summit.
Molly Presley from Hammer Space. Thanks again. I'm Dave Nicholson.
Stay tuned for more on infrastructure and other topics here at the six five Summit.


