UT07x07: Accelerating Storage Infrastructure using GPUs with Graid Technology – Utilizing Tech
Modern AI infrastructure has exposed the importance of reliability and predictability of storage in addition to performance. This episode of Utilizing Tech, presented by Solidigm, features Kelley Osburn of Graid Technology discussing the challenges of maximizing performance and resiliency of storage for AI with Jeniece Wnorowski and Stephen Foskett. AI servers are optimized for machine learning processing, and Graid Technology SupremeRAID offloads processing to GPUs similarly to the way these massively-parallel processors offload ML processing. They also have a peer-to-peer DMA feature to direct the data directly to the processor rather than forcing all data to pass through a single processor or channel. There is a need for RAID software at many spots in the data pipeline, from ingestion and preparation to processing and consolidation, and each requires performance and availability. There are many applications that require maximum performance and capacity without impacting the host CPU, including military, medical research and diagnostics, and financial, in addition to AI processing.
Transcript
Modern AI infrastructure has exposed the importance of reliability and predictability of storage, in addition to performance, this episode of Utilizing Tech presented by soddy features Kelly Osborne of grade technology, discussing the challenges of maximizing performance and resiliency of storage with ai. Welcome to Utilizing Tech, the podcast about emerging technology from Tech Field Day, part of the Futurum Group. This season is presented by soy and focuses on the question of AI data infrastructure.
I'm your host, Steven Foskett, organizer of the Tech Field Day event series, and joining me today from SOY is my co-host Janice Kowski. Thank you for joining us on the show. Oh, thank you, Steven.
It's so nice to be back. So Janice, uh, we have been talking quite a lot on this season about the varying requirements for data underneath ai, and one of the things that keeps coming up is, I guess the practical question of how these AI servers are built and what they look like and what the infrastructure around them looks like. And it seems that a lot of the focus of the industry, especially when it comes to AI, has been on optimizing the, uh, the processing, uh, aspects that, and the data movement aspects, uh, without as much consideration for storage and, and we're here to, to set that record straight, right?
That's why we're doing this. That's right. That's exactly right.
So it's all about the infrastructure, infrastructure matters, uh, but the underpinnings of that infrastructure, particularly storage, right, as we're dealing with loads and loads of data storage, is really being put back on the map in a big way. And we're excited to talk to various guests about our large capacity, high density storage, and how does all of this, uh, become a game changer for those AI workloads that everyone's dealing with? Yeah, I think that's the, the key there, as you say.
I mean, storage has become sort of a critical path, a a critical point. Uh, you have to have high performance storage certainly, but you also have to have reliability and, uh, storage features. And, and one of the questions about reliability too is, is if, if there was some kind of failure somewhere in the data path, um, you know, how would that affect everything?
Well, the answer is pretty badly. You know, you wanna make sure that, uh, that everything you're building all the way from the, uh, CPUs to the, the GPUs, to the network, to the storage layer, everything has to be redundant, has to be reliable, and has to be predictable. So that's what we're gonna talk about today.
Um, we have a guest on, uh, Kelly Osborne, rep representing, uh, grade technology. Uh, Kelly, welcome to the show. Thanks.
Uh, as you mentioned, I'm Kelly Osborne. I'm, uh, senior director of OM and channel business Development here at Grade Technology. And, uh, we've been in business for about three years based out of California, and I appreciate the opportunity to, uh, chat with you guys.
So tell us a little bit, I guess to start off, um, you know, where in the, uh, AI data infrastructure stack does grade technology fit? So, um, grade technology and our product Supreme Raid is actually a, uh, GPU accelerated RAID stack. So we actually dedicated GPU to do all of the infrastructure processing for rate calculations, for parity, um, and we also feature a peer-to-peer DMA technology that allows movement of data from the drives to the applications directly across PCIE, uh, eliminating bottlenecks that we traditionally see with, uh, older hardware rate technologies and the scaling problems of CPU utilization with traditional software rate technologies.
So Kelly, uh, we've worked a lot together over the last year here. Um, been to a couple of shows together and, and certainly very excited about your Supreme rate product, and I know our teams have been doing some work there as well. 44 terabyte specifically.
Can you talk a little bit about, you know, the value of, you know, QLC SSDs in conjunction with Supreme Rate and kind of what you're seeing? Sure. And I think it's not just about capacity, it's about performance.
And when you look at dense server environments, and we're talking about, uh, servers that have 12, 16, even 24 or 32, uh, NVME SSDs, uh, and now with that kind of a capacity point, how do you deliver the per potential performance of those drives? So if I said I'm gonna do read zero stripe across all those drives and then do read, I'm gonna get very high read performance, unfortunately, I'm not gonna have any data protection. So as you layer in data protection, uh, for an environment like that, uh, the traditional methodologies will cause bottlenecks and will, you'll never be able to get to that performance that you would have with Raise zero.
What we have been trying to do here at Great Technology with Supreme Raid is build a stack that can deliver very close to raid zero or theoretical performance on read and write of those big drives you have and yet provide all of the protection you need for that data that you're storing on those. So let's talk a little bit about how you're doing that. Um, I guess some people listening might be scratching their heads thinking, wait, what do GPUs have to do with storage?
And, uh, I think that's the, the clever thing here, right? In that just like GPUs and their massive, uh, parallel architectures can be used to accelerate, uh, the kind of, uh, matrix math used for machine learning, uh, those same, uh, that same kind of hardware can be used to accelerate the calculations that are used in, uh, storage, uh, data protection. Right.
Uh, why don't you talk us through a little bit about how GPUs are useful for storage? Absolutely. So if you, if you take a look at the rate technologies that we've always had, the, uh, it's nothing really new in terms of what we call Read Solomon algorithms.
The issue is generating parity information using those kinds of algorithms. Uh, we'll take a heavy toll on your CPU. So if your CPU is very heavily utilized because you have a large number of drives that are really, really fast, it doesn't leave much room for your application to run.
So if we can take that and put it onto A GPU, um, in our case, we work with Nvidia GPUs, we've taken, uh, and written our supreme rate driver as a cuda based application. So we are taking advantage of those really, really, uh, dense cuda cores on those GPUs to do that mathematical calculation much more quickly than you could on A CPU. Very similar to graphic, uh, rendering and things.
GPUs are much better at that than a general, general purpose CPU. Um, if you tried to do that on traditional hardware rate controller, the, the flip side of that is in that environment, you, you create this problem of a multi-line highway feeding down into a single card. And so the problem there is if you've got, say, 10 or 20 of these, uh, big solid I drives each taking four PCIE lanes, uh, if you had 20 of those, that's 80 lanes.
If you feed that into one card, you're funneling it down just like you would've toll booth with a a, a super highway. So the other part of our technology is to use a peer-to-peer DMA feature, um, that allows us to act as a traffic cop. So in addition to offloading from the CPU, we can tell the data on the drives, or we can tell the drives, excuse me, to directly send the data across PCIE, kind of like a bypass road around that toll booth.
Um, so that is the other side of our technology that solves kind of a hardware rate bottleneck. Yeah, and a lot of this too, Kelly Wright is, is really plug and play. I think we talked about this, uh, when we were at nav, but there's like effortless installation, no cabling required, no motherboard, you know, relay out, uh, uh, required the SSDs actually just plug right into the, uh, you know, PCIE interconnect, and then everything kind of just works together.
Um, is that, is that unique, uh, or is that similar to others in your, in your space? We really don't have a competitor in our space, um, that all the traditional hardware rate controllers would require them to run cables from your drive direct straight to a card. As you mentioned.
We wanna see the drives plugged right into the motherboard, and so we see them across that, uh, PCIE, you know, uh, root complex, and we can communicate and control the drives without forcing the data through our card, which eliminates that bottleneck. And what's interesting to me is because of your simplistic design and nature, and you guys, you know, work with a lot of really interesting customers, um, we'd love to just talk a little bit about the work you'd been doing with, uh, cheetah rate as an example. I know at, um, the NAB show, we had a very interesting demo there.
Do you wanna just talk to us a little bit about the uniqueness and, and what we showed there at the, on the show floor? Sure. And I think, I think, um, being part of a solution is a much more interesting thing to discuss than mathematical calculations on A GPU.
I think, you know, you kinda understand what we're doing, how do you apply that in the real world? And, um, in the broadcast world, you're dealing with really, really large video files in, in many cases, and, uh, which are perfect for the soy. Uh, 61 terabyte drives the, the type of thing that you see in that market is, uh, if you're doing digital recording at a studio or on a film site, for example, that data will fill up those drives.
And at the end of the day, you want to edit those, you want to edit that data, so you've gotta get them over to a post-processing house or a, a graphics, uh, you know, CGI kind of house and things like that. You're not gonna take that data when you've got hundreds of terabytes and transmit it to the cloud and then let your, uh, video editing, uh, shop go access those files. It becomes very expensive and time consuming.
So what Cheetah does is they build a server that can take four of your drives and put it into a, a removable cartridge, and then they have three of those in the, in the server, so they can pull those out, and then you could ship those somewhere, uh, put new cartridges in and then start filming again the next day. And then back at your post-production house, you could plug those in and then access all of that data. We're providing the rate protection and high performance read and write access to make sure you can get that digital video down fast.
And at the same time, in an editing environment, if you're gonna have five or six, uh, windows, for example, based editing, uh, stations, you need very, very fast access. And many times, four or five people are gonna want to edit that file at the same time, and they might be working on different, uh, parts of that file. And so the final piece of our, uh, of our solution, uh, was, uh, a company called Tux.
And Tux is a company that makes Fusion file share. And this is, uh, basically a SMB protocol file sharing technology that allows those multiple, uh, video editors to access the same file at the same time, but also with extremely high performance. And so this really solves a lot of problem when you have, uh, on waiting for data to, you know, to scroll forward and things like that.
You can move back and forth very, very quickly. You can have multiple people attacking, uh, that data at the same time instead of having to do it sequentially. So it saves time.
And so that was very, very cool, uh, you know, very cool solution that we were showing, and it was fun to be part of that. It was amazing. Yeah, thank you for being a part of that, Kelly.
And just downstream all of the kind of AI workloads that are being run, you know, within the studios, there's software out there now that actually works with Tech Sarah, where they're able to take that, that video footage in real time and then start editing and looking for, um, sequential B roll that supports that video. So, uh, you know, it's kind of to, you know, taking a look at the entire pipeline of how this stuff actually becomes, um, a movie or a video, you know, within five minutes, which could ultimately be five days. So, uh, the work that we've done here, I think has really helped these guys, which is really exciting.
I just wanna ask, Kelly, can you tell us a little bit more about some of your other partners and customers? I know we have like NetApp as an example, we've done some work there, um, with those guys and I, I believe you're a part of that solution, but can you talk a little bit more about some of the mission critical workloads, uh, you work with, and how high density storage could potentially benefit, uh, some of those workloads? Sure.
So, um, in the machine learning and AI space, we're seeing more and more, uh, customers purchasing very, very expensive GPUs. Um, I was at GTC, I got to see you out there, it was a great show. Um, got to see Jensen walking around with his leather jacket on.
But, um, we're, we're seeing servers that have eight H one hundreds and now the blackwells are coming. These are, these are servers that are gonna be, you know, four or $500,000 machines by the time they're finished being built, and they're very, very data hungry. And so I'm, I'm sure you've talked about it before, but, uh, one of the biggest issues is how do I keep those very expensive devices that I've purchased fully utilized?
Uh, why buy eight of 'em if I'm only gonna have them, uh, 50% utilized because my data is too slow in getting to the GPUs, I might as well just do four. So, you know, very, very fast read performance from very large machine learning and training data sets is a big part of what we're capable of doing with a large number of drives. And so that's an area that, uh, we're involved in.
Uh, we also, uh, partner with, uh, several companies that build parallel file systems. And, um, and the primary use for those are high performance compute environments, so think very large super compute environments. Uh, so we, we work with Think Park and they, uh, represent a product called BGFS, and we partner with them.
Um, in fact, uh, we were recently at the ISC conference in Hamburg, Germany, and we'll be at the, uh, SC 24 conference coming up in November. I think that's one of the interesting things about this technology, Kelly, is that, um, while people like to focus on really kind of the tip of the arrow, you know, the, the, the processing and, and storage of data when doing ML training, there's a whole broad pipeline, a data pipeline here all the way, as you talked about a minute ago, from ingesting and preparing data in the field through the actual practical inferencing of data with applications. And it's, you know, I think it's easy to get tied up thinking about the importance of, uh, of storage right there when you're doing your, you know, you're, you're, you're building your ML models, but there's a million other things that need to be addressed in order to have an effective AI solution.
And, and it seems to me that, um, you know, solutions like this, like yours are, uh, part of a broader perspective of AI data infrastructure. Is that how you see it as well? Yeah, absolutely.
You know, um, AI is not an event, it's a workflow and you have to do injust first, obviously, and that could be multiple sources of data coming in all at the same time. How do you write that really fast? That's a very right intensive operation.
Um, and you still want that data to be protected as well, um, in case there's corruption and other things. And then you move to kind of a, a tagging model, um, and a training model, which generally is much more read intensive because I'm now re running back through that data with multiple GPUs in parallel to, uh, to try to build my trained data sets. And then you move into inference and inference really becomes a whole nother animal, um, much more read intensive once again, but not quite as GPU intensive.
And so each stage of that, and, and they can be different depending on which frameworks you're, you're using. Um, so I think that, you know, the experience that we see is that high speed access to data, both read and writes, smooths out those workflows and, um, and makes, makes them, you know, much more efficient. I'm curious, Kelly, can you tell us a little bit about like, just an overall, um, uh, TCO model here in terms of, of utilizing, you know, QLC SSDs with, uh, Supreme Raid?
Is there, uh, a cost benefit to any of your customers as they're kind of looking at their overall infrastructure? I believe so. Um, when we take a look at a, uh, those, those, uh, drives are gen four drives today, and they're generating seven gigabytes a second, um, capable of, of read performance.
So if I took four of those drives, that's 28 gigabytes a second of theoretical read performance, if I plug those into a traditional hardware rate controller, I'm basically maxed out on that read controller. So if I had eight drives, if I plugged all eight into that one controller, I'm only gonna get 50% of the throughput because that PCIE Gen four by 16 slots only capable of doing 32. So I'm right at 28, 32, 2 overhead of that controller.
So the only way I'm gonna get better performance is to buy another controller and another controller and another controller. So if I have 16 drives, I need four controllers. If I have 24 drives, I need six controllers.
Now I'm, now I'm worried about how many PCIE slots do I have, where do I put my GPUs if I've got all these RAID controllers, right? So the ROI for us is we can handle up to 32 drives with one slot. Uh, we have an HA feature if you want to have two of our devices, um, two NVIDIA cars with our software, we do offer that HA failover, so maximum two slots.
Um, so that actually saves you on PCIE slot and it saves you on cost from having to purchase all those physical controllers. The flip side of that is if you said, well, I could just do 24 drives and run software rate, if you want to have any room left on that CPU to run any applications like your ai, you know, you know, a PyTorch or something like that, um, you're gonna have to have a lot more expensive, uh, CPUs to be able to handle the overhead of the infrastructure. So by offloading that to the GPU, uh, you could theoretically buy a cheaper CPU potentially.
Um, and then that can also save you money and still give you the performance you're looking for. Yeah. I know you mentioned it's, you know, our QLC drives have, you know, seven gigabytes per second with reads, uh, which is great, right?
4 gigabytes per second, which is amazing for QLC. So, you know, in a nutshell, QLC can really handle some of these higher performant, you know, workloads. And, uh, a lot of folks out there, there's debate around whether or not QLC can hold up to it, right?
But, um, it's actually not only a benefit of performance, but cost and it's seamless every time, which is really great. Yeah, the, the point of the, uh, you know, CPU resources is an interesting one because of course, um, there are a lot of software rate solutions that use CPU, uh, to do the same processing that you're doing on GPU. And many of those use, um, accelerator instructions, uh, built into the CPU, um, are, are those still, uh, I guess using up resources that could otherwise be useful for machine learning processing?
They can be. Um, some of the instructions that we know about with those are instructions that aren't gonna be around much longer. Um, um, and they may be intel only features.
So we see, you know, um, that kind of a problem where if they decide not to keep those vector instructions, you could be, um, hampered by that. The other question is, where do you run your application? Um, is it in user space or kernel space?
We are, uh, we employ a kernel driver as well as the CUDA driver. So there is a piece that runs on your CPU that works in concert with the, uh, kuda based driver that we put on the GPU. Um, some, some companies out there will advertise that they get the best performance by running in user space, but that has vulnerabilities along with it.
Well, and, and I think that that's, uh, these are all considerations that people would, uh, would need to address in their individual environments as well, right? That, uh, they would need to look at the, uh, capabilities of the servers that they're deploying in various spots along their data pipeline and decide whether it would be appropriate to use something that relies on CPU instructions, uh, versus, uh, A GPU or even a raid card at various levels. Could you see a place for all the different, uh, storage solutions in the same data pipeline?
Very possibly. Um, the traditional hardware rate controllers that have been around for a long time, uh, they feature batteries and caching battery back cache, for example, if you're talking about storing large amounts of data over time, there are still places where hard drives are being used, and that's an area that we don't particularly work with. So that traditional harder rate controller that has caching and other things is a perfect fit for that kind of environment.
Um, if you're talking about a lightweight server with two to four drives in IT, software rate's probably adequate. Um, so where the grade, uh, solution really fits is when we're talking about much higher densities of NVME in a single machine. Uh, so, you know, we kinda understand where our fit is and, and that's our market is to focus on these machines that have more than four NVME drives, where you have to start figuring out how do I deliver the, the maximum performance from those drives to the application As drives get bigger and faster, uh, is there a greater processing requirement to do the RAID calculations or is there more nuance to that?
It's, uh, it, it's a processing in terms of which rate level you want to use. So you have rate five and six, which are, um, more, typically we're seeing rate six once you have more than four drives, uh, because people want a higher level of protection. And the, I think the biggest issue is when you are talking about a large number of drives that are so fast now with NVME that have so much data, it just consumes your CPU.
If it's software rate, we already know the hardware rate controller's not gonna handle that at all. If you just dedicate software rate, you'll never get to the performance, um, that you want because you're gonna be consuming that CPU more and more, and then you're gonna have to spend more on, on greater or more powerful CPUs to try to keep up with that processing. Um, in the testing, we've done even software rate at rate 10, uh, which is how a lot of people will try to deploy software rate to minimize that CPU involvement because it eliminates the mathematical calculations.
The problem is you sacrifice half your capacity to do R 10 because you're mirroring and striping, and also your right performance goes down because you have to make two rights, uh, to get the mirror before you're acknowledged. Um, we offer basically what we call very close to RIG zero performance with raid five or R six usable capacity. And once again, free that CPU up.
Yeah, Kelly, so, uh, let's go back to use cases for a minute. Um, just wanna hear a little bit more about what other type of, of partners and, and who really needs this type of performance. Can you give some specific examples?
Sure. So, um, we're working with the military on lots of edge deployments where, uh, high performance computing is still mandatory, but they wanna be very compact, very low power, very low heat, and they need to e as much performance as they can to, to capture the data that they're capturing in those types of environments. Uh, we have research hospitals, uh, one in particular that most people know of in Memphis that actually has a new type of microscopy.
So this is a microscope that is an atomic microscope, and it generates a huge amount of data, and it has to be written really, really quickly. And if you think about it, the faster I can write that data, the more, the more quickly they can move on to another patient, uh, study. And so if you can help more people with that same device by not having to wait for the data to write down to disc, um, that's valuable.
Uh, similar, uh, area in the medical device world, we're working with companies that build CAT scan, CT scan, MRI type systems, and similar situation, a huge amount of data is, has created very, very quickly in a very short amount of time. And today they have to wait a long time for that to get written. If they could double that performance, they can have handle twice as many patients in a day, that helps more people.
It helps the clinic pay that equipment off. That's very expensive, much more quickly because they're seeing more efficiency. Um, database, high performance database.
We work with accounts that are working with Oracle and Postgres, um, Redis, other high performance in memory compute type databases, um, Splunk servers. I have a large credit card company that's, uh, using us, um, in that kind of environment. Um, high performance compute supercomputers with parallel file systems like BGFS.
So we could go on and on and on. Um, and where, where this makes sense, um, we have gamers call us all the time, but we're too expensive for the gamers. Well, that's great.
Thank you so much for this. Um, it's, it's interesting to consider these aspects because again, we, as I said at the top, I feel like so often people focus only on the, uh, sort of the, the signature, uh, data center full of GPUs, and they don't realize that there's just a lot more to the question of AI data infrastructure than, than that. And even that has requirements for high performance and as I said, high rate reliable and highly predictable storage that comes from, uh, raid.
So thank you so much for joining us. Kelly, before we go, uh, where can people learn more about grade technology and Supreme Raid and where can they, uh, connect with you? com and, uh, upcoming shows that you might wanna come see us in August.
August will be at FMS, which is formerly Flash Memory Summit. It's now the future of Memory and Storage. That'll be in, uh, Santa Clara, and then again in November, uh, this Super Compute 24 SC 24 show, which is in Atlanta this year.
So, uh, I'll be personally at both shows along with, uh, other folks from my company. And, uh, we'd love to hear from you. And, uh, Janice, I imagine that folks will see solid I at some of those shows too, right?
Soy will be there almost at all of those shows, and then some. And, uh, yeah, thank you again for, for hosting us today, Steven. Well, thank you very much for joining us.
It's nice to see you. And, uh, thank you everyone for listening, uh, to this episode of Utilizing Tech, uh, our special AI data infrastructure series presented by soy. You can find this podcast in your favorite podcast application.
Just look for utilizing tech, and please do consider giving us a rating or a review. This podcast is brought to you by Tech Field, a home to IT experts from across the enterprise. Now part of the futurum Group, as well as, as I said, soy.
com or find us on X, Twitter and Mastodon at utilizing Tech. Thanks for listening, and we will see you next week.