UT08x12: Revolutionizing Data Infrastructure for AI with WEKA
Storage software running on modern hardware can deliver incredible performance and capability to support AI applications. This episode of Utilizing Tech wraps up our season with a discussion of WEKA’s data platform for AI with Alan McSeveney, Scott Shadley of Solidigm, and host Stephen Foskett. Modern hardware is capable of incredible performance, but bottlenecks remain. The limiting factor for AI processors is memory capacity: GPUs are hungry for data and must be refreshed from storage quickly enough to keep them running at scale. Storage can also be used to share data between GPUs across the data center and to cache working data to accelerate calculation. The secret to scalability, from storage to applications to AI, is distribution and parallel processing. Modern software runs at incredible scale, and all elements of the stack must match. Technologies like Kubernetes allow applications to use huge clusters of workers all contributing to scale and performance. WEKA runs this way, matching the GPU clusters and web applications we rely on today.
Transcript
Storage software running on modern hardware can deliver incredible performance and capability to support AI applications. This episode of Utilizing Tech wraps up our season with a discussion of Becca's data platform for AI with Alan Sev of cca, as well as Scott Shaley of Soddy and myself, learn how Modern Hardware is transforming storage for AI in this episode. Welcome to Utilizing Tech, the podcast about emerging technology from Tech Field Day part of the Futurum Group.
This season is presented by Solid Im and focuses on AI at the edge in related technologies. I'm your host, Steven Foskett, organizer of the Tech Field Day events series, and joining me from Soy for this final episode of our season once again as my co-host and old friend Scott Shaley. Welcome to the show, Scott.
Hey, Steven. It's great to have fun doing the season with you. And it, it's sad and, and also great to know that we've, uh, made it through another season of these, uh, wonderful, uh, episodes and some amazing conversations that we've had over the last few episodes.
Uh, and working with you guys has always been so much fun. So Yeah, it's, it's been really great. Uh, you know, it's been great welcoming you as a co-host.
I, I knew you could do it. Uh, glad to have you. Yeah.
You know, I like to talk about things, you know, that kind of stuff. So when it comes to talking tech, it's, it's a lot of fun and, uh, the guests and yourself make it a lot of entertaining, so. Well, that's what I was just gonna say.
I mean, the guests are incredible. Um, you know, we get so much great insight from them and just so much perspective on how, uh, this AI thing is being implemented around the world. Uh, you know, I think that people have this feeling that AI is somehow kind of a big iron thing, that it's some supercomputer in a, in a big data center that's sucking down gigawatts of power.
And it is. It is, but it's more than that. AI is being implemented outside the data center in smaller environments at the edge.
Um, you know, maybe, uh, let's say, uh, interesting venues, entertainment venues, all sorts of things. Exactly. It's not just a, the, the, the home of, of Big Iron, right?
Yeah. That, that's the wonderful thing about this is this is kind of the convergence of a whole bunch of different technologies at once, and the ability to generate data in the way that we can generate data and then actually do something with it in a more meaningful way, uh, as we talked about in a couple of the previous episodes, of what people are doing to go back in time and bring that forward with the, the technologies that we have available today. So, today's a, a fun one too, because we have a, a literal convert, uh, joining us today from, uh, being a customer to an employee.
And so that's kind of fun as we bring a Alan from WCA along. Hi, I'm Alan MCs. I'm the field CTO for Media and Entertainment, uh, and related AI at, at wca.
Um, I, I recently joined only in, in December, but I, I was a customer of w for four years prior to that, uh, in a, uh, a cloud-based visual effects studio. Um, most of my, my career has been related to creative industries. Um, actually I was a professional musician for the first 10 years of my life and really enjoyed marrying, um, creativity and, and technology and really pushing the boundaries of what could be done with technology back when actually audio and music was a quite a challenging thing you could do on a computer as opposed to now when you could run an entire studio and a laptop.
Um, but transitioned out of there into the visual effects world and large scale playback systems. Um, and that, that's been an immensely rewarding career. Um, just using technology and being able to push boundaries is really been, uh, a place where I'm kind of happy.
Well, it's interesting that, um, in audio, yeah, audio visual. Yeah. Old Atari st.
Guy here. So I know a little thing about using personal computers for music. Um, you know, what happened there was, you know, basically specialized hardware gave way to software running on commoditized hardware.
And the same thing Scott and I, in our career in enterprise storage, have seen the same thing happen where, uh, what was once the domain of literally special boards, special processors, special, everything has now become the domain of software. And that's really the story of CCA too. I mean, my understanding is that essentially the, uh, the, the founding team and the origin of the product was what if this was done in software, and what if we took advantage of the latest, um, you know, incredible advances in more commodity hardware, especially NVME, but also all the things that you can do now on the processor.
And, and, and it worked, you know, I mean, this has been, uh, something where the software based storage solution from CCA has literally become the bedrock of, of ai, right? Yeah. I mean, the, the marriage of three things really allowed WCA to, to become, uh, not just the product, a vision in the first place, which was someone soldering an SSD onto A-P-C-I-E cart, um, someone coming up with the concept of containerization, and then the network teams out there far surpassing everyone's expectations and blowing away what anybody considered could be achieved with network performance.
Um, so you put those three things together, and now you can access, uh, NVME storage across an array of servers all over network while orchestrating all of this through containers. That that is essentially what WCA is. Um, gives you a huge expandable data platform, uh, that is all NVME based, that is addressable over network that will at times not only outperform local NVME within, uh, a client, but sometimes even dra.
It's, it's a creative innovation where you guys are literally transforming that ecosystem to allow the underlying hardware that, you know, comes from someone like ourselves into something that peoples just see as right next door. It, it's kind of unique that the software platforms and the data platform you guys have has that ability to transition data from point A to point B as if it were just sitting there and not have to worry about all that transition time. 'cause to your point about networking, we, we keep seeing networks go up and down and faster and slower, and as we start getting further away from the, the main processing, that ability to see that localized data is something unique that you guys are working on.
Yeah. I mean, there's a bottleneck always somewhere, right? And it, you know, every, every so often it gets moved to somewhere else.
But, um, network performance today, um, is astonishing. You know, we're, we're seeing ethernet networks up to, you know, you know, 400 gigabit, um, we're able to push data into a single host at, you know, hundreds of gigabytes a second. It's, um, it's not something that I thought we would see, uh, so quickly, but this is where we are today.
Um, I think you're right, the, the idea of having data local to whatever your processes are, whether it's on a laptop or, um, on whatever compute you're using, and then expecting to have to transfer that someplace and there'll be some penalty or some time spent or some transfer process, that would just make you sigh. Um, it is most of our, most of our collective experience and history, um, today, that is just far from the case. Uh, actually having all data centralized and accessible or network can be more performant than having it local.
And that's one of the challenges I think when it comes to ai. 'cause this, this hardware is incredibly capable, but I don't know that every system can take advantage of those capabilities in order to, uh, you know, kind of move the bottlenecks out of the way. I mean, the entire history of technology is all about moving bottlenecks.
It's, it's, it's, you know, this, you know, we, we eliminated this one and then it pops up over here. We eliminated that one, it pops up over there. And if you look at this kind of classic computer system hierarchy with processors and memory and storage, uh, you know, storage for a long time was just the ultimate bottleneck with SSD that has been, um, dramatically reduced, as you say, with things like NVME and now with, uh, you know, ethernet networking, uh, a lot not to mention proprietary networks, it's been reduced further, but the demand for data from these AI processors is just absolutely off the charts.
It's insatiable. And, you know, one of the things we've heard about repeatedly on this season in the last, here on utilizing tech is the, basically the need to feed the beast. If you are not keeping your expensive GPUs fed, then you're essentially wasting money every minute, every hour that they're, that they're not working at maximum capacity.
That's pretty much what companies are looking at CCA to solve with software, right? Yeah, I mean, the, the limiting factor with, uh, a large data center enterprise skill GPU today is memory capacity. Um, the amount of parallel compute available in, uh, a cutting edge GPU is mind blowing, but the, the memory footprint on each car is just not where the processes that are running today, um, needs to be.
And our aim is to be able to augment that memory with a place where essentially you can tear the data that should be in memory off to WCA at such a data rate, that it can also be retrieved so fast that it becomes practical to, to now scale your memory footprint into petabytes, uh, of space. Um, there's obviously, you know, we're talking tiers of performance here, but when we are able to outperform in some cases what DRAM could provide to GPU memory, um, at that type of scale, you know, hundreds of petabytes if you like, um, it really starts to be a paradigm change. The, um, the amount of, the amount of time spent, for example, in LLM processes, during pre-fill calculating, uh, key values and creating KV cash data.
A lot of the time this will either just be cached to, uh, DRAM or to local NVME as KV cash, but this could only be used by other GPUs inside the same server within, or within the same, uh, envy link. And now we can drop all of that KV cash out to, to WCA and make it available not just to other GPUs in the same server, but to every single GPU in the entire data center. Um, and at rates where, for example, to calculate around 105,000 tokens, uh, of prefilled takes around 25 seconds on GPU, we've able to, we've been able to take that same KB cash, place it back into GPU memory in around half a second.
So we're talking 40 more than 40, almost 50 x, uh, speed up in some cases. Um, and, and every time there's any query that's performed on an LLM and this KB cash is generated, we can just keep that for as long as that model is around. Um, so it never has to be recalculated again.
So the more and more and more that queries are common across, um, multiple processes or customers, we just don't ever have to calculate it again. Um, and then the GPU can get on with the, the meaty part, which is, um, deco. Yeah, you bring up an interesting point about the, the idea of the tearing, right?
Because we, we all looked at it as kind of, if you think of the hype cycle and all this kinda stuff, I brought this up with some of the other examples that we've gone through in this season, but we're at the point now where people realize they need more of something, and that something is really being able to offload and, and shift the performance tier into an aspect of a larger footprint. Like, for example, the, the massive drives that we can provide give you those petabytes of storage that can look like that fast memory just because if you overload the memory, again, moving bottleneck to bottleneck to bottleneck, the, the eliminating of those bottlenecks is really kind of the key here. And, you know, fast delivery to your point.
Um, it's really cool you guys have had the, the ability to highlight the performance characteristics in real time of what you guys are up to with some of the work you've done in some of the recent venues that have come to light. So it, it's really interesting to see how you guys have been able to transform the idea that it's really more about the data and not where the data is sitting and being able to ma let the user maximize their hardware configuration by way, what you can do with your software. Yeah, I mean, for example, the, probably the most prominent place that WCA could be seen in action would be the Las Vegas sphere, right?
So it, it feeds data to, to that screen. It's involved in rendering, it's involved in carving, um, at this point it's touching pretty much every aspect of, um, content creation and, um, and delivery to the screen. Um, and, and we, we've seen enough interest off this where other large scale venues are, are looking to, to do the same.
Um, it's just being able to deliver this type of performance over a network is, I mean, I don't wanna say that we don't have any competition, but there's, there's nothing else right now that is able to hit these numbers that, um, that are also built that can scale up to the correct size, not just in performance, but in capacity. Um, this is obviously the trade off, right? I mean, you look at systems that historically have been extremely performant.
They're usually direct attached. If you want 'em to be a little bigger, you would switch to something like san, um, that's not, was never quite as performant as direct attached storage, but it could be much bigger. And then you could go bigger still and reduce some of the complexity by deploying na, which would be slower still, but could go larger.
And then if you wanted stupid scale, you could go to object. And also your performance goes through the floor. So the place where WCA sits really is beating the director type storage performance and also scaling all the way up to object.
Um, and customers that have these massive high performance requirements and large scale, um, are coming and trying our product and just can't really find anything else that can hit those, those metrics. So, um, I'm sure people will catch up, but today it's a good place to be. If, if I can provide a little background in there from a long time storage nerd.
Um, you know, it's funny that people talk about, uh, you know, what you just talked about San and Nas and object. The, the, the scalability and performance of those is really a, a function not of the intrinsic element or nature of the storage or the storage protocol. It's about basically the modernization of the, uh, delivery mechanism and the software that's being, that's constructed to build, to, to deliver those.
Um, and I think that that's sort of the insight that some of these companies recently have had is that, you know, the reason that, uh, direct attached storage was high performance was because it was dedicated. Uh, and the reason that, uh, object storage scaled so, so, so well was because it was distributed. And the idea that you could build a massive scale solution that would kind of combine the best of all possible worlds with, uh, with software is really the reason that so many of these modern systems are able to scale.
Frankly, it reminds me a lot of Kubernetes and the cloud, and frankly, AI itself. I mean, the reason that ai, uh, processing is so incredibly power consuming and high performing is because of this whole idea of distributing it, breaking it up into small tasks and distributing it massively among multiple nodes and parallel. That's exactly what's been going on in the leading, uh, software for storage.
And, and that's the reason I think that CCA is able to scale the way it does. It's not because, um, of some, you know, specialized little trick in there, it's simply because the system scales to just incredible levels, just like the cloud does, just like AI does. And, and I think that that makes it uniquely suited for this AI application because it's such a scalable platform because everything is just completely distributed.
There aren't, there isn't some, some, you know, monolith somewhere that says, you know, this is only how fast it can run. Everything is distributed, everything is run in software. I think that's how people think that things work, but not everything works that way.
And, and, and yours certainly does. Yeah, I mean, you touched on one technology that has been instrumental in achieving planetary scale in anything, and that's Kubernetes. Um, and that the ability to orchestrate containerization in, in a way where as long as you can provide the resources behind it, you can scale horizontally in a, in an extremely resilient and redundant fashion, um, is, is phenomenal.
Um, so WCA released, released recently, uh, is on WCA operator where you can actually provision an entire W cluster or multiple wacker clusters, uh, deployed fully in Kubernetes. So if, if a Kubernetes based shop, um, has compute that already has NVME available in it, and that is all siloed per server, installing, installing WCA via Kubernetes now allows you to bundle all of this NVME in a one giant file system that is available to every single server. Um, and if you, if you're in a multi-tenancy environment, you could actually compose more than one cluster in this environment shared across that, that infrastructure with each customer having its own, um, its own entire dedicated cluster with cluster admin privileges per custom, um, I don't know another product that's doing that to date, but it's, it's pretty wild.
I mean, normally, normally you would treat storage like, um, like pets and everything else could be treated like cattle. Um, but today actually, actually being able to run, uh, storage in a cattle ranch is, um, it's pretty interesting. It's pretty wild.
I really do appreciate that you're, uh, putting some focus on storage. I mean, we've been the overlooked, uh, you know, pet, uh, for quite some time. I'm not sure if I like being a pet or on the cattle ranch, but at least maybe I'm the, the dog managing the cattle.
I like that idea. Uh, but it's interesting 'cause you, you talk about these ability to shift, uh, access to information across multiple points of physical location, which plays well into our conversations of this kind of season around edge, how far away some of those platforms can be from the user or the operator, right? Because we all have this different definition of the word edge, and we've talked about those definitions all season.
Uh, but realistically, I mean, how far out there are you guys seeing that the future of what would be classified as your ability to reach closer and closer to where the data generation point is? Are there certain platforms or, or solutions that you're kind of investigating or already working on? Yeah, it's funny.
Uh, I think the, uh, the definition of edge moves probably more often than the bottleneck moves, right? Is, um, you know, and also one, one person's, uh, edge infrastructure could be larger than another's core infrastructure. Um, and certainly some organizations, the amount of edge infrastructure they may have, uh, can vastly outweigh what they have as core.
Um, but I think, I think as we see more robotics, uh, appearing in the world, as we see more, um, more healthcare being more tech driven, um, the, the ability to feed data to all of these compute processes that are literally out there in the world, not, um, not not sitting in a data center, is gonna become more important. Um, that will require different, different stages, different tiers, fantastic data movement between all these tiers. Um, so yeah, I mean, edge Network is gonna be really important.
Uh, edge data centers, um, edge storage within them, data tiering from there, back to much larger data centers. All the orchestration of this is, um, you know, it's an intense focus for, for cca. And, uh, you know, we, we wanna make sure that as that whole world gets more complex, that we, we stay at the forefront of it.
And I, I think the nature of the solution too, kind of matches the needs there too. Because, because it's built up of sort of this parallel architecture, you can scale up and scale down very effectively. So you can use it at smaller scale in, well, comparatively smaller scale, uh, for, you know, AI processing outside the, the, the data center.
And then you can ramp it right up, uh, to massive scale, and then you can use your tools to enable data to make that leap from, you know, location to location from size. And, and I think that that's, again, ma that, that matches the way that people wish that software worked, but it doesn't always work that way. Yeah, I mean, so we have, um, we have customers right now, for example, running, um, autonomous vehicles all around the world who are generating massive amounts of metrics.
Um, and trying to send all that home, um, is not very efficient. So deploying WCA in many, many different data centers all around the world that can be as close to these vehicles as possible to collect all their metrics, um, clean them and, you know, reduce their size and then set all of that, um, back to a, back to a core for additional research, for additional training and modeling is, um, is a place where Weck has been really successful. As, as, uh, that type of autonomy moves into additional places within the edge, um, with more robotics, I think we're gonna see, we're gonna see a lot, a lot more of that.
Another place that we've been successful is within media attainment as well, where the edge can serve to, to provide, uh, tool sets to talent that can be all around the world, because talent is something you can't really scale. You know, you have to find talent where it resides. Many, many companies have to set up infrastructure where you might, um, you might make a, a tool set available to, to people for either a permanent or a temporary amount of time, but they need huge performance within the, the compute, within the storage, within rendering.
Um, and all of this has to be able to seamlessly communicate with all of the other parties that are participating in the same project workflow, for example. Um, so we've, we've had great success there. Um, particularly in cloud, we've, we've watched customers being able to deploy temporary setups in different countries where you wouldn't even have normally any footprint, um, employ talent in that area, tear it all down when the project's finished, and while you're bringing up more someplace else for a, for another project.
Um, that's been, uh, there's bit of a game changer actually, When people think of ai, especially nowadays. Uh, I think a lot of them are just focused on chatbots and chatbots and more chatbots. But of course, there's a lot more being done with this technology, whether it is using AI and ML in different ways or using HPC, uh, for other related applications.
I know that you all are involved in some of that. Can you, can you tell us a little bit about some other, uh, applications for this technology? Yeah, I mean, for example, we have, um, a few companies in, in, in health and life sciences who are trying to solve some of the, the hardest problems in the world here that really matter to people.
Uh, Memorial Sloan Cantering, for example, deployed wacka to help speed up their, their modeling, um, in the pursuit to solve many cancers. Um, and they have managed to massively contract the time it takes to coherence for a model and massively reduce energy footprint in the same time just by being able to achieve more miles per gallon on the exact same hardware in a shorter time. Um, so this, this is gonna be a game changer if, if Memorial Sloan Kitten can actually achieve what they think they can in the next few years, um, which really excites me.
I mean, it's fun to work on media entertainment, it's fun to work on, uh, cars, you know, many things, but when you actually see life changing, um, work being done, it's quite humbling. Absolutely. And, um, and it's always fun to hear about technology, as I said, that's not just, um, you know, not just the same old thing that people are, that people are thinking of and, and using AI in, in new and exciting ways.
Um, thanks so much, uh, for this incredible conversation. Um, Scott, this is our last episode of the season. com, they'll find both of those seasons along with, uh, six other seasons, uh, previously.
Um, I, I guess before we go, Scott, uh, sum up a little bit about season eight ai, uh, AI data infrastructure, AI at the edge. How exactly, um, should people be thinking about, uh, data infrastructure and storage for ai? Yeah, I appreciate that, and it has been a, it's been a lot of fun this season, and I know Janice has had fun over a couple of seasons as well as my coworker Ace.
Um, from our perspective and from my personal perspective, it's just AI is, is a shiny object, right? It, it is something that's very real. It's very true.
But the fact that we're combining AI and now where we're generating the data and we're generating so much data nowadays, it, it's unique to think that people don't tend to realize as much that you have to put that data somewhere. And a lot of this season, it, whether we intentional or not, has been focused on the advent and benefit of storage. And so it's kind of cool as a, as a longtime storage guy to see the value and the benefits of what we see and our daily lives coming through to everyone else's as something of value.
'cause you, you spend so much time working on data, and that data always seems to be, you know, the, the star of the show in certain different processing and things like that. But as you saw through the season, if you go back and look at it, we've talked to a whole bunch of different ways of looking at managing data, focusing on data, and all of that revolves around where the data sits. And the data doesn't always just sit in a CP or dram, which are wonderful toys and tools, but it does have to, you know, have a long time, uh, placement of that.
So I, I see that as kinda one of the bigger nuts of this whole season is just, it's, it's cool to know that storage is really getting a, a play in the space, and all these companies are doing so many cool innovations to again, shift those bottlenecks and, and talk about it, whether it's the industry standards bodies, the software platforms, the hardware platforms, a combination of all that. So it's been a great season. I've had a lot of fun, and I've learned a lot myself.
Well, thanks a lot. Yeah, it, it has been a great season for me as well. Uh, obviously an old time storage nerd here.
It's fun to see where this, uh, industry is headed. And, and it is fun to see just how all of those things that we wished we could do have, in many cases come true with modern software. So just, just incredible overall.
Um, Alan, again, thank you for joining us and representing Waka here on Utilizing Tech. Um, as we wrap up this episode, where can people connect with you and continue this conversation? So actually, uh, this week, um, 18 to the 19th, WCA will be presenting, um, at the three big AI conferences, uh, San Francisco, London, and Singapore.
So yeah, if you can make it down to those, please, uh, come in here, what we're all about. Great. Uh, Scott, uh, I guess going forward, uh, where can people continue speaking with you and your colleagues and, and learning more about Soy?
Yeah, for Soy, it's pretty straightforward. com/ai/ai, and, uh, also you can find me on LinkedIn, blue Sky, and, uh, Twitter, formerly known as or ex, formerly known as Twitter at SM Chadley, uh, I'd tend to spend a lot of time having fun, sharing insights and just being a little bit social. So, Yep.
And you'll find me as, as FoST on most of the socials, including, uh, blue Sky and Mastodon as well, uh, and of course on LinkedIn. Thanks for listening to this episode of Utilizing Tech. Uh, you can find this podcast in your favorite podcast application as well as on YouTube.
As mentioned. This is the last episode of season eight. Uh, yes, that's right.
There are eight seasons of this, and you can go back and listen to those, uh, all the way back to, to the pre-chat GPT era. If you enjoyed this discussion, please do leave us a rating and review. It's really nice to see those.
This podcast was brought to you by soine this season, as well as Tech Field Day, which is now part of the Futurum Group. com or find us on X Twitter, blue sky, or Mastodon at utilizing Tech. Thanks for listening, and we will catch you next season on utilizing Tech.