UT07x06: Connecting Ceph Storage to AI with Clyso – Utilizing Tech
Many of the largest-scale data storage environments use Ceph, an open source storage system, and are now connecting this to AI. This episode of Utilizing Tech, sponsored by Solidigm, features Dan van der Ster, CTO of Clyso, discussing Ceph for AI Data with Jeniece Wnorowski and Stephen Foskett. Ceph began in research and education but today is widely used as well in finance, entertainment, and commerce. All of these use cases require massive scalability and extreme reliability despite using commodity storage components, but Ceph is increasingly able to deliver high performance as well. AI workloads require scalable metadata performance as well, which is an area that Ceph developers are making great strides. The software has also proved itself adaptable to advanced hardware, including today’s large NVMe SSDs. As data infrastructure development has expanded from academia to HPC to the cloud and now AI, it’s important to see how the community is embracing and improving the software that underpins today’s compute stack.
Transcript
Many of the largest scale data storage environments use cef, an open source storage system, and are now connecting this to ai. This episode of utilizing Tech sponsored by soddy features Dan Vander Stir CT of cly o, discussing CEF for AI data. Welcome to Utilizing Tech, the podcast about emerging technology from Tech Field Day, part of the future and group.
This season is presented by soy and focuses on the questions of AI data infrastructure. I'm your host, Steven Foskett, organizer of the Tech Field Day event series. And joining me from SOY is my co-host, Janice Roski.
Welcome to the show. Thank you, Steven. It's great to be back.
Well, it's good to have you here. So, as we've spoken about many times, uh, in the past, especially here on this whole season of utilizing tech, there's a lot of data out there. Um, there's a lot of existing data, a lot of existing data sources, and a lot of the existing data platforms and all of that is going to eventually need to be integrated into the AI data pipeline and into the whole AI picture.
Yeah, absolutely. And, uh, you know, there's lots of ways to go about it, right? A lot of different hardware out there, software, um, and, and folks are really trying to figure it all out.
How do I make all of this work together? Um, and there's one, um, software tool out there that's open sourced, uh, that many, many companies have been using for years, right? And now with the advent of ai, they're kind of like, how do I use this tool I've been using for a long time that's open source and free.
Uh, how do I make this work for my, my workloads going on today and, and the ones that are ever evolving into the future? So we're excited to talk with, um, somebody from Seth. So I, you know, I'll turn it back to you, Steven, to introduce him, but excited to dive into this topic.
Absolutely. Yeah. And, and, and Seth is one of those things, um, storage nerds like me have seen for many, many years.
Um, I, uh, I've watched this project, um, for, for Grow. I've watched it become absolutely a critical component. Um, many people might not have heard of it, but, um, as our guests said, um, it, it kind of is the Linux of storage.
It is everywhere, and it is used especially in environments that have lots and lots of data. And those are the environments that are going to need to be integrated into the AI data pipeline. So let's welcome, uh, Dan Vander Stir from, uh, uh, lyo here to talk a little bit about the importance of SEF in ai.
Thanks. Thanks Steve and Janice? Yeah, I'm Dan Vander.
I'm CTO at Lyo. I'm coming from, from cern. I spent about, about, about a decade at CERN working in the IT department, working in storage.
And, uh, I'm also like wearing multiple hats. I've been, have the pleasure of working with STEP for around a decade as well, um, early, uh, testing at scales that hadn't been seen before. And, uh, yeah, participating, getting more active in the community, community and eventually, um, now, uh, acting actually in the project, the open source project as a member of the executive council leading the overall open source project.
Yeah, So many people haven't heard of s like I said, but many, I mean, I'm sure it affects almost everybody now because it's, it's, it really is everywhere. I mean, essentially this is, um, a software based storage solution that uses, um, as, so full disclosure, I was there at the beginning when this was originally announced, and I was excited about it because of what it is. It uses unreliable components to build reliable, high performance, scalable storage solutions.
So essentially it is massive scale, massively distributed, and it's designed to not just be able to adapt to failure, but to be ready for failure. And that's one reason that it's become so successful. So tell us a little bit more, where is SEF today in the overall picture of, of the, the world's information?
Right, so that's, uh, I mean that you highlighted a lot of the points that actually attracted us to s early on and made us one of the early adopters. Um, you know, what organizations are looking for, and also like the, the, the people operating the storage infrastructure are looking for is something that's like reliable and can be built out of like, you know, low cost, um, commodity components. But then you can build something that's reliable and scalable.
And one of the, like, main things is that if you're building a large scale storage system, you don't want to have to like every four or five years lift and shift, migrate data from one, from one appliance to another. You want something like an organic storage system. And this is how we, this is how I think we wrote an internal memo at cern, like toward an organic storage system for, for our cloud.
And that's how we got it started. And yeah, it really proved, it proved to deliver what it promised, which was that kind of scalable, reliable system where the operators can, you know, sleep at night and things fail all the time, but you end up with a reliable system that that can grow and evolve with the, with the organization. So on the, on that notion of, of creating something that's scalable and reliable and kind of always on, um, could you dive a little bit into, um, how you are working with AI with some of your partners today and, and, um, how you're bringing through some of that Seth goodness and, and bringing that to ai?
Yeah, I mean, ai, AI use cases, certainly the hot topic with Seth these days, because SEF is a, like, in addition to the whole reliability aspect of sef, it's also very flexible. It's, um, it's a low level object store internally, but actually on top it's like very familiar storage components. It's like block storage for a private cloud, or it's object storage compatible with, you know, the public clouds.
And it's also a file system, a normal POS file system, uh, like a typical NFS what, what you would expect from NFS. So like, because it's so flexible, um, it already has made, like, it's already used quite massively in those, you know, in, in cloud environments for object storage, especially like self-hosted or hybrid cloud or multi-cloud environments. So there's really just like a lot of data out there that now in the AI context, um, you know, organizations want to process that data that's inef more quickly, more rapidly, and it always puts pressure on the project to like deliver more and more features and performance, uh, needed for, you know, the, the cont the exponentially, uh, expanding like, uh, pro like processing capacity that we have in the AI world.
So are those, um, I, I guess what are the industries and use cases that are predominantly using CEF today? And, and let's, you know, let's talk about that and that's, then let's talk about how those industries are gonna be using AI with this dataset. Right.
I mean, so s had its beginnings in the academic sphere, so it was really a lot of initially universities research centers, research labs building and using it like they would use, uh, hp. Like it's, it started out as a HPC file system, and then it evolved, these other flexible use cases. Um, today it's used by all sorts of industries, every industry from, let's say, financial industries, um, trading, you know, you know, high, high frequency trading or, you know, complicated algorithms that are, that are processed through, you know, crunching through lots of financial data.
Um, of course still still like super computer centers doing any kind of research, let's say biotech research, things like that. Uh, and also like, you know, on the entertainment side, like media storage, like SEF is massive in media storage, like backing, backing, you know, some of the largest, uh, media companies as well as like video games and things like that. You know, it's really, it's, it might, you, you say, it may not be well known, but it really is backing the largest infrastructures that we have out there.
You know, SEF is behind the scenes, powering it all. And you mentioned a moment ago too that, uh, you know, s isn't always seen as being maybe a, you know, high performant, uh, solution. Right.
Um, can you talk a little bit more about, you know, how SEF does bring forth that performance? Yeah, I mean, it's important to, it's important to understand when Seth was created it, like the idea was to do something better than previous storage systems, and that was to put a very strong priority on the durability and consistency and reliability of the storage. So to not make any sacrifices, uh, for performance or for any other reason on that, on that data consistency, um, CEF has a lot of, there's a lot of smarts internally to enable that to happen, like at, at, at high performance speeds.
io, we've done recent tests to show and to try to make a little bang and make the other file systems a little bit, you know, pay attention a little more. We did a one terabyte per second demo showing that you can get, I mean, I'm not sure another file system is demonstrated to one terabyte per second recently. Um, but this was with, you know, just with open source, open source software on commodity components, few hundred n VMEs standard standard servers that anyone can, can purchase and build, you can build like the best AI platform for your, for your organization with it.
Yeah, It seems like that is really true to the, the philosophy of SEF two, because right from the very beginning, it was all about using, you know, mundane components. You know, it's not extreme specialized components. It's about using ordinary components that are accessible and affordable and, and, and, and varied and combining that into a unified system that offers, you know, massive scale.
And one of the fundamental architecture concepts of C is distributing everything, not just in not having a single point of failure, a single bottleneck. So it makes sense to me that that approach would be able to deliver high performance if that was one of the design, one of the goals that that, that folks had when kind of re building it out, tuning the, tuning the system, because that is very similar to how, you know, AI clusters and HPC clusters are built. They're built out of, well in, in some cases, uh, more exotic components, but they're still built out of massively scalable components, and they're massively distributed because anytime you have any kind of bottleneck, well, that's a bottleneck that's gonna cause a performance hit.
Um, so do you find that you are able, uh, it sounds like you're able to change, I guess, some of the tuning or some of the, the, the, the way that Seth was designed to really focus on performance and distribute that workload, is that, is that how it works? I mean, in the earlier days of sef let, I don't know, maybe let's say eight years ago or 10 years ago, I think we all in the community had the idea that, yeah, we call it like horizontal scalability. If you want more performance, just buy more servers or buy more devices, and you just add them and then you, you know, scale linearly the performance and for the very lowest levels of cef and for things like object storage, that's really true.
Um, and we had, we had the, the ambition and the like goal that like, okay, we can like see beyond NFS, we won't have psic file systems anymore. We don't need them anymore. We can just do everything with object storage.
But I think we didn't really, I, it still seems that psic file systems are, are popular and needed, and they're just the normal expectation for, for, for a variety of different reasons. And pause. Excel systems are not as easily horizontally scalable.
That's the issue. Scaling AI workloads brings new challenges to metadata performance. Um, really opening really like millions of files per second becomes a i is, you know, that's a very highly, highly metadata intensive task.
And, you know, so we work hard to, to, to have new ways to make the metadata the set for s the c metadata features actually scale as well. They're pretty good at it already. I, I, I would say it's like we anticipated this few years ago, and it's already quite good, but, you know, we're always trying to push to the next level and be, and be ahead of the, ahead of the curve a bit.
Yeah. Well, when it comes to AI training, I would think that that would be a problem because of the scalability of the clusters that are being used. You have a ton of clients, they're accessing the same data, uh, like you said, they're opening lots and lots of files.
Is that what you're seeing that it, it tends to be a very, very massively parallel access pattern. That's right. Yeah, exactly.
And repeating the same exact, the same ones again and again, again. And it's really this like, um, it's, it's the number of files that's the, that's the thing. It's very easy to make a file system or a storage system, which deals with large objects or large storage, and you just stream everything in parallel.
And, you know, in the HPC space, we call this embarrassingly parallel. So it's easy to make an embarrassingly parallel storage system, but when you have contention and you have clients, you know, all looking at the same files, maybe modifying files in the same directories, and because SEF has this very strict view of never allowing any client to have an outdated view on the current situation, um, that's where Seth is like, actually, you know, and with, with our developer hat on, we're like, Hmm, maybe for real life that kind of strict consistency is not always needed. So maybe we should consider relaxing some of those, those, um, to, to, to behave more like the, the rest of the, the, the file systems that we, that we compete against.
Yeah. 44 terabyte Q-L-C-S-S-D? Yeah, I mean the, so of course there's a benefit, right?
Because you get, so when, when N VMEs arrived on the scene, SEF was not prepared for this because SEF is a very, it's a very smart, there's a lot of soft, a lot of code, a lot of lines of code that go into, into like doing the, the storing the data. And, you know, it was WR written at a time when I don't, we had, we probably had early days of flash, but we, but it was written at the time of eight of spinning discs. And the, of course, there was a lot of work done in those early days of, of NVME to really make sure that you can extract all the high performance out of that.
And there were some early workarounds done as well. Yeah. One of the like, low level hacks that, that, like the practitioners learned early on was that they should take like large N VMEs and split them into many vir different virtual devices and then like treat them as separate discs.
And then that way, that way Seth could get more performance and get the native performance because they're really just, they perf they have too many iops and they have too much bandwidth. And, and the, the software couldn't keep up, but that's like maybe five years ago. Now, in the last two, three years, the SEF developers, the SEF project has spent a lot of time getting reworking the internals to make sure that they can extract the, all of the performance.
And there's a, there's still an ongoing, so we've achieved quite a lot already. And there's like an ongoing major project called Project Crimson, um, which is like a major rewrite of the storage of the low level storage demons, which is a hundred percent focused on, uh, extracting performance of large n VMEs like you're talking about now. So there isn't, in your opinion, any sort of, um, issue with, say the endurance, right?
Uh, when it comes to working with QLC, I mean, this is, this is definitely a relevant issue that comes up when we are designing systems for, for, for customers. We pay attention to the endurance, we pay attention to things like right amplification that might be relevant for their exact use case. So we, so when we're, when we're, you know, when we're working with a, a customer or a user, we try to understand exactly how they're using it and then guide them.
SEF is one, one of the things about SEF is it has about 10,000 tuning options. And so you can really manipulate anything about the sector sizes, the block sizes, how data is chopped, chopped, diced, and sliced and distributed across the cluster in a way that optimizes the usage of the underlying devices, right? So if you have that, you know, full like holistic view on the, on the hardware, you can tune the software to, you can tune stuff to extract, like maximum performance.
And so maybe out of the box it might not always work optimally, but with a little insight and a little expertise and, and guidance, you can really get, that's how you achieve results, like one tart per second. It's like paying attention to all those stacks we see a lot of, we see a lot of, a lot of CEF users making, making use of those large devices. Yeah.
So, um, Dan, you are out there helping people build, um, CEF systems and adapt their CEF systems and, and basically bring them into the AI world on a daily basis, I assume. Um, talk to us a little bit about some of the real world here. I guess, uh, one of the questions that immediately springs to mind is, you know, something you brought up, which is the, uh, the object store versus file system versus block question.
Um, you know, some of the performance questions, I'm just excited to hear how real world users are using SEF storage with ai. It's such a wide variety. Um, it's hard to find one classical explanation, like one classical model.
I mean, so if you look at, if you look at the typical, like very large supercomputer right now, um, they're often, I think most, many of the recent designed supercomputer centers, like let's say top 500 or IO 500 stuff, you'll find a C tier on the outside, on the other layer as a sort of data lake. Um, that's very common because it's really the lowest with erasure coating. It has very efficient erasure coating.
You can, you can get a low cost, um, you know, standards compliant. It speaks S3 and Swift. So it's, you got, you got a compliant normal object storage that you can deploy and, and run it, um, and run at a scale that makes sense to, to different organizations, right?
Um, but you know, we, we, I mean there's no, there's no, that's the beauty of cef. It can do whatever you need. It can, it can, it can work In that use case.
We also have AI environments that are building their solutions on top of block storage. So like, oh, let's prepare some images, some, some block device images that have data sets, and we'll just like mount those on the fly and deploy the data like that. Um, or of course the file system is, is, is like always the standard.
Let's, let's have a file system, let's mount it and let's put all the data there and then, and then throw our, our massive, um, uh, GPU clusters at that file system. And then we get the call like, oh, it's like, it's, I think, did we overload it? Like is it broken?
I mean, it's like, we got, so we can help, we help with those kind of issues too, because it's like, it's one of those things that just, I don't know if I, I don't, it is a good thing about Seth, but one of the things that happens again and again is that it works very well that it like absorbs more and more of the use cases AI here, oh, let's store the pension fund on it too. And then like, eventually you have this, like, you have this thing that, that like this becomes the core critical backend of the whole company or the whole university. And that's like, then you get, okay, let's, let's re-architect this, let's, let's move things around.
Let's do like, let's, we can do stretch clusters across multiple data centers. We can, um, we can take that part out. That's really like a analytics thing and AI focus.
Let's move that out into a separate system. 'cause it's not appropriate to have that with this other system. And you know, so then you can talk about like third or fourth generation, like after you under, after they understand their, their, their use case a bit more.
Yeah. This is, that's a repeated thing that that happens, happens quite often in the sef in the SEF world. Yeah.
Well, I do love that, you know, uh, Seth is just incredibly, as you said, flexible, right? And because it's been around for so long, there's so many different organizations that can take advantage of it and evolve with it, especially as we are in the world of AI where everything seems to be evolving every second of every day. Um, can you give us, as you know, back to Steven's Point, give us an example of a particular partner, um, say, uh, gosh, I was just reading an article.
Hugh had something about, um, you know, CEF on, uh, Abuntu or, you know, any, any partners that, um, you can speak to specifics on, around how, you know you're using SEF for ai? Uh, I mean, the, the, IM, the important thing is that, is that, you know, clayo and we participate actively in the SEF foundation, that's a, the SEF foundation is, it was established. It's, it's part of under the umbrella of the Linux Foundation.
So that's like the open forum where the organizations, uh, can coordinate their, their, their activities. We have different hardware vendors, software vendors, and you know, everyone, we find those projects to work together. So Dan, you have a lot of experience in the HPC space, and I think that many of us have seen a ver very much a similarity between HPC and ai, both in terms of architecture, but also in terms of, you know, the, the technical solutions that are used as well as the community that's supporting it.
Obviously a lot of, um, the, uh, AI space is built on open source tools and open source projects. A lot of, uh, the, uh, HPC um, customers are out there deploying AI or, or experimenting with ai. It certainly is, uh, a lot going on in academia.
Um, talk to us a little bit about the community and, and the ways that that AI and SEF and open source are working together. Yeah, so I mean, it comes back to the roots of SEF and then like how it evolved. It started in, like you said, in that HPC targeted use case because of its flexibility.
It was the right technology on the spot with private clouds coming. Then we went, you know, we went from private clouds to containerized and Kubernetes. It was the right, it was the right technology for Kubernetes and remains today with, um, with the CCSI, uh, plugin.
And, you know, each time, you know, you grow the community. And now, so then other than also Data Lakes, you know, with object storage, the massive large scale data lakes, it was the relevant technology. And now with ai, it's the same.
It's like this ever growing community. And, you know, for those, for those of us that have been participating sit for, for a while, like, like you also, um, had the pleasure. I wish I had been there in the, in the initial announcements, but I was maybe like a year later.
Um, like the strength of s has always been, it's like awesome community. Like there's, with the mailing lists, with the events that we organize, the, we have CEF Day events around the world, always organized. We have a Cephalic con, uh, that's our yearly large event this year.
It'll be in December at, at my former organization, cern. Um, you know, we, we bring people together and it's like this friendly open source community where everyone's sharing experience, contributing valuable feedback to developers to make it better. And it's like, you know, it's really, it's just like, it's the Linux of storage.
Like everyone work together and make something that is gonna like, solve real problems and make, like, make, make lives better and make organizations like more efficient. And it really works. io, follow the links, join the mailing list, join the slack, check out our YouTube, watch the videos, come to one of the events.
I'll be, I'll be giving, if you check on the cly O LinkedIn next month, I'll be giving a webinar on, uh, related to, to mastering some kind of CEF operations tasks. So some kind of tools that, that we have to, to simplify the, the operations of some of the, some of the more maintenance tasks that might be like tricky for some operators, but we have some tools to make it a lot easier. Um, yeah, so that's it.
Yeah. Yeah, it really seems that, um, you know, the, the whole open source environment is, is still quite vibrant, quite alive and well, and, um, and I'm glad to see that because as I said, I mean, it's, it's important in HBC, it's important in ai, and it's important in all of this. And, and like you've noticed, I, I think Dan, the, one of the most important points you made was that, uh, tools like this tend to find their way outside of their intended use cases.
And, and if they're useful and if they're, uh, they prove themselves in one area, they tend to absorb other use cases as well. And as companies are deploying AI applications, the, the question is, how do we get data into this? Well, in many cases, the customers already have data in something like sef, and they could absolutely think about using that as part of their AI data pipeline, just like they would as part of their analytics pipeline or their research.
And I think that it's, uh, you know, if there's one message that comes out of this, it's that Seth is absolutely useful and will be useful in the, uh, AI space as well. So thank you so much for, for that overview. And definitely, uh, I, I echo what you said.
Check out some of the, uh, open source work that's being done and, uh, get involved. Thanks very much for having me. It's been a pleasure.
You. Well, thanks for being here. And uh, Janice, thank you for joining me, uh, for this episode of Utilizing Tech.
Uh, thank you also to the listeners. Um, if you enjoyed this podcast, uh, please do leave us a rating, a review, a comment. You'll find it on your favorite podcast application, and you'll also find us on YouTube.
This podcast is brought to you by Soy and by Tech Field Day, part of the Futureum Group. com, or find us on X Twitter and mastodon it utilizing tech. Thanks for listening, and we will see you next week with another episode of utilizing Tech focused on AI data infrastructure with solid.