UT07x08 – Deploying AI Data Infrastructure in the Datacenter with Ariel Pisetzky of Taboola – Utilizing Tech
Transcript
As practical applications of AI are rolled out, they're increasingly being deployed on-premises at scale. We're wrapping up this season of utilizing Tech with soy focused on AI data infrastructure by discussing practical deployment considerations with Ariel Preki of Taboola. Listen in and get some practical tips on how to deploy AI data infrastructure on-prem.
Welcome to Utilizing Tech, the podcast about emerging technology from Tech Field Day part of the Futureum Group. This season is presented by soy and focuses on the question of AI data infrastructure. I'm your host, Steven Foskett, organizer of the Tech Field Day events series, and joining me from Soy today is my co-host, Janice Roski.
Welcome to the show, Janice. Hi, Steven. Thanks for having me back.
You know, this is the end, this is the end of this season of, uh, utilizing tech, and I thought that it would be sensible, and I think you thought the same thing to end with a practical application end with somebody who is really building AI data infrastructure out there in, uh, in the real world. Yeah, I couldn't agree more. We've talked to lots of different, you know, ISVs, hardware vendors and the like.
Uh, but it's really exciting to talk to an organization that hasn't just been looking at AI over the past couple of years, but really been looking at it for the past 10 years, and how are they doing things differently Now, I'm excited to talk about or understand how they're looking at, uh, you know, on-prem versus off-prem Cloud solutions. And, uh, yeah, we have an actual end customer, so excited about this. Cool.
Well, let's, uh, let's waste no time and, and, and, and bring 'em in. Um, so Ariel pki, um, VP of Information and Technology and Cyber at, uh, tabula is our guest today. Ariel, welcome to the show.
Tell us a little bit about yourself. Thank you very much for hosting me. Well, I've been in it for many a years now.
Um, I think my bragging rights go back to as 1 6 8 0, which is one of those first 2000 ISPs on the, on the worldwide web as we called it back then. Uh, we even said security. We didn't say cyber when, when I started out and now, uh, really many years later with, with physical data centers cloud, so many things that one could not have imagined back in the day and having a lot of fun providing services to many, many readers out there.
Yeah. Uh, so Ariel, let's just dive right in on this. So, um, you're obviously a content delivery network provider, and we wanna talk about ai.
Can you just give us a little bit of insight about how you are utilizing AI today in, in your workloads and applications? Sure thing. So, uh, I'll give a few words about Tabula, because Tabula is not a brand name.
We're not a consumer facing product. So many people might use our products on a daily basis, but not be aware of them. We are a content discovery platform, and that means that we reside on many of the publishers that you read on a daily basis and love and receive content from.
And we are that place where advertisers go to turn users into paying customers. So we actually provide that matching service between advertisers and publishers and content. Now, bringing all that together, we eventually serve over 4 billion webpages a day, and that means that we recommend over 40 billion different articles and kind of, I'd say, pieces of content for you to consume at any given moment.
Now, we do this with high levels of personalization. What does that mean? That means that when you are on any given website and you are browsing reading, you are immersed in the content, then we will provide those content recommendations for the next thing that you can read, the next article, the next content that you may like, bringing that content to you without actually knowing who you are, without you logging into our service, without you providing us any specifics about yourself, not your name, not your age, nothing about your gender.
That is the kind of essence of using ai, specifically training and inferencing in large scale and in real time. Over the past year, with the emergence of generative ai, even a bit more than a year, we have been also taking advantage of LLMs and different generative AI technologies to provide additional tools for editors and for advertisers to curate the, I'd say article, name some of the article itself and imagery that you might get on any given day within your, uh, beloved websites. It is interesting that, um, you know, people when they think of ai, they think of gen ai, they think of like everything that's happened in the last, well, 12 months, basically.
But that's not the whole picture of ai. In fact, uh, companies have been deploying systems based on artificial intelligence concepts for a long, long time. I'm a particular proponent, uh, for example of expert systems, which have existed literally since I was born.
Um, and the whole world of, of leveraging data to build productive applications. I mean, that is the story of the 21st century really. Um, everything that you've talked about in terms of what Bulah does, uh, in terms of how, how it impacts its customers, how, uh, those of us out there in the, in the real world interact with it is basically the story of it in modern times.
And I'm glad to hear you say and recognize and claim, you know, yeah, this is an AI application as well, and generative AI is a new direction, but it's not the only direction. Um, and yet it is a interesting way to go. So I, I guess talk to us a little bit more about the ways that data feeds your businesses, because I think that your business is all about data.
That's, that's so true. So data is actually the basis for training. What you get when you see 4 billion web pages a day is these massive amounts of data that are written somewhere, in our case to physical servers and physical drives, solid Im drives that we utilize on-prem, and we'll talk a bit more about that and the additional tech at the low level in a moment, keeping at the AI level, that means that this data, if you don't do anything with it, if you don't try and I'd say infer any knowledge from it or create knowledge from it, it's just inert data.
It's just inert bites or bits even that just sit there with, with no value. So the place that it comes to create value for the business is taking data, even if it's vast amounts of data and turning it into, into knowledge and turning it into business relevant knowledge. So creating that understanding of, oh, wait, we've been providing recommendations for 15 years now, somewhere, uh, seven, eight years ago, we said, wait, there is a better way to provide content recommendation by trying to understand what is actually interesting for people.
And we say, we're sitting on all this data. How do we train? How do we look?
How do we put this data into different boxes? So that's like the beginning of the training. What are the signals that we are getting from this data and how can it impact and affect our kind of service moving forward?
Then of course, there is natural language understanding and natural language processing, which is a big part of AI today with LLMs, it's a totally new field, but even seven years ago and moving from then and until last year when LLMs kind of came out to our lives, that was something that was a hard problem to solve. You needed to understand as a service provider for us, as a service service provider for publishers, how do we recognize the article that we reside on? How do we recognize the user coming into that specific article, and what is the relevant article where that user kind of browsing arc is, is going to end?
And that is been do, that has been done with AI on different levels. And today with more and more with LLMs, all of that happening on-prem, because storing those vast amounts of data and actually processing them and creating value out of them is something that when you own the data and data has a certain level of gravity to it, when you own the data and you can process it, then you have a lot of advantages over putting it somewhere in the cloud where it might be very expensive to even run operations on it, reminding us all that. When you run your own drives, you don't pay per 10,000 operations.
You don't pay for deletes, you don't, I mean, you pay maybe in in performance, you pay, there is a payment there, but it's not hard cash once you bought your drive, that is it. And then the use of that drive over time is a bonus. Thinking about accounting, sorry, to bring the, uh, pencil, pencil pushes, pushes into the room.
But if you use the drive for over three years, oh, it's even not on the books anymore, and it's free. So free compute, who can't love that. Who can't love that.
And you know, you talked a little bit earlier, Ariel, when we first started, you know, uh, discussing this on-prem versus off-prem, and there's a lot of organizations out there. We're having the same kind of debate. So can you tell us a little bit more about how you are, you know, working with AI on-prem and you know, how that might be different from what others might be doing in this, in this space?
Yes, yes. That's a, that's a, well, the story there is, is actually fun. I came to Tabula years ago to cloudify our operations, and, uh, somewhere in that Cloudification story, we started noticing costs and how we as it impact those costs and how we as it can do better for the business, and what does on, on-prem actually mean.
And when you optimize and look at the use of data at the use of CPUs, at the merging of those two, and over the past year at, or even a few years, the merging of CPUs, GPUs and data, you suddenly find that working on-Prem has its extreme advantages. So when I talk about optimization, I'm talking about multiple layers of optimization. First, understanding that when you're looking at a data center, you're looking at compute, you're looking at data or storage, and you're looking at networking data itself can be, of course, data at rest or can be data within a database service.
Now, in any one of those levels, you have multiple layers of optimization using cpu, better using CPUs in a, I'd say higher capacity, uh, making sure that all of your, uh, GPUs and CPUs are fully utilized. Understanding what LE layers of optimization exists there. One great story is that we looked at our CPUs and we found that we can better utilize our CPUs by updating our code in various ways from, uh, different libraries coming from Intel or coming from NVIDIA, or well, those that for GPUs, of course, or coming in from, uh, other vendors.
And then really better utilizing our CPUs then when we better utilize our CPUs. We of course have the drives themselves, the drives themselves today, the NVME SSDs, the NVME interface, and the SSDs provide really so much performance. And when you understand the geometry of the drive and start to think not only, I'm not only buying space, but I'm buying specific types of space, space with different drive geometries that fits in different places.
If you are able to optimize that as well for your read size, you suddenly get this boost of performance where you can do so much more with your on-prem hardware, your on-prem investment in CapEx. So we really love to see how we year over year optimize our use cases for storage, for CPUs and for network, and bring them together to a place where our developers now have so much raw power at their fingertips that they just do not want to go to the cloud for many of the day-to-day operations at the age of GPUs, where you need to push a whole lot more data into the GPUs, controlling your storage layer and controlling your network layer also provides you with multiple opportunities for, um, optimization. One great example for anyone out there listening, think of data compression.
Data compression costs you with CPUs and it costs you with network, meaning the network layer will work harder. But the CPUs, let me just, uh, clarify there, if you compress CPUs will work harder and the network will work, uh, lighter and you'll have more storage space on your drives. The flip side there, if your CPUs are the most expensive component in your data center, why send CPU cycles to something that isn't business oriented?
Why compressed? Get the right type of storage that you need, get enough bandwidth, the bandwidth in the data center is literally free, is literally free. When you have that level of free usage of bandwidth, you can really, really utilize your drives and push a whole lot more data into GPUs.
It really is an interesting, um, use case situation because I think that it pushes back on a lot of the story that people have bought into over the years about cloud. Um, you know, I think that there is certainly a valuable, uh, use case for cloud. I think that, that everybody would agree to that, but I am hearing more and more companies that are looking at deploying, uh, cloud technologies on-prem and deploying now in the future AI technologies on-prem.
Uh, once, uh, applications are rolled out, once they're fleshed out, once they're, you know, sort of have a, an understanding of the level of hardware that they're going to need to support those applications. Because as you say, you know, you can buy equipment, you can deploy equipment in internally, and I think there's another aspect that you kind of, uh, mentioned there too. Once things are depreciated, you can continue to use those past their expected lifespan.
Maybe they're not in production, maybe they are. Um, are you seeing that there's a much longer lifespan with, uh, modern hardware and specifically modern storage than you expected? So drives have become so much more reliable over the years.
And again, when you work with a good vendor, and dyme is a great example of that, when you work with a good vendor that provides you drives that do not fail beyond the expected MTBF, you're getting a good bargain for drives that maintain their value beyond their, their three year depreciation. So you get this really, I'd say great deal out of using hardware on-prem. The cloud, of course, is relevant.
We in it have not done a good, a good service to many of our kind of business counterparts over the years. We have provided services too slow, it took too long to install things. We did not embrace automation on time.
The hyperscalers have taught us so much about how we, as the regular Joes of it should be providing services within our data centers. So the cloud is great for dr, the cloud is great for testing. The cloud is great for production loads that vary and have peaks that are maybe short-lived, but for businesses that are always on, they should probably be always off the cloud.
So if you have something, a workload that is always on, it should be always on-prem. That is where I would, uh, that that's where I would take it. So Ariel, uh, given that that notion and that thought, uh, is there any benefit to efficiency by being always on-prem and are is, you know, tabula doing anything specific to address some of the, you know, energy issues that are coming about with ai?
Oh, yes. So when, when you're on-prem, of course efficiency suddenly becomes, well, your problem when you're in the cloud, then you get the great, um, carbon footprint of the clouds that are in renewables. You get wonderful e-waste management by the clouds and and so on and so forth.
So when you are on-prem, you need to kind of control your own destiny there. Anything and everything on the e-waste side and anything and everything on the kind of, um, energy side, that that is clear. But then when you look at AI specifically, which is a huge energy kind of hog, because they, these GPUs and these CPUs are running way hotter than than they would normally for other, uh, activities.
You want your storage at least to be as efficient as possible. So when you have SSDs instead of spinning drives, that's of course the no-brainer. But when you look at the larger capacities out there that still provide amazing kind of, uh, performance in terms of the level of IOPS that you can get and their thermal footprint doesn't warm up your data center and their energy levels are in use only when you are at full right mode and so on.
But you have varying levels of energy that you can also manage through the NVME interface, should you so choose all of those, provide you with the ability to A, optimize and b, provide a great service, yet still keeping a very, very echo friendly footprint. One of the things I've heard from, uh, infrastructure managers that are considering buying equipment for on-prem is that it makes sense to buy the biggest, baddest, you know, hardware because then it'll have the longest lifespan. In other words, you buy the bigger drives, not the smaller ones you buy, the faster CPUs, the faster GPUs, even though they cost more today, they'll have a much longer lifespan and you'll be able to get much more out of it.
Do you agree with that approach? No. So a short answer, but here I'll give you the longer version.
Thank you. That's a wonderful question. Uh, thinking of what you need is really what the business needs.
Are there big, bad, amazing servers that we buy? Yes. Do we always buy the biggest, baddest, um, servers out there?
Absolutely not. We try to balance a lot of our servers into something that is a, in, I, I'd like to call them Lego block mode, where we buy big servers, but not very big servers. And then these servers have a life cycle, so they can be a server that is driveless at the beginning and it's CPU intensive.
Maybe it even has the GPUs in it two years later. I have now a better CPU or a better GPU. And in terms of energy and optimization, this server is not a good AI server for me anymore.
But lo and behold, it has drive base. Now I can use it as a storage server and I can get that CPU that isn't the latest and greatest to do great things around managing all the data that is in that server. So Lego blocks really the, which is also why I, I'd say so many of us love Lego, their versatility, the fact that they, that, that their lifespan is so long and that they can, you can use them in different models and in different, uh, scenarios and build with them different things.
That is what is so enticing for us to buy servers that are somewhere in the middle, not too big, not too small, and then for specific use cases, go wild. Ariel, one of the things I was gonna ask was being on-prem and having to deploy across, you know, multiple environments, uh, it's pretty taxing, right, to have some of your IT guys go out and have to rip and replace drives, right? And, you know, on the notion of buying, you know, not just the biggest drives or the fastest or, um, or even the, you know, that same case with your servers.
How, how do you feel, uh, solid state technology helps to, um, improve how, how you do business, you know, at the edge or, you know, within your, in your on-prem infrastructure, is there, is there a benefit of solid state storage versus hard disk drives in your opinion? Okay, of course, yes. So having solid state drives at the edge is super important.
I'll talk about that in in a moment. And then of course, just to talk a moment about Tabula, we have our frontend edge edge data centers, and we have a backend data center where we run all the training in the backend. Yet the inferencing that I spoke of, the, um, hyper-personalization of our service happens at the edge.
It cannot happen without data. It happens on data on SSDs at the edge. So the closer you are to the edge, the faster you will be able to provide service and the faster you need the drives, the servers and the service to be putting all that in perspective in terms of operations and IT operations.
So I, I'll go back to the Lego blocks I spoke of earlier. The idea, if you don't go biggest baddest, but you go good and exactly what you need and good being the solid dime seven terabyte drives as an example where you're not going all 64 terabytes, 60 terabytes, uh, 30 terabytes, you're going at a, at a nice, uh, moderate size and then you spread them in the servers, but maybe you don't use them because there are Lego blocks and this server is now a CPU server. But once it's installed, because the installation cost might be your prohibitive cost, you can then down the road change roles for that server and the drive is already in there.
So that is just one example. Um, we also, of course try to optimize by holding drives in our kind of front end operation center, and then we ship them out when we need them and have remote hands install them. So there are multiple options to be had here if you have the automation stack to manage it.
And that is a big part. We wrote our own automation stack in terms of data center automation, not in terms of, uh, uh, compute automation. We use Kubernetes for that.
So we have the, uh, agility to provide storage where we need it, when we need it. And again, the beautiful thing with, with soddy is the connection to the OS and the tooling that is provided, and we can manage the drives remotely through the OS providing us with all the serial numbers and asset management information that we need to do this in a responsible way. It's interesting that you bring that up because you know, now that you mention it, um, storage is one of the only data center components that's actually upgradeable and replaceable in the field easily.
I mean, certainly you could take the server apart and swap out the memory, but then what are you gonna do with the old memory? You know, same with the CPUs, same with the network cards. I mean, I guess network cards can be replaced, but storage is by far the easiest thing to replace and upgrade because in many cases you can even hot swap.
I don't know if you would want to do that, but you could probably do that in a server. And, um, you know, and, and, and as you said, that would let you upgrade things on the fly and match the capability to the, to the requirement. There's another aspect too, and that I want to emphasize here on this because I mean, I'm a storage nerd.
Um, one of the coolest things about storage too is that capacity influences longevity to a great extent. And that if you have sufficient capacity, especially with flash, if you have sufficient capacity, then that can greatly extend the lifespan of that drive because of wear leveling. Essentially, the drives are going to, um, you know, they can only handle so many rights, but if the drive is big, then so many rights spreads out across a lot more cells, which means that that drives lifespan can be a lot longer.
And I think that we've seen that in production. Um, as you mentioned, once drives get into the terabyte, multi terabyte range, I mean, certainly I imagine a 60 terabyte drive is gonna have a very long lifespan, but once drives get into the, you know, seven terabyte or something like that, that's an awful lot of data in order to quote, wear that drive out. And most people are never gonna hit it quite that hard because of course reads don't affect it.
It's only right. Um, are you seeing this as well that, uh, that bigger drives are, are more reliable? So amazing, amazing question.
I have multiple things to say. Nonpolitical. I'll start with taking a leaf out of, uh, yes we can.
Hot swappable. Yes we can. Yes, we do.
It is amazing to see when you control the drives and you have good reliable drives and you're buying drives with the level of endurance that you need. As you mentioned, it doesn't have to be crazy levels of endurance because you're at the multi terabyte size and you are at a, even at, at just one, uh, at like endurance one and not, not crazy high endurance levels. You can write a lot of data and if you have the software and you have the automation in place, you can monitor the media state and you can monitor where that drive is.
And we saw over the years that we've been using the drives that we are not wearing them out, and we should not be afraid. We can really use drives for a long time. They are sustainable.
They are not prone to failures and doing good. It also means that you have clustering and you have n plus kind of, I'd say architectures. It's all today.
It's almost no-brainers. You don't have to pay, uh, like exorbitant amounts of, of cost to different fancy licenses or storage providers. You can do your own storage and software defined storage today has built in a redundancy that is really good enough and provides it's, it's excellent, it's beyond good enough and provides you with a solutions that are cheap, effective and provide a ton of performance for the business.
So let's move up the stack a little bit from storage. So, uh, clearly, uh, storage is something that we are all passionate about, but also something that gives a lot of flexibility. But talk a little bit more about the rest of the infrastructure that you're building there on-prem.
Uh, you know, what do these servers look like? What does the G-P-U-C-P-U network and all that? So when you're looking at, uh, GPU servers, just as an example, you, you spoke earlier about these big servers.
So GPU servers, some of them are these nice slim two 4G PU servers, which are, I'd say specialized for data crunching. You can talk about them in a moment. Um, spark specifically, we'll talk about them in a moment.
And then you have these huge humongous six eight U servers that host eight 10 really large Nvidia drive NVIDIA drives really that that host, uh, eight 10 GPUs. And these GPUs are data hungry. So if you look at the two of GPU servers, let's start with the really big ones.
When you're looking at a huge GPU server hosting multiple GPUs, it really needs to pull in a lot of data to fill in all of that data. It's gonna come with the kind of huge network cards, the 100 gig network cards, and multiple ports of that going into the top of the rack, uh, switch. And it's gonna be pulling a whole lot of data coming in from your HDFS or other storage solutions, uh, sifts or whatever you are using.
And then you would need some scratch space locally to be able to cache it locally, because bringing it off the network will get you to one level of performance. But if you really want to harness those GPUs to their fullest and use all of that capital that you spent on GPUs and making sure that those GPUs are fully utilized a hundred percent of the time, then you would like to see a, I'd say local cash fully flash with a, with a lot of iops that can serve the, those, um, NVME, like NVME to NV link or NVME to PCIE, depending on what type of NVIDIA you're using and, and really feed those GPUs and then bring that data back and feed it out. When we're talking about our data processing GPUs, um, we've, as Tabula, we've spoken about this, uh, quite vocally, you can find it on our engineering blog.
We have switched from using in many cases, CPUs to GPUs, the smaller GPUs actually, um, let's say the A 30 and, and the 40 family to crunch data in spark format. So when you're looking at crunching data at large quantities, uh, instead of having 10 CPU U servers running separately, having maybe much less scratch space and doing each one their own thing, you suddenly have one server doing the job of 10 modern servers pulling in a whole lot of data from, uh, data storage, crunching it and preparing it for some type of report or some other type of, um, I'd say spark slash sql, uh, processing. And that we do on GPUs as well, having only 25 gig connections for those servers.
But again, having NVME locally so that while the GPUs are active, you can pre-read as much as possible from the central storage and then provide space that the GPUs can work with locally to optimize their utilization. And that reminds me of one of the questions we, or one of the discussions we had, uh, earlier this season. Um, talking about keeping these hungry GPUs fed.
Uh, that's really the whole ball game. 'cause these are the, that's the expensive, uh, or the most expensive component, uh, out there. And of course, as you mentioned as well, uh, GPUs can be used in other areas of infrastructure.
We've talked about that this season. Uh, you know, it's interesting here, uh, Ariel, we, uh, everything you've said, I think really does kind of summarize what we've talked about all season long here on utilizing tech. And it's been so great to hear a practical, real world use case for all the things we're talking about, especially an on-prem use case.
Because, you know, I think that increasingly people are gonna be looking to, that. They're looking to deploying their own AI data infrastructure on-prem. And, uh, the lessons that you've brought here, I think are very valuable to them.
So given that, uh, thank you for joining us. Um, where can people connect, connect with you? Where can they continue this conversation?
Where can they learn more about the lessons that you've learned? You can find me on LinkedIn. I'm always happy to continue this conversation.
And there's the Tabula engineering blog where we provide additional information about our infrastructure and what we do. It's much more down to the technical details. So if you want to know about running, let's say volcano on GPUs and MIGS on GPUs, so virtualizing our GPUs to really optimize and squeeze more out of them.
Or if you want to read more about how we're doing, uh, optimization around MySQL and SQL servers in general with storage, with, uh, CPUs, it's all there. And we are always happy to share. Well, thank you so much for that, and I, uh, really appreciate your, uh, everything you've had to say here today.
Um, thanks for joining us. Uh, Janice. Uh, I'm sad to say that we're wrapping up the season with this episode.
Uh, let me get your call to action. Where can people learn more about soy? Where can they continue the conversation of how, uh, flash storage can be transformational for their IT infrastructure?
Thank you, Steven. It's been such a pleasure being able to sponsor this series, and we're excited to continue, uh, throughout the year, um, sponsoring other activities that you're going to be doing. You know, things like, uh, AI Field day and storage Field day and, and many other things you do.
Just a great job, um, tapping into to our audience. com/ai where we have, uh, a multitude of, um, stories, uh, solutions, and also details around, um, the myriad of, uh, drives that we have to offer. So thank you so much for the opportunity.
Excellent. And actually I'll call out too, uh, soy presented at our AI Field Day, uh, brought in partners. Um, just a tremendous, uh, thing.
Use your favorite search engine, look for field day, look for soy. You'll see great presentations, a lot more information about that as well. And of course, uh, we've had many other kinds of companies present at Field Day over the years too.
And so you can, uh, look for those too. Thank you so much for listening to this episode of Utilizing Tech. Uh, you can find this podcast in your favorite podcast application as well as on YouTube.
If you enjoyed this discussion, please do leave us a rating, a nice review, maybe send us some feedback. We'd love to hear from you. This season was brought to you by Soy and of course by Tech Field Day, part of the Futurum Group.
com, where you can find the entire season. In order, uh, you can also, uh, contact us, uh, on x Twitter and Mastodon at utilizing tech. Thanks for listening to the season and, uh, thank you Ace and Janice for co-hosting as well.
We look forward to our AI data Infrastructure Field Day event, along with our AI Field Day event coming soon. com to learn more information about those, and maybe if you'd like to participate, reach out to me, Steven Foskett. I'd love to hear from you.
Thanks for listening, and we will catch you next season with another exciting episode of utilizing Tech.