AI Workloads: Data Performance to the Test with Aniket Khosla
Aniket Khosla, vice president of Wireline Product Management at Spirent Communications, is at the forefront of addressing the challenges posed by the rapid growth of AI workloads. Khosla highlights how AI is driving an unprecedented demand for low-latency, high-bandwidth connectivity between servers, storage, and GPUs. As data centers evolve to support the needs of cloud-based 5G networks and edge computing, Khosla emphasizes the necessity for intelligent transformation.
Transcript
This is Textron tv. Hey everyone, welcome back here to techron tv. It's great to have you on.
Um, we've got a new guest for, to introduce you to, uh, from a company that I, I was telling them off camera. I haven't had, uh, much updates for him in a while, so I'm really happy to have him on here. His name is Anit sla, and Anit is the Vice President of Wireline Product Management over Aspirant Communications, one of the leaders in testing and performance, performance testing, et cetera.
Um, but we'll let him tell us about what pyron s the leader are on, I guess at it. Welcome to Tech Drug tv. It's great to have you here.
Thanks for having me, Alan. I appreciate, I appreciate you taking the time to have this chat. Um, I appreci just, you know, yeah.
Before we get started, a little bit about, you know, my background. Um, I've been in the core networking test industry for over 20 years. Um, mm-Hmm.
Of those 20, I've been at Spiron now for about three and a half years. Uh, before that I worked for Ixia that eventually got acquired by Keysight. I used to run the product teams there as well.
So my background over the past 20 years has been purely in the networking space. Um, you know, more, uh, routing, switching layer two three, basically testing the infrastructure and the plumbing of whether it's a carrier network, whether it's a data center network. Um, so that's, that's been my background essentially for the last 20 years.
I moved over to Spire, like I mentioned about four years ago, getting close to four now, and in my responsibility I manage. Um, so, so Pyron as a whole builds products to test the network infrastructure and the applications that run on that infrastructure. So think of it as us building products to test, you know, the lowest layer one all the way up through layer seven, and even potentially the services offered on top of that.
So we test core data center networking technologies. We test 5G uh, related technologies. We test satellite positioning technologies.
Um, so essentially the entire gamut of tests, what I am specifically responsible for is the core, uh, network infrastructure test products. 6 terabit. So, you know, we build the hardware platforms that a lot of, you know, the, the biggest names in the industry, the big chip set vendors, nms.
They essentially use the product line that I manage to validate and test the performance of their routers, switches, applications running on top of that infrastructure. So that's essentially, you know, my purview in, uh, and my responsibility for the product line. Inspiring.
Absolutely. You know, there's a big product. Pyron does a lot of different things though, right?
And I don't wanna spend a lot of time on it, but for our audience out there, just so they know, you know, maybe when they can go or what they should, you know, the way the kind of things that they should go through, they'll, I even say, Hey, SPI might be able to help me with this. Can you give us a little bit of a broader Sure. So, so you could, the, the easiest answer there is your equipment that your network equipment manufacturer who's building a new router, is building a new switch, right?
With a bleeding edge technology. Like a hundred gig, uh, sorry, like, like 800 gig, let's just say, right? Mm-Hmm.
I think the first thing you wanna do is make sure that the equipment that you've built is up to spec, uh, with an inter with different vendors. It's giving you the performance that, you know, you expect it to give. And the best way to do that is through a third party like pirate, right?
We're an independent test vendor. So that's the, that's the most basic use case. If you are a big financial enterprise, as an example, who's building their own data center out, okay?
Um, you want to test and, you know, stress test the performance of your, um, data center, right? Are you getting the performance you want to get out of it? We also have security applications.
So if you are somebody who's deploying a new security service, deploying a new firewall, let's just say in your network, you want to make sure that the firewall, you know, is performing the way it needs to, it's blocking the right attacks, it's not letting malware through. So we can test the performance of the port networking pieces. We can test the performance of the, the security posture of the, uh, data center that's essentially been deployed, right?
So those are just some really simple examples of the things we can do in the 5G realm. We are used by all the top carriers in the world to test the performance of their 5G rollouts, right? So we can actually sit in the production network as well and say, look, you know, this is how your 5G service is doing, so on and so forth.
So, you know, think of us as being able to test, um, the core, um, networking infrastructure and the security applications that sit on top of it, and any other applications might that, that might sit on top of that. So, you know, if, if you are a customer, you know, for us who's either at enterprise, we can test how well your network's performing before you deploy. We can test the security efficacy of it.
And if you're somebody who's a switch vendor or a router vendor, chances are we already work with you. And, you know, we test the performance and scale of your routes, uh, of your, uh, routers and your switches. Excellent.
And, um, and really, I mean, it, it's, it, it is, it starts at the network and builds all the way up through the application layers and pyron tests, all it. And I've used pyron in the, you know, past lives at companies I've, I've helped, uh, found. And, uh, it's always been, you know, it was always a market leader.
Still is now. And I, if it's okay, I want to turn to kind of the topic of our discussion today. Mm-Hmm.
Which is a, you know, pretty familiar topic here on text, AI's changing everything, right? AI is, uh, I mean, we did our Textron gang recording this morning, and we're talking about technical rioting and, and just so many different you things you don't really think about, but I gotta imagine that it's bringing about or writing, you know, some big changes within the, uh, data performance testing. Yes.
Arena it Testing in general, but data performance, It, it absolutely is, right? I mean, it starts with, you know, added score. The AI data center is built differently than your traditional data center, right?
It's called what, what you call the AI backend data center. It's more, you know, GPUs. It's more either InfiniBand or ethernet.
Um, and it's, it's very differently built than your traditional data center. I mean, if you talk to the big hyperscalers today, most of them actually physically have different locations for their, what they call the traditional front end data center and their AI backend data center, because the, the way they're built and the expectations of how that data center perform are very, very different, right? In a traditional data center, you can throw more compute at a problem and, you know, through high performance computing, you can, you know, solve a task, if you will, in an AI data center.
Given the size of these large language models and things like that. Everything is distributed across a cluster of GPUs, right? So when it comes to being able to validate and test that the old way of, you know, pyron creating workloads is not good enough, right?
We have to think of ourselves now as a tester who can simulate and emulate these new traffic patterns that the GPUs in an AI cluster are generating, which are fundamentally very different than the traditional data center. So we've had to go back to the drawing board and rethink, you know, how we, how can we emulate these new traffic patterns with the solutions that we have. So, you know, the way we've gone about doing it is fundamentally very different than, you know, some things spiral has done for the last 20 years, which is, you know, just generate, you know, standard workloads.
So the AI workloads are very different. So that's, that's kind of how we looked at the problem. The AI data center is very different.
The way they operate and they're built is very different. Uh, so we also have to do something very different with our products so that we can test these new AI data centers. Yep.
So let's talk about what you're doing different. So, yeah, so, so I'll get into some of my technical detail, maybe not too much. So, um, the, the big question right now in a lot of AI backend data centers is ethernet or fin, right?
infin, clearly right now is the dominant technology. You know, it is a lossless architecture, low latency. Um, but ether the ethernet share in that AI backend data center is, is expected to grow.
I think all analysts will tell you over the next 3, 4, 5 years, the ethernet share will grow. Okay? Um, but ethernet has a bunch of inherent problems in it.
Uh, it's not lossless, it's gonna have packet loss, it's gonna be high latency. Um, and in an AI data center, packet loss and latency is disastrous because if you have packet loss and latency in an AI data center, the GPUs are not utilized to their full capacity, right? In, in a training environment, if you lose packets or you have high latency, the GPUs could sit there idling doing nothing, and, and think about it, right?
People are spending billions of dollars building out these AI data centers with GPUs. Um, and very often what you'll find is the bottleneck in packet loss and latency is the performance of the network itself. Okay?
So we looked at that problem statement and said, look, you know, you, you know, you Mr. Customer who's building out an AI data center, want to get the most bang for your buck by getting the most efficiency, uh, uh, out of your GPUs. So you need to test in the lab before you deploy to production.
So, you know, if you have these network bottlenecks, you can eliminate them. So what we set out to do with our products is give the hardware that we have an AI personality. So what we mean by that is we now take the hardware that we have, and each, you know, physical hardware port that we have can evaluate A GPU.
It creates traffic patterns based on what is widely known in the industry as a communications collective. So the, the way we transmit packets over the network is different. The way we report statistics is completely different, right?
People care about latency, but they care more about job completion times. They care, they, they care about tail latency. So we've had to rethink the way we generate these workloads from our hardware to mimic GPUs so people can test the performance of their fabric before they roll out, right?
The end goal for the customer is, I need to get the most efficiency outta my investment in these GPUs, right? So that's, that's where we come in. Uh, we've come in with our hardware to evaluate GPUs in the lab.
'cause very often what we found is customers want to test before they deploy, but it's Expensive. Get GPUs, right? Who could get GPUs?
That's the issue. I mean, Yeah. And even if they can, they're incredibly expensive not just to buy Alan, but in a test environment, the power consumption requirements for these GP just through the roof, right?
Yeah. So, so people, we found people were testing, but they were running into all kinds of issues when they deploy. I mean, I've spoken to a financial enterprise customer who told us that 90% of the time their GPUs are setting idle before.
Yeah. And right now, unfortunately, what most people are doing, instead of optimizing the network, they're throwing more GPUs at the problem. And that's only going to be, you know, sustainable and terrible for that long.
Yeah. Well, but, but as wasteful as that sounds, that's probably been the model. Throw more hardware at it, throw more, you know, processors at it.
And that's what we've done in tech for a long time. Yeah. I, I think it was, I I wanna say it was in the New York Times, but it might have been the post this week, an article about, you know, what they call that data center corridor in, in Northern Virginia.
Mm-Hmm. I forget how it's 50 plus data centers, like right there. But they were talking about, you know, the new AI data centers that are optimized for ai.
I mean, you, you, you really, you can't cool 'em down with air. They need to be water cooled. Right?
And the, and the networking, right? And that's a very connected area in northern Virginia, right? All the knocks, uh, uh, the kns and stuff were there, right?
But even with all that, the, the, the bandwidth requirements and then the sustainability of these data centers and everything that goes with it is tax. And that's an area that's probably more prepared for that than any Mm-Hmm. Other concentration that we know of in the world.
But it's really, it's taxing it, right? It, it's, it's kind of coming apart at the seams almost. So, um, this, this is a real world problem as we move to the, you know, quote unquote as you call it, the AI data center.
Yeah. I I think that same article also referenced that given the extreme consumption of powers by these AI data centers, residential prices are going up. Yeah.
Residential prices are going up and it's, you know, that's real world impact to, you know, someone like you and me. I remember. Absolutely.
Yeah. It was a great article. Um, and, and just so our audience knows, right, this, this is a problem that affects not just the hyperscalers, you know, the big cloud centers, public cloud providers, but it's, it's wherever you're running, you know, whether it's a private data center, a co-located data center, the these are, they're issues that are in here.
And you know, as you said, we're throwing more GPUs at it instead of saying, Hey, what can we intelligently do to make better use of, of the resources we have there in place? Um, let me ask you an off the wall question. You've see the, do you foresee spiron coming out with test equipment that actually has real GP news in it?
Or you get the same bang for the buck kind of just emulating it? Yeah, so I'll, I'll, I'll say this, right, this, the, the product launch that we just had a couple of months ago, I think we put it out in July. The GPU ation solution, that's our first step at first foray into this AI world, right?
Um, I think we're going to look at a myriad of products that is going, that are going to help test the AI data center. Whether it's pieces of software, whether it's new pieces of hardware. This is our first foray in, can I rule out a product that Pyron will build using a real GPU?
No, I can't. But I will tell you that the problem statement that people have that they're coming to us with more now is, is there a way for them to test the performance of their GPU? The thing is, right now the GPU market is pretty much owned by one vendor, right?
It's, it's essentially Nvidia at this stage. So not a lot of GPU performance testing happening, but as the market opens up, I, I do expect other players come in, right? And then, which GPU is better than which GPUI think will become For what?
For what job too, right? More Job two, right? Yeah, exactly.
So I mean, this is definitely our first foray in, uh, we are learning a lot through the engagements that we've had, problem statements that, you know, I don't think any test vendor has encountered in the past. And it's, it's, it's an exciting time to be in the networking industry right now because the pace of innovation is also unlike anything at least I've ever seen in my 20 year career piece of tremendous. I think AI's making an exciting time to be in tech in general.
It's developing. Anyway, Anique, we're about outta time. I want to thank you for coming on.
You know what, think ofpi, think of testing for AI data centers. Think about what AI data flows mean. I think that's a great message.
Carrying it forward, continued success. Keep us posted on this. 'cause this is obviously something we discuss a lot here.
It's extra I right about and cover and, uh, keep up the great work at sp. Thank you, Alan. I appreciate the time again.
Now you're happy to stay in touch and we'll, we'll chat. So again, thank you. All righty.
Etiquette co Kla V vp, wireline product management, Sping Communications, talking about AI data centers here on Techron tv. We're gonna take a break. We'll be back in a minute.