Cisco AI Networking Cluster Operations Deep Dive
Paresh Gupta’s deep dive on AI cluster operations focused on the extreme and unique challenges of high-performance backend networks. He explained that these networks, which primarily use RDMA over Converged Ethernet (ROCE), are exceptionally sensitive to both packet loss and network delay. Because ROCE is UDP-based, it lacks TCP’s native congestion control, meaning a single dropped packet can stall an entire collective communication operation, forcing a costly retransmission and wasting expensive GPU cycles. This problem is compounded by AI traffic patterns, such as checkpointing, where all GPUs write to storage simultaneously, creating massive incast congestion. Gupta emphasized that in these environments, where every nanosecond of delay matters, traditional network designs and operational practices are no longer sufficient.
Cisco’s strategy to solve these problems is built on prescriptive, end-to-end validated reference architectures, which are tested with NVIDIA, AMD, Intel Gaudi, and all major storage vendors. Gupta detailed the critical importance of a specific Rail-Optimized Design, a non-blocking topology engineered to ensure single-hop connectivity between all GPUs within a scalable unit. This design minimizes latency by keeping traffic off the spine switches, but its performance is entirely dependent on perfect physical cabling. He explained that these architectures are built on Cisco’s smart switches, which use Silicon One ASICs and are optimized with fine-tuned thresholds for congestion-notification protocols like ECN and PFC.
The most critical innovations, however, are in operational simplicity, delivered via Nexus Dashboard and HyperFabric AI. These platforms automate and hide the underlying network complexity. Gupta highlighted the automated cabling check feature. The system generates a precise cabling plan for the rail-optimized design and provides a task list to on-site technicians; the management UI will only show a port as green when it is cabled to the exact correct port, solving the pervasive and performance-crippling problem of miscabling. This feature, which customers reported reduced deployment time by 90%, is combined with job scheduler integration to detect and flag performance-degrading anomalies, such as a single job being inefficiently spread across multiple scalable units.
Presented by Paresh Gupta, Principal Technical Marketing Engineer. Recorded live at Networking Field Day 39 in Silicon Valley on November 6, 2025. Watch the entire presentation at https://techfieldday.com/appearance/cisco-presents-at-networking-field-day-39/ or visit https://techfieldday.com/event/nfd39/ or https://Cisco.com for more information.
Transcript
Okay, so what are some of the common challenges we get to hear from customers is like, number one, look, there's a lot of moving components here, so we want this whole to be validated, right? So reference designs, simplicity. There's not just one network involved, right?
There's enough com complexity just with storage, just with compute. And then there are networks, not just one many type of networks. So we want simplicity, then we wanna make it secure.
I think all of us keep on hearing news articles about ai, uh, issues here and there, uh, right? So we wanna keep it secure. And then there are specific challenges that I'll deep dive into in a few minutes about, you know, inter g backend communication.
Now, look, there are some basic challenges about network delay congestion, uh, things like that we have heard from forever. But what makes inter GPA network very unique is that all these GPUs send and receive traffic at the same time, which is unprecedented in the industry, right? So if you have 800 gig quotes, and they are eight of them on a single server, all of them send traffic at line rate at the same time.
And then if there are many of the servers, all of them are doing the same similar challenges of the frontend too, although, you know, storage traffic is there, right? So storage causes a lot of InCast. So if you're reading data, let's say you have 32 servers, they want to checkpoint some data to the storage.
Now, checkpoint is a state during training where, you know, all the GP nodes say that this much I've learned, so let me store it, right? That's what Checkpointing is. When it's writing Checkpointing or doing checkpointing, all these GPU nodes are writing to the same storage appliances at the same time.
So when you run training, your backend traffic literally drops and your storage traffic spikes. And if there's congestion in the frontend network, that CHECKPOINTING will take a lot of time. And because of that, your training gets delayed.
That's again, uh, wasting time of GPU cycles or the money put on those GPUs. The other challenge is on front end network, all this traffic must coexist, right? There's not, not only, uh, storage traffic there, remember there's all kind of traffic flows there.
So all this traffic must go exist and honor each other. So I've got a question for you in, in looking at this and looking at, you said the, these are the challenges you see from customers. Normally on this, if you have large traffic bursts and you have, you know, packet loss or any kind of network delay as you're doing this, um, you know, it, it seems counterintuitive to converge on the same network, or is it just because of the increase in capacity in going to that next generation architecture that, that solves these problems?
Is that right? Yes. So you, you are right.
So for those customers who have those massive traffic flows, even on front end, they stay with the isolated networks, right? Okay. So there are some that are that, that have these challenges that are gonna go with that architecture.
Yes. Okay. Well, for anything I'm saying, you will find an exception out there.
Sure. Yeah, that's fair. That's fair, Right?
Yeah. You know, you will find some of the customers doing something different because, and the other thing is, I talked about hyperscalers. One of the things that I'm seeing about hyperscalers is that all of them are highly opinionated for the right reasons, and all of them do things differently because, you know, massive r and d budgets.
So as per their needs, as per their models, they build infrastructure from scratch, right? So that's the reason people do different things, but this is what I'm, uh, trying to convey the most common view what I'm seeing. Right, right.
Maybe Arun would disagree with me on some points, but this is me talking. No, Gotcha. No, that's cool.
I was just curious because having talked about converging them and then looking at common problems like, you know, packet loss, congestion, burst traffic, things like that, that would be, are you, I mean, is that just basic quality of service, or how do you solve those problems? We, we'll talk about that. So before next, okay, before solution, I'll explain you.
The problem was, the problem is, you know, problem short answer is, again, it has to do with the load balancing problems and e ethernet network. So there is network capacity, but it's unused. So before adding more capacity to the network, I wanna make sure I'm using all the capacity.
Got it. Right. So how do you add, uh, or make use of the network to the full, full of its potential?
That's where we have some features. Okay? So that's short answer, but I'll, I'll explain the detail.
But before that, let's go step by step and let's talk about the reference architecture. And of course, NVIDIA is the largest name in the field of ai. It's the largest company in the human history.
And in this photo, you see a lot of my peers, a lot of them are famous except one who, who's a little bit more famous. Uh, This photo is coming from last week. This is where, um, the NVIDIAs conference, GTC DC took place in Washington, DC Uh, Before this GTC that happened earlier this year in March.
There's another GTC right here in San Jose. We announced partnership with Nvidia and we promised a couple of points to our customers. One of the points was offering joint reference architectures with Nvidia.
So we offered the enterprise reference architecture that applies up to a thousand GPUs that was published in March, uh, this year. Now last week we announced two more architectures. The first one is the cloud reference architecture.
This is based on NVIDIA's own cloud reference architecture. But this is, this is Cisco's reference architecture. It applies to, you know, all the way to 32,000 GPUs is what we have published, but we can also have larger designs.
The third announcement is we also took NVIDIA's Spectrum silicon and productize a Cisco switch using that. And then we plan to run Cisco's own, own operating system on it, NXOS and Sonic. So the photo that you see on right Jensen signing that box, that's the switch.
And I'm kind of have a mixed feelings because it's awesome. You see in right next to Jensen, we have Will Etherton, who is Cisco's, uh, who leads Cisco engineering and then uc for us. He's my awesome teammate.
I was thinking to get that switch in my lab and use it, but now after the sign, I cannot use that anymore because the switch is going to museum, right? But the point that I'm trying to convey, say that again, send on Vacation and keep the switch. Okay.
I don't think even I'll, they'll let me even touch it. So anyway, with all these reference architectures, the point that I'm trying to convey is that if you think about why the basics foundation, I think yesterday somebody was asking why did we do it? Because look, Cisco is already massively deployed already.
All the force of network engineers, they find the operational simplicity of using the same operating system and the operating model out there, right? So regardless of whether it's a hyperscaler or enterprise, we have a solution. Now, of course, Nvidia is the largest name, but not just Nvidia.
We are also working with others. Like you see all the full list on the slide and there are some photos. In fact, Cisco is building one AI cluster every quarter, right?
So I have some photos of that. All this information you'll find in some social media feed or some information. But we recently validated A MD GPUs last year.
You see photo on the left. I personally validated Gaudi two with Cisco network, recently completed Intel Gaudi three validation as well. We have validated all the storage vendors.
Um, DDM recently completed. Vast Waca was done before. We are going to the next level.
And most importantly, we have also an in-house dedicated benchmarking team. And that benchmarking team has been a massive help in, in making these customers feel confident about the entire stack. Because it's not just about how much traffic can, you can send on a link.
It's all about can you run a distributor job? Can you run a training? And can you achieve that performance really with that entire, uh, you know, slideware that you're showing.
So yes, that team is delivering that. We have submitted results on ML Commons. We also have customers coming in-house and running their models and feeling confident before going live.
Next point about operational simplicity. Now we do realize there are customers of different demands out there. So first model for operations is an on premises running network management platform.
That's Nexus dashboard. So essentially the idea is a customer can bring, uh, any type of server, they can bring any storage, use Cisco network and manage that using Nexus dashboard. I'll talk about some examples of what features we have built within Nexus dashboard that simplifies the operations of AI clusters, right?
But then some customers also appreciate the simplicity of a cloud controller, right? They don't have to install anything on-prem, it's running on cloud. com and the portal is there, right?
There's a variant of hyper fabric called hypo fabric ai. That's what I have mentioned here. Hypo fabric AI is kind of a prescriptive model.
This is really designed for people and there are a lot of enterprises out there. They say like, you know what, uh, we have not done this before. Uh, can you help us out?
And our answer is, look, you're deploying for the first time we have validated this architecture end-to-end, use the servers, use the network and use the storage. And we offer them into multiple t-shirt sizes, like small, medium, large, based on the enterprise reference architecture that we wrote before, right? So all the variables go away.
And when people decide second cluster or third cluster or fourth cluster they can do after that, they are capable of making their own decisions. I have a question, Rita Younger, um, the Nexus hyper fabric will now work with the Nexus nine K or will be able to work with the nine K. Will hyper fabric AI also be compatible with the nine K or is it gonna require the 6,000?
Yeah, That's happening soon. Okay. So it is the nine K also.
Yes. Perfect. Yes, You already know the future.
Okay, talking about Nexus dashboard, uh, again, look talking about each of these platforms, it's like a multi art conversation. I just wanna give you one example that when people deploy this, it's not just one network, right? So in this example, look, if you're talking about a b inter bg, uh, inter GPU network, that maybe it's a routed network, the front end network might be a VXLAN network, then you have management network, you have other kind of, uh, services running, guess what?
In Nexus dashboard, we integrate all these different kind of, uh, networks and we allow provisioning all these different kind of networks and also provide a consistent API platform. We are also bringing the end-to-end visibility all the way into GPUs and Nicks, even the distributed jobs that are running on the GPUs and then correlate their network traffic pattern. Uh, just to give you an idea, this is a zoomed out view, but you can take a look on all the list of all different kind of networks that you can provision, right?
So essentially all different kind of networks that I showed you, single management platform delivers the full simplicity of the operational model. So we talked about two challenges. The third one is security, right?
And again, security is a massive topic by itself, but this is what we are doing. We realized what are the entry points and exit points to an AI cluster. So what we did, we designed smart switches in Cisco.
These smart switches essentially use Cisco Silicon one ASIC for line rate non-blocking switching performance. But we also put a MD DPU on these switches. These dpu are able to perform a lot of servicing typically, which is done by some dedicated platforms by itself, right?
So the first use case that we are building or the the use case that we are planning to build for AI environments is provide a layer for zone based segmentation using the architecture of hyper shield. So hyper shield designs the policies, but the enforcement happens on the switches. And by the way, because this is network field day, I'm just showing you a network view of it, the hyper shield architecture also has enforcement happening at the workload level, right?
So consistent policy plane, but the enforcement might be different. Alright, let's focus on the network piece more, right? We are going down or deep one step at a time.
Now let's focus on frontend and backend network. And if you think about what exactly is the type of traffic going, so now we become network geeks and start looking at protocols. Look, most of the user traffic will be TCP based, but a lot of storage io to your question, uses Rocky.
Like Rocky, um, it's RDMA or converge ethernet, uh, people. If you put Rocky in ai, it says Rossi, check it out. Okay?
As is, you know, you know, say it, it say it will say Rawi. It didn't realize, you know, uh, somebody really named it really nicely. Uh, even GPUs talk to each other using Rocky y because when Rocky using Rocky implementations, the whole kernel is avoided.
The CPU path is avoided and memory to memory data transfer happen without, with minimal delay, right? This is a topic by itself, but the point that I'm trying to convey, again, because this is network field A, is that if you look at the network RDMA or converge ethernet version two is really U-D-P-I-P traffic over ethernet, right? Which means it has UDP inside it, right?
So it goes back to the classic differences between TCP and UDP. That TCP has a built in congestion control, uh, protocols, right? That uses either packet drops as a signaling or uses RTT as well.
Uh, guess what? When Rocky was done the very first time, it did not have any, but, well, you guys have been around for a while. You, you, you already seen that.
We take one part of the, you know, stack and then we learn then the other part of the stack and then we transfer the same information. So essentially over time, Rocky came up with similar congestion control algorithms. One is the very famous name called D-C-Q-C-N.
Essentially it means the network detects congestion, inform the end devices and the end device adjust the rate quite similar to what TCP does, right? And then recently people realized, you know what this D-C-Q-C-N is working well for storage traffic, but because these GPUs are able to send and receive at the same time, by the time the notification goes to the end device and end device does something, guess what? Congestion is gone.
So we want something more proactive. So the end device is now, because this whole, uh, RDM operations are mostly done by the Nicks, them, the super nick or smart nicks or AI nick are different names for them. These nicks are able to calculate the round trip time of every connection.
And if there's a slow down, they throttle the way. So I've got a question for you. This is, um, Kevin Meyers the, um, so Rocky as I understand it as you first explained it.
So that was memory to memory. So we're talking RAM a node that's working and has something in working memory is able to transfer that to another node via the Rocky protocol. That's one of its use cases.
Yes. And then it looks like according to your architecture diagram, that's also also to storage also to some kind of long-term storage. It's used for both of those.
And that that is a, and that's a storage protocol that would replace contemporary storage protocols for this. Um, you know, as far as the, the storage protocols that we have been using, I mean, I'm not a storage guy, I'm just curious. So let's say NFS is one like net Yeah, Like NFSI think you have like what NVME over IP now?
Yeah. Things like that. Yeah.
So NFS you can run our TCP. Okay. And you wanna make NFSA bit more efficient.
You run it over Rocky. Got it. Okay.
So almost every storage protocol, and I have a massive diagram written on it, like, uh, uh, in my book, which I published last year on storage network congestion, the point is that every network, uh, or storage protocol that you can think about, there's A-R-D-M-A implementation to it. Okay? And NFS or RDM is one classic example.
NVME or there's A-N-V-M-E or TCP, there's NVME or uh, Rocky. Okay, there's NVME or IB two. There's all kind of variants out there.
And What, what runs the, what runs the IP layer? The IP four layer, the IP V six layer. Is that EVPN vxlan?
Is that something else? What does that architecture look like? Similar.
Okay. Similar. I, I had the previous slide, if you remember.
I had like different BGP routed on backend Front. Okay. I saw I think EV VP MDX line on the store, on the front end network, but I didn't see it on this one.
I must have missed it backend. So you're gonna run Yeah, a similar network, but that's a separate, that's considered like a separate network On the backend For this part? Yeah, for the backend network.
Yeah. Well, the reason, uh, look, protocol is not the reason to separate the networks. The reason is network capacity.
Okay. On the backend network. If GPUs are consuming the entire capacity of the network, then you probably cannot put anything else on it.
Sure. Right? Okay.
Can you merge it? Yes. I mean, there are some smaller design.
In fact, if you look at the Cisco reference architecture, the Cisco enterprise reference architecture for smaller deployments, we just merge all these networks. Got it. They like, there's not enough requirement.
'cause there isn't a reason to to run them separately. Yes. Okay.
Now, TCP also has the built-in, uh, packet loss re transmission. Rocky doesn't do that because it's UDP. And because of that, packet loss leads to significant performance degradation for RDM and traffic.
Now, why packet loss happens two major reasons, big errors, right? CRC errors happen and the packet will drop. The other major reason is network congestion.
But why exactly performance degrades so much? And I have a, you know, kind of a geeky example here. If you think about when GPUs talk to each other, there's a concept of collective communication.
Essentially all these GPUs are talking to each other at the same time, right? They are actually running the same job. And when they are running the same job, and let's say one GPU does some calculation, the other GPU does some calculation, they must sync the states.
Otherwise, you know, these are split brain scenarios, right? So that's why there's a collective communication going on among all the GPUs. And to convey the point, I have a very simple example.
Imagine if all these GPUs want to transfer four megabytes of, of, uh, uh, data to each other. Now, Rocky, MTU, typically we take packet is 4K. So if you divide that leads to roughly a thousand packets.
Now imagine in this example, all these four GPUs are transferring a thousand packets to each other, right? The way RDMA operation happens, it's typically RDMA right operation where one end just writes all the data to the memory of the other device. That's a remote direct memory access.
But on the network layer, it leads to a thousand packets. And this is, these are, this is how packets look that there would be first packet, there would be middle packets, and then there would be last packet. Each packet carries a packet sequence number.
Now, a couple of years back, what used to happen is that if one packet gets dropped, guess what? The entire IO happens again. The entire operation happens.
Again, people realize is not very efficient. They ca came up with another approach, say that go back to N and they said that let's say 1500 packet gets dropped, then you know what, let's send everything after 1500. Very similar to what you know, you can argue TCP does too.
Other protocols do too, right? Regardless of what you do, this induces delay, right? Even if the packet is remitted, there will be delay.
And what happens, most of the operations complete fast, but very small percentage of operations get delayed. But remember, it's a collective communication. All the GPUs are talking together.
So unless all of them are finished talking, none of them are finished talking. And because of that, the entire collective communication may time out or may significantly get delayed. So this is the basic reason of why, you know, when people say, why can't I run with a normal ethernet because of this reason?
And we have run benchmarks, you'll find multiple. If you do a web search, you'll find a lot of, uh, research. Mia, even industries have done this benchmarking in their lives.
No, it's not delivering performance. Although there are some exceptions if you catch me later saying that, right? People have done their own things nicely.
Okay? Now, one of the ways we can avoid packet loss is using priority based flow control. This has its route in convergence.
FCOE. Yeah, that Was gonna be my, the PFC is gonna be my my question. Now look, uh, PFC configuration takes some effort.
It must be consistent across the network. And also when you bring PFC with network congestion, there's a whole concept of ECN, which is notifying the end devices. And remember a few minutes earlier I said that the network must inform the end devices in a fine, in a, you know, timely fashion, otherwise it's already delayed.
So there are some ECN thresholds required. These thresholds we have learned by building all these GPU benchmarking environments and inside Cisco, and we ship these fine tuned thresholds within the built-in templates of Nexus dashboard as well as hyper fabric. In fact, on these first screens, you don't even see a way to enable PFC or ECN Because It's hidden for sake of simplicity.
Of course, if you are a power user, you will like to edit it, your field group, you know, edit your templates and like push that values consistently across all the switches in the network. Similar concept would apply to network delays too. Right?
Now, let's say imagine we don't drop any packets, but if there's some kind of delay, the similar performance degradation will be seen with, with, um, the collective communication or RDMA traffic. If you go back to the basics, what are the reasons of network delay? I mean, you know, if you go back to basics, it takes a while for a packet to transfer on the link, depending on the speed of the link, how fast you can really serialize bits of a packet, Right?
Honestly, Uh, When the network is up, there's nothing much we can do about it, Right? It It's governed by the physics Laws, right? But can we redesign our networks in a better way, right?
The only variable factor that you see here is the queuing delay, right? Because let's say if there's a packet that's trying to go out of a switch port, and if it's stuck behind a hundred of the packets, Right? That would be a massive congestion problem.
I mean, the packet has not been dropped yet. It's, it's seeing significant amount of congestion. But except the queing delay, there's nothing much we can do except few scenarios, which is we can redesign a network, something called a rail optimized design.
We can have our networks with no over subscription, right? And then we can also bring some simplicity. So let me give you some specific examples.
And this is where we are talking about the backend GPU to GPU network design to be precise Now, right? You're probably seeing that we are going one level deep each every time. Now, when we connect these GPU nodes, we connect them in a very specific way.
This is called a rail optimized design. There are are two major attributes to it. Now, before I talk about rail optimized design, let me talk about the little bit easier concept first, which is these networks are typically, again, typically designed in a one is to one over subscription or no subscription or a non-blocking fashion.
So if I talk about this spec specific example, every nick, every GPU can send and receive traffic as eight at 800 gig. And then the switches that we are using in this example have 64 ports of 800 gig. So if I want to have a non-blocking network or a 64 port switch, I take 32 ports and connect them to the host, and then I take rest of the 32 ports and connect them to those points, right?
No over subscription in the network. And because there are 32 connections per per switch, per for every host, I can have 2 56 connection going down every server. If it has eight GPUs, that means 32 servers or 2 56 GPUs in a kind of like a, you know, a scalable unit.
That's a term, you know, people typically use in a scalable unit with eight leaf switches, four spine switches, right? So this network design allows all these GPUs to talk to each other at line weight without any network bottleneck. There's a caveat to it, I'll come to that, but question Oh, okay.
Hi Denise Donna here. Um, okay, but then the serialization, um, to serialize the, the actual, you know, it's on the, on the wire into out the port. Um, did these do line rate serialization?
Is there no, no, no delay, no, you know, the whole, the old whole old problem of a big, a big packet, a small packet stuck, you know, But that problem's still there. I mean, I I just said a few minutes earlier, right? Like let's say if a packet stuck behind, yeah, a hundred of the packets, But this is a still not able to, we're still not in this world able to address that, right?
Oh, Well look, there are some technologies, uh, I think last time I was here I talked about that, where we have the capability to put the traffic in a prq, right? I mean, if you go back a few slides earlier, I had this RDMA operations showing a hundred packets, right? Thousand packets, first packet, last packet.
Now imagine all the packets go through through a switchboard, but the last packet stuck is the operation complete until all the packets are received? The operation is not complete. So what we allow in our switches is like allow putting those last packets in a different cube that at least finish the open operations before, you know, taking more operations, okay?
Right? So to an earlier point, yes, all these systems are able to send and receive at line weight, right? But there will be congestion scenarios.
I'll cover that. Like, you know, you can ask that if we dedicate NF capacity to, to the network, where, where's the scope for congestion, right? I think last week we were talking about what 600 access points, like we were having some, all of them open the laptop and like, uh, access point, like getting into bottleneck or this is kind of analogous to that, right?
When all these GPUs talk to each other at the same time. So we cannot have an oversubscribed network, but to reduce the serialization delay and the, and the transmission delay and the propagation delay, what we do, we connect these systems in a very specific way called as a rail optimize design. Essentially take first GPU of every node, connect to the same leaf switch and take second GPU of every server, connect to the same leaf switch and so on.
The benefit of that is the way these GPUs talk to each other, all of them now benefit with the single pop forwarding. And in this specific example, traffic does not go to the pines. So if you follow that, uh, bowl line, uh, between the first GPO, both the servers, I'm trying to signify or convey the traffic flow.
And if the traffic does not go to spine, it helps a little bit with the serialization delay or, or the net overall delay from the network perspective, right? So this is a benefit, but every benefit, you know, in real life you solve one problem, you run into another problem, right? So we solve this problem and see that.
Now it leads to a lot of operational challenges. And to understand these operation challenges, think about a typical data center. The way we deploy, we have rack, and within the rack we have maybe 24 servers, 30 servers.
We put a top of the rack switch. We use cable, copper cable, two meter, three meter, the shortest possible, the most affordable one. If you want to connect that cable to the next rack, it's not even possible.
The cable is not long enough. So really cabling goes wrong. Right?
Now, think about the AI deployment, a AI factory deployment. Imagine it's a massively powered rack, maybe 60 kilowatt plus rack every server. By the way, this is like 16 kilowatt server, if I remember correctly.
You put four server in each rack. This is how scalable, scalable unit is built. And then at the end of scalable unit, there would be dedicated rack for switches.
And this rack, all it has switches. Now imagine somebody has to pluck 2 56 cables. In some environments it might be double the number five and 12 cables, and then, you know, somebody's connecting them for three hours ready for a coffee break.
You know what happens? Like instead of the top switch, one cable goes to the bottom switch. I've run into this condition many times, by the way.
Yeah. Yes. That's why I'm, I'm so well connected with this problem.
I thought it's problem with me. But then there are other people who ran the same problem. As I said, Cisco, I is building one cluster every quarter and I see people running there too.
It's just a human thing, right? So how can we solve this problem? And by the way, one point that I wanna convey is that what happens if the cable is connected to the wrong, wrong switch instead of leaf one?
If the cable is connected to leaf two, what happens? I mean the, the LED still comes up, the link still comes up and I, I guess most of you have spent more time in the industry than I have. And you tell me if the connection is totally down, if the network is down, it's relatively easy to solve, then it's up, but it's crippling, it's it's sick, not dead.
And this condition miscalling runs into a massive problem. And imagine it's such an embarrassing answer to the management saying that, you know what? We had bad performance because of miscalling.
Uh, nobody likes saying that, right? Massive embarrassment. So the way we can avoid is like, this is one example that I'm bringing forward from hyper fabric ai.
What we do in hyper fabric for ai, because these are prescriptive designs based on our enterprise reference architecture, we offer a feature called run cabling, and then it generates the entire cabling map, right? So on the left you see it towards the bottom there's an entire cabling map based on the rail optimized design and not just rail optimized design. We also go up to a level where we recommend that if you connect this connection to this specific port on the switch, your performance might be slightly better because of the internal architecture of our asics.
So that's one, that's the first help we provide before, uh, you know, started cabling. But then if you look on the right, that's the task assignment, that's the onsite task assignment to the people who are cabling it. And once they connect, the cable light does not turn on green here, the light will turn green.
Let's say if I connect one cable to any other port, and if both the ports are okay, the light the LED will turn on, on the devices here on the ui, the light does not turn green until it's correctly connected as per that cabling plan, as per the rail optimized connection plan, right? So if somebody is physically looking at the device and say the light is green, nope, until it's green here, not, okay. Right?
So before we help, and by the way, as I said, uh, we have beta customers Cisco it using that in their own words, they reduce their deployment time by 90% using this feature. So this is a two person job then, or where? Well, if you connect so many cables by one person, depends on how many cables you can connect in a day.
But the cabling plan is all automated. There's no, you need to, uh, there's a run cabling button. I don't know if I can zoom into, But the, but the person on the floor doesn't see this.
There's somebody else they're in communication with during the, during the process. So We offer multiple options. Uh, so the onsite task can be assigned, if you look at the screen right now, it say assign the task.
So you have users delegate the task to users. That's one way, right? Some people just like use using or Excel way of doing things.
So we allow exporting this in Excel format and you take a look and when you are done, you need to go to the ui. Let's say if I'm cabling, I will go, I'll connect one cable said, done. And when I say done, uh, it must be green until the US says it's green.
The job is not done. Okay? So they can do the, the whole job and then come and check.
Yes. After, After the network is going be checked, type fabric AI doing this checks that actually as per rail optimized design, the cable has been connected to the right leaf switch. Okay?
Right? You mentioned, uh, Ethan, uh, you mentioned that depending on what port you plug into, you might gain a little bit of performance because of what front panel port is mapped to what ASIC on the inside, are we talking nanoseconds or microseconds? Nano That matters?
Uh, Yes. Mm-hmm. Everything matters in AI clusters.
Uh, And it matters in high frequency trading for sure. It matters there too. Yeah.
Uh, it matters up to an extent in storage environments too. Wherever we need performance ev everything matters, right? Because remember, um, that's why I had that diagram, right?
Like four GPUs, slight delay at one place, everything is getting delayed so it adds up, right? There's not even single operation adding up. So you keep on adding every delay point.
So nanoseconds will add it to micro, micro will show up somewhere. Yeah. And you know, I'm talking about the beach, the the GPU benchmarking environments.
We have customers come there and say that no, this is the completion time of the job and until you really match it to millisecond to millisecond level, it's, we are not gonna accept it because it's real money in millions sitting on the table. And if you delay the job, you know, million millions getting lost, like high frequency trading, you know, we all understand if you are a trader and if the stock price gets delayed by even one second on your laptop, you lose, lose a lot of of significant dollars there. So very similar concepts here too.
Okay, I'll take the next example from Nexus dashboard. Now, what I did at had the example of a 2 66 GPU cluster. Now let's say I wanna have a larger cluster of five and 12 GPUs, right?
And of course people have a lot larger GPU cluster. I'm just taking example to convey the point. And let's say there's a user who wanna run a distributed job that requires 16 GPUs, eight GPUs per server.
So the question is, how would you allocate your GPU servers to the user? First way is you take the first scalable unit and then you take second scalable unit. You just randomly assign any GPUs to the user because of that traffic.
When these GPUs talk to each other or traffic goes to the spines, we wanna eliminate that or at least reduce that to an extent. So what we prefer is that run your jobs within a scalable unit. And of course there's a job scheduler out there, very, very famous one called slum.
Slum has this feature called Slum Block Scheduler. It allows doing that, but then we have seen some misconfigurations there. So what we do in next is dashboard.
Because of the slum integration, because of the visibility, Nicks and GPUs, we are able to detect based on the network traffic path. So on the left you see the whole topology of a 2000 GPU cluster. And on the right there's a screenshot that shows, um, that this, the GPUs of the job are spread on two scalable units, which is an anomaly, right?
Nexus dashboard flags it, people should correct that. And let's say after a careful investigation, if I lies, oh no, this is exactly what I wanted, fine, acknowledge it, let it go. But Nexus Dashboard detects this anomaly and tells the user essentially goes back to the basic point of reducing the network delay by looking at the network traffic pattern and job integration with the Nexus dashboard.
I have a question that came from, comes from Mike Whitty who's watching remote. He wants to know if there's any plans for adding the cabling check functionality from Hyper Fabric to Nexus dashboard. We so look, I'm just giving examples.
You don't know. These features probably already exist. So You could ask the same question, is this feature available in hyper fabric?
Yeah. 2, which is coming out in a few months. Okay?
Right? So yes, that's the answer. All right.
Just examples. Of course, like covering all the features of all platforms is probably dedicated session, right?