Cisco AI Cluster Design and Operational Strategies
Arun Anavarpu, Director of Product Management for Cisco’s Data Center Networking Group, opened the presentation by framing the massive industry shift towards AI. He noted that the evolution from LLMs to agentic AI and edge inferencing creates an AI continuum that places unprecedented demands on the underlying infrastructure. The network is the key component, tasked with supporting new scale-up, scale-out, and even scale-across fabrics that connect data centers across geographies. Anavarpu emphasized that the network is no longer just a pipe. It must be available, lossless, resilient, and secure. He stressed that any network problems will directly correlate to poor GPU utilization, making network reliability essential for protecting the significant financial investment in AI infrastructure.
Cisco’s strategy to meet these challenges is to provide a complete, end-to-end solution that spans from its custom silicon and optics to the hardware, software, and the operational model. A critical piece of this strategy is simplifying the operating model for these complex AI networks. This model is designed to provide easy day-zero provisioning, allowing operators to deploy entire AI fabrics with a few clicks rather than pages of configuration. This is complemented by deep day-two visibility through telemetry, analytics, and proactive remediation, all managed from a single pane of glass that provides a unified view across all fabric types.
To deliver this operational model, Cisco offers two primary form factors. The first is the Nexus Dashboard, a unified, on-premises solution that allows customers to manage their own provisioning, security, and analytics for AI fabrics. The second option is HyperFabric AI, a SaaS-based platform where Cisco manages the management software, offering a more hands-off, cloud-driven experience. Anavarpu explained that both of these solutions can feed data into higher-level aggregation layers like AI Canvas and Splunk. These tools provide cross-product correlation and advanced analytics, enabling the faster troubleshooting and operational excellence required by the new age of AI.
Presented by Arun Anavarpu, Director of Product Management. Recorded live at Networking Field Day 39 in Silicon Valley on November 6, 2025. Watch the entire presentation at https://techfieldday.com/appearance/cisco-presents-at-networking-field-day-39/ or visit https://techfieldday.com/event/nfd39/ or https://Cisco.com for more information.
Transcript
I'm Arun Pu and I'm representing Cisco. And uh, I'm the director of product Management for Data Center Networking Group and I manage the Nexus dashboard suite of products. And today we are going to talk about how big shifts in the industry are making data centers prepared for ai.
We'll also talk about how Cisco is addressing those requirements and challenges with its end-to-end solutions. I'll briefly touch upon the operational readiness of what we can do with, uh, the AI networks and then I'll hand it off to my partner here who will talk about more of networking for AI. To begin with, I can't even get out of my room without hearing the word AI or seeing the word AI on any of the objects that's sitting out there in the room.
And better yet, I'm here talking about and contributing to that the shifts in the AI world, starting from LLMs that we soon, um, see it as obsolete agent applications, agentic ai, and then the inferencing at the edge. All these technologies are creating that AI continuum and demanding more of the infrastructure that supports that AI applications out there. The underlying infrastructure component, which is the key for connectivity, not just for traffic that is flowing within the data center, but also traffic that leaves the data center is the network.
So the AI network demand is such that you have scale up networks and scale out networks, which are front-ending the, um, data center fabrics. And then we also have this new scale across network fabrics that are the new demand where you're connecting cross geography of data centers together to serve those agent AI applications that are serving users needs. So the demand and the requirement on the network is that networking is no longer just a pipe.
It needs to be available, it needs to be lossless, it needs to be high bandwidth. Most importantly, your GPU that you've paid monies for should be better utilized. So the network problems will directly correlate to GPU utilization.
So the network being available, resilient, reliable, and most importantly, secure is the key for having a great foundational architecture for in AI infrastructure. What is Cisco doing to contribute to that infrastructure with its end-to-end story, starting from silicon to the hardware that supports with that silicon and the software that runs on it, along with the operating model that allows people to use Cisco products seamlessly. We are trying to provide that end-to-end solution even with the optics that connect those hardware as well.
I'll briefly touch upon the operating model for AI networks. The ability to provision your networks on day zero with one click, two click, three clicks of a button instead of going through pages of configuration. The ability to have visibility into that network you just deployed using telemetry to do analytics and remediating issues, sometimes reactively, sometimes proactively, and then having that in a single pane of glass across your estates, whether it's scale out, scale across, or your front end or backend fabrics is the key for success.
Cisco has a couple of options and a couple of form factors to provide that end-to-end story for you. All of you have heard about Nexus dashboard and its evolution of how it became a unified solution, allowing you to do provisioning management, security infusion into the network fabric as well as your analytics is now able to manage with a single click of a button and configure your AI fabrics. My partner Perish will go through more details about what it can do, but this is one available option and one form factor where in your on-premises network, nexus dashboard sits in the on-prem network.
It also has the ability to sit in the cloud on AWS, but it allows you to manage your network from where your data center is. We have another option, which is the hyper fabric ai, where if you want Cisco to do the management of that management software hands off for you. We have the SaaS option of hyper fabric ai, which does the same thing, easy deployment of your AI fabric, easy visibility into the fabric and easy remediation of problems that occur in the AI infrastructure layer.
So a couple of form factors there for you to think about your operational models for AI infrastructure. It's not just these two elements. We also have something called as AI Canvas and of course Splunk, which are the higher order aggregation layers for sending traffic from Nexus dashboard or hyper fabric AI into those layers.
Do analytics across cross products because these two products serve your cross product correlations and analytics use cases with AI and it'll enable you to do faster troubleshooting, proactive insights and give you that operational excellence that is required and demanding for AI infrastructure needs.