Cisco Nexus Hyperfabric: A New Cloud-Controlled Data Center Fabric as-a-Service
In this session, Cisco discusses new cloud-driven Data Center Network as-a-service. They cover how the product works, key use-cases and how this also relates to new Hyperfabric for AI offerings from Cisco.
Presented by Dan Backman, Product Architect and Distinguished Technical Marketing Engineer, and Minako Higuchi, Principal Technical Marketing Engineer. Recorded live at Tech Field Day Extra at Cisco Live EMEA 2025 in Amsterdam, Netherlands on February 11, 2025. Watch the entire presentation at hhttps://techfieldday.com/appearance/cisco-presents-day-1-at-tech-field-day-extra-at-cisco-live-emea-2025/ or visit https://techfieldday.com/event/clemea25/ or https://Cisco.com/ for more information.
Transcript
Thank you much. Uh, first of all, uh, welcome everybody. Uh, my name is Dan Backman.
I work for Cisco Systems. Um, I've been at Cisco for about seven years now. Before that, I've been few other places.
I've worked in data center technology, a couple startups, and before then I worked for a company that makes routers and gin flavoring. Uh, and I have enjoyed a few sessions here at TFD before, but I'm super excited to talk about, uh, what we've been doing for hyper fabric. Just as a little bit of background, uh, this is a product I've been working on for about two years now.
We've spent a lot of time working with real network engineers, people like yourselves in the trenches who are building out large networks, and we've literally started looking at how do we rethink the operational model of running a data center. I'm also joined by my colleague Manco. And Manco is one of our awesome tme, just one of the best group of people that I've worked with.
We've worked very closely with our DC group as we've been putting this together over the past years to make sure that this is an awesome session. So, Manco is gonna walk us through, uh, a lot of the functionality of this. And likewise, we are also gonna show you a live demo because in the end that's what matters.
You can listen to me all you want, but let's actually see exactly what this looks like. So first, let's take a look at what hyper fabric is. You can see clearly what's in the title, but first of all, what exactly is it?
I really have to start by saying I can't believe Cisco let us build this. This is awesome. This is a cloud controlled data center fabric running on switches that are purpose built and hardened for deploying data center environments.
Running Sonic. Mm-hmm. No, just Sonic.
Ah, not the gene. That's also, that's also a Sonic the Hedgehog. But actually no, we're actually running a completely new operating system.
When we built this, we had a chance to go back and look at if we were gonna solve a lot of the operational problems that we had running networks, what would we do? Number one, if we start looking at data center fabrics, data centers should be a science and not an art. How do you make a reliable data center?
It should work the same way every time we're here at Cisco lives. If you go out on the show floor and you were to go pick up five CCIS and put them in a lab and give 'em a whole bunch of switches and tell 'em, I want you to build the data center fabric, what do you think is the probability that the configurations they put together, even though they are, they're all awesome, they're gonna make it work perfectly. What is the probability of those configurations being even close?
Zero. Probably they might choose the same protocols, but even then you're not sure, right? So one of the things we looked at is there are best practices for how you run large scalable data centers.
And we learned a lot from the, from the hyperscalers and said, if you're going to run a highly stable large scale data center, what would you do? You would make it very repeatable. You would wanna make sure that none of these deployments are a snowflake.
And we want to build an architecture to actually help our customers build these types of fabrics. The other things we would wanna do are make a fabric not fragile. Data centers are critical infrastructure and one of the problems we often see is once you build that network, people are afraid to touch it.
If you're gonna build a data center fabric and let's say you want to go expand it, if you're building it on a typical fabric technology, in many cases you may actually have to go back and re-architect that because you have to think about all the different layers of signaling protocols, reachability that you have to rebuild. How do we actually change the experience of deploying this? And that's what hyper fabric really is all about.
So if you were gonna build something like this, what would you do? You would start by looking at how do I change the control plane that I can make it plug and play? How do I look at the operational model and maybe even look a little further than just running a fabric and make it something that we could actually start to eliminate?
A lot of the frictions you get when you typically deploy a fabric and then you would wanna make it repeatable, you'd want hardware that always works. You want something that could actually solve this problem of running large numbers of data centers in distributed environments. Solve that edge compute data center problem.
And if you were gonna do that, what would be all the pieces you would put together to build that? That's what we had the opportunity to do two years ago and we've been working on this. We are super excited to talk about it.
And with that, let's go a little further and take a look at the actual product. com. It's already public.
You can reach out to the ULL already, but unfortunately you cannot log in yet. It's gonna be available next month. So the first thing we are gonna do is by using Cisco account to log in to the cloud-based controller.
The customer doesn't have to prepare on-premises controllers so that they can easily try out. The first thing we are gonna do on the cloud controller is creating fabric. So imagine you have a new project deploying fabric.
What is the typical first step? So let's walk through how it would look like in hyper fabric. Hey, done.
How big earth fabric would it be? How many did we need? Let's make a leaf spine.
Mm-hmm About two spines and eight leaves. Eight leaves. Notice this Manco is building this.
The whole point of the designer is we want customers to actually interact with this. We want people to actually get guided into what is a valid design. The whole idea of hyper fabric is also, there is no Cisco validated design for this.
This is actually built into the designer process. Now what you see is we've created a blueprint for that fabric. Now that blueprint has the switches that you've selected.
We now know the size and shape of it, but where else would we go with this? What's the next thing you would do? Manco, if you're gonna build a data center fabric, Uh, I probably need to define how to connect them each other by checking what's the speed and what's the optics supported on the devices.
So probably people need to take a look at data sheets of which vendor, switch network devices and also open up some famous application called Microsoft Excel or some sheet kind of application to list up the covering plan because I'm probably not a person who gonna cable it here. The interaction happens and then sometimes we might make some mistake, right? So how many people here have built that Excel spreadsheet?
That sheet that shows every port that goes to every port? We've all done it, right? This is actually the joke that happened.
We've spent a lot of time with real network engineers as we've been building this product. And the joke that kept on happening is what's the number one most used tool in network engineering? Nobody liked that Precisely.
I thought it would be SSH, but no, everything comes out to the Excel spreadsheet. So what this highlights though is that as you're building networks, what we do find is that there's a bunch of places where you run into friction because the people who design a fabric are sometimes not the same ones who actually deploy the fabric and then they're not always the same ones who actually operate the fabric. Later we rebuilt this controller and with a web presence, what we can do is start to rethink how we can get these different groups to collaborate together and get the right answer the first time.
And that became the backbone of how we actually built and managed the fabric. So let's go ahead and build a cabling plan. So how many cable do you need for each switch pair between two by eight?
Yep. So let's say one link for each pair. Mm-hmm.
And then what's the speed do we need? So let's pick 400 G then what's the cabling type do you prefer? Let's, depending on the data center, I mean facility, there might be a requirement of multimode or single mode.
The system gonna tell us what's the available optic based on your requirement. So let's pick this R photo two. Then boom, you don't have to deal with excel sheet, but it's gonna create a list of cabling plan.
But let's take a look at the faster example. How would you connect by the lift topology? It's most likely you're gonna pick the fast fabric port of the spine is going to fast fabric port leaf one and then continuously doing this and then managing this in the Excel sheet.
And as you may probably notice that that is a cabling plan that is exportable. And if your field technician has an access to the system, you guys can share that information together. So fabric blueprint is a source of truth, not only just a logical configuration, but also cabling plan.
So why did we do this? Who here has gone into a deployment where you've built a design, you've figured out all the pieces and parts you need, you order the gear and then maybe a month later you're sitting in the data center and you're plugging in all the gear and you figure out, oh these leaf switches are rear mounted in rack. I need the airflow to go the other way.
Does that ever happen? Never happened. Or I ordered a whole bunch of multi-mode optics and as soon as I go plug it in, they're oddly blue colored and realize, oh no, that that's single mode.
So that's a problem that happens a lot, believe it or not. And it doesn't matter how good network engineers and partners are in as part of this. Mistakes like this happen.
So we all spend lots of time on the internet. What is the right way to get the correct answer to any question on the internet, The chat bot, uh, with AI or Not? The 2025 answer before that, the right answer is you post the wrong answer on a chat board and the entire world will CR will correct you and tell you the right answer.
We can do the same thing here. So what we showed you is actually that Excel spreadsheet that you would normally build, but we're populating it with all the information about the optics that you need to know and is maybe not that you need to know it, but now you have a sheet that you can hand to the physical layer team in your data center that says, I'm gonna build a fabric from these racks to these racks and I'm assuming that I'm gonna use multimode fiber. And we even print out what is the max supported distance on that type of fiber with that optic so that your physical air team now at the planning stage can tell you, oh, that one rack is actually over there and it's more than 200 meters away.
And maybe you need single mode for that one rack, but now you can actually get the right answer for the physical air team 'cause you've already given them the plan before you order the gear. This is why we built this. The other thing that we're doing is this now becomes source of truth for the controller, which we'll talk about in a second.
The next part is you need to order the gear too. So why do we have to go call up somebody and have them do a quote for us and go do a configuration where we already know the right answer, we know it switches you picked, you pick the airflow, we know which airflow direction it needs to take and we've already picked the optics. Why can't we just click print?
Well we can. And we're also adding a button that's gonna connect directly back to CCW so that we can actually generate an estimate for you right away. Now if you're working with a partner, the other part of this is this is a website, this is a web-based controller.
You can invite other people into this. So you could actually take your SE or your partner or your trusted partner and actually invite them in and collaborate on this design as part of this. This is a slightly different take on running networks, but we, as we started looking at a lot of these headaches that you often run into, there's real issues of when you show up and you've got the wrong optics.
A 400 gig optics, first of all, they get kind of interesting once you go past a hundred gig, there's a lot of options you even need to start worrying about whether it's angle polish or not. It's signal mode. And then at the same time you need to worry about a whole lot of other factors.
So it gets complex. So how do we actually get the right answer and get all this built right the first time so that you don't actually have to wait if you've got the wrong optics and order new ones that can delay a data center deployment by literally weeks to months. I know you have heard this before, but the order button to the Cisco web shop would be nice, uh, beside the print.
Yeah, That would be great. Cisco doesn't make it that easy to do business. We wish we could do it connect.
The best we can give you is an estimate where you can connect directly. But yeah, we would love to be able to do that as well. Um, the other point is, uh, we, we also have this information.
So now we can start taking this and give you different views of this. This is an assembly list, like when you get the gear on prem, how do you know you got all the right pieces? How do you know where you're gonna plug in all the optics?
We're actually putting that information forward here as part of that same view. So we're gonna pause there for a second. We're talking about a data center network.
We haven't got into any gory details, but it sure looks like a pretty cool configurator, doesn't it? Well, it's actually the controller. This isn't a configurator, it's actually the source of truth for the controller that will actually be building and running your fabric as well.
So what did we cover here? What is different about hyper fabric? We've looked at how do we build the network experience but take a step back and actually try to incorporate lifecycle 'cause it matters.
This is where you run into real problems and if we already know the right answers, how do we connect the people that need to know so that we can get the right answers all the way through? The next piece is right now we talked about day zero, sort of the planning stage and the ordering stage. But how do we continue this?
Wouldn't it be great if we also took this information and helped you on day one when you're physically installing it? We'll talk about that in a second. The other piece is we also, to extend it further is this is not just about deployment, this is also about lifecycle of running a fabric.
That means we also have to do things like software upgrade and maintenance and changes to the fabric. We want to help you actually deploy it but also make changes in the future and make those painless and seamless as well. So this is really what's different about hyper fabric is we're looking at this end-to-end life cycle of the fabric and trying to leverage this as a source of truth throughout the entire system.
So what are the other pieces that we need? We need to actually also take this and build a fabric that now becomes much more easy to deploy and much more flexible. Let's take a look at a little bit more of this.
We talked about the bill of materials. We are going to build a connection to CCW so you can get the estimate. So How's it, how's it done into the CCW?
Is it like for the export CSV and then you import in the CCW or Oh we're gonna make an API connection. We're exporting A CSV today. We haven't enabled the button yet.
That will actually generate the estimate. We're still, we made some changes in PIs so we have to make some updates before we release it. Uh, but it's actually gonna be a button that says generate estimate on the backend.
We will take that bomb that we have, including all the software licenses that attach to this, which is actually a very simplified model. I think you'll like how we're doing licensing. All of that gets automatically sent into CCW and it generates an estimate.
Mm-hmm. And the estimate is something that you can share with a partner where they can go figure out what the pricing structure is you to offer. So as a customer, if you're not a partner, what you would get is an estimate id.
We can email that to you or you can give it directly to your partner. If you're a partner, then you can get direct access to the estimate and generate a quote directly. Mm-hmm.
Perfect. Thank you. The other thing that came up here with all the network engineers that we've been working with is there's a real desire to start to make these networks self-documenting as well.
If we can establish the controller as a source of truth, why can't we just take that and emit that information? You start to see that with things like selection of the pluggables where we're automatically deriving information for the pluggables that you need. Likewise the actual interconnect matrix by the way, it's smart enough to also deal with things like breakout cables.
We spend a lot of time making sure that this actually works in scales for the different deployment types you need. But we also add something else, which is a day one feature. The one of the biggest problems that you have when you start deploying this is somebody's gotta physically plug this in.
Who's been on that call where you've got remote hands in the data center, they're on a phone, there's really loud fans behind them and they're sitting there plugging in cables and you're sitting there looking at the console and the first thing you say is what all you hear is, what about now? Nope, didn't work. How about now?
Who's been on that call? Yeah. Lucky you.
If they are speaking the same language like the Exactly. So if we have a cloud presence and we have a source of truth of where every connection should go and if we have switches that call home to cloud, why can't we make this easier and just tell people what to do? So we've also built in a remote hands interface into hyper fabric.
This is something where you can take a cell phone and we assume the person whose remote hands doesn't work for you. You can invite them in, you can give them a link to actually join this fabric and it will give them a list of tasks. Here's all the things I need you to do.
I need you to get these switches plugged in connected to cloud and we'll talk about how that works as well. And then here's all the things we need you to cable in. Day one is a lot of plugging things in and cabling it, but that's where a lot of mistakes happen.
We'll talk about in the future why this gets even worse even for the fabric we just built. That's 32 separate links and when you consider each one of these is probably going through structured cabling and you could potentially have polarity problems all the way through. There's a lot of trial and error here.
But wouldn't it be great if they had an interface when they plugged in a cable and they saw link go up. If we tell 'em did you plug into the right port? Did you plug in the right pluggable?
If we know that's supposed to be a 400 gig single mode interface, but you plugged in a multimode interface, we should tell 'em immediately. 'cause maybe that was intended for a host instead of a fabric link. If link came up, the next thing you wanna know is did it come up to the right place?
Is that fabric link actually between the right two switches and did it land on the right ports? And even if it comes up and lands on the right ports, we can then give feedback to say no, this came up but this is wrong. I need you to move that from this port to this port.
Why can't we do that? It shouldn't have to be this hard, but this is also built natively into hyper fabric. One question for the screenshot.
Do the smart hands interface also show the physical view of a switch and highlight part in the fashion? Yep. We're still built.
We're still doing some updates. That's a mockup. One of the things we are actually adding as a physical view, because you're right, why is it that every switch, sometimes the ports number top down left right, sometimes left right, top down.
It's a little infuriating. Yeah. Often the the interface names for networking person it makes sense for smart and person may be not.
Yeah. I'll tell you a little secret as we are developing this, um, we we're using hardware from MCG effort, a couple different business units to put this all together. And what we learned is that in some parts of Cisco ports start with a zero in in other parts of Cisco ports start with a one.
Yeah. Best is bring some kids. Everybody brings their six, 7-year-old you tried out, it's working with them then it's good.
Yeah, Absolutely. I love it. And the other thing here is that this is not an app.
We don't wanna force anybody to install an app. This is a mobile responsive webpage, which means that we can make updates to this. You don't have to, you just have to hit a URL, you've got access to this and then it's, there's only limited things that you're allowed to do as a smart hand.
One of the things you can do for instance, is um, there's a portion of sort of claiming a switch and binding to a fabric. We're enabling limited read-write access. So you can accomplish those tasks as smart hands as well.
So we spent a lot of time on this. We're really proud of this feature. We think this is gonna be a really important thing.
I think Dominic had a really interesting point. Um, and I definitely seen that. So, you know, sending switches to remote sites where there are no available English speakers.
Like if you would translate this into like, you know, uh, local languages, that would be a huge, because I can imagine like, you know, we, we've had that so many times where you're literally on the phone with someone who doesn't speak English. Um, so, you know, if this would be in a local language, this would be a huge, huge help. That's a really good point.
What's your top three to five if you had to pick? Uh, aside from the obvious ones. 'cause I think like, I think what you're calling out is there are a bunch of, there's a bunch of remote compute going in to now probably countries that may not be speaking sort of your standard English, Spanish, German, Et cetera.
Exactly. Exactly. So, so I don't know like what would be the top three, but like, you know, I think when we were like for instance shipping stuff into China, um, you know, that that was a challenge.
Mm-hmm. Um, but even in countries like France, like, you know, you just like getting someone who, who's locally in some like small village, uh, not necessarily in Paris but like, you know, some sort small village in south of France, you know, and there's no, no like the remote hands who are available in there are not, you know, speaking language. So, um, speaking English.
So you know, even French Spanish as well, The toughest one are the ones where the language is different and all the sites are different. Like Japan where you have both different, then it becomes challenging. Yeah, yeah, Absolutely.
Well I think you actually bring up an interesting point because the whole assumption of this is that these are not network people. Mm-hmm. And that's the other thing.
Yeah, yeah, yeah, yeah. Like you know, when you're deploying switch in like in a local branch somewhere, uh, there's like someone not technical and someone who doesn't speak English. Absolutely.
That's a really good, that's a really good call out. We absolutely do need to localize that. So let's go to the next step.
If you're gonna build a fabric, what would you need to do to it? So we have really powerful fabric technologies. Manko you, you know more about a CI than about anybody.
I know you've worked on it for about a decade now, right? Likewise, we also have solutions that are deeply built on EVPN vxlan. There's a lot of tools in the toolkit that we can pull from to build this, but when we built it, we needed to do a few things.
We needed to, to be something that you can easily deploy remotely. So I need it to be hardened. I need that switch that I ship out to always call home to cloud and not do anything else.
I can't have a switch that pops up in a RAM bond mode where you have to go search the file system and find the right boot string to magically make it work. That doesn't work if you're shipping to a remote location and maybe you don't have A-C-C-I-E on staff. I also need something that I can automate easily.
Right now. Everything, especially a lot of these workloads that are coming back from cloud, they're coming back with an automation scheme from cloud. So being able to automate this by a and by an engineering team that knows how to automate cloud but may not be a set of ccis IMP is important.
If you're automating EVPN fabrics today, even if you have controllers and tools to help you, you have to think about how does this port map to a local VLAN to an MVE interface to a VT E. What's reachability of VT E? What's A VNI?
How does that happen on the other side? Where does IT gateway that's exposed to the automation team that's asking a lot of the automation team. If you configure a network in AWS, it's just a network.
Why does it have to be hard? So we did need to think about the programming surface here as well. And then likewise, when you're building the fabric itself, one of the things we see is fabric technologies today are incredibly powerful.
Every discussion you've had about them is about scaling them up though. What happens when you want to take advantage of a fabric but maybe you don't need that many switches. Wouldn't it be nice to have the ability to deploy an EVPN fabric where we can make it much more plug and play and flexible?
Absolutely. These are the things that we've considered as we started building this. So we had to build a lot of new components.
I need a new switch that we'll call home to cloud. I also, that switch has to be hardened. By the way, if this is going into remote sites, maybe it's not a fully secure environment.
Maybe this is a shipping container sitting outside with a whole bunch of compute in it. I need that fabric to be something that is relatively hardened against attack. So I need integrity validation, I need hardware root of trust inside these boxes.
I need to make sure that they always call home to cloud. Likewise, it may be even be an untrusted environment. What happens if these switches are outside of firewall?
I need to harden the surface of this device as well. So we've done that as well. That's what we've done on the hardware side.
We've also taken standard EVPN vxlan, but we've changed the way the control plane works so that it's actually plug and play. And I'm sure a few, I'm sure a few people out there that have built this by hand are calling BS on me right now. But we actually did it and there's some key things we did to make this work.
But the net effect is that what you can do is now deploy an EVPN VXLAN fabric that is plug and play and you can even change the shape of that fabric in service without having to re-architect it. Fabrics today are too fragile. That's one of the key things we wanted to fix.
Let's look a little bit more at what this looks like. So we've built a fabric though. We've gone through the configurator, we have a blueprint design.
What's the next thing you wanna do on a fabric back in stuff? There you go. And then configure connectivity across it.
Right? I Stop Oh sorry. Sorry, go ahead.
Sorry. Yeah, let's, let's build a network. Yeah, Let's build a network.
So the actual data center network deployment doesn't finish without a logical configuration such as we RF logical networks, any guest gateway and so on. In hyper fabric we don't have to care about the physical topology that much after our initial physical cabling the plant, let me hide this and then lemme create one logical network across these two spine eight switches. Lemme create one logical network and then it's pretty easy and it might be too fast to create one.
So let me create a second one and then let me pause here. What configuration I just put, I just, Who here has done this by hand? Is anybody enough of a masochist to actually build an EVPN VX lane fabric without controller by hand in lu?
Yeah, I think we've all done it there. There some of us are a little weird and yeah, you gotta, you gotta want to go through it at least once. Um, in order to get A-A-E-V-P-N VXLAN switch, it doesn't matter what vendor it is, you have to build out an underlying routing network.
You have to sort out reachability, you have to set up a PGP mesh for EVPN. Then you have to start to start to configure services. You have to sort out VT e reachability.
You have to define a profile for what this network looks like, what VNI it is and mapping. It's about 400 lines of config per switch before you forward a packet. How many lines of config is that?
Actually it's not even a line, it's just a text and then you can specify BNI but similar to a CI. We don't necessarily care about it because it's automated. If you want to edit it, you can still edit it then lemme actually make one of them as a layer three.
Meaning any case to distribute the gateway. Of course we support multiple B-R-F-B-F. But let me just select D for VRF then I'm gonna just make it, you are stuck.
So I think it's 2025. Been a rough year already. It's pretty short.
So far flies. There's not really an excuse to not be fully V six compatible. So as we were building this fabric, the directive was 100% V six dual stack everywhere.
Ann, can you repeat that for posterity? Just say the first part again. It's been already Strange though.
It's Crazy. Five, it's been a long year. There's no reason for this not to be.
Thank you. Yes. There's no reason for things to not be fully dual stack V four V six.
We have that on film now. So we're gonna be showing this to a lot of other people in the industry. Don't seem to get that.
I await the darts. Uh, But I think the main reason was there was a lot of additional lines of code. People were kind of struggling.
Mm-hmm And then if it is that easy to add dual stack Yeah. Then there's no reason to not do it. Yeah, Absolutely.
I think in the past mostly lazy admins in too much complexity maybe. Yeah. That stopped us.
Um, Out of a new protocol. Yeah, that's fair. 'cause it works differently compared to number of Yeah.
What is the single best thing about V six by the way? If you had to pick one thing they did in that protocol that's better than anything else, what would you pick? They got rid of broadcast.
What's the other thing? No nut, No net. So address space.
My favorite Would be my actually preferred thing because actual protocol works sensible. Y'all are too smart i's probably go for it. Easier joke.
Um, my favorite thing is actually link local addressing. I hate it. I, it's easy to hate, but what it does do is it's always there.
It's self configuring and if you wanna build something plug and play, wouldn't it be nice to not have to worry about an underlay addressing model ever again Resulting in a network having four different addresses who I cannot control If you had to see it and deal with it. So what we've done in this fabric in order to make it plug and play is all that underlay addressing is actually V six link local. We do that so we can facilitate auto discovery and our protocols are actually building all the reachability based on that.
In fact, we went through and we rebuilt the entire EVPN stack so that it's V six native. In fact there is no V four packet on the wire in hyper fabric until you configure a V four SBI. Beautiful.
If you have environments that have a mandate for V six, it's done. It's ready. You can already prove that it's fully V six.
We did it mainly because it saved a lot of work in the, in the provisioning. We didn't want it to be fragile. So if we have a numbered addresses on each individual link, I can move a link from any port to another port and we just dynamically discover an adjacency and I don't care what the address is 'cause it's just link local And it's Absolutely, and the other key part to making this work is, uh, by the way we've configured svs L two reachability.
Have you seen any IP address configuration other than what you're exposing to the user at all? Have you seen BGP config yet? Is the EVPN fabric?
I assure you there's plenty of BGP in there. Have you seen as numbers yet? If you wanted a fabric to just work, especially if you're deploying in a remote location, you can, all you have to do is literally plug in switches.
You have a plug and play EVPN fabric. And my abstraction model here is this looks like a big layer three switch. You created VLAN layer two network.
We called it layer two network. 'cause technically there's a VLAN in the switch and we didn't want people who knew too much to get really confused by what we're doing. But it's a layer two network, got an SVI on it automatically configured, distributed any cascade gateway.
It's done. I need a fabric that gives you all the benefits of a fabric. But maybe in that remote location you don't have A-C-C-I-E ready to pre-configure it for you.
And, and I think this lowers the burden for entry level people. Yeah. It makes it more streamlined.
And at the end of the day, they just care about the services that they expose. Yeah. The underlay is maybe attack engineer care about that, but the, let's say average, uh, lazy admin, as long as it works, he doesn't care.
Yeah, Exactly. The other benefit here is when you design an EVPN fabric, how many we, we, we talked about if you took five ccis, how many different choices are the, are there of how you would build that? You could choose different IGPs, you could have different addressing strategies.
You have different strategies on what you use for reachability or VT taps. There's lots of ways to configure this. What I want to do is make fabric technology something that's pervasive and always works the same way.
Because you don't, in your data center, you don't want your fabric to be a snowflake. You don't wanna be the only one that has that one combination of features that nobody else actually has. So out of the box.
Out of the box. Mm-hmm. Out of the box.
Yep. The way it works is, you can think of the underlay as a bubble. What's inside the bubble is ours.
We auto provision it and we don't leak it to the outside world. In fact, you peer the outside of the fabric with the rest of the world. You peer your VFS to the outside world.
You attach hosts into these networks, you can then peer this with other, uh, fabrics at layer two. All of those features and functions work the same way, except what we're doing is we're hiding that underlay and we are able to dynamically provision it. And we did actually have to write a whole new equivalent of an IGP underneath this to make this work.
There's another magic trick here. I can take you from a leaf spine down to a full mesh with no spines in the same fabric. And I can do that in service.
Try that with a regular fabric. So now comes the controversial question. Yeah, I guess customers will love this, but the value added consultancies.
Yeah, they're kind of, you have done their job. Yeah. So Not really.
This is the this is the crap work. This isn't a fun, this isn't the fun stuff. The fun stuff.
All the pro, all the interesting network problems, as you well know, are at the boundaries. Yeah. You're still configuring boundaries.
You still integrate this with your network. What's different is now you're not actually integrating the underlay into the rest of your routing plan. Yeah, that makes sense.
But I would say a lot of, let's say enterprise businesses can live with that and is happy and their daily production work is totally fine. Without buying every day consultants, which has BGP special knowledge to uh, to let's say configure little detailed in the fabric. Yeah.
So Certainly, I mean to be als also, to be perfectly honest, I work at Cisco. We make lots of great products. We are really proud of what we built in hyper fabric.
But this may also not be your fabric. This may not be your product or your use case. There are plenty of things that maybe you don't wanna simplify it and you want those, those amount of knobs.
So this is not here to replace existing products that we have. This is a new operational thought model for where this makes sense though. There are plenty of places where you do need low level controls.
Yeah. That are still there. Maybe there are customers that have a large number of ces which are doing this full time all the time.
And they are the shops where you have this one person that needs to manage everything and has also, uh, other responsibilities. Um, You know, we were actually really worried about, we were worried about two things as we designed this fabric over the last two years. And believe me, we spent an amazing amount of time with real network architects.
We have, we have people that were, that basically ran all of routing for very large SaaS companies. Uh, people that ran routing and networking for large commercial banks inside the United States energy companies. The surprising thing was that was what we, we were expecting to hear two problems.
We're worried about reducing the complexity and we're worried about cloud. Those two problems were not coming up. We have energy companies that are looking to do this even on their OT networks because simplifying the networks actually solves their bigger headache right now.
I think right now what we're seeing is there is so much complexity in networking that sometimes maybe it does help to reduce some of it. And we were, to be perfectly honest, it's a legit question. I was worried about the same thing and that's not what we were hearing from people in the trenches.
Agreed. Often a more simple uh, design is also more robust. If you turn on all the no knobs, then you have to manage these details all the time.
Yeah, Absolutely. Oh, about interoperability. Interoperability with other fabrics like NSXT, Perfect example.
So what's the fabric like? Anything else? So interoperability is designed to work just like another fabric.
What would you do with another faab? NSXT is a special case because that's an overlay. So I'll get back to that one in a second.
But what would you do with a normal fabric? Well, normal fabric, you'd have northbound pairing. Typically you do layer three ports.
You'd speak some dynamic routing protocol or static if it's a smaller, uh, a smaller environment. So we have that. We do the same thing.
You just do northbound peering over L three hosts. You'd attach hosts, you could single attach 'em. Ideally you're dual attaching everything unless you're running VMware and LAC P is kind of a pain.
But you'd want something like multi-channel link aggregation. We have that. We built it in.
So we've got LAC P port channels. It's multi-channel today by the way. It's not VPC.
We built it using native EVPN. There's more modern technologies that we took advantage of because we had to chance to rebuild this. Then you also need to do east west peering.
So you, how would you do that? Well, one of the common models, we spend a lot of time with our data center engineering TME with tac, what do you do? Build back-to-back VPC model, build a port channel, peer them, tag those links as a trunk and use that for coexistence and migration.
We're not that different. Now the overlays is a question of where you start peering, but again, I mean that's, it's just a standard fabric. Sorry, I interrupted your phone.
Sorry. Too many good questions. Actually, I have another step to complete the configuration.
So let me show you. So I just define logical network, but I haven to decide, I happen to define who is going to belong to this logical network. It's most likely we'll have Frank.
So we need we membership configuration. So let's say we 10 is belonging to logical network one or something like this. So it's pretty much similar to existing we to with external ID mapping.
But you don't have to deal with underlay. What you just need is selecting a logical network and you can simply define who is belonging to this logical network. So if you have just standardized switches, you can say all of switches is ANet one slash one.
If we tag eight 1011, then it's gonna go to net one. But in this fabric, let me actually select leaf one, leaf two, ethernet one slash one to one slash three. If Tropic is coming on these ports with this V tag, it's gonna belong to net one.
So probably people still need to un understand what is real mean, but it's pretty easy to configurate because they don't have to deal with underlay and so on. So let me just create second things. So as Monaco does this, think about automation when you're asking your automation team the right automated provisioning inside your fabric, what do you see here?
You saw how easy it was to create a layer two network. It's, it's just a layer two network. You don't have to give a lot of information.
In fact, you could just give it a name, you map ports to it and that's it. And even the way that we're mapping ports to it is designed to be very automation friendly. You're, you can basically program that data structure directly.
Everything you see here, by the way, is programmed directly and exposed by an API. But we have 100% API completeness with everything that you see. And as we ship this, we're gonna ship this with terraform providers and we'll have Ansible uh, playbooks coming shortly after.
So this is also designed to be automation friendly because what we're also seeing is that as these workloads come back from cloud, especially some of these edge compute workloads, some of the people automating this are more familiar with the cloud model. So it was really important for us to think very carefully about the APIs and the object models that we expose so that they are familiar to people and we don't force them to understand gory details of how the network works underneath. In fact, there's two levels of abstraction that we built.
The first one is at the network level. You can see that just through what NACO provision in just the basic, uh, networks. And then there's another level of abstraction at the physical layer that's where we auto provision the fabric.
So you as a customer provision services on top, plug in cables underneath and we automatically fill in the glue between the two. So can like, you know, not that it would be like so important, but I'm just curious, like, so how is that communication like happening between this uh, dashboard and your switches? Is there like, uh, net con for Oh God no.
No. Oh God No. Yeah, See I'm just, I'm just curious.
Like it's, I don't think it's important. That's the beauty of this point, right? Like it's not really important, but I'm just curious I guess.
Yeah, no, it's, it's, it's a really important thing. So let's actually cover a couple things. You're asking the harder question, which I'll answer in a second.
Let's ask the obvious question. Okay. Okay, let's do It.
How do we connect to cloud? So first of all, how does it work? So the switches have an agent that calls home to cloud.
So this is pretty familiar. Anybody here used a Meraki device? Kind of similar, right?
What you want is you want a device that you plug in, you feed it internet and it calls home to cloud. That's a short answer. It's very survivable.
The other thing is the protocol itself. We've designed to look exactly as much as possible as an actual web client. Why does this matter?
Because the entire internet is designed to connect web browsers to servers. So everything that you would go through, any type of nat, double nat carrier grade, nat, um, proxy servers that could be required to go outside on on a management network. All of these are designed to work with browsers.
We've designed our protocol to look at act and smell just like a browser. But it also means is that there is no more hardcoded IP addresses. Everything is in FQDN.
We do global lab balancing. We have a globally distributed front end that will automatically handle global resiliency for that connection for these upstream servers. That is all exposed basically as a TLS client and it's a restful interface.
Everything is just rest, rest with a published API model. And it's already published. We can share it with you And how you do the telemetry for instance.
We do. But why would you have to configure telemetry? I don't know, like, you know, for the troubleshooting, like, you know, why do You have to configure?
We know we need it. Why do we make you have to do it? Yeah, no, I don't want to configure it.
I'm more curious again, like, you know, at the back end, like how do you, how how do you do it? Just me being curious. Yeah, it's actually part of the protocol.
So if you're gonna build a system like this, um, one of the reasons why a lot of people will ask us why is this brand new? Like you built a new cloud controller, we have new set of switches, new operating system, new agents. Because if you make this work, the way it works is you need the ability to create a feature in four places at the same time.
I need, when I create a new feature on a switch, maybe like we started off building like a SVI for layer three that has a data plane definition, I have to program the A six to four it, I have a control plane definition. I have to basically, uh, program the side layer to make that actually happen in the switch. So there's actual, there's kind of a data plane and control plane Kubota on the box.
I also need an intent model in my controller to describe what is an SVI and it needs an address and it has a certain schema. And then I also need to have a provisioning model that's common between upstairs and cloud and downstairs on the switch. And if I'm gonna make this not a dumb product, I need to have a feedback loop with automatic telemetry and I need to actually build all of those parts of the feature at the same time so that they all work together because I need that full closed loop to build this control model.
Mm-hmm. All of that is automatically built into the, to the client. So we're able to, our engineers that are designing these features have control both upstairs and downstairs for doing that all at the same time.
Mm-hmm. And because of that, telemetry is not an afterthought. Telemetry is actually a key part of every feature we built.
And the controller itself has control to authorize the agent to go extract the type of telemetry it needs. And in fact, it's dynamic when you actually go look at telemetry, it's actually dynamically turning on a real time channel to uplift telemetry faster. 'cause you're looking at it.
But that's our opinion of how we should do telemetry is make it an intrinsic part of the product and actually have it controlled from above. There is no telemetry config because we should be doing telemetry, not you. Uh, the other obvious question by the way, um, and, and I'm looking at the clock here and I wanna be, uh, mindful of that as well.
What happens if a data center network gets disconnected from cloud? I'm surprised nobody's asked me that one yet. That was, that was my fault that I wanted to ask.
Yeah, the short answer is the only obvious one it keeps on working. So what we've designed here is that to get the behaviors we want, we want to enable this through a cloud controller. The reality is we can never allow your data center network to stop working.
But what happens first, the, the actual cloud connection itself is highly resilient. If it gets disconnected from cloud, any one switch can connect through another switch connected to cloud, which means my disaster recovery is I bring in a cell phone. I, it was actually we were, we were worried about wireless outside.
That was our disaster recovery plan here too. Just tether on your cell phone and plug into one switch any and your entire fabric is up and running. That's number one.
Always be greedy about getting back to cloud. The other part though is there is no critical path into the controller for operating your fabric. All of the control plane protocols run on the switch and you can even unplug switches and rebuild your topology if you're disconnected from cloud and it keeps on running.
That's what survivability should look like. The only thing that happens if you get disconnected is you can't, you queue up configuration changes, telemetry pauses until it reconnects. I guess it keeps on working for as long as you have to find it a policy.
And I apologize because otherwise you plug stuff and it's not just gonna work magically, but it assumes there is a topology already defined. Absolutely, yeah. And for instance, you have to have switches admitted to the fabric first before they can connect.
But in general the whole point is it does need to keep on working. And we spent a lot of time focusing And you have to be a bit mad if you want to introduce major config changes during an outage. Yes, Absolutely.
But in the end you always have a get out of jail free card. It's just a web connection going to a public front end. You can tether on your phone and as long as you plug into one switch your whole fabric gets online.
How much bandwidth is needed for this control plane. That Was gonna be my next question. We show you it's actually pretty low.
We spent a lot of time reducing the amount of bandwidth. It's kilobits. It's kilobits per second.
Hi. It's high. Yeah.
So I'm gonna end here because I wanna be respectful of your time. You guys have asked me so many great questions. Look, there's a couple slides we haven't got to.
Um, it's not a Cisco presentation if I don't show you hardware. So here it is. Um, we are, so we talked earlier about we have to have purpose-built switches for this.
So there is a high density spine oriented device as well as a leaf oriented device. Um, these are actually running the same asic and I don't care where you put each one. You can just as easily use that as a leaf, as a spine.
You can even put host ports, gateway ports, anywhere you want on the fabric. And the thing I will leave you with is we added one more magic trick, which is we scale down. You can support spineless fabrics.
So the one thing we haven't talked about that I wanna cover before we finish is where would you use this? Think of a place where to get started. We can scale up, we can build huge, huge networks.
We've got, we'll do the math here for, you can build a big network, but if you're trying to plug in and you've got 300 ports in a remote location, you probably, if you're deploying a typical fabric, you probably have more r use of servers than you do switches. Now I can deploy a spineless fabric and I can even do it with a single box. Think about that as an operational model.
Drop ship it, plug it in, remotely manage it, and no additional hardware around there required. I'm gonna stop there. Thank you for the opportunity to present.
It's a Meraki data center actually.