Multi-Tenancy & Network Automation for AI Infrastructure Operators Demonstrated with Netris
Netris CEO Alex Soroyan demonstrated the multi-tenancy and network automation solution in AI infrastructure. The presentation began with a live demonstration of the Netris controller, showcasing how it facilitates the setup and management of AI infrastructure networking. Utilizing Terraform modules and a “CloudSim” simulation, Soroyan illustrated the process of initializing the controller, generating network configurations based on user-defined parameters, and creating a digital twin of the network for validation.
The core of the presentation focused on day-2 operations, specifically the creation and management of tenants and network isolation. Using templates, Soroyan showed how easy it is to establish isolated clusters (VPCs) for different tenants. These templates translate high-level server assignments into low-level switch port configurations, enabling a cloud-native approach to network management. The demo also highlighted the integration of Elastic IPs to expose the internal clusters to the outside world.
Finally, Soroyan discussed monitoring features, which automate the configuration of monitoring tools and provide network health checks, including link validation. The presentation also touched on InfiniBand networking, demonstrating Netris’s capability to manage InfiniBand fabrics and integrate them with Ethernet networks. The key takeaways were automating network tasks, simplifying complex configurations through templates, and comprehensive monitoring capabilities, all contributing to a more efficient and manageable AI infrastructure environment.
Presented by Alex Soroyan,CEO and co-founder, Netris. Recorded live in Santa Clara, California, on April 24, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/netris-presents-at-ai-infrastructure-field-day-2/, https://techfieldday.com/event/aiifd2/, or https://www.netris.io/demo for more information.
Transcript
Uh, I'm Alex Soan, CEO, and co-founder of naris. Uh, and, uh, in this section, um, I'll show small demo of some of these things that, uh, I just explained. So this is Naris controller.
Uh, it is blank. Uh, this is basically day zero situation. Um, so we can see that the inventory is blank.
IIPM topology, nothing in there. Now, the first step or, uh, first step is to, you know, initialize the controller. And of course, it is possible to initialize, uh, manually by defining your switches, your ips.
Um, but we're gonna use, uh, more, uh, more simplified method. Uh, we'll use, uh, one of our Terraform modules that are, you know, designed for particular use case. Uh, one second.
So, on left, on left terminal, you, you see this is, uh, mattress controller, uh, CLI. So I'm, I'm on that machine, and there's this Terraform folder. Uh, now this Terraform module, this is not something that customers need to write themselves.
It, it's part, it's, we ship this with the product. It is open, it is written in HCL language. Customers are welcome to make all kinds of changes.
We will be more than happy to help them with that now, but in most cases, there's no need to make changes, because what you need to do is primarily to edit this file variables file, which is, you can think about this like a, like a questionnaire where it's asking some of the questions that, that, like, professional services would ask, like, how many servers are you planning to deploy in your, in your network? Do you have a questionnaire that populates this file? Uh, that, that, that happens manually?
Okay. Yeah. So this, this particular one is designed around, um, NVIDIA's, uh, reference architecture for ethernet based, uh, clusters.
Spectrum X, what they call. And, uh, in spectrum X, your EastWest fabric, uh, has a certain, uh, prescriptive design. This module knows that design.
So for backend fabric, we're only asking these questions. Basically, we, we just need to know how many servers are you planning to have? I have 64 in this case.
Now, this part is for north south fabric. That's because north, south fabric, the front end fabric, there's, it is less prescriptive. There's more freedom.
Customer may want to do all kinds of, uh, modifications. So that's why we ask more questions. How many lift switches do you want to have?
How many spine switches? What kind of redundancy you want to have with, with them between them? Now we answer these questions.
We apply this module, uh, it's applying. So what is happening, this module is basically based on that inputs is generating, uh, generating objects in network controller, like IP addresses, like topology information. You'll, you'll see this in, in a sec.
Uh, we can see that these numbers are changing. So controller is being populated, and now it's done. Now, uh, uh, I'm going to start.
Uh, so one second. Okay. So in this demo, we're not using physical hardware.
We're using, uh, a simulation. So now, now that I've populated network controller, and I explain a second what, what I've done here, I'm going to, to create a digital twin of the network that I've designed in the network. That process takes about five minutes.
So I will start that. Um, this digital twin in this example is, um, is using our own technology we call Cloud sim. We've developed it in-house.
Uh, we're using PMI for this. It's a, it's a basically, you know, smart script that reads all kinds of parameters from controller and creates lots of, lots of VMs in, in a very powerful hypervisor, uh, basically in a set of powerful hypervisors. So this way we were able to simulate really, really large networks, like 40,000 links, networks.
Like we cannot create that in the lab. Okay? So that simulation is in the process.
Now, let me walk you through what we've got in, in the controller. So we've populated inventory, and now we've got entries for servers, HGX. Uh, these are, uh, you know, simulated, uh, GPO servers for servers.
We have a little bit of data, their name. And, uh, basically this JSON file. This JSO file basically describes the host configuration on the server.
This is for the case. If customer opts for using our plugin to also automate network configuration on the server server. It is optional, but it's very useful, uh, for switches, uh, because we do a lot more with switches.
We, we have more information like what operating system to install, what is number, the switch should have I IP address management IP for the switch. Uh, then we've populated ipam, and this becomes kind of source of truth for, uh, IPM. Obviously we can have, uh, you know, additional outside, uh, source of truth and you can connect with APIs.
But within net, this is your source of truth. So these ips describe Which kind of, so, so tool you can connect. So are you able to, with GI with, uh, so many control version?
So, uh, you, you can, uh, you can, uh, manage net through web, rest a and Terraform, and every source of truth has something, something similar. So for DevOps engineer, it's really easy to, to write like four, four lines of code. Okay.
That takes from their end and generates. Yeah. But let's suppose that someone is touching the, the interface.
Is there a kind of alignment between the interface of same kind of drift detection from the social truth to the, uh, the infrastructure element that is supplied? Uh, no. No, not really.
This is more like this source of truth. So idea here is not to, to, to replace your existing source of truth, but idea here is that net needs to know what are your IP addresses. Okay?
Like we, we, we like have to know that. Okay? Um, so this is where you provide, if you don't have source of truth, this is it.
If you have another one, perfect. Just, just let us know what is it? Great question.
So, and then we have this place called topology. Now this topology is, is the topology that we've just created with that module. And, uh, this top part is backend network.
It's kind of reversed. Vacant is in the top, and the front end is in the bottom. But, so, and you can see that this part has like lots of links connected, like crazy number of links.
This is, this is how this ai, um, beacon networks are lots of, lots of network connections. This bottom part is front end fabric or north south fabric. Uh, so there's less switches.
Uh, these four nodes here, these are our soft gate nodes. They are connected to lift switches, uh, like a router on a stick. Uh, these small boxes in the bottom, these are GPO servers.
Uh, so they have eight nicks. You know, each nick kind of connected to HGPU, and, uh, each nick connected to backend fabric. Uh, so Can the users select whether they want a rail network or a regular fat tree, or do you default to something?
Yeah, that's a, that's a great question. So there are different rail optimized designs in the world, right? So in the, uh, you, you need to describe your design here in this topology.
So Naris knows how to configure, but I, I, I've generated this in, in just like in a couple minutes using that module. Now, idea of module is that it, some designs are very prescriptive. Like Nvidia, spectrum X is very prescriptive by design.
It doesn't allow you to, to like connect in, in a different way. And because of that, that's, that's how our module works. Mm-hmm.
If, if you're, if you're doing spectrum X, you have to do like this way, uh, but if you're deploying a design which is less prescriptive, uh, y your, your questionnaire will be bigger. Uh, but you can d define whatever you want. Um, yeah, you can s This is Tom.
I've got a quick question. You mentioned the Spectrum X module. Um, I know that you probably have modules for the different, uh, switch manufacturers and, and systems that you work with.
Are those ones that net develops and then the only place to get them is from you and your team? Or do you take submissions from the community for supporting potentially devices with different functionalities that you incorporate in there? Is it kind of a, a closed system, or are you more open to other people kinda helping you figure these things out?
Yeah, great question. Uh, so technically the, the, the code, uh, of the agent that runs on, on the switch, is it, it is closed code? Uh, it's, it's not open source.
Uh, but we, we are very, very open in general. So when it comes to spectrum X, uh, because it is prescriptive, uh, the, the way we've built a lot of things is based on, uh, collaboration with nvidia. So there, there is number of specifications coming from nvidia.
It's, it's like hundreds of pages specs that are available to us as a partner. So this agent is doing like literally that. Uh, now in, when it is less prescriptive, when there's more freedom, we, we are very open-minded.
We, we are happy to listen to our customers, and we have cases when customers said, Hey, in my particular use case, this thing should happen in, in, in a different way, like x, Y, Z way. Uh, we, we are open to make this modifications in the agent if required, but customer themselves usually cannot do. But there's other part of it.
If, if we, if it's kind of a question that makes sense to kind of open up to, to a customer for modification, we do that too. In some cases, we're saying, you know, this is matter of preference, not, uh, best practice. So we will give you a way, uh, to make this modifications.
That makes sense. Thank you. Yeah, thanks.
Great question. So, okay. Uh, by now, you know, the, the network has converged.
So what happened? The switches came alive. We installed operating system ulus in this case and installed our agent.
And here we can see, uh, that monitoring, uh, started automatically. So we're, we're not only automatically confi generating configuration on the switches, but we also, uh, configure monitoring, uh, and we monitor for like, like typical things that engineers would love to have monitored. Like is, is wiring consistent with a topology?
Like, or is there any miswiring, uh, is like, are your fans spinning, right? Are your power supplies okay? Are your, you know, this, this BGP sessions that we have configured on switch to switch links, are they up, are they getting prefixes?
So we monitor for things like that, and if something is wrong, we let you know. Question. Thank you.
Um, do you look at every individual link? I see 4,732 interfaces means about half that many links. Yeah.
Yeah. You each one and to, to verify that it is where it's supposed to be. Yeah.
Yeah. We do that and we've tested this on, uh, on like 40,000 links. Okay.
And when it's not, do I get a picture? Uh, yeah. So you, you, you, you are getting a message saying like, switch port one happens to be connected to switch nine, switch Port X, where it is supposed to be connected.
So how it should be and how it is, it, it's okay, I can simulate that in a, in a simulation, because simulation always makes connections, right? But Yeah. But, but with real customers, typically, when, when we launch a new customer, when, when the physical is done and we run this process, uh, we, we, we usually discover that approximately like 15% of cables are mis wired, and then they go and yeah.
When, when it's like thousands of cables, right? 15 percent's a lot. I mean, was just gonna say that.
And isn't, isn't that one you can't get good help. It's hard to do on your own, right? I'm going through a declarative approach.
I'm trying to push code and I don't get any drift information back. Mm-hmm. Part I have to go fix it, then rewrite the code, push it again, and see if anything's broken.
So the, the approach is very, seems like you're solving the approach problem, right? For the customer. Yeah.
Thank you. Uh, yeah. I know myself and most, uh, people on leadership are coming from network engineering background, and we've tried to solve this problem for, you know, our prior employers, and we were like, it's, it's a hard problem.
Like it's, it's a, it's not a, it's, it's like entire team should focus on this one thing to solve it. Okay? We've got the fabric up and running and, uh, let's connect to some of these GPO servers and see what, what's happening there.
Now, here on my left window, this is my controller, and because I'm running the simulation from the controller, uh, the way this simulation, uh, platform is designed, I can use kind of a backdoor to, to connect to the simulated servers. Like in production, this backdoor would, backdoor would not exist. And this server, this host zero in production would not have operating system by now, right?
Normally we provision network, then your compute platform provisions operating system. But here, for the sake of demo, we, we, we bring up machines with operating systems. So you can SSH and ping here and there.
Now, on this machine, we have this script we call cluster pink. It takes U ID and the, and the host ID as, uh, as input. And it calculates IP addresses, uh, based on no well-known IP address plan and executes parallel pinks.
So in this case, I'm pinging every rail, I'm pinging production interface, and I'm pinging IPMI interface host zero. Pinging host zero. I'm pinging myself.
Obviously it works, but when I ping my neighbors, it doesn't work. Why? Because I don't have any tenants onboarded, I don't have any overlay network.
What I've had, what I've got so far is just fabric underlay up and running. That's it. Now, let's say, you know, a tenant comes in, maybe that's a tenant coke.
And this is how I create a, you know, VPC or isolated cluster. So I select a template and then I will show soon what this template is. I will explain all this how what hap happens, uh, on underneath.
But basically for first tenant Koch, I'm selecting these four servers, host zero to host three. And that's basically the list of all machines. We have add, add, and, uh, for tenant Pepsi, uh, I'll do the same, similar, uh, same template.
That's a template for, uh, AI clusters for VPC, I'm saying create new, I could have said use 10, uh, VPC for Coke, but that's not the intention. So I'm think create new here in this list. I don't have servers from zero to four 'cause they're busy, uh, by Koch.
And I will give four servers to this next tenant. Now, this, both clusters are in provisioning. While it's in provisioning, which takes like a minute or two.
Uh, I'll explain what's, what's happening underneath. So I've just used this template called template one. Now what is this?
See what's happening in networking, typically, you, everything this vlan, VR Fs, all network configuration is tied to switch ports. And we, we've seen switch ports on this topology, right? This is network engineers world is switch ports.
Network engineers work and deal with switch ports. Now, cloud operators, they don't want switch ports. They don't wanna say for this cluster includes switch port ones and, and like 500 switchboards.
They wanna, they wanna operate in terms of servers, server one, server two, server three. And we have construct for switch ports too. Not a problem.
Uh, this is kind of more advanced version of that construct. Now, this is the template, which is basically template between, between platform engineers and network engineers. It, it translates server one, server two, server three into WA switch ports.
Uh, these templates are fully editable by customers. Uh, this is just an example one where I'm saying basically go and find switch ports that are connected to EH one to ETH eight of any server in this cluster. Basically for this cluster.
I said server first for servers. Basically the template is saying, go find which switch ports are these cables connected? 'cause you know, the topology and go and configure all the right VR fs via excellence on that, uh, appropriate, uh, switchboards.
The next section says, you know, for this other ports connect another vxlan and then connect, create this third VXLAN for this other ports for management, uh, based on that template. Net creates kind of sub objects inside. This is object called vnet, which is virtual network.
It can be vxlan, can be vlan, can be InfiniBand, can be any kind of virtual network. So we create three of those for each. In this case, it's a vxlan.
Uh, in one case it's L three vxlan. In one case it's L two vxlan, uh, um, anyways, three, four Coke and three four Pepsi. And you can see that they are tied to different VPCs, 41 and 42.
That's what makes the isolation possible. Now, if we go back to server cluster, we'll see that both provisioned. Let's see if my host one can p its neighbor.
See every, yeah, so it, it works. Uh, I can p uh, host two, host three, but the next one belongs in, in a different cluster. I cannot ping it.
And, uh, to, to prove the point, I will exit this one. I'll connect, uh, to that other cluster. Uh, let me go and find that cluster.
Uh, uh, basically one more. Yeah, host four. So if I do cluster ping here, uh, I cannot ping host zero.
Obviously that belongs to Coke. And now I'm Pepsi, but I can p uh, host three, no, no, Four, Yeah, sorry. Host three is also cook.
I can ping host five. Yeah. Uh, 6, 7, 7.
One more, right? Oh, no, one more is my myself. So host four.
So basically Pepsi's host four, host five, six, and seven. So 4, 5, 6, 7, all works. Okay.
Uh, so, okay. Uh, so this part, uh, uh, this, this part was for isolation. Now we've got two tenants, Coke and Pepsi.
They are isolated. One, you know, internally, training can happen. Nickel can run coal backend, front end, everything connected.
But how do I connect to this cluster? Cluster works. Congratulations.
I'm data scientist. How do I connect to this cluster? So let's say I'm data scientist of cluster K, and, uh, cluster K has this, uh, this private IP address for, for basically connecting to this machine.
How do I connect to it? So, uh, we'll create, uh, a construct that typical cloud call calls, uh, you know, elastic ip. Um, so we'll select the data center.
'cause this thing can manage multiple data centers. We need to select the VPC because IPS can overlap across different VPCs. So in this case, we select VPCK.
This defines who can access. So in this case, entire world. And here, this is basically pool of public ips that network engineers of this AI cloud defined for this purpose.
A pool of IP is for elastic ip that's defined in our ipam. And here I will paste this, uh, the actual private IP of the server. Click add.
Now I need to copy the public ip. Okay, come on. Copy.
Okay, so I'm pinging it and, uh, timeout, timeout. Provisioning is in process, and once provisioning was done, I can see that this IP is responding. So left hand side is my controller.
I'm using, uh, backdoor right hand side. This is pure internet. I'm using wifi connected here.
So this pink is, goes through this internet. Uh, I was wondering why it was 30 milliseconds, right? As evidenced by its Delay.
Yeah, actually, this simulation runs in Santa Clara. It is, it is located on, in our data centers is, uh, basically in this zip code. Mm-hmm.
Of course, it goes to like bunch of places before it makes to the data center, It has to go through the field day. That adds bunch of latency as well. Yeah, it's wifi.
Yeah. So, so just to prove the point that I'm not just pinging one ip, but I'm can, I can connect to this machine. I will run, uh, netcat, I will run like basic service here and will say hi from AI field infrastructure field day.
And, uh, so hopefully I'll be able to connect from my laptop to that TCP port. So see, this basic chat kind of proves that I'm not just pinging an ip, but when I connect to that IP that traffic goes to, into like literally that machine, um, that's, um, so any questions about this? Uh, two questions.
So I did hear you correctly where you can use the same address space in multiple domains. Mm-hmm. Okay.
And, uh, since you've just demonstrated that you can talk on an arbitrary port, you actually have some sort of fire wheel walling capability in it. Uh, YYY yeah. You know, because the, the way the, that soft gate note talks to the fabric, it, it, it can distinguish between different VPCs and that way we're able to kind of extend the notion of VPC into the soft gate itself that allows us to achieve multiple things.
One is different. VPCs can have overlapping ips, and soft gate is not getting confused. Another important thing is scalability, because this is designed for, for like potentially scaling to like millions of VPCs.
And, and we, uh, allow customers to run like multiple soft gates like horizontally 8, 16, 12, any number. And, um, scale, scale scaling with that, uh, VPC awareness is a, is a pretty challenging task. And we've tried lots of things until we found the formula.
So, you know, we we're kind of, you know, grouping different VPCs into like different groups. So, so that way we can manage the, the scalability. Okay.
But I, I'm sorry. Maybe can you also do firewalling at the port level to per to, yeah. Okay.
Yeah, yeah, that's a great question. So we also have, uh, acls where we can do just, just like permit deny. I wanna chew Off, you've answered my question.
I didn't wanna chew up your very few remaining minutes. Uh, so we, we also have, uh, elastic load balancer as I explained, which is again, um, VPC aware. Uh, so that's for like, this is typically used in inferencing use case when you want traffic coming from the internet to be balanced between multiple, uh, servers.
Now, all this example was, uh, was based on ethernet, ethernet network backend and frontend ethernet. Uh, I wanted to also quickly show how we do this for Infin bank networks. So this other controller here, uh, has, see this has half of that topology only ethernet.
And we have UFM here. So, so this, this is not a simulation, this is connected to physical UFM, you know, physical infin bend, which is we can simulate Infin bend. And, um, uh, I have some pings running.
Uh, see this top part is pinging from host zero. It's neighbors, uh, on left side. This is ethernet, this is pinging on InfiniBand, and this is host three pinging neighbors, ethernet and InfiniBand.
Now, they, again, each host can ping itself, cannot ping others because there's no tenancy, uh, there's no server clusters. So let's create first, uh, let's create Coke here. This is different cluster that Coke and Pepsi have showed that that was in another cluster.
Uh, in this case, I'm using different template called IB template for InfiniBand. So this is one. And, uh, yeah, just to show you in, in UFM, you, you can see here, I don't, we don't have any pickies yet.
Uh, Pepsi second, uh, tenant, uh, again, same template. I'm picking two other servers, two, three, both are being provisioned while it's provisioning. I'll show, uh, the template.
So I've explained the con, uh, the idea of this template in previous example, and you can see that ethernet part is the same. But here we have different part that is specific to Infinity Bend. It is saying for East West, previously, east west was ethernet in that example, right?
In previous example, east west was all ethernet. Here, east west is in ban. So we have different setup.
In this template we're saying, Hey, talk to UFM, which has, um, you know, identifier called UFM lab and use pickies automatically. So basically Netflix, you figure out what IES to use. And, um, once it's provisioned, we can go and refresh UFM and we can see that net just created two IES and net provision gds because net learns these GDS automatically and keeps in, in its database.
And if we go here, we can see that host zero can ping host one on ethernet on InfiniBand, but cannot ping others. And same from the Pepsi perspective, ethernet operational and InfiniBand operational. Uh, that's pretty much it.
We've got minute and a half almost Question. Uh, when you just deployed the servers and you were showing the mapping in the template to E one through E seven or E eight. Mm-hmm.
Uh, were those discovered by an agent running on the host? So you can map the ETH number to the MAC address for proper connectivity? Great question.
So EETH 1, 2, 8. These are logical numbers significant to, to net. Mm-hmm.
So if you go to your server, you may type IF config. You, you may see different names. I will.
Now, for modeling your obstruction, you need this, these normalized numbers. E th one to eight, you need this. But, uh, see when it comes to, uh, topology cabling, uh, validation, and we have EH one to eight, and on the server you have EMP 47 or whatever, right?
Uh, there, there's a, there's a place in net you provide mapping, uh, in, in, in the server profile. Uh, i, I can show documentation for this, but, uh, you, you provide mapping between like ENP 47, your real format and your logical format. And net net.
Based on that, NARIS continues doing LLDP and seeing real interface names. And based on that, naris will tell you, Hey, uh, logical, e th h one happens to be connected to logical ET H three and, you know, on the server physically, which cable two change with, with which one.