Introduction to Multi-Tenancy & Network Automation for AI Infrastructure Operators with Netris
Netris helps GPU-based AI infrastructure operators automate their networks, provide multi-tenancy and isolation, and offer essential cloud networking features like VPCs, internet gateways, and load balancers. Netris focuses on network software designed for AI and cloud infrastructure operators because the growing popularity of AI necessitates specialized networking solutions to handle demanding AI workloads. Netris’s technology is particularly well-aligned with NVIDIA’s networking offerings, which are based on the foundation of Mellanox and Cumulus networks.
The presentation highlights the importance of dynamic multi-tenancy for maximizing the utilization of expensive GPUs. Netris provides “cloud provider grade network automation software” that allows AI infrastructure operators to achieve security levels comparable to physical isolation while maintaining software-driven speed. This solves the problem of manual network configuration, which is time-consuming, error-prone, and doesn’t scale. Furthermore, Netris supports cloud networking functions like Internet gateways, NAT gateways, and load balancers, offering a complete solution that addresses the need for secure and flexible network management in AI environments.
Netris’s solution is built on three key pillars: VPCs for isolation, cloud networking functions for connectivity, and fabric management for network operations. They manage both Ethernet and InfiniBand fabrics, providing operators with a single pane of glass. For InfiniBand fabrics, Netris integrates with NVIDIA’s UFM controllers. On the Ethernet side, Netris acts as the fabric manager for several vendors, including NVIDIA, Dell, and Arista, automating the management of network switches and streamlining operations. The goal is to offer a comprehensive, integrated network automation platform tailored for the demands of AI infrastructure.
Presented by Alex Soroyan,CEO and co-founder, Netris. Recorded live in Santa Clara, California, on April 24, 2025, as part of AI Infrastructure Field Day. Watch the entire presentation at https://techfieldday.com/appearance/netris-presents-at-ai-infrastructure-field-day-2/, https://techfieldday.com/event/aiifd2/, or https://www.netris.io/demo for more information.
Transcript
I'm Alex Soan. I'm, uh, CEO and co-founder of net. Uh, so net we're, uh, we're based here in Santa Clara, California.
Uh, we're, we're for, um, you know, our customers are everywhere. So we also have, uh, team members in Europe, in India, and, uh, Sydney. Uh, we're focusing on network software for AI and infrastructure of, uh, for AI and cloud infrastructure operators.
Um, we've been doing this since 2018. Uh, you know, in 2018, AI wasn't as big as today, and most our customers back then were tier, tier two, tier three cloud providers, uh, private cloud operators, enterprise cloud operators. But some of the technologies that were use were, were being used for cloud providers in networking are, uh, also relevant and applicable to AI clouds.
And, you know, with, with AI becoming very, very popular and, uh, lots of organizations are now building infrastructure for hosting ai, uh, you know, Nvidia, uh, NVIDIA Networking is basically former, former Mellanox. They acquired Mellanox, they acquired Ulus Networks. And, uh, we've, we've been working with Mellanox and Ulus for a long time.
And that way when, you know, technology is created by Mellanox, and Ulus became basically the foundation for NVIDIA's networking, we appear to be kind of very much aligned with the, with the stack. Uh, that created a good opportunity for us, uh, to work closer with Nvidia and, uh, uh, you know, refine our technology for their AI infrastructure use cases. Uh, you know, primarily around, uh, multi-tenancy and automation for cloud providers.
I will tell more about this. So, you know, although, uh, organizations that are running AI infrastructures that are coming from different verticals, we work with, you know, neo clouds. We work with infrastructure as a service providers pass with sovereign AI operators, with telcos.
They, these organizations, they have different business models, but from our perspective, from from networking, from technical perspective, what they're building is very similar, uh, from the stack perspective. Um, and, you know, they all are, you know, heavily investing in infrastructure to run AI workloads. And, uh, AI workloads are very demanding.
So the entire stack required to be improved, starting from power, you know, cooling, compute, hardware compute software networking, hardware networking became way more complex in ai. Uh, so is networking software. Now, this, this AI infrastructures, uh, they require, they have many components.
We focus on just one component, the networking. This is area where a lot of organizations are having the most challenges. Uh, and, uh, you know, one company cannot solve all this.
Uh, all, all these components, all these pieces of the puzzle, even Nvidia being this amazing large organization, even them, they don't have all the components. That's why they, that's why it creates opportunity for partners like us and others. So, um, these GPUs that most organizations are willing to host and, uh, use are expensive.
And because of that, the path to return of investment is through dynamically sharing this hardware resources with internal and external tenants. So that requires dynamic multi-tenancy, what we call, and, uh, you know, because of that, networking becomes kind of complicated. Uh, but to understand the problem and to understand what we're, uh, how we're solving the problem, uh, I wanna do a little bit of kind of a simple history class here.
So, on this, on this, you know, business school 1 0 1 kind of comparison. Um, one, one X is ba basically utilization of GPOs. And the other axis is basically safety.
Now, why this matters is because when these organizations create AI infrastructures, uh, they need to share hardware with internal tenants or external tenants. If they are cloud providers, obviously they're businesses to share with external tenants. And some of the earlier generations of kind of new wave of AI cloud operators, were, were do, were doing physical segregation.
So basically different clusters were separated from each other physically, of course, that provides maximum security. There is no way to break in from one cluster into another one if they're physically separated, right? But the problem is that it takes a lot of time.
And when you're building a cloud business, sometimes you want to isolate resources for one hour, one day, one week, and some manual method doesn't scale. And if it doesn't scale, this very expensive G ps are basically idling. So what, what what we've seen, uh, no, historically, some of these, uh, operators, they started to share the network, basically build one infrastructure and manually configure the infrastructure, manually reconfigure the infrastructure to make it shareable across different tenants.
Now, that does the job, yes, but the problem with that is manual. It takes time. Uh, it, it takes some times, weeks to, to configure, you know, hundreds of switches.
And, you know, the AI infrastructures have lots of switches, uh, and there's high risk of human error. You make one mistake while you are creating a new tenant and you misconfigure something, you can potentially break network for other tenants, and their training workloads may suffer and the customers will not be happy. So cloud providers realized that, obviously, and they started, you know, you know, uh, automating things in-house.
Now, in-house automation is very challenging. You know, we've built a product of, for network automation. We know how challenging it is.
It, it took us years to build, and we were focused on this one thing. But also, the big problem with the building automation in-house is that the organization that is building in-house automation, they learn from their own errors versus a specialized product is learning from the data, coming from lots of lots of customers. So these are obviously not great solutions, and, um, they're risky.
They're taking a lot of time, not a solution for a cloud provider. Now, there's this other extreme, you know, using Kubernetes and containers and namespace to, to make isolation, which is very fast. That is API driven, that is software driven.
It is fast, but the problem with that is that there's no isolation on a networking layer. So that creates a security problem. And, you know, it takes just one container escapee vulnerability to jeopardize the reputation of a cloud provider.
Security is very important in the, in AI because these customers are training, are dealing with a data where, which is highly, highly sensitive. Now, basically, AI infrastructure operators want something that provides a security level that is comparable to physical isolation, but is available at the software speed. So basically, some sort of, you know, uh, we, we call this cloud provider grade network automation software.
This is what, uh, AI infrastructure operators need. And this is basically what net is focusing on. We're a, we're a software company that provides network automation, obstruction, and multi-tenancy software for those who are building AI and cloud infrastructures.
Now, uh, I like to explain this using this method of three pillars, which are kind of, what are some critical pillars when you are building an AI infrastructure. So pillar one is this notion of VPC. You know, anyone who used any public cloud, like AWS or anything knows what VPC is.
It's a unit of isolation. And when AI infrastructure operator is building a cloud, you need this, the same obstruction model that works in the cloud. But, but you need, so these cloud providers, they've built law lot of software in-house.
They're not using any, uh, of available network automation techniques. They have built this in-house, right? But it took them many years.
Like they, this first generation of providers, they had that time. The new generation of providers, they wanna go live today. Like when we work with customers, they say, we wanna go live this, this month.
And, uh, it's, it's like every month someone is going live, the competition is growing, and new cloud providers want to want to deploy fast. So notion of VPC is critically important. Pillar, uh, you know, networking, traditionally, everything in networking was tied to switch ports or in Infiniti Bank to GUID.
Now, as a cloud provider, when you, uh, you know, when you are creating, when you are defining your tenants, you don't want to deal with switch ports. Your, your, your customer facing portal, which talks to your network automation backend doesn't, doesn't know what are switchboards. Your portal knows what servers are, what GPO servers are.
So basically, when your customer facing portal creates that, uh, backend request, it wants to say, Hey, let's create a VPC and let's include not one, not two, not three. In this VPC, you, you want to create a basic, you want to make a basic request and get, you know, all the details configured. So that's notion of VPC, that's pillar number one.
Uh, and by the way, any questions feel free to, uh, pillar number two is what we call cloud networking functions. See, with, with previous pillar, previous pillar allowed to create this isolated environments, uh, you know, isolate on ethernet networking, finna bank network everywhere. But these are islands.
Now, now, once you have that islands isolated, protected from other tenants, you, you, you have the need to, to let this islands to talk to internet, for example, maybe you need to download a package from the internet. So you need something like an internet gateway to provide internet connectivity. If you are, um, you are as a, as a cloud provider, you're probably using a shared storage system, kind of very common for cloud providers.
And you wanna a ability, you wanna, you need constructs that allow to say, Hey, my tenant, number one, my tenant maybe is tenant. Koch, for example, wants to access storage. So let's create access to storage.
Let's, let's let this other tenant talk to storage. But without two tenants talking to each other, um, internet gateway was another example. And if you're running inference workloads in this AI cluster inference workloads are used by users in the internet, right?
Think, think charge GPT and their customers. So you need, uh, some sort of load balancer, which will take traffic from the internet and will balance between these different machines, like in the cloud, in the cloud, inside your, your VPC, you can create internet gateways, not gateways, load balancers, elastic load balancers. And if you are a cloud provider yourself, you need methods to provide this constructs yourself.
So pillar number two is cloud networking functions. And, um, pillar number three is, look, these first two pillars are kind of nice and fancy. They, they provide cloud providers these constructs, which help them drive their business.
But let's not forget that all this beautiful cloud is running, uh, on, on the physical network. And there are network engineers who need to take care of the network. They need to deploy the fabric, upgrade the fabric, downgrade the fabric, perform maintenance, learn if something is not good with the health of the fabric.
So for that, you need a system that does your fabric management. Now, if these three critical systems are, you know, different components differently developed or like, I don't know, coming from different vendors, what can happen? All these three components are trying to configure some of the same things.
What can happen if two systems are trying to make changes in the same front end network or in the same backend network, right? It's, it's, it's, it's a, you know, it creates a risk for logical disconnect and, uh, it can create error risk. Now, in our case, these three components are not really three components.
It's a one product where the three components are tightly integrated into it, into each other. They're using same database, same data structures, and there's no risk of, uh, creating conflicts. Uh, you know, it's like when you drive a car and your wife is saying, take a right, and JU is saying, take left, but you go straight because you know that you're running out of gas and you need a gas station.
So, so that's what NET does for, uh, for network automation for ai, uh, cloud operators. Now, um, uh, so, you know, AI came from the HHPC world and InfiniBand, uh, networking was big in HPC, even though there's a lot of developments in ethernet. Uh, and, and we see a lot of ethernet developments.
It a lot of ethernet based, uh, clusters. Um, in our experience, uh, based on what kind of customers we work with, approximately half of the customers are using InfiniBand in their backend fabric, and they're using ethernet in their frontend fabric. Uh, where, uh, for InfiniBand fabrics, we work through Nvidia UFM controller.
So, like I explained in previous example, two systems should not edit the same system. We, we don't want break our own rules. We don't want to go and touch the switches directly.
So we talk to UFM and UFM talks to infin and switches, versus on ethernet side ethernet doesn't have something like UFM, uh, and on ethernet side net is the fabric manager on ethernet. We, we manage the switches, we load agent on every switch, and that way we auto automate management of the switches. So from, from customer perspective, there is single pane of glass that takes care of entire stack, but, but behind the scenes, we manage ethernet and we manage Infiniti banks through U fm.
Uh, question A question on the, the ethernet, uh, management plan. Are you making an assumption that it's a multi-vendor ethernet or is it a single vendor where, because some of those vendors do have fabric management. So would Nitris be talking to one of the proprietary networking vendor fabric managers, or you're assuming multi-vendor?
Yeah, that's a, that's a great question. Uh, we, we support multiple vendors as, as of today we support Nvidia, uh, we support Dell, we support H Core and arisa. Now how we support, uh, we, in, in this four cases, we are the fabric manager.
Nvidia H Core doesn't have fabric manager. We are the fabric manager. We are NVIDIA certified for this job, for other vendors.
There's no certification. But we've been doing this for a long time with live customers. It's not just theory, uh, with some vendors.
They also have fabric managers. Like if we take Arista, Arista has, uh, cloud vision, which is a very decent product. Uh, we do have some customers who use our technology in combination with Arista instead of Cloud Vision.
Why? Uh, because we, we bring, first of all, we bring this, you know, independence. So these customers may not only have our study and have multiple vendors, right?
And we bring some of this cloud functionality that are kind of more convenient for customers. Um, and, uh, so we, we also do lot, lot of e ethernet, uh, ethernet cluster. So this, this case, um, uh, we in, it's, it's a customer case study where, you know, we worked with Scott Data, uh, their, their AI cluster is using ethernet both on backend and front net networks.
Uh, backend network is based on NVIDIA's, what, what Nvidia calls Spectrum X. It is NVIDIA's ethernet based, uh, fabric. It's kind of alternative for InfiniBand.
It is more scalable than InfiniBand. Uh, and, uh, you know, it's ethernet. It, it has some, uh, benefits.
We can talk later. Uh, front end is also ethernet and, uh, you know, both fabrics. They don't have fabric manager other than net.
So we manage the fabric. We also manage this layer, uh, I will explain later. We have this cloud, uh, net networking functionality layer we call Soft Gate, uh, which provides, which makes, uh, possible all that not load balancing, uh, functions that switches do not provide.
And customers get this single pane of glass for entire management. Uh, similarly, uh, we, we have customers, uh, so for example, if we look into a solution for, uh, for this customer use case, um, so Tensor Wave, uh, is running a MD based cluster. Uh, they're one of a MD based neo clouds.
Um, their backend network is based on Edge Core and Arista switches. So it's a mix. Uh, and the front end is based on edge core switches.
Uh, any, any questions around this? Not yet. Sure.
Is that a, so, so from an Edge core perspective, what would be the network operating system on the Edge core? Is that specifically Sonic or something else? So that's a, that's a great question.
For, for every switch we integrate with one operating system that works the best on that particular vendor on Edge Core, that's Sonic. Mm-hmm. Uh, for our restart, it would be obviously EOS.
And uh, for Nvidia it would be humus. Some vendors, they support multiple operating systems, but we do not integrate with all of them. We just pick one, which is usually there is one that works the best with that vendor.