How the Retirement of Ingress NGINX Signals a Shift to AI-Native Gateways
Lin Sun, head of open source at Solo.io and a member of the CNCF Technical Oversight Committee, explains why the retirement of the Ingress NGINX controller marks more than the end of a widely deployed project. As cloud-native architectures evolve, IT teams are being pushed toward next-generation gateways purpose-built to support AI agents, real-time inference workloads and increasingly dynamic application traffic patterns.
Transcript
Hey guys, thanks for the throw. io, and we're having a little chat about, well, this NX ingress controller that's about to be retired and what we should all maybe be doing about that next. Lynn, welcome to the show.
Thank you. Thank you so much for having me, Mike. I think at first glance when people looked at this announcement, they're pretty much like, well, I should get a different Ingers controller and I'll go from there.
But there's other ways of thinking about this thing, right? The whole space has evolved a little bit. There are gateways now and other things that we might use.
So what are my options to go forward here if I'm gonna have to replace that controller? That's a great question. I, I think you're ly right.
I actually first heard about this, uh, at Cube calling Atlanta, right? So the team kind of set, uh, uh, uh, archive date, remember it's approaching end of next month, uh, right, and they said they're not going to maintain it after that date. So, uh, timing is everything.
So, as user are looking at replacement of ingress Gin X, I believe they want to look at a couple of categories. So first, uh, the maintainers, if I remember correctly when I was reading the announcement, the maintenance definitely recommends stay on the standards, which is the Kubernete gateway, API, right? So that's the recommendation, uh, from the community, the cloud native community, the maintainers, um, C network, and even the maintainers, I think, uh, who put out that post out, uh, uh, ingress Next, uh, retirement.
Um, so I do think you want to pick out a implementation of Kubernetes gateway APIs. So first of all, also that's important check marks because you wanna make sure you're adopting open standards. I think the second check marks is you want to look at, um, the features of what you need with your, um, what you have today with whether the new one you're gonna pick, supporting those features.
Does it have the conversion tool right, to help you easily convert from your existing, um, configuration to the new one with Kubernetes Gateway, API, and potentially some projects specifically in, um, uh, extensions. The third thing I think is the performance, uh, is always a big thing, right? As people looking at, uh, another option for replacement for INE X ingress, you know, how is the gateway performing in terms of the throughput CPU memory latency in terms of the time to translate the configuration and make it live, uh, in the proxy.
So those are the important metrics people should look at. Uh, the first thing I would say is the community momentum, right? You certainly want to pick a project that's in a foundation, uh, which has a better chance to success, and also check out the diversity, right?
Um, whether the project have maintainance from multiple different companies, you know, whether the project is growing steadily and healthy, whether the project team, uh, are responsive to your questions, if you ever have a question. Those are the four things. I think it's, uh, it's important to check it out.
Mm-hmm. I also think, you know, it's maybe not always like for like, I mean, I had a controller, but if there's something called a gateway that's a little more elaborate, maybe I go that route because the gateway does more than just the controller, and I have to think about it a little more broadly. Yeah, I mean, there's really two pieces of the puzzle, right?
So one is the, the proxy that's actually doing the work, uh, which is Nginx proxy, right? Uh, so the other piece of the puzzle is the control plane who is programming that proxy. So whether you stay with Nginx as your proxy, or whether you pick Envoy as your proxy, or whether you pick a new high performance proxy like Asian gateway, for instance, which is completely written from scratch, um, based on rust for AI to adopt the rapid change of AI agents and MCP, uh, all these new protocols to allow us to innovate and iterate a lot faster.
So, um, so the, that's another important piece. As people looking at their requirements, are they looking at, you know, adopting, doing the same thing as they were doing in the past 10 years? Or are they looking at adopting MCP, you know, adding agent conversation into some of their, um, front facing, uh, API with their users?
So that's going to be critical. Are they consuming large language models and adding intelligence to their applications? Is this also maybe a good time to start thinking about things like, I don't know, IPV six?
I mean, should we migrate as part of this thing? Or how, how ambitious should we get? That's A great question.
You know, honestly, I do think, uh, you want to be, uh, you want to be thoughtful as your migration and probably not taking too many risks instead of one time. It's like you don't put your X in one basket, right? In a sense, um, the IP V six migration, I do think you could potentially migrate, uh, relative smoothly once you decide what is your replacement project, right?
Like, you don't have to migrate to IP V six in day one once you decide this is the new project you're going with. The challenging, though is you don't want to have too many variations as part of your migration. Um, you want to identify what is the critical parts of your migration and then face it out.
I think that's probably a more viable approach because we're not talking at, uh, about you have a lot of time to migrate. Um, 'cause the, 'cause the end of March is approaching right after that. You will not see like fixed packs or, um, CVE releases, uh, for ingress nginx.
So there is a present timeline approach. So I personally wouldn't recommend unless you are really like adventurous and willing to, you know, like eat all your sandwich together. There you go.
It might be too ambitious. It's good to have a long-term plan, but maybe not all at once. That's right.
Yeah. I think we have seen, at least in the last year, a lot more clusters and fleets of clusters. So have we reached some sort of tipping point now where the networking between those clusters has become more critical?
Because I think a lot of people, you know, if I had one or two clusters and they were running stuff in isolation, I didn't think too hard about it, but I think now we've crossed some sort of Rubicon there. Yeah, I, I definitely agree with that. And the gateway is a critical piece of that puzzle, right?
Um, so we've seen, uh, at least, uh, working at solo and also some of our open source users, we've seen tremendous interest around beyond one cluster, right? 'cause typically people can roll one single cluster themself, uh, follow open source documentation where they need enterprise help a lot of times is when they have more than one cluster, and when they actually having cluster talking to each other, they want to either do failover or secure communication or join their, um, different trust domain together and still have that boundary of trust. Um, so that's, uh, where we think Gateway can play a huge role in this type of scenarios where the gateway can help you to apply policies, connect traffic across different clusters seamlessly, and provide that visibility observability as the traffic travels through, uh, multiple, uh, different clusters.
Um, what is your best advice for managing all this stuff today? And I'm asking this question because historically in a lot of larger enterprises, there's been this networking team and sometimes a separate security team, and then there's the DevOps teams and the guys who build the applications. Should that be all centralized or, you know, do we need to figure out how to kind of manage all these things in isolation and the networking people need to take more responsibility for the network controllers and gateways and all the things that connect those Kubernetes clusters, not only to each other, but the rest of the enterprise?
Yeah, I think it really depends on how many, uh, how large is the organization. I've seen sometimes when the organization are not super big, the network team and the DevOp team, uh, the platform team could emerge, right? But if your organization, the number of cluster, the number of the teams are big, I can see a clear separation between the two.
I do think what's important though is, uh, we probably going to change how people work together, uh, soon if it hasn't happened already, right? Same as right now. Uh, I remember a year ago, um, I just learned about, okay, I can use natural language to ask, um, ask cursor or, you know, cloud to actually help me develop my applications for me, write code for me, right?
So, um, which is great, and I'm really impressed to see the speed of the innovation in software engineer with AI agents and large language models. Um, now if we think about networking engineer, developer engineer and application development or deployer, right? That experience I've constantly heard from network engineer and platform engineer is what they had to put the guard rail in place.
They had to constantly educating, onboarding, uh, new application developer who is, you know, developing things in the environment. They have to constantly educating them. These are the rules, uh, for these are the ports can be open, these are the resources constraints.
Uh, you have to go through this repository and do this and do X, Y, and Z before you can deploy in production, right? So all these rule books that may be on paper in the past, I think we're mo we're seeing moving that more into a agent flow where AI agents along with MCP and agent skills can potentially play a really key role here to be that God real person, uh, to manually guide you instead of having a person manually guide you through either, whether it's, um, you know, a cookbook, um, but it can be more programing through natural language, uh, through, uh, agents. I think that's going to be, uh, tremendously helpful to change how people are going to interact, uh, between different type of roles.
Mm-hmm. Um, well you mentioned platform teams, and I guess I wonder if, you know, I I get past four or five Kubernetes clusters, does that force the platform engineering conversation because teams will go, we gotta rethink how we manage all this stuff? I think so, right?
Because, uh, having that skill of platform engineer, you don't need everybody to gain that skill, right? To be able to stand up the Kubernetes cluster, to manage the health of the Kubernetes cluster to, you know, watch and having the right governance in place, that's a skill you don't necessarily want whoever developing application focused on business innovation to learn, you want to kind of abstract that, um, for the application developer so they can focus on ship code faster so they can focus on how can they make their application more user friendly if it's the business needs. Um, so I think that line is certainly very useful.
What I've, we've seen though in the past, uh, with application developer can iterate a lot faster now, and I see this with myself. Uh, I could, you know, rip off my application and redo it in a different language and to see if it performs better. Um, so the ability to ship is a lot faster now.
Uh, so I think that abstraction provided by the platform engineer is critical and also can potentially enable platform engineer to support way more application developers, uh, with, uh, with their guardrail. We talk about through AI agents and MCP and agent skills. What do you see people doing today when it comes to Kubernetes and networking?
That kind of just makes you shake your head a little bit and go, folks, maybe we should be a little bit smarter than that. That's a really cool, uh, interesting question. Uh, I would say people are definitely frustrated when they are problems and problems do exist a lot of times.
Unfortunately. I think things works really well on paper and when you try it in the POC environment, but a lot of times in production, uh, things can get a little bit wired. And I think as a human, you know, our brain is, uh, focused on one thing at a time.
Uh, it's harder for us to, you know, think completely and think very fast. Like when you're developing, uh, depo debugging a problem, right? You might be thinking about X but not thinking about the impact of Y on X, and here you finish looking at X and then start to looking at a, at a different angle to see if that could have been a problem.
So, um, one of the example I run into myself was, I remember I was deeply debugging a problem, a networking problem in my cluster, and it was, uh, it has like 10 million pieces. Like I was loading a new version of software. I was noting a new version of my gateway stack.
I also de was developing, uh, deploy MCP servers. It has like two moving pieces, and I got frustrated with my debugging. Um, but at the end it turned out to be a very simple problem, but somehow I didn't look at that problem.
I didn't thought it would be wrong. And, uh, in hindsight, I actually thought that it would be really nice if I actually had a debugging agent looking at the most fundamental problems I had and could potentially actually spot out the problem for me. But because I was frustrated myself, I didn't think about looking for the agent.
I think I overestimate what I could potentially do. Uh, I went, I bugged a lot of people for help. I did find out what the issue, it took me a few hours, but I know, um, my agent, um, could potentially spot that a lot faster because the fact that the agent is, uh, it can check things a lot faster and it doesn't, it wouldn't be afraid of check some of the most dumb mistake, which we all are human and we could potentially make mistakes too.
So, So I also can help but wonder if developers are kinda ignoring the physics of networking. And I asked the question because they're building more distributed applications and there's latency, but I think that they somehow think that magically, you know, all these distributed clusters and networks will just somehow or other, you know, account for all that latency when in fact maybe we need to think about the network more as we build and deploy our applications. Yeah, that's absolutely right.
I think the goal is, uh, the network can be disappearing for developers, right? But that's a aspiration. I don't think we will be getting there.
Uh, I mean, it, there will be latency for sure. I think the question is, is the latency small or tiny enough that we don't need to talk about it, right? I think, uh, this is also why we started the agent gateway from ground up because we've been amazed with some of the performance number from agent gateway, and we're talking about like, uh, way less than milliseconds of latency, uh, when the proxy is on the load, right?
'cause to a lot of people with concurrency traffic, when the latency is like zero point, uh, one milliseconds, a lot of time it is, uh, it's, uh, it's so minor that you may not think any latency at all, right? So that, that's kind of the goal we always wanted to achieve. But in reality, there is that hub.
You have to, um, think about how to minimize the hub as little as possible. There's also like the networking, the cable, um, bandwidth, if it's across region, right? If it's within the zone and region, it could be faster.
So that, um, networking latency across region is always going to be existing. And on top of that depends on how many proxy hops you have and how performing your proxy is. That's when it's getting really important, which is also why I would definitely encourage people to look at the Asian gateway.
It's one of the best, it's actually the best, uh, performing, uh, in terms of network speed latency in terms of, uh, CPU and memory utilization in terms of the moment you programming your gateway, API resources into the controller and how fast it actually program the gateway proxy of gateway, it's like the fastest in the industry. So, uh, that, that that's the way to, you know, minimize your network latency as much as possible is choosing the best performing, uh, gateway. Uh, in between your network hub.
You mentioned ai, and I'm gonna ask the question, um, is AI gonna ultimately be the means where we achieve that goal? You described where we make the network kind of transparent to the developer, and will the AI agent kind of just magically take care of all these latency issues? No, I don't think AI would magically take care of the latency issue.
The latency will always be there. Like I mentioned, it's just how much latency it is, is it, uh, tiny enough that's, that's not worth talking about, right? I think the AI could potentially play a huge role in how help us to be more intelligent, help us debugging problems, uh, help us overlook issues we didn't thought about looking or maybe faster than us to look out those issues, helping us onboarding new members that help us delegate the works we don't want to do because it's repetitive work.
I'd rather have somebody else who, um, can do it and onboarding new members, uh, you know, so that I can focus my energy on learning, doing new stuff, and for the stuff I already know how to do. Um, I want to build that as recipes to my agent skills and con, you know, config my agents to do some of that work for me. I think that's going to be, um, an important way to make us more efficient.
But I really don't think it's going to, you know, making our, um, network of gateway faster. Um, unless, unless, you know, AI can beat some of our best engineers, uh, improving infrastructure. Maybe we'll get there one day, but, uh, but right now I don't see it yet.
Well, folks, you heard it here. Even in the age of ai, the laws of physics have not been suspended, and we still need to figure out how to make all those networks work. Lynn, thanks for being on the show.
Well, thank you so much, Mike, for the chat. I appreciate it. All right, and back to you guys in the studio.