Microservices Orchestration – Cloud Native Podcast EP 25
Mike Vizard chats with Orkes CTO Viren Baraiya about the challenges organizations will encounter as they look to orchestrate the next wave of microservices in the age of artficial intelligence (AI).
Transcript
Hello, and welcome to the latest edition of the Cloud Native Now podcast. I'm your host, Mike Bazar. Today we have a guest, Byron Biray is the CTO for Orca.
So we're gonna be talking about, well, how AI agents and orchestration and microservices is all gonna come together because they're on a collision course. It's just a question of when now. Hey, Byron, welcome to show.
Hey, thanks. Uh, thanks for having me here. As we've seen so far, microservices are not easy to build and maintain.
A lot of folks are challenged by it, and as a result, maybe they sometimes even go back to monolithic applications. We are seeing, though, the rise of these AI agents that are increasingly starting to be able to take on more complex tasks. I guess the issue is how would we manage all these AI agents and create something that feels like, um, some order in what could be a potentially chaotic scenario, um, and apply that to microservices.
So you're at the forefront of this. What's going on? Yeah, I think, I think that's a great question, first of all, and I mean, if you kind of look back and about it, right?
There are a lot of similarities as well. Like when we look at, when we look back and, and look at the microservices world, what started to happen was, uh, you know, people started building microservices, uh, you know, single responsibility functions. And once you had a, you know, a host of microservices available in your kind of, uh, infrastructure, the next question came was that, okay, how am I gonna manage all of those things?
Uh, how am I going to essentially put all of them together to achieve my business goals? And what we are starting to see with the AI agents is kind of a similar story that, you know, the agent is only as good as the tools it has and the autonomy you give it to them, right? Otherwise, essentially, you know, you are not really solving the fundamental problem in terms of how they are able to plan, execute, and deliver the results that you want.
Um, so I think the way I at least see is, you know, uh, lot of learnings that we had with microservices, you can potentially apply them to how you build agents, how you operate and run them. And I think there are two critical pieces there, right? One is visibility into, you know, what is happening.
Um, ML explainability, explainability has been a subject of research, um, for quite some time, uh, with the adoption and a broader option of ai, and especially LLMs, it has even gotten a little bit more mainstream. Uh, just imagine a case where you are leveraging AI agents, which are making fully autonomous decisions. How do you get visibility into why that decision was taken?
Um, and we are starting to see some, you know, uh, kind of news about, uh, agents doing things that doesn't make sense or, you know, got company into hot soap and things like that. So this is where I think, you know, the learnings from microservices in terms of like, you know, it's not just about kind of how you put things together, but more importantly, understanding why it did that. Uh, being able to get visibility into it and, and be able to kind of do that in a safe way is gonna be very critical.
And I think there's a lot to learn from, you know, the adoption of microservices, how we could evolve, how it matured and apply them to, you know, agents. Mm-Hmm. I, in a lot of sense with microservices, one of the things about it is that, um, it creates redundant paths for API calls.
So the application is more resilient, and in a similar way, we may have to route different tasks around the different AI agents, many of which may be checking on each other to ensure that the quality of the overall application or whatever route we're building, doesn't go off the rails. Yeah, absolutely. And I, I think, um, I, I would say like, you know, in microservices world, you had the concept of edging where you would send the same request to two different API endpoint, uh, so that like, you know, if one of them is, is slow to respond or is not available, your overall request still continues to operate.
And I think you could apply, and we are starting to see that as well, right? Uh, with the AI agent as well, that, you know, you are sending the same requests to two different models, um, and, uh, maybe using a human to evaluate the response or using a third model to again, evaluate the response before actually, you know, responding back, right? Um, and, and leveraging that kind of patterns to kind of, uh, put some more safeguards, some more checks and balances.
Um, and sometimes it's purely for evaluation. Like, you know, if you build a new model and you want to test it out, um, before kind of rolling it out, you are kind of sending a shadow, uh, set of requests over there, evaluating them. And once you are confident, you know, switching that, uh, path, um, and especially the way models are evolving, like we are seeing a new version coming up pretty much every year.
Um, and, uh, you know, oftentimes what was working in the previous version, the behavior changes. So, you know, up upgrading to the newer version also requires, uh, testing. And the biggest problem with, um, ai, um, and language models is that they are non-deterministic, which means you can't really have a set of use cases and say, oh, I'm gonna run through this set of use cases, do a Q testing at their pass.
It's all good to go. That may not apply here. So now we have to kind of do this in production environment with real world, you know, user inputs.
So this becomes even more critical to be able to orchestrate the requests across different agents, different versions of the same agent and, and things like that. Yeah. Mm-Hmm.
So I may have agents that help me write code, and then I'll have an agent that helps me test, and then there'll probably be a set of agents that are doing security reviews and, uh, uh, governance management kind of tasks. Um, what will be the role of the human developer in all of that? How will they kind of be involved?
Um, software engineers will orchestrate this using what? That's a great question. Uh, what are we going to do if AI does everything?
I think the more important thing there is like, you know, um, somebody has to still kind of figure out, um, how these agents are going to orchestrate the task amongst each other. Um, also be able to kind of, uh, do that in a safe way, uh, which means there is gonna be some human, um, overview intervention required every now and then. I think that's one.
And like, you know, and this kind of brilliantly applies to software development, for example, right? Like a lot of software developers today are using co-pilots to write code, but you know, they are still responsible to, um, ensure that the code that is written kind of matches the requirement. It, it, it is safe to run.
Um, there are no bugs. Um, and, and like a lot of people are also using AI to kind of generate automated tests, but, you know, somebody still kind of orchestrating all of those things, right? I, I don't think we are there yet where we can say that everything is gonna happen, uh, by AI automatically.
Um, and I, I think, you know, essentially, you know, we are shipping ourselves slightly at a higher order in terms of what we deliver as a value, uh, which is what we are best at it, right? Um, being able to understand the business, being able to find the right set of agents to orchestrate and where, and how we orchestrate. I think that's, that's the field that I feel like, you know, is not fully developed yet.
That's one area where I think, uh, Orca, what we are working on is, is gonna be very critical in terms of, you know, you have all of these agents, you need a run time for this agent so that you can run them safely, you can inspect, visualize, um, uh, log them and see what's going on, right? That's, that's I think one area where I, I I see a lot of opportunities and, and I need, So in effect, the AI agents still need a boss, and that would be called us, right? Yeah.
Right. Um, you know, one of the things you see in microservices applications is there's just a lot of containers coming and going. And if you listen to folks, we might be building more software in the next few years than we've built in the past decade.
And so are we prepared to deal with the amount of ripping and replacing of containers that we're gonna see at that level of scale? Because, well, we can't throw more bodies at the equation, so we gotta find a way to kind of keep track of these containers that might only live for, you know, what, 30 seconds. Yeah, I think, yeah, and I, I think this is where the orchestration plays a very critical role, because you need a system that has a complete overview of what's going on.
Um, because let's imagine a case where you know, somebody, um, start, uh, spinning of containers and you end up spinning thousands of containers and you have no checks and balances, and, and you will find out when you get your next cloud bill. Um, so that, uh, that's definitely kind of the, the case. So, you know, this is where like, you know, the orchestration systems, systems which are essentially responsible and have kind of right set of limits, um, and, um, checks and balances in place, uh, becomes more critical to understand, you know, what's happening with my system.
And I think, as you rightly pointed out, right, it is, it is gonna become easier to write a lot of code, uh, build a lot of systems, where we will start to see, um, is how well are you able to run this? So runtime aspect of it is, is gonna be a lot more critical now than than ever because, you know, we can just do a lot more now. Um, It also seems to me at least that, um, we're reaching a level of complexity.
Maybe we're already there and we're just struggling with it. That, and the application environment is just too difficult for humans to manage by themselves. I, I think we are already there.
I think we are already there. Like, even without, like, even if you take the AI out of the equation, um, applications are complex. Um, what, and especially as you kind of move things over to the cloud, uh, in a, in a more distributed, uh, uh, world, um, applications are increasingly complex in terms of its interconnectivity to other applications, different systems, um, microservices, um, events, um, that are kind of, you know, being exchanged across different applications.
So things are definitely a lot more complex. Um, there is no longer the case where your application is a very simple, you know, you have a database, you a form, and somebody fills out the forms and service in your database, right? That typically never tends to be the case.
Now, there are a lot more kind of business logic that is implemented. There's a lot of orchestration that is happening across multiple applications, including third party systems and vendors. Um, so they are definitely a lot more complex.
And, and guess what happens when you have a very complex systems? It becomes increasingly complex and difficult when there are failures or things that you have to reason about as to why certain things happen, right? Doesn't have to be necessarily failures only.
Um, now I would give you an example of, um, a system that takes some decisions, right? Um, why did it arrive at a particular decision? You need to be able to know the entire graph of, uh, you know, steps it executed to understand why he derived at that decision.
Now, if I were to build a graph in my mind, you know, that's a lot more cognitive overload, the amount of time it takes. Um, and, and that's where, you know, orchestration systems, systems like conductor and orcas, uh, becomes very critical. Where, you know, you are able to understand what is happening, why it is happening, and if there are failures, uh, be able to kind of, first of all resolve the failures automatically, and if not, uh, be able to kind of have a, a simple API call or a one click button to, you know, resolve them.
Um, but yes, they, they, they, they in in somebody, they are already complex. Um, On the opposite end of that, are we in danger of becoming maybe too dependent on ai and we let the machines kind of figure out everything that can be done and we may not understand how it was done. And then when somebody calls us up like a compliance auditor and says, can you explain this?
We may not be able to. And I, I think that's a extremely valid point, and I'm, I'm kinda, I hear that a lot, uh, and this is one of the common themes that I kind of also hear from our customers, the users that I've spoken to is, yes, AI is great, but I need to be able to explain the decisions it took. And this is why I believe that like, you know, being able to explain the decisions that AI made is gonna be extremely critical, non-negotiable in some sense, uh, or especially in some industries where you have to be able to justify and understand why that decision was taken.
So, you know, instead of AI becoming a black box, the way I would kind of think about it is that like, you know, it, we are starting to see, and uh, that is one area where we, we also kind of, uh, position ourselves is, you know, you think about AI as tools that can help you, uh, automate things, um, and take, um, automated decisions, but you are still in control and you have to have the complete visibility into, you know, why this happened, what was the input given to it, um, what was the rational? And maybe have a parallel kind of, uh, execution made, uh, to another model or an algorithm to validate that like, you know, even if I use a different AI system or an algorithm, I would arrive at the same situation. So now when my compliance officer calls me and says, you know, why did you do this?
I have the full explainability as to why, um, Do you think we'll also be able to use that capability to a degree to, uh, determine what is causing a, a performance degradation in my application? 'cause one of the issues with microservices is, well, they may not necessarily completely fall over the way a monolithic app. Well, I can spend a long time trying to figure out what it is that is causing a certain performance issue.
And the maddening thing about it is it might take me two days to figure that out, and it's about a 32nd fix. Oh, yes, absolutely. And I think this, this has been kind of already, um, you might already see some of these systems similar to that, right?
Like doing auto tuning. Um, so, you know, you put an AI agent, um, um, or as a sidecar in your application deployment, which is constantly monitoring, um, how your application is performing, checking logs, um, and other systems, CPU memory and kind of understanding, you know, the application behavior. And in today's world, uh, it could start with giving recommendations, but I can totally see in the future world, um, and I have seen those systems, uh, at play as well at large companies like Google, where it'll just automatically tune it for you, um, to the best of its capabilities.
Uh, so, you know, and, and the whole idea is, you know, if you think about it, right? Uh, where we are best as humans is being able to understand the business, um, implement the business logic and drive the business forward. Everything that we have to do on the infrastructure side.
How do I run my application? How do I get visibility into my application? As you mentioned, I queue my code for performance and everything that can be very well done by machines.
Um, and, and yeah. Um, so you, you kind of coexist and, and leverage them as your tools rather than like, you know, think about it as kind of replacement, uh, makes it more productive. How will the software engineering teams be organized?
'cause I, in my mind, it looks like we're gonna have a small army of AI agents and working alongside humans, and, but you know, it still takes a team of folks to build something. So how will the AI agents that I created work with the AI agents you created? I think that's, that's an interesting question.
Um, I would, um, I mean, I, I, I don't think I, I have the kind of answer that I, um, because right now it's such an ascent field, um, and, and we are starting and, and we are seeing frameworks, right? Frameworks, like, for example, conductor, where, uh, we are allowing people to kind of, uh, orchestrate across multiple agents through API calls. Um, there are frameworks like auto gen, uh, that allows you to kind of do the same thing why, where you can have conversational AI agents, uh, talking to each other and to achieve a goal, um, or collaborate on a particular task and things like that.
But I think that's another field that is, is evolving quite rapidly. Uh, and we will start to see maybe some amount of convergence there in terms of frameworks and, and, um, how those things are kind of managed. And, uh, yeah, Right, you mentioned the Netflix, uh, framework that you guys are using at the base of your approach, but, um, I'm trying to figure out where the cart and the horse is here.
'cause sometimes I think people are gonna try to create all these AI agents and then kind of add the framework to kind of manage it after the fact. So maybe we should be putting the frameworks in first and then figuring out what's going to attach to them second Maybe, uh, maybe, but I mean, then I think it's, it's the kind of dilemma, right? In terms of like, where do you invest in focus your time on, right?
Building infrastructure is always more expensive, tricky. It requires very specialized skills overall. Um, ai uh, agents, fortunately what has happened over the last few years is it they have become almost ized, right?
Like you have a plethora of choices in terms of the models, uh, their cost. Um, so it's much easier to use them to build a POC, um, or at least, you know, put together a simple business use case, um, which is what everybody is doing today. Uh, where I think, as you rightly pointed out, is there's a need for infrastructure, there's a need for a runtime for all of these things, which does not exist yet, um, uh, outside of, um, you know, few kind of initiatives like ours.
Um, but I, I would say, you know, as people kind of realize, um, and, and try to find out, they will both converge, um, eventually, uh, as in pretty soon probably. Do you think ultimately we hear a lot about the phrase platform engineering, and I wonder if this transition to AI agencies gonna force us down that path, the more organizations are gonna have to have some, uh, centralized approach. But I guess one of the joys of DevOps was that, you know, we embraced it in the first place so we could get out from underneath centralized it.
And so how do we strike a balance? I think platform engineering is gonna be the commonplace, and it's not already, you know, I mean, we are already seeing a lot more effort, emphasis on platform engineering. And if you think about it, right?
Like, um, you take your, um, developers who understands business, uh, who is able to kind of, um, implement complex, uh, business solutions. You take AI who is able to kind of take care of heavy undifferentiated kinda work, um, like, you know, helping them write code, um, and things like that. Then what's missing part there is that platform where they can all put everything together and run it so that now they don't have to worry about kind of that.
So in some sense, I think, you know, it's, it's a complimentary thing. And you know, when you put all these three things together, you, you get the most efficient. Uh, you know, what I would like to call it, like, you know, at 10 x developer side, um, 10 x developers are less about developers and more about the frameworks and platforms that they operate on.
Clearly we're kind of at some sort of crossroads when it comes to software development. So what's your best advice to folks about how to get ready for all this? Like, if you look at software development, right?
Like, it, there is a lot of hype, um, and, um, interest in, in AI today, but software development has always been similar to this, right? The, it has never been the case where you learn one thing and then you keep on working on it for next decade or so, right? Like it's constantly evolving in terms of the framework.
Um, in terms of the architecture, um, you know, we went from mainframe to PCs to, um, data centers to, you know, cloud, um, now hybrid cloud ai, uh, microservices events like this has been constantly evolving. So I think the advice is, is the same, right? Like, um, it's, it's about kind of constantly learning, um, understanding where you add the value, um, and, and I think in the end, um, be close to kind of the foundations, right?
Um, I, I think that's, that's the hardest part. Um, and, and that's where we add value. Yeah.
All right. And folks, you heard it here. Hey, even in the age of ai, don't forget the fundamentals because that's what's gonna help you get through all this.
Hey, Barron, thanks for being on the show. Yeah, No worries. Thank you so much for having me here.
All right. And thank you all for listening to the LA or watching the latest edition of the Cloud Native Now podcast. You can find this in other episodes on not just our website, but on Spotify or Apple or any place else that you listen to podcasts.
Once again, we invite you to check out the entire library of podcasts that we have. And until then, we'll see you next time.
