The Evolution of AI Agents with Cognizant’s Babak Hodjat
Babak Hodjat, CTO for artificial intelligence (AI) for Cognizant, traces the evolution of AI agents from the initial launch of Apple Siri to a new era of agentic AI.
Transcript
Hello, and welcome to the latest edition of the Techstrong AI video series. I'm your host, Mike Fazar. Today we're with Baik Hojat, who's CTO for AI at Cognizant, and we're talking about this whole shift to agentic AI and how we got here.
'cause it started all the way back with things like Siri, and it's a continuing journey. Baik, welcome to the show. Thank you.
Thank you for having me Put some perspective around this. I think everybody kinda understands Siri or whatever, um, the folks in Google and Android provide, I think it's Bixby, but these things were the precursors to AI agents that we see today. And it seems like there's a straight line here, but I don't think everybody appreciates how exactly these things are all connected.
Yeah, well, uh, AI agents have been around, actually even before Siri. Um, uh, they are conceptually the same thing. When you ask Siri to call your wife and it actually dials a number and gets the call done, that's an agent working on your behalf that understands natural language.
Um, the difference today is that we actually have these Gen AI models, uh, to help with the understanding of the natural language and to help with the reasoning of the agents. So they're much more powerful than what we used to have in the past. But even before that, you know, AI is, is really, really hard.
So having a single AI system that does everything and is generally intelligent has always been very difficult. So in the late eighties and in the nineties, people started, uh, thinking, okay, what if we simplify the environment within which the AI is operating and let's call that an agent. And back then we did have a pretty simple environment.
It was the web, it was much simpler than it is today. Uh, just some texts and hyper texts and, and links. Um, and so that's kind of where it all started.
Uh, but you know, the, the aspect of, uh, agents that's really in interesting to me, which also started in the nineties, was multi-agency. So you have one agent, and it might be collaborating with another agent or communicating with another agent to get something done for you As we go forward. As I look at it right now anyway, it seems like there's gonna be, I don't know, hundreds, thousands of these agents.
How will we kind of connect them all together to kind of drive something and will the agents know of each other? Not to mention who not only am I, but who am I working with or who's part of my extended family? Exactly.
That is, um, a very important question. And I think it's something that in my opinion too, is inevitable, uh, with people, marketing agents right now for various different tasks. Um, you would want these agents to know of each other and actually collaborate and talk to each other.
Uh, and so that we do need, uh, some way to make these agents interoperable. Uh, we do need to make sure that, uh, when they're aligned, we're we have less of a problem when they're not aligned with one another. How do you deal with that?
How do they negotiate, uh, you know, to get something done for you? These are all problems we will be running into the future, it running into in the future. And, uh, I, I think we are already seeing that.
Uh, so companies have been starting off by creating these knowledge extraction chat bots, uh, using genai. And, uh, they've created them, you know, large companies like ours have been creating them for various different use cases. And very naturally, uh, it just doesn't make sense for you to talk to one agent and the agent saying, well, I can't really help you here and you gotta go talk to this other agent.
And then for you to actually repeat yourself to the other agent that's responsible for it, because you started off talking to the HR agent. Um, and so that interconnectivity, that interoperability is really important. And many companies are actually bringing in agents from third parties.
You know, we have agent force, we have, um, uh, you know, SAP has its own agent. You know, the various different companies are coming up with their own agents. There's agent space by, uh, by Google that, um, in fact has adapters into many different backends.
And, uh, so yeah, I think the, the future, the way we see it is this, um, uh, coming together of agents representing various different functions in an enterprise. Uh, you could think of it initially as really looking very similar to the microservices that we've created and even the organizational structure of, of a business. Um, and you should be able to, uh, get stuff done, uh, as long as you have authorization for a task, uh, across multiple agents, and in fact have agents talk to each other and work out collaboratively, um, and consolidate their responses.
So you're only dealing with one, uh, interface, but in the background, you have a team of agents working on your, on your behalf. Um, I think we need to make that this can't be created in a centralized manner. So we have to come up with a way to actually nurture this organic, uh, emergence of these multi-agent systems.
So for example, I want to see business units be able to create their own agent sub networks and have a process for, you know, sandboxing them, validating them, and then being able to plug them in into a larger organizational network. And by virtue of doing that, expanding the domain of discourse, uh, and and capabilities for, for that particular organization, Will these agents negotiate with each other? And, and how might that be accomplished?
Because sometimes I wonder if I have an agent that is optimized to go buy something at the lowest cost, and you have an agent that's gonna have something that is optimized to sell something at the highest margin, won't they kind of just meet in the internet somewhere and beat each other up to a standstill and then call us for help? Yeah, in fact, that is something, uh, I I think that will happen in the future. It's not the case right now.
So we're really building from the ground up, and we're only at the stage where the assumption that the agents are all working for the same organization, uh, is is is probably gonna work for us for a while now, but we will very quickly get to that world, Mike, that you're, uh, you're describing where, for example, I have agents that on my behalf might want to broker some, uh, uh, you know, uh, service from a third party, and that third party, um, has created an agent, uh, to receive orders, for example. Um, and, uh, and, uh, you know, there might be a negotiation going on. In fact, um, we did a research, um, uh, with Oxford economics just recently that showed that while consumers like you and I, um, are not comfortable with having AI agents actually, you know, do the transaction itself, we are comfortable with them giving, you know, do the, doing the research and giving us the options.
And so that inspired me to create an agent network that helped a consumer like myself make decisions. And it was a number of different agents that came together to help me with various different decisions from, you know, what hobby to choose, to financial decisions, to, you know, buying stuff. And then I actually had a few agent networks that represented different, uh, third party, uh, B2C kind of businesses, uh, for example, for, um, uh, you know, hotels and, and accommodation and so forth.
And I had more than once. So I actually had a lot of fun, uh, having the agent that is responsible for, uh, you know, uh, coming up with my vacation, uh, agenda, negotiate and talk and, and I would tell it like, you know, yeah, I want vacation, uh, and I have a budget of whatever X dollars and see it go back and forth, uh, with these third party agents and in fact do a negotiation. And as part of their prompt was, don't trust the other side, the other stride is not fully aligned with you.
So it would actually get some results and go to the other hotel provider, for example, or accom accommodation provider and say, Hey, I have a better deal than you, and can you beat this? And, um, it was amazing just to read the transcript of how these agents were kind of trying to, uh, uh, you know, meet their own KPI, uh, while trying to, you know, close the deal on something. And ultimately the, the agent representing me came back with some a list of options like, here are the options that you have with these different providers.
So the final decision is up to me, uh, the user, but a lot of that negotiation, the headache of trying to, uh, check and find these things out, uh, took place, um, without me involved, uh, it takes a while. These, these systems might, uh, run for a while. One of the things that might be an issue, again, this is a future state we're talking about, is that these, uh, large language models are fine tuned to be very kind, uh, you know, very positive, very, uh, ethical, moral, uh, systems.
So when you actually follow the dialogue between two agents, they tend to agree with each other very, very quickly, much quicker than you would like actually in these types of situations. So that might call for a different kind of fine tuning when it comes to non-aligned agents talking to each other, for example. So there's, there's a whole world of research that has to go into, um, uh, kind of, uh, uh, coming up with that kind of world, uh, that, that kind of setup.
Um, I must say that, um, again, currently the, the, the overarching assumption for everyone is that we're building agents from multi-agent systems that are all aligned and in the service of either our company or our consumer. So that's an easier problem than, you know, we have a bunch of agents that might be talking and negotiating with some other agents that are probably not aligned with us. I've been trying to figure this out and, and I imagine, you know, the answer and maybe it's not, uh, obvious to everybody, or maybe the techie folks understand it, but so Siri more or less ran on my phone, tablet, whatever, is there enough horsepower to run these AI inference engines out at the edge because we're trying to create an interactive experience.
So the round trip to the cloud, the latency might be too long. So, uh, how much of the AI runs locally and how much of the AI runs somewhere in the network and how much of it runs in the cloud? Well, uh, let me correct you, uh, on, on the premise here.
The initial Siri was a hosted system. In fact, it was one of the first, uh, when, when Siri got acquired by Apple, they actually built the data center for Apple. Apple didn't have a data center before that.
I mean, uh, so, so that was actually hosted, but for the processing capacity back then that was required, um, uh, not anymore for that, uh, type of, uh, Siri, uh, system. So you're very right that, that, uh, kind of, uh, processing didn't require what we require today, like our most powerful large language models have to be hosted. They actually run in very large data centers and take a lot of processing capacity.
Today, we cannot even, uh, start to imagine what it would look like to run a GPT-4 oh on your phone. That just doesn't, it, it won't work. However, having said this, you could today think of systems that are running in a hybrid mode, where some, um, you, you even see that actually today in Apple Intelligence.
Uh, when it came out, I was thinking, okay, they have the right idea here because, um, much of Apple intelligence is actually running on your phone, and that is for data security purposes and so forth. It's actually really, really good. Um, but it's, it's not that powerful like what it does, what Apple Intelligence does on your phone.
It's a small, uh, large language model. So it's not that powerful and there's not a lot of functionality. You can, you can run just on your phone alone, uh, for agents that can understand natural language, do reasoning, and, um, be kind of a higher order type of agent, you will still today need to, uh, run them on the cloud and data centers.
But you could imagine a hybrid model where some agents are hosted and some agents that are much more specialized and therefore don't need a large, that large, uh, uh, a large language model to be running, uh, locally. And what that does for you is it, it it does, uh, save money, uh, um, um, uh, because it kind of, depending on the use case and depending on the agent, you're using a different, um, model that, that might not have the same, um, processing capacity requirement. It also gives you some data security aspects for those particular, uh, parts of your multi-agent system where it's closer to the data.
Just imagine if you have an agent sitting in the cloud that generically knows what you're talking about, and then defers it to the agent that's sitting on your phone that is actually responsible for the data on your phone. That way your data isn't really moving off of your phone. Uh, and that functionality, um, is, is resolved through that, um, interaction between the agents without the cloud agent even getting that data.
So, um, that we can do, even today, the general trajectory of the technology today is, uh, these large language models are getting smaller, uh, at the same time, more powerful, um, deep seek an obvious, uh, example of that where a much smaller, large language model could do things at the scale, uh, that wasn't possible in the past. You can run a 14 billion parameter, um, deep seek model on your, um, M three laptop and, and, and it will work, and it's at, you know, it's, it's on par with the GT four oh mini, for example. So it's actually pretty powerful and it, in my mind, for many agentic use cases, it pa it passes the bar.
Um, and so, you know, we're, we're still in the early days, but I think in general, we will move to a point where these, um, uh, running more and more of these agents, uh, locally is gonna be viable. The other thing people seem to be struggling a little bit with is, is an agent, should we just treat them like digital labor or are they essentially a new type of employee or is it just software that we're using and we just need to invoke it through an API, but we don't need to think about it any differently than we do any other application. There is a distinction in that the moment we talk about an agent, we are assuming some level of autonomy, because if there were no autonomy needed, then, you know, it doesn't need to be an agent.
It's just pure software. You just write, write some code and some rules, and you have it do whatever, um, it needs to do. But the reason why you're using an agent beyond the fact that it understands natural language is to defer to it, to make a decision as to what, uh, which one of its tools it needs to use and in what order and synchronously or asynchronously, and perhaps in some cases just go off and use its tools and review and go back and, you know, so to maybe make multiple calls and then come back to you.
So I don't know if we can treat them exactly like software for, for starters, we have to accept a world in which systems do exhibit some level of autonomy, so we have to design for that and make sure it's responsible, uh, in its operations. The second is we need to get used to a world in which things are not deterministic. Uh, you know, the same inquiry is not gonna get you the exact same response every single time.
Um, and it takes some getting used to, uh, but, uh, yeah, i, i, I don't know if I answered the question, but generally I'm just, just kind of highlighting the distinctions here. As you look forward and now that you're working over at Cognizant, you're touching more customers, um, at least among the early adopters, what do you see them doing right, that you kinda wish other people would kind of cribble a little bit? Yes.
Um, I think, uh, uh, those who do recognize the fact that, you know, identification, AI enablement is not a one-stop shop. You're not looking at a single be all end all model that will do everything for you. I think that that is the right kind of thinking, uh, those who are actually making interoperability a requirement for the agents that they, um, utilize, uh, I, I think are taking the right steps and also recognizing the fact that this age identification is an incremental process.
It's not a lift and shift. You don't have to go off and, uh, you know, say, okay, I'm not even gonna embark on this 'cause my data shop isn't in order. Or, oh, it's a huge undertaking, like identifying my entire, um, business is gonna take forever or culturally, I'm not ready for that.
Actually. I think because of the nature, the modular nature of multi-agent system, you can start small and grow. And as long as you set the framework in a manner that allows for that kind of incremental, uh, adoption, um, you're in a good place.
So being an early adopter, uh, in this case doesn't mean being completely disruptive to your, to your business. It's, it's actually a smoother, um, less painful process than, for example, I don't know, back when we were migrating to the cloud, uh, for instance. And, um, so those of our clients that recognize this, those early adopters that are doing that, I think, I think those, uh, that, uh, I'm just highlighting what, what, um, I think keeps them ahead of everybody else.
And I've been surprised that, uh, some of our clients who are from domains that are traditionally quite conservative, um, are recognizing that and are actually taking, taking steps there. So it's unlike what we've seen in the past where, you know, a conservative bank or an insurance company that would say, yeah, we'll wait until that, that this new technology is, has been out there in the risks been, um, uh, mitigated. Um, no, many of them have embarked because they recognize the incrementality and the, the importance of, um, endorsing and managing this in a safe, uh, manner versus just letting it happen organically because it's happening organically and you just don't want that.
It just, uh, it's not the responsible way to go. We haven't talked about security and we are struggling with just securing the humans who are a part of our environment. And if we have all these agents and they're essentially entities, our endpoints, much like people, how will we secure thousands and thousands of agents that the bad guys are gonna go try to steal the credentials and manipulate?
Yeah, that's a, that's a really important question. I think, um, uh, I think it's important for us to, uh, recognize that if an agent is hosted by a commercial entity and it's a complete black box, especially if it's hosted by maybe a third party country, um, you know, that, that, that there's a trust element there that has to be, um, you know, taken into consideration. Um, I always say if, if, if the functionality, uh, that you're, um, so delegating to this agent is sensitive or it's dealing with sensitive data, I think it makes sense for you to, um, host it, uh, and run it internally and secure it.
Um, but there are ways, there are mitigations, um, uh, uh, around authentication authorization. Uh, and agent by definition is not just a la large language model, it's a large language model plus code. And the code is actually superseding the large language model.
The code is what is making the calls. The large language model only takes input and, and, and, uh, output, uh, produces output. It's the code outside of that, which is, uh, a deterministic human design code that decides what to do, like, which tool to call what API to call.
So if you want to secure your agent, that's where you have to be writing your code. And, um, the other thing is this acknowledgement that the agent is going to have some level of autonomy means that we have to always think very explicitly about what is the line that we draw. We're like, okay, up until this point, with this sort of level of certainty, with this kind of functionality, we allow the agent the autonomy, uh, because it gives, makes everything more efficient and productive and everything.
Um, but here's the line below which I'm gonna take over. I'm gonna write the rules, I'm gonna make sure there's a deterministic methodology for what happens. And we always talk about a fallback, like if all hell breaks loose, I want to be able to, you know, disconnect all LLMs in my enterprise and fall back into a rule-based model or a human-driven model.
So having that in mind I think is, is very important, especially with critical processes. Um, yeah, I, I think, uh, you're, you're touching on a very, very important point here. All right folks, you heard it here.
As you listen to this conversation, it becomes pretty apparent that despite what AI agents may automate, the missing ingredient is still gonna be human intelligence. Hey, Baik, thanks for being on the show. Thank you very much, Mike.
And thank you all for watching the latest edition of the Techstrong AI video series. You can watch this episode and others on our website. We invite you to check them all out.
And until then, we'll see you next time.