AI Agents in Business with Tredence’s Unmesh Kulkarni
Unmesh Kulkarni, head of generative artificial intelligence (AI) practices for Tredence, dives into the impact AI agents will have on organizations once they are more widely relied on to make actual business decisions.
Transcript
Hello, and welcome to the latest edition of the Textron AI video series. I'm your host, Mike Bezu. Today we're with Umesh Karney, senior Vice President for generative AI at trades, and we're talking about the rise of AI agents.
They seem to be everywhere, all of a sudden, or soon will be. Esh, welcome to Shah. Thank you.
Glad to be here. We've been experimenting with generative AI for a while now, and people have been playing an our prompts, and some people are better at it than others. But how far does AI agents take us to another level?
I mean, how automated will automated get? Yeah, it's a great question, Michael. And you know what I see?
It is not a brand new technology that came out of the blue, but really the evolution of how we have been working with AI for the last two, three decades, right? We had the statistical models that led to deep learning and machine learning, and then, you know, large language models with generative ai. And really, agents are now the defacto means to actually consume all of these AI agents and work with all of the IT backend systems that we have today.
So a lot of automation to come, and I believe this is the next wave, but we are fairly at the beginning of this next wave right now. Um, how many AI agents might there ultimately be in an enterprise? Because if there's an AI agent for every task, well, there are thousands of tasks that does every one of them need an AI agent or eventually, will AI agents handle multiple tasks?
How do you think that might come together? Yeah, I, based on the agent take AI implementations that we have done at treatin, and we are, uh, a very specialized data AI services provider. So we have been actually very, very focused on generative AI, doing some key agent TKI implementations.
What we have seen is the best practice to think about AI agents is to not think of them as a task agent, right? Think of them as a, as a role agent. So you may have an agent for your software developer, right?
An agent that helps your quality assurance engineer or your data scientist, right? So that's for IT and engineering. But on the business side, you may have an agent role that manages your supply chain or looks at your inventory and logistics, right?
And these agents are actually going to work together, just like we form teams of, you know, people in our companies to actually perform functions. I see this evolving into teams of agents working together to actually perform tasks and really propel the human teams that are performing these tasks today. So many, many agents, but agents not mapping to tasks, but to roles in the organizations that you have today, How will those activities therefore be orchestrated across an end-to-end workflow?
Because, well, some outcome usually involves multiple applications, multiple processes. So how will I organize that in a way that creates the desired outcome? Yeah, fantastic question.
And I think that question is at the crux of why agents are so much more powerful than a lot of the technology that we have seen come out in the last 10 years, right? So agents have, they're not just large language models, right? They actually have reasoning capabilities.
A lot of that reasoning comes out of the language models or the prompts that we provide to the, to the models. But agents are able to think, Hey, what am I specialized at? What are my tasks?
What is the desired outcome? And when I say think it is still us prompting and, and setting up and configuring these agents, but now think of these agents then saying, okay, if I want to perform this task, how do I actually plan about doing the tasks? Large language models, just think of next word or the next token.
But agents think about long-term outputs, they think of the goals for them, and then they come back and say, okay, now I have a plan to get there in five steps and I also need to work with this other agent, or I need to get input from the human expert in this case. And that's how they are much more powerful in actually making things happen because yes, these agents will then very, very dynamically be able to collaborate with one another, get back to humans, or give feedback to humans or collect feedback from humans, and then collectively process these tasks for automating a lot of common work. And I, I see this more as a productivity gain right now, the, the currently being, Hey, let me actually make the human expert 20, 30, 40, 50% faster by automating and taking care of a lot of the run of the mill things that you and I do on jobs today on, on a daily basis.
Mm-hmm. And of course, everybody's talking about, you know, whether their job will be impacted and to what degree, and everybody has a certain sense of, uh, fear of somebody moving their cheese, as it were. But as I look at it, a lot of these tasks are things that maybe we don't like doing in the first place and in the second place, we don't do them all so well anyway.
Yeah. And that is clearly the goal. I I do feel that today there is little bit of uncertainty and it's fair for people to be, you know, concerned.
And I, I think there are some indications that some of the routine low-end tasks will be taken away by, by agents and will be done by agents. But if you look at the history of technology, haven't we be be doing this all the while, right? There are cars that kind of do a lot large part of driving today, but has autonomous driving replace the human who needs to go somewhere, right?
Has, does the car decide where you go? No. Right?
So it's a human that still needs to take the agent system and say, what do I need to do with it today? Right? How do I guide it the right way?
So we, we, as, as, as experts in the industry or as leaders in organizations, we need to be very careful about how we message the rollouts of ai. I think we need to start thinking human first. We need to start thinking of how AI can actually aid our companies move faster, how to set the right goals.
And that those are all human tasks. The, the ability to communicate, the ability to understand and perceive, right? The ability to make key decisions is still going to be a, a very much a, a human domain, a human endeavor in for a long time to come.
So I'm not worried about the jobs, but if you are doing run of the mill regular tasks all day long, I think you should be worried. And the way to get out of it is to really upskill yourself to really use the human wisdom, the human decision making powers that we are all kind of bond with, right? And can develop further.
And then AI doesn't become a substitute for us. AI actually becomes something of a tailwind for us to do what we want to do in a much, much faster way. Now, I've seen some early implementations and one of the things that strikes me is sometimes there's too much of a good thing and there's too many of these agents popping up asking to handle a certain task.
And, um, and as they wind up getting in the way, as much as they are helpful or they're trying to be helpful, so how do I kinda tamper down the enthusiasm of these AI agents? 'cause every time I'm doing something, one of them seems to pop up and say, can I help? Yeah, I do.
Uh, I I, I actually feel that's a good problem to have, right? When any new technology blossoms, you are going to have a little bit of exuberance, right? You're gonna have just a little bit of too much of enthusiasm all around too many products trying to do stuff, and they're not doing it exactly the way you want.
And, you know, when mobile came up, it was similar. When cloud came up, it was similar. When computers came out, it was similar.
So it's fine, right? We are gonna come out of that phase. It's, it's perfectly okay.
What I feel is, is gonna happen is there will be, um, a a lot of these trials, like the thousand flowers Blossom or whatever, and then we are gonna see some ecosystems emerge, right? And that's happening with like, for example, Google kind of really consolidating a lot of what they did with Gemini, with what, what they call now as agent space, right? How they actually build around manage agents, right?
Microsoft has been doing a lot of interesting with copilots and auto gen and, and TIC frameworks, right? So there are these frameworks evolving, which will now also form ways of interacting with each other, right? And then a lot of the complexity will go away.
And yes, still that time you will have these agents power up. You just have to figure out what works for you and also keep a watch on where the industry is going. So you're not left with some of the laggards, you're actually moving where there is the critical mass or development and, and further progress and, and traction, particularly on the market side.
Um, these agents are based on those reasoning engines that we find in the LLMs. How smart are those reasoning engines getting? And at what rate?
Because, um, it seemed like initially at least some of the ones I saw had the roughly the reasoning capabilities of a five-year-old. But now they seem to be, I don't know, college students. What, where are we on this adventure?
Yes. I, I think most of the language models are, you know, late teenage kind of capabilities in my view. But One way to think about them is a very, very smart teenager who has learned a few tricks, right?
And still doesn't have the experience or the wisdom for how to apply, um, those skills, right? So that's where I said human orchestration is still, still very critical. I'll give you a few examples, right?
So kinda just ground it in some facts. We have, um, an a benchmark called SWE Bench that stands for Software engineering, SWE Software Engineering Benchmark. And it, it has hundreds of really complex software engineering tasks that need to be performed.
You can give these two software agent tick systems and agent tick system is fairly Reliably able to solve 50 to 60% of those tasks. Today, when we started, we were at 2%. Now we are not, well not of 50% already, right?
The agent tick systems are able to collect information from all your database table structured data. They can merge the information from documents and unstructured data and make sense out of it, right? I, I have obviously you probably heard of this PhD agent that's gonna charge B be, you know, coming out for $25,000 a year, right?
So there will be agents which will cost tens of thousands of dollars and potentially have the ability in a narrow field, presumably to actually go fairly, go deep and do deep research into areas. But I don't think we are at a PT level yet, although there are some benchmarks that say, oh, these language models themselves are able to solve fairly complicated math problems to going to the level of, let's say, an American Math Olympiad or, um, American Invitational Math. But that, again, is a very, very specific narrow field in terms of the ability to do concrete work.
Think of it as a bunch of interns that are available to you if you can guide them, they're very committed. They work 24 7, they're diligent and they love to work. So how can you guide them?
That's the, that's the right mindset in, in my experience. Um, as we kind of move forward all of this, um, How Will we insert these AI agents that are based on probabilistic outcomes or guesses, as they might say, and they get better over time, but the workflows tend to be deterministic and they're supposed to be done the same way every time a hundred percent of the time. And the AI agent's gonna do it maybe differently every other time.
So how do I kind of reconcile those things? I am laughing because we had exactly the opposite problem asked a few years back. I was the head of AI and automation at a, at a public, uh, company that did customer engagement.
And obviously customer engagement is where like agents, human agents in contact centers are talking to their customers, right? But the problem was those hardcoded workflows didn't work because humans cannot be bossed into standard workflows, right? So when we build those software systems, humans would always go and ask a question that the system didn't know how to answer.
So we actually wanted more flexibility. Now people are saying, Hey, is it too flexible and how can I, so there is a little bit of balance there, and we will go a little bit from guardrail to guardrail. There are ways for enterprise systems to be fairly become deterministic and reliable.
For example, within language models, you can set temperature settings, you can do top grade, top K, top p and there are different ways to kind of really lock down certain things. You can also actually build a lot of reliability in the way you build those agents. The way you define the prompts, you define the roles and the job descriptions of the agents.
You can instruct the agents to follow steps or to come back clean and say, I don't know. Right? But these are not, I believe very, very different problems than fundamentally what we deal with with humans.
Sometimes we humans tend to not actively say, I don't know, right? We try to figure things out. And language models are similar, right?
So you can actually build those systems to avoid those problems of hallucination. We can do what's called grounding, which actually tracks you back to where does the data come from. Give me a concrete reference of where you pick this data point from, right?
So those approaches are where kind of professionals actually go. When we build these systems, we make sure that the systems are not hallucinating. That if they are hallucinating, there is detection and confidence scores that the user gets back, right?
Or the answer does not ever reach the human. So there are ways to, to mitigate the problem. And we are getting better every week in terms of how we can drive a much more reliable deterministic automation outta the systems.
But the bottom line still is that flexible workflows are better, especially when you're dealing with humans at the other end. Mm-hmm. Um, you mentioned hallucinations.
Do you think we have a new appreciation for data that despite 40 or 50 years of computing, we never really had, and now in the age of ai we're starting to realize that, you know, the way we manage the data, describe the data store, the data matters. Yeah. Uh, um, it is interesting that many of our projects, Michael, we we start obviously with, okay, here is the ROI and ROI is now proven, right?
It's not like, if it is, how can I do this right now? The question has moved to, okay, do I have the right data in terms of the, the quality? And by quality I don't, I don't mean like, does ca stand for California or Canada?
Right? Those, those problems have been largely kind of solved. But I, I think the world is world of data is excluded because we are now bringing in the, the 70% of untapped, unstructured data sources into the mix, right?
So the structured data with data warehouses, lake houses is the, that world seems to be under control. There's still a lot of work of to be done with data engineering, but we have now added complexity with documents, images, videos, right? So more and more content is now available to extract business value out of.
And yes, that does create interesting challenges, interesting pre-work where before you can start extracting value, you may have to do some data engineering work. The good thing is you don't have to wait for all your data to be in one place in your house or whatever to, to kind of get started. You can get started with a fairly small amount of structured and unstructured data, start ruling out your agent systems and then add more and more data sources to get richer insights, recommendations, and actions over a period of time.
So it's a journey. It's not one and done. And you can start expecting business ROI from from D one, which which means in, in about a few weeks from, from the, from the time you start.
Mm-hmm. Um, so if we play this out to the nth degree, we keep talking about AI and aging as a technology problem, but how much of this is really almost a sociology cultural challenge as we go forward together? Yeah, I am not sure I'm that expert to kind of really talk about the sociology angle of it, but let me give you my 2 cents.
I, I do feel it is a fairly pervasive big, big issue that all of us really need to, to look at, right? And it does mean just like Industrial Revolution did quite a few social changes, right? In terms of the way we work, the way we interact, the way we actually go to work and do stuff every day, right?
And that does require, um, a, a lot of adjustments from our side and I see two bookends, right? There are people who buddy their head in the sand and say, I'm not gonna even allow Chad GT in my organization, right? And then there is like, let's go and do, you know, like let's open up all the, uh, all the paps and just run multiple projects.
I think there is a fine balance that particularly the CIOs, the cd, aos of the world have to kind of really reach in their organizations and they have to figure out how to roll out the AI solutions with governance, with data safety, with security. And that's where I think professional help, like someone like NCE coming in or getting your people trained in the right technologies helps. On the other side of the coin, as individuals especially, I was at a, a career fair last weekend volunteering and a lot of like high schoolers and college kids came to me and said, how should I think about my careers?
Because I'm thinking software engineering will not be as attractive or should I be doing X or Y right? And I, I do feel we have to start thinking about some of the core skills that will survive, that will actually flourish and be in demand five, 10 years from now. Especially if you are early in your career, you have to pause, give it some thought and not rush into what was hot four years back.
We have to really think through the, the change is as a society come together and, and think through it. Alright folks, Shanan here, change a coming as the song says. And who knows, they might be entirely new careers that no one ever thought of, but it's gotta be nice to figure out how to absorb all this stuff.
Hey, you mesh. Thanks for being on the show. Thank you.
It's my pleasure. ai video series. You can find this episode and others on our website.
We invite you to check them all out. Until then, we'll see you next time.