Agentic AI Is Ready – Are Enterprises?
The rush to automate is hitting a major speed bump as enterprises realize that bolting a high-performance AI engine onto a legacy horse-and-buggy process is a recipe for a pilot program disaster. Harshil Shah, Director of Engineering for R Systems, warns that the real challenge isn’t building the agents themselves, but rather evolving our governance models from simple software checks to role-based oversight that treats AI like a new member of the engineering team. To survive this transition, organizations must move beyond the hype and prioritize an “observability-first” approach that builds evaluation and security guardrails into the foundation long before the first agent goes live.
Transcript
Hello and welcome to the latest edition of the Techron AI Leadership Insight series. I'm your host, Mike Bazaar. Today we're with her Shaw, who's director of Engineering for our Systems, and also heads up there, Gente ai Center of Excellence, our show.
Welcome to the show. Thanks. Thanks.
Thanks Mike for having me. Really excited to be here. And we are talking about agents, So we're gonna have a little chat about AgTech AI and the workflows and how it fits into the enterprise.
'cause it seems like we're running into some challenges. Everything from governance to well, uh, change management and even cultural issues. From what you've seen so far, what are kind of the biggest issues that we're kind of struggling with now?
'cause I think building the AI agents themselves is not even half about. Exactly. Right.
So I think, uh, the, the top two or three issues that we see when we work with a lot of enterprises is first is that enterprises are quick to jump to implementation. They want to, they'll identify a process in their existing, uh, ecosystem, and they'll, they'll have the mindset that let's introduce an agent to actually, um, help me with this process right now. That what that that approach does is the, the analy, the analogy I like to use is like the horse carriage and the analogy, right?
That you have a horse carriage today and you figure out you have, uh, uh, you know, you have a really great car engine waiting for you and just go and put it to the horse carriage. It won't work, right? Because horse carriages were never meant to be driven by, uh, a, a Ferrari engine, right?
So there's a lot of, um, issues that they run into when they realize that, oh, okay, if when I tried to bring agents into a certain ecosystem and I tried to jump to an implementation, a lot of, um, consequential issues came up, which I had to deal with, and, you know, then I had to scrap the whole pilot. That's one where readiness of, uh, the data, the ecosystem to accept agents was not done before the jumping to implementation. Second, uh, the governance framework around agents is something which is not a very well established, um, uh, framework out there in the world today because this world is new, right?
So enterprises, they think of agents as software. They think of agents as something which, um, I need to consider as a piece of code that is running in my system. But that, that, that never sits well, because when you look at a certain, um, addition to your ecosystem as software, you kind of change your mentor model to governing the software.
But this is not a piece of software. You are adding someone as good as you, you're adding a piece of software that's as good as adding a new engineer or adding a new role to your team, right? So the governance needs to be as per, you know, the governance model needs to change to actually embed role-based governance for agents rather than software based governance for agents.
So these are like the two top issues that we see with enterprises when they start with agents. And the third largest issue is they don't have visibility into how the agents are performing because the first step of building a concrete evaluation and observability pipeline for an agent was never done, right? That that doesn't happen, uh, typically with enterprises today.
So being able to go back and check why the agents are fail fa failing is an, is an answer that enterprises fail to on is, is a question that enterprises failed to answer when they go into this journey. So how do I determine if I am ready for AI agents? Are there a set of, I don't know, tests or frameworks or something that I should be applying here?
Or how does that kind of work? Yes, Yes, absolutely. I think first very important PO point is, is your data ready for to be consumed by agents, right?
Today, um, a lot of datas, a lot of the data or the system of records for the enterprises sits in a lot of different silos. Um, it sits in, um, it, it sits, it sits as tribal knowledge within existing team members, right? It is very important to understand that today agents are not, do I have a unified view of what my product does as an enterprise, right?
Do I have a unified place for me to go and check what all, what, where all, where all my data actually lives, right? That part is step one. Where do you, is is your data readiness in place so that agents can consume the data to take actions or decisions on top of it?
That's step one. Step two is, are you ready to put, uh, like, do, do you have the right infrastructure in place to track and evaluate and observe the agents that are being run in your ecosystem? Right?
Do you have the right logging mechanism in place? Do you have the right tracing mechanism in place? Right?
All of these questions need to be answered first before you jump onto bringing agents on board, right? So that second piece is very, very important to make sure that any kind of agents that get deployed in the enterprise they need there, there needs to be a strong evaluation and observability layer that supports it so that at any point in time I can go and check how the agents are doing. How are the last 10 runs of the agent?
Have then been very major errors? Have we been seeing hallucinations? As long as you are able to answer these questions by going back and looking at a dashboard, then I think that's a good sign of readiness.
And third is with respect to the organizational mindset, right? The organizational mindset needs to evolve from being, um, you know, today enterprise employees are people who are doing certain processes themselves. That is their sense of that is the, their sense of value they bring to the organization.
So for when you bring agents on board, a lot of friction that we've seen in the past is adoption doesn't happen because the employees tend to think that this is gonna take away my sense of value to the organization. But that evolution or that messaging across the organization needs to be passed on, that you are evolving from the role of the person doing the process to evolving to a role, to a person who is orchestrating the process with agents or is supervising the process with agents. So that organizational change management in terms of the mindset of employees, that's very, that's very important to establish readiness of the, uh, enterprise to start using agents.
One of the other things that comes up is security. And it feels like once again, we are deploying some emerging technology without thinking through the security implications, but, um, AI agents have or should have some sort of identity and some sort of permission. Yes, I think we have a tendency to let them inherit the permissions of the humans, but maybe that's not a good idea because well, shall we say, there tend to be a little overly aggressive in terms of what data they go look for.
So, yeah. Um, how do we kind of navigate this, Right? So I think it's necessary to understand it's necessary that even before you build the agents, uh, you put a simple, um, um, you put a simple list of responsibilities that are tagged to each agent, right?
And, uh, also not having the mental model that one agent is doing the job of a human. It could be several agents that are doing the job of a human. It could be, uh, several agents doing 30% of the existing job that the human is doing, right?
So it's essential that when you're go building from the ground up, attach responsibilities to the agents, at least tag them after that's done, after you, after you define the responsibilities of what the agent is supposed to do, then you need to audit what systems is the agent gonna talk to, right? And for each system the agent is gonna talk to, as for the responsibilities that you've developed in step one, it is necessary that in step two, you de you decide, um, what kind of permissions is are needed for this agent to perform this responsibility in the most safest way possible, right? So inheriting what, inheriting all the roles and inheriting all the access controls that a human had doesn't really add up.
So you have to build from the ground up that you attach responsibilities, you figure out the ecosystem that the agent is gonna talk to and only allow certain level of accesses or to the agent to interact or take responsibilities only for that specific level. That way that is the new way of setting up, um, agents, uh, identity and access management for agents where you define a clear framework of how much access does the agent have, right? So that naturally will ensure that even if there are prompt injection attacks, even if there is the agent goes, you know, out of his way to try to access or try to read some data that it's not supposed to, your access control layer that was tied back to the responsibilities is covering, or, you know, adding the security layer for it.
It seems like, and understandably c-level execs are obsessed with, you know, well, what's the ROI here? Um, you know, are we gonna be more efficient? Et cetera, et cetera.
And there's two things that come to mind about that conversation that I wonder if people are working their heads around. But the first is, well, not everything that we're gonna apply AI to is gonna deliver a competitive advantage, right? Because on a certain level, it's the new table stakes.
If I wanna remain competitive, I just have to be able to have this capability 'cause well, everybody will have this capability. Is that just kind of one of the fundamentals? I think what Henry Ford had said a lot of years back is that it is useless to innovate on something which is already working, right?
So if let's say you have a very concrete workflow that is doing a very, very good job for you, it could be a piece of code, it could be just an established framework. Then introducing agents there which have, which are non-deterministic in nature, they have their own sense of reasoning is where you will start seeing the friction points of trying to adopt AI for the sake of adopting ai, right? So what typically C-level executives that we work with, um, you know, and we, we, we consult them, we tell them that you need to identify the top five use cases in your organization that require reasoning, that require high level reasoning, that require high level autonomy, and then the most deterministic way of building that either through code, either through introducing LLMs at different places.
Try to do that, right? Try to make it as deterministic as possible, and if it still doesn't work, then try to bring the agents on board who are trained to do the multi-level reasoning to be able to achieve that task. But you have to start from the bottom up where you try to make it as deterministic as possible and then get to a stage where the agent is doing the reasoning and doing the job for you.
So I think for C-level executives, typically they start from the top down, but it actually needs to start from the bottom up. To your point about that, in some conversations I've had with folks, they're saying that the process that they're managing using AI is now taking longer, but the output, the quality is better. So, you know, what they're doing is working through multiple AI agents to orchestrate something, and it's taking them longer to do the task, but they are creating something that is better at the end of the day, but it's not something that necessarily the company can charge more for.
It's just more something about, you know, it gives the end customer a better experience, but it's not, shall we say, monetizable. Yeah. Got it.
No. So see, I think, um, those kind of cases where your time has your time to do orchestration more is, uh, actually more than what it would've taken to actually do the job yourselves. That the, the value add that does is that you have to figure out that, is my agent providing quality?
Am I charging for quality or am I charging for speed? Am I charging for acceleration? Right?
In most cases you are charging for acceleration. In some cases you are charging for quality. But yeah, definitely being able to, um, frame your narrative about what value add this agent is introducing and then trying to monetize it is what is what, you know, setting the narrative straight from the scratch, like from the get go when you start interacting with your customers that, okay, this is the value add that my agent will provide, goes a long way, right?
So a very good example is, um, let's say, uh, project management, right? Today, project management tools out of the box do a really good job, right? Um, if I introduce AI agents there, and we've done it for a couple of customers, typically the, the, the time it'll take for me to orchestrate agents to create dashboards or to, uh, run, run my sprint boards, et cetera, will take longer.
But eventually, if, let's say the quality it introduces is that each ticket has more descriptions, has more informations, uh, my, my cadences of making sure that my project management views are all quality and they're up to the mark, anytime my C-level executors wants to come and check my project management boards, they're always updated. That level of quality, if it is bringing on board, then that is definitely something enterprises are willing to pay. So it's just about figuring out in terms of monetization, you have to think about what is it that you're offering through AI as a quality of life improvement.
So it is either quality or is it acceleration? Is it speed? So what's your best advice to folks?
What are you seeing them doing today that just makes you shake your head a little bit and go, folks, we need to be just a little bit smarter than that. Yeah. Uh, I spoke about this earlier, but definitely I think building agents without an evaluation and an observability layer doesn't work, right?
Because, uh, the LLMs are getting so good that everyone wants to jump onto attaching your ecosystems, attaching MCP servers start asking questions to the l lms run agents on your data. They wanna do that because that's exciting. But what people don't do is the other side of it where you establish scores, you establish metrics, you establish accuracy scores, you establish semantic similarity scores, which will help you understand how your agents are performing, right?
That is something which people do it after the agent's implementation is done or after the agent is live, but ideal, in an ideal world, you have to define how will you score the performance of the agent beforehand, and then you have to jump to the implementation only then you are in that 10 or 20% bucket of successful agent tech pilots that enterprises do that will enable it. Second thing, when you jump to implementations, one thing that I've seen is people do not take guardrails and security vectors very seriously. They will to a certain extent add some level of security, but there is a term called red teaming.
Red teaming of agencies trying to break your agent system to the fullest, right? Trying to do prompt injections, trying to poison your knowledge store, trying to do agent, agent to agent escalations, where I tell one agent that, okay, you know, try to do this and another agent. So trying to red team your agent first and then deciding the security architecture of the agent to reduce your attack vectors, that will go a very, very long way because especially for, um, consumer facing or B two f B2C facing conversational agents, the attack vector will blow up very significantly if your user base blows up very significantly.
So these are some of the bits that we, that I've seen that you have to take an evals and observability first approach, and you have to take a security first approach to, uh, building AI agents or even using AI agents in your ecosystem. All right? Hey, folks, you heard it here.
Even in the age of ai, there's no substitute for the fundamentals, and if you skip them, you're just gonna pay for it harder later on. R thanks for being on the show. All right, thank you, Mike.
Thank you all for watching the latest episode of the Techstrong AI Leadership Series. You can watch this episode and others on our website. We invite you to check them all out.
Until then, we'll see you next time.