Powering Autonomous IT with ignio AI Agents from Digitate
At AI Field Day 7, Rahul Kelkar, Chief Product Officer at Digitate, presented on powering autonomous IT with ignio, Digitate’s agentic AI platform designed for IT operations. Kelkar began by outlining the industry’s evolution from manual IT operations and cognitive automation toward modern AIOps and agentic AI, framing the journey towards fully autonomous IT as a progression through stages of manual, task-automated, and augmented operations. He described how ignio leverages a unified three-pillar approach: unified (or business) observability for comprehensive monitoring of both technical and business processes, AI-driven insights using traditional and agentic AI including machine learning and generative models, and closed-loop automation that not only provides recommendations but executes prescriptive actions with high confidence. This architecture aims to proactively eliminate business disruptions due to IT, identify issues before they impact business productivity, and reduce incident resolution times.
Ignio operates on what Digitate calls the “Enterprise Blueprint,” essentially a digital twin or knowledge graph that captures both structural and behavioral aspects of enterprise IT. The platform integrates with common monitoring and IT management tools, ingesting metrics, events, logs, and traces to provide a layered view of health across infrastructure, application stacks, and business value streams. Observability data is enriched with AI-based noise filtering, anomaly detection, correlation, and dynamic thresholding, automatically triaging and suppressing redundant alerts. Kelkar highlighted ignio’s “composite AI” approach, combining logical reasoning (rule-based and machine learning models), analogical reasoning (using generative AI and large language models for contextualization where knowledge is incomplete), and assisted reasoning (bringing domain experts into the loop to validate and tune recommendations). The workflow encompasses agent-based management of event and incident handling, automated root cause analysis, and remediation actions, all while learning from human validation to continuously improve performance.
The platform is designed to address complex, applied use cases in large-scale, modern environments, such as event reduction, proactive incident management, observability, patch management, and cost optimization across multi-cloud and containerized workloads. ignio supports integrations through out-of-the-box adapters for 30-40 major tools (with customization for specialized environments) and specific modules for SAP applications, batch scheduling, digital workspaces, and procurement processes. Its agentic capability is extended with Ignio Studio, empowering SREs and IT operations teams to continuously extend and customize workflows. As demonstrated, ignio’s AI agents interact via conversational interfaces, notifications, and dashboards, enabling a shift to smaller, cross-functional SRE teams—supported by autonomous agents handling the bulk of monitoring, triage, and remediation, with humans focusing on governance, validation, and improvement. This supports a vision of truly autonomous, resilient IT operations that adapt rapidly to changing workloads and technologies, minimizing disruptions and keeping business-critical systems running smoothly.
Recorded live in Santa Clara, California on October 30, 2025 as part of AI Field Day 7. Watch the entire presentation at https://techfieldday.com/appearance/digitate-presents-at-ai-field-day-7/ or visit https://TechFieldDay.com/event/aifd7/ or https://Digitate.com/products/ignio-aiops/ for more information.
Transcript
Hi, my name is Rahul. I'll be talking about autonomous IT today. Uh, we are Digitate.
We are all about agent tech, ai AIOps. The name keeps on changing, but our objective remains the same. It is autonomous IT operations.
So we have been around for 10 years. We have seen this transition from cognitive automation to AIOps to now Agent TKI. Uh, right now we are agent TI platform for IT operations around 800 people, two 50 customers globally, most of them globally, 2000 customers.
How we are doing is publicly available on G two. You can go there and look us up. Um, we have fantastic ROI for most of our customers.
All of what we have built is our own ip. So let me just walk you through what, what we are doing. Autonomous enterprise evokes a lot of emotions depending upon where, who you are, what your experience has been.
Uh, the way we look at it, it's a journey starts with enterprise IT operations being done manually. And this is largely the state of the art where there are people who are having a lot of tacit knowledge with years of experience. They kind of know if something goes wrong, why this is going wrong in their sleep, and they do things manually.
The challenge with that is it's not very efficient because whenever the issue happens, the right person has to be available at the right time at the right place, and there are associated challenges from there. Some enterprise take up task automation, and that's the assisted mode where there are machine assisted task automations that happen with somewhat lesser input. Still, pretty much the operation is show is run by people having tacit knowledge.
It's slightly better, uh, marginally better efficiency, uh, but it's still way off the objective of autonomous enterprise. From there, the maturity moves on to augmented. This is where now there is eradicated educated actions that trigger.
It may be not just tasks, but it, it may be compound actions, handling situations and things like that so that it starts mimicking the way people do things where they're thinking and acting. Uh, together actions are based on predicate actions pre, uh, predefined heuristics. And this is much better in terms of consistency, much less prone to errors.
And the nirvana state for us is autonomous. It where the IT or the brain that is operating the, IT is able to learn, adapt to changing environments, changing workloads, changing technologies. Uh, it has the ability to reason, think, and act so that situations can be dealt with and not actions.
Tasks can be performed only more reliable, lot more sophisticated decision making, not based on rules. A lot of ai machine learning context needs to be there, obviously for this to work. And of course there is a continuous learning feedback, which is what makes this interesting.
And that's really the north star for us. That's very, that's what we are aiming for. Yes, our purpose, we are committed to help enterprises become autonomous.
It. How do we do that? So we have a three pillar approach.
We call it unified observability, or sometimes we call it business observability. It is all about learning what is happening, observing what is happening, uh, plum, uh, integrating into the plumbing, which may be monitoring parts of the estate, but bringing the unified view of vertical stack observability, horizontal data flows, transaction observability, as well as experience observability outside in. Then there's that observability feeds into AI driven insights.
This is where decision making reasoning, insights, observations, lot of data mining, machine learning, uh, traditional AI agent ai, generative ai, all of this kicks in. And finally, we don't stop at just giving recommendations, but we actually try and close the loop by giving prescriptive recommendations with high enough confidence so that the value actually gets accrued and it's just not an advisory system. So this, this three pillars, they create a virtuous cycle and that's what we believe will eventually lead to autonomous IT operations.
So what does this help us with? There are three simple values that this provides. First is eliminating business disruption that happens due to it.
This everybody likes every, every IT operations wants to get this done. Second one is ensuring IT finds issues before business does so that you don't get a barrage of calls and then you are reacting to it. Third one is even if the issue happens, self feel fast enough so that the issues feel more like a blip and there are a lot, not a lot of pain minutes associated with the issue.
That's really what the three values that we are trying to derive using the three pillar approach. I'll explain how. So let's take a look at the prevent A prevalent IT operations model.
Right there is monitoring, monitoring generates alerts. Alerts go to the level one team. Sometimes they're called command center team.
They will assess, register, and assign that incident to somebody from there on the level two technology teams, which are organized in technology silos, pick it up. They try and triage and resolve the ticket. If they can't, they will escalate.
It goes to the level three teams who are again organized in similar silos. They will validate if the issue has been resolved at level two, otherwise they themselves will roll up their sleeves and triage and resolve and so on. There is governance, which is done by operations management team, which sits outside this silos.
And then of course there is problem management, which is supposed to bring in continuous improvement. What happens is a lot of activities are happening, all of these are done by a lot of people divided into teams and silos and layers. As a result, there is a fair bit of disconnect in the way this prevalent IT operation works.
So we did a survey for about 300, uh, 600 delegates who, who have been using AIOps for last two years roughly. And we looked at, uh, their responses on what kind of use cases they're trying to leverage in the future immediately or within the next 12 months. And it kind of, uh, brings to the fore the fact that all of these use cases that these practitioners who are responsible for IT operations want to get enabled.
They have the liability for it, are all applied use cases. They are taking care of, let's say, event management, which is going to reduce a lot of grief and effort that is spent in dealing with noise. Then incident management, which is going to take out a lot of errors and effort that is going into fixing the incidents reactively day and night, 24 by seven.
Proactive problem management, as you can see, is fairly high on the radar because this is an area which is often deprioritized, uh, for some burning issue, and there are always going to be burning issues. Uh, tickets, observability, patch management. These are some of the areas that came out of this survey.
Uh, in addition to this, uh, there are two more driving forces, which is in, in some sense fueling us on this journey. The first, first force actually comes from technology itself. Uh, there is a lot of improvement in the general term ai in the last couple of years, the veracity has improved, uh, accessibility has improved.
It is much easier. It is a lot more trustable. It's no more random, it's no more tie.
Applied use cases can be actually done consistently with it. So there is a lot of adoption happening. So whether it is gen AI or AI agents, agent, TKI or the more traditional machine learning data mining, uh, predictive ai, causal ai, all of it, right?
So this entire spectrum. Can I, can I ask you Ray Lu, Casey Silverton Consulting? Sure.
Can I go back to the prior slide? Uh, can you explain what, what's the vertical access here? So, um, in some sense they, their focus areas weighted.
So this is not a, a, a partition of the incident types that no, no. This is a survey on where what use cases Customers are seeing Practitioners 600 people are seeing right now and want to do it using agent TKI or AIOps. So number of respondents or percentage of respondents for answering In some sense, in some sense.
And the expectation here is that 12 months later with AI agents, this is how it's gonna look. Okay, so I will answer it slightly differently. Uh, this is where they feel confident in 12 months.
If I rerun the survey, they might give a different set of use cases as what they will be aiming for, But this is their expectations today As of now. Yes. Yeah.
Okay. So continuing, uh, so I talked about push from technology. As a result, it is almost mandate mandatory to have AI driving anything and everything that anybody does, and it's almost you get disqualified if you don't have AI in anything that you are trying to do afresh.
So it's a tremendous push and we like it. We are able to leverage a validated tremendously. The second driving force that we see is pull from businesses, and this pool is actually very strong.
And when I say business, it is the enterprise, its of the world. Um, most modern applications that are getting built and deployed are cloud native. So they run on microservices containers and things like that.
As a result, the, the type of management observability required on such architectures and such type of applications is very different. All of these applications inherently use a lot of AI as part of the business application stack because they are using ai. Uh, that's on the technology part.
On the operations part, there is this traditional pyramid you see, right? There is a large level one command center service desk team, smaller level two teams for smaller L three teams and our whole, Uh, we've seen various versions of MIT's report. So it's, you know, 95% of ai, uh, proof of concepts fail.
Um, so you're seeing that, that the businesses are actually implementing AI adoption in their, uh, container workloads and things of that nature. I mean, it seems like it's contrary to what we're seeing. Okay.
So, uh, I wouldn't say that all the workloads have ai, but the situation three years back was, okay, AI is in the lab. Yeah. And there is nothing in production, but that's no more the truth.
There are AI workloads in production Coming out. Okay. That's what I mean.
Yeah. Okay. So going back to the pyramid, a smaller number of architects, very few.
That pyramid is definitely getting disrupted to a more nimble SRE kind of model. Uh, people have different ways of calling it, but that, that's really a smaller number of people who are, who have multiple talents, multiple skills, if you will, and they have the end-to-end responsibility of dealing with something happening, continuous improvement at the same time to improve the overall efficacy, efficiency and experience outside in. Uh, that's the transformation we see, and that is another important pool that we have.
Uh, that's the driving force behind what we have. So you're saying that's all, can architects are the least most likely to recommend genic ai? Is that what you're saying?
No, That's not what I'm saying. Today's, if you take a more traditional IT operations setup, there are very few architects in that mix. Oh, okay.
The largest number of people are level two, uh, engineers who are very siloed. They just have one technology view. And whereas the newer model is not like that, there are no silos.
The people have a comprehensive view. There might be a view of, okay, I'm responsible for these set of applications, but there is no further division or no, I will only look at database. That's what I meant.
And uh, as a result, uh, there is a lot more emphasis being given on experience and not SLAs, because SLAs is a nice way of governing distributed, siloed teams that you set SLA at every level, level one SLA, level two SLA level three response time, SLA and resolution SLA, and then the whole game becomes reporting governance on are the s ls green or red and things like that. So there is this famous thing called watermelon effect where all SLAs are green, but if you look inside, it is all red. So that's what happens typically and fairly common occurrence instead of that.
Now, there are few people, I mean there are like 50 SREs, and thou shall be responsible for everything that happens. So there is no more watermelon there. They will be red in some sense.
So what we foresee, and we see a tremendous transition happening to this model. It's not like overnight it is changing to this model. There is still going to be monitoring tools.
They will still raise alerts. Now instead of these alerts going to level one team, the first right of refusal goes to an AI agent That is, that needs to deal with that alert. So the AI agent is now going to look at that alert, assess it, filter, suppress, aggregate, validate, write, assign it if it needs to be resolved, and then it'll register an incident once it registers, another AI agent will pick it up, try and see if you, the agent can do a root cause analysis, automatic healing of that issue, if not possible, because there are always going to be exceptional situations, new situations, uh, which have not been encountered in the past that AI is unaware of.
Then it needs to be escalated. And this is where the smaller nimbler SRE teams come into picture. The escalation point for a agency is going to be the SRET, much smaller in size.
Of course there is going to be operations management. It's required to govern the whole setup. What what is happening in this processes, same set of activities that are being done are now re aggregated into a fewer roles.
There is only going to be an SRE team and the operations management team, but then the scale is going to come through AI agents. That's really the model lot of activities in the earlier picture, all of those are still required to be done, but they are done by a fewer set of people performing more activities augmented by AI agents, which they are by AI agents are going to provide the scale. That's really the shift in the model that we foresee and we see a lot of our customers on this journey.
I wouldn't say that they're there and that will be foolish of me. Um, so finally, I mean, what are we trying to transform, right? If you take a very simplistic view of enterprise IT setup at the top, right, it starts with the infrastructure cloud, a mix of that at the bottom.
Then you have containers sometimes, then you have other platforms, database databases and other things. You have end user devices. Then you start getting into the appli, traditionally called application stacks, where you have a batch streams that are running, uh, batch processes, and then you get these applications which are bespoke.
It could be cost applications like SAP. And finally you have the business value streams. This is where the business function executes across multiple applications.
Transactions are flowing to get something done, whether it is ordered to cash, procure to pay, all those transactions are flowing across a multitude of these applications. And that's, that's really the stack we are trying to manage. When I say enterprise it, of course there is security which cuts through all of these layers.
Now, what does Igneo do? What does igneo bring to the table, right? We talked about three values.
So the first value ensure that it finds the issues before business does. It is enabled by two critical modules, functions, capabilities of IO if you will. First one is observe, the second one is analyze.
So what does observe mean? Of course, every enterprise has a multitude of monitoring tools. Uh, there is no point in bringing yet another, but we also have seen with 250 customers and 10 years of experience that there are always gaps in the monitoring.
There is always, uh, scope for improving the quality and precision of the monitoring. That is where this metrics, events, logs, trace kind of thing. We augment the existing plumbing.
Uh, there is always scope for, uh, hierarchical business specific health checks. These are not, uh, single point in time checks that you are not checking CPU, but you are checking a larger business level outcome, which may be dependent on a bunch of things going together. Hierarchically then there is business health because the health check is not just restricted to business application, but it needs to get into business functions.
That is this, let's say I have a merchandising business function in my retail operations. Is that business function healthy? That question can be answered only when all the applications used by teams which are delivering that business functions are healthy.
And then that question can be answered only if all the vertical stacks are healthy, and that's how the vertical stack kind of becomes green, but that's not the end of it. While all the applications may be fine, there, there are still going to be issues in horizontal business transactions that actually flow from application to application. And there could be various reasons for it.
Data issues, uh, network issues, many other issues, which vertical stack still may be green. Now, from this observability, there is a need to now leverage a lot of machine learning and AI to analyze, right? So there is this event management where there is a need to autonomously remove whatever 80 90% noise that is there in every anomaly that gets detected and raised as an alert, right?
Otherwise, the teams are going to spend, even if they spend 10 minutes on every alarm or alert, you get 10,000 in such alarms in a day, they're going to be busy forever. So you don't want the noise to reach those teams as much as possible. Once that happens, and let's say you remove 80% noise and 20% incidents remain, again, it's a lot of work because every time recurring is issues keep on happening.
There needs to be a way to automatically triage for perform root cause analysis and figure it out. There are scenarios where the IT metrics and KPIs need to be correlated to business I KPIs and metrics. For example, in a retail store, footfall is a important metric that the business measures and there is a direct correlation between if some past terminal on lane three is not working, footfall goes down.
And so it is, there is a necessity to correlate that context of business issues with IT issues so that the people have the sense of urgency to do it. Last, but not the least change that is almost single handedly the largest culprit for all issues that happen in it. Uh, things are stable when there is no change.
Suddenly you introduce change every day, every week, every month, something breaks, uh, because of the side effects from there. Second value, once it is observed, analyzed, then it needs to be self field fast enough so that it feels like a blip. You don't have a lot of pain minutes associated for the business, whether it is selfing incidents, auto fulfilling requests, change request, service request, or performing other lifecycle activities like certificate management, patching, and many other things that happen in a typical IT operational setup.
Last but not the least is eliminating business disruption that happened due to it. This isn't, like I said, an area that is often deprioritized because some priority one issue, some priority two issue is permanently. Where does, where does ignio, what, what is ignio here is that whole bottom section.
Yes. iOS, That is what I'm saying, Capabilities that you're bringing to the play. Exactly.
It is trying to manage the stack at the top, and I'll get into more, uh, details of how it exactly does that on this slide only. So, uh, the improved piece comes from ability to systematically eliminate tickets. We call it ticketless.
The idea is not, not logging tickets by idea is systematically not generating tickets, fixing, uh, putting fixes, elimination fixes so that you don't get the issues in first place. Capacity has been a perpetual issue for last decades. Cloud cost is becoming another dimension apart from response availability.
All of that cost is becoming critical because most business cases are not met when it comes to, uh, moving to a multi hybrid cloud kind of setup. Uh, that's the improve piece. All of these set of features are running on what we call as ign new enterprise blueprint.
It is nothing but a, you can call it a digital twin of your enterprise. It captures both structural aspects as well as behavioral aspects of the it, uh, the modern name for it in the industry's knowledge graph. All of these, of course, uh, there are a bunch of AI agents and I will go through that in the subsequent slides.
These AI agents are purpose you. They are augmenting day in the life of specific people. I talked about a lot of roles there on the earlier slides.
These agents are doing lifecycle activities, whether it is event management, incident resolution, problem management, business, SLA, predictions, cost optimization, and they are collaborating and augmenting with some of those personas on earlier slide. Of course, the ways to collaborate and uh, uh, collaborate are through conversational assist notifications, dashboards. Many of these modern sari kind of personas want to be hands-on.
So they want the ability to extend existing functionality on a continuous basis. So that's why we have Ignio Studio, which allows them to sort of extend functionality. Now in this case, are you using standard LLMs to drive these, uh, agent uh, workflows or are you doing some special services, special models that are associated with, uh, enterprise ops or, or, so We have a, a three pillar approach there.
Somehow. We like three pillars. The first pillar is what we call as logical reasoning, um, which is based on machine learning.
Traditional AI models, a lot more deterministic. Uh, the second pillar is analog reasoning, which is dependent on generative models, uh, which, uh, work on a general book of knowledge. Third pillar is assisted reasoning, which is where we bring in the experts that have the tacit knowledge so that the generative analogies that are coming out of, uh, general book of knowledge can be specialized personalized for that specific enterprise situation.
So we use this what we call as full spectrum or composite ai. There are other names for it, and that's what really it is. It's not one answer that we are using this open, this or that, and I'll get into next slide, um, into details.
I'll give some examples. Can I, so let me frame what I've heard so far. Scotton al, um, interesting architecture makes sense.
I feel like you've made claims that I can't tell if you mean they are present state or our future. I think the best way would be to see sure what you're doing and with some real customer examples Perfect. That demonstrate auto autonom, auto autonomous genic behavior.
Absolutely. I don't mean you have to fast forward through everything. Okay.
But I think that would be the best way to answer I think some of the questions we have. Perfect. So we have, we have organized this session in three parts.
I will go through the product, then I'll go through demos for 25 minutes. Uh, all of the agents I will demonstrate, and then Rajiv will cover the customer, uh, scenarios. Okay.
You're right. I remember five customer from last night. Yes.
Okay. Okay. So moving on, uh, let's skip this.
So we have, like I said, right, we have been around for 10 years and it's a field that is changing, uh, quite rapidly in the last two, three years because of whether you call it gen AI or any other thing. Traditionally, AI ops, the way the industry has defined, and it is also evolved over a period of time, it is all about contextualization, correlation, anomaly detection, prediction if possible. We as Digitate have always looked at AI ops as not just these three pillars, but we have always, uh, looked at it in conjunction with closed loop automation.
The point of detecting anomalies, point of predicting and prescribing fixes needs to take the last step, last mile and apply automation and resolve those issues. That has always been our definition. From there, we have matured into leveraging gene AI to bring in efficiencies, right?
With previous releases, whether it is chat bots or code accelerators here and there, so that the overall value journey for a customer becomes accelerated with the latest release that we have out in the market. That is the agent T care release. So to answer your question, I'm talking about what is out there generally available and not coming out in the future.
When does your, what does your timeline start? So what year would you put on your first AIOps Bubble? So we have, uh, okay, so our journey 2015, okay, we called ourselves cognitive automation.
So the idea was, I mean, there always has been automation, runbook, automation, whatever you want to call it. We wanted to make it a little bit cognitive so that it is able to sense and understand in what scenario, what automation to use, and that that was really the play for first couple of years. At that time, there was no category called AIOps.
Nothing was defined as we started doing that, deploying it across customers this category emergent, this was great because we were naturally born into that category. The category was saying almost 80% of what we were doing. We were saying 20% additional, and that's where we started calling ourselves AIOps for, um, IT operations with the three pillar story, okay?
And very recently, we have now that scaled in, into this modern stack that I showed modern IT operations way of operating agent DKI platform for IT operations. So that's really the three, if you will, three phases of evolution. Okay, So this is the slide.
So the ai composite ai, right? I talked about logical reasoning. Uh, this logical reasoning is based on models, so it is a lot more deterministic for a given model.
Uh, uh, it leverages a lot of enterprise blueprint, which of course needs to be populated by doing a lot of data mining, machine learning, it captures structure, dependencies, behavior in terms of workloads and monitoring data, event data, incident data. Then model kicks in. There is of course the domain IT model that we bring into the table.
Uh, combined with the enterprise blueprint, which is situational model. Uh, there is model based reasoning, case based reasoning. A lot of that is enabled by machine learning and traditional, uh, causal and predictive ai closed loop automation.
Of course, this adapts to changes and it frees up people's time in the, uh, prevalent IT operations model. That's really the first pillar. The second pillar is now the challenge with this is this is assuming that, uh, sufficient information is available for the machine learning to kick in and take decisions.
In many scenarios, dependencies are partially known. Monitoring data is thoroughly populated. There are gaps in that data.
As a result, one has to make assumptions about both the dependencies as well as the behavioral model in a typical IT setup. And that's where it starts becoming, uh, time to value starts becoming longer because then you spend time in populating that blueprint so that the answers to the model are more trustable or you go on, right? So that is where we started, uh, bringing in this, what we call as analog reasoning.
Uh, it augments this existing knowledge of enterprise blueprint and the domain models with analog reasoning, right? This is the general book of knowledge that I was talking about. The combination of this worked great because, uh, it, it solves this problem of, uh, I don't have complete information in all parts of my IT operational setup.
Wherever you don't have, you don't have to wait for populating that, all of that. So you start getting value right of the bat. And then the advantage of using en analogical reasoning is there is this inherent ability to learn, adapt, and generalize so that you are able to now apply some of these learning into situation, which may not be exactly same when it happens subsequently.
The last pillar is assisted reasoning. This is very important because neurological reasoning is going to get you generic advice, which is general book of knowledge, right? It needs to be contextualized and, uh, specialized for tacit understanding of the, that specific understanding.
So there is going to be collaboration always between neurological reasoning and assisted reasoning, and that is where we see that the experts come into picture and overall effectiveness of the system becomes very, very sharp, if that's the term. Advantages are great. I mean, you don't have to populate, spend a lot of time in populating context before you start delivering value, uh, lower variance, because now this is an approach which is augmenting different personas that are playing that role.
As a result, there are going to be experts who are super smart, right? There's no issue with that, but there are going to be people who have just joined a couple of years or three years sooner or even lesser than that. They're not as effective in dealing with situations, but this augmented approach kind of augments everybody.
So the overall effectiveness of everybody is going to go up as, as a result, the variance will drop between the elite versus the average. Uh, of course this is a great way of dealing with complex unknown scenarios because in the first column, if there is an unknown scenario, I mean, there is nothing that model will do, but this is a systematic way of dealing with that so that next time a similar situation happens, the whole composite is able to deal with it. This is really the crux of what we are doing, and let me now, uh, explain this a little bit more in details.
Sorry. Let's take this example of event and incident management. I talked about enterprise blueprint that needs to be populated, whether it is the application stack.
I mean, this application runs on this server and it is connected to this database, and there are a bunch of dependencies like that. Uh, in addition to that, there is a behavioral aspect of it. That monitoring data is that there at each level, and all of that blueprint needs to be populated using a multitude of data mining and other techniques.
You bring in this information that is available in a lot of enterprise IT data sources, whether you, it is monitoring tools, ticketing system, excel sheets that people maintain when they're managing their own databases or servers. All of that we, uh, bring together, of course there are conflicts and we figure that out. Now, what do we do with it on that?
We run a bunch of machine learning algorithms to reduce expected normal behavior. So this is going to tell us this server, uh, this application is going to behave at this level between 10 and 12 on a Monday morning. And the subsequent servers on which this is running the utilizations levels are going to be 74% to 79%, right?
And this is what is derived using a lot of machine learning from the data. There is graph data dependencies. There is time series data, which is all this metric information.
There are few more things that happen, uh, um, as part of logical reasoning, right? Uh, using the same information, we also mine a lot of filtering rules. These are rules or association, we do association rule mining on that data to figure out are there scenarios where people are just not acting on certain types of alerts, having certain characteristics.
And these are great candidates for filtering because anyone, nobody's going to act on it. Uh, then similarly, we figure out co-occurrences of alerts that are happening together, either, uh, one leading to another or they are co-occurring together as bunch of alerts. We also figure out dynamic thresholds I talked about on a Tuesday morning between 10 and 12.
This is the level of utilization expected based on, um, significant persistent changes that have been observed recently. Classic machine learning. Um, and this is all done through the logical models.
This is where now assisted learning comes into picture. There are going to be experts who come into play, uh, play. They look at the suggestions for filtering rules.
They will look at the evidence provided. They will confirm those rules. They look at the co-occurrences, they will promote those co-occurrences as correlations when it comes to alerts, and they will look at the dynamic thresholds that are suggested and they will confirm those thresholds as a result.
Now, the machine is now more tuned for filtering irrelevant alerts. Again, logical reasoning, suppressing redundant alerts based on the correlation signatures and suppressing noise based on the dynamic thresholds. Now, what happens is, as part of this exercise, there is always going to be, uh, some parts of the enterprise blueprint that are not fairly populated currently.
So naturally, that shows up as errors in this process. So naturally, and our side effect is suggestions for blueprint enhancement so that the overall effectiveness goes up. These enhancements, again, are again, are done by assisted reasoning to enhance the blueprint through conversational interfaces.
Now, let's say this happens and over a period of two, three months, this becomes more and more efficient. 70, 80, 80 5% of the noise is now reduced from your alerting system. Now all the incidents remain 15% they need to be now dealt with, right?
So there is automated root cause analysis. If the root cause analysis using model-based reasoning is confident enough and there is high enough support, then it goes into auto resolutions because the prescription is good enough, uh, with high enough confidence in scenarios where there is not high enough confidence or not sufficient support data. Then again, experts come in, they confirm the root causes, identified, the fixes prescribed, and then the machine keeps on learning from it, right?
All of this can be done through conversations. Now, in scenarios where none of this is working, the machine simply doesn't understand what to do with this scenario because it has not happened before or not sufficient information is available of dependencies around that. This is where now analogical reasoning comes into picture by now suggesting summary, providing historic context, debugging context automatically suggesting fixes through general book of knowledge.
And these fixes now need to be assisted again with tat because there is this risk of something, uh, being said, which is not right. Once that confirms, then that fixed case gets operationalized the next time. So the idea here is it is always, it has always been composite AI for us.
We are not dependent on a specific LLM or anything like that. We always try and do everything logically first, if not analogically and then assist it. That's really the crux of the basic idea behind Ignio.
Let me take an example. Let's say I'm an SRE person and uh, I come in the morning and naturally I will ask this question that is there anything that needs my attention? And now the perception part of the agent kicks in.
It'll monitor the, the incidents that have been happening in the last few hours. It knows who I am, my context, what is my responsibility. It is now going to infer, um, give it to the reasoning agent, which is going to now assess impact of each of those incidents, prioritize those incidents, and then present me a list.
So now, as an SRE person, otherwise I would've to type shoulder, tap three, four people, look at some Excel sheets dashboards and figure this out so I can simply ask this. Okay, great. Then I pick one of those and say that, can you help me diagnose this issue?
Now, perception agent is going to kick in. It is going to find out, let's say this application is slow. That's the incident.
It is going to find out, okay, what are the influencers in that hierarchy of that application? Is it running on a, uh, Linux server Oracle database? What is the app app server?
Is it running on a virtual machine? Is it on cloud? Is it on-premise?
All of that, right? This is the dependency hierarchy. Then, uh, it is going to infer the expected normal behavior.
I talked about machine learning, defining the envelope of workloads. Then it is going to perform a bunch of realtime health checks. It's almost like a blood test, if you will, right?
You perform a blood test, you compare, uh, results of various metrics of that blood against ranges, right? The range is defined by the expected normal behavior in this case, which is dynamic in nature. And then whatever is out of that range actually is anomaly.
So the diagnosis becomes simple, okay? The, I checked at 14 parameters across three dependencies, and two of them are out of the expected normal range, and that's the result. I get that.
Okay, maybe those are the probable causes. So now I ask the, uh, agent that, okay, give me a historic context of this. Have similar things being seen in the past and what happened when this happened?
So then again, perception agent kicks in. It tries to identify a cohort of incidents that are similar along multiple dimensions. It then tries to infer common causes that were happened that happened during that incident.
Common fixes that were applied, it worked, didn't work. Whatever it is, if not possible, then it'll go back to analogical reasoning, reach out to an nl, LLM, bring that in, bring human in the loop, and that's how it continues, right? Okay, I'm happy with the fix that has been suggested analysis that has been done.
Go ahead and apply this fix and tell me if it fixes the issue. The agent is going to then check for safety and conformance of that fix, validate that fix in some sense, apply that fix. It may be one fix, it may be multiple fixes at multiple layers in that hierarchy.
And of course, it is going to learn from this investigation. And this is exactly how the way I explained earlier, the three logical reasoning and logical reasoning assisted, this is how it is going to work together. I will of course show a live demo, uh, subsequently.
Now, uh, I talked about lot of personas that are playing different, different roles. We want them to be only SRE teams or other enterprise. It wants them to be only SRE teams, which are smaller and nimble and many enterprises are there.
But otherwise it is going to be a mix of traditional operations teams. Some SRE, there is the CIO it management layers. All of these people actually perform a lot of day in a life jobs, right?
For example, a or operations team. They will monitor, they will investigate, they will resolve, they will validate, they will improve, they will govern. Whereas the IT management or CIO kind of, uh, personas, they're largely going to focus on improvement, transformation, and governance.
Right Now, I talked about these agents and I just used five agents as example, who are going to augment the day in the life of many of these personas. You can imagine from the example, given most augmentation is going to start from left hand side. So the agents are going to first do a lot of augmentation on monitoring investigation resolution, and job of people will slowly remain as validation, improvement and governance.
And slowly the agents will also start suggesting improvement opportunities. As a result, the people are going to only handle exceptions eventually. That's really the dream.
But we believe strongly that that's where we are going, uh, as part of this autonomous enterprise. Yeah, Sorry, Carl Fugate. Can, can you go back just one second?
So I'm just curious, um, and we can talk about it later if, if it makes sense. So from the monitor and investigate kind of phase, um, you know, especially monitoring, uh, there's, most teams have tremendous amounts of logs, um, which are very expensive to, to move. Yes.
Um, but I would imagine your tool is going to need access to those in order to, to provide value. How do you, how do we go about that? Yeah, great question.
So, uh, the more the merrier, that's the simple answer, right? So let's say if you have a hierarchy of an application, there is going to be monitoring data, some metrics at application level, application server level, database level O waste operating system level network storage. And those are relatively easier to get because it's not that expensive to get that much data.
The challenge there is, uh, what metrics are getting monitored, and that's where I talked about ignio coming in and augmenting this thing because unless the model is not sufficient to derive conclusions and reason about it, it is going to not produce good results. Having logs is going to be great, but you don't really need entire log. That is our experience.
All you need is anomalies from the logs, which is going to be a lot lesser than transmitting terabytes of data. So our approach is only deal with log anomalies and not complete log. So if you already have log at some place, we have ability to just extract anomalies and deal with that.
Okay. Do you, so do you, do you need just kind of access to raw metrics or do you have plugins for, to support other, other tools that you, that you might already have? Yes.
So I mean, yeah, I mean, if we always run into yet another tool, uh, but more or less the standard 30 40 tools that are out there, we cover out of box adapters for those. So that's not an issue. But like I said, uh, there is always yet another tool that is used by somebody so that in that case we build that adapter.
Okay. Many a times the tools are supremely customized, so our adapter doesn't work out of box, so we will have to click it. Okay.
Okay. Going back to the enterprise autonomous suite, uh, of course AIOps is our main offering because it covers infrastructure cloud, application stacks, but there are specific nuances related to parts applications. Like S-A-P-S-A-P by itself is an interesting, I don't know whether to call it animal, but that's the word.
And it has these ECCS four HANA and suit on hana. I mean, there are those variations that are there. So we have a specific offering where it is trained on specific SAP architectures and models because it's a different, different game.
Uh, then we have a specific module, similarly trained for batch schedulers, right? So lots of enterprises run literally tens of thousands, if not hundreds of thousands of bad jobs, and they just keep on running, they keep on firefighting. So we have a module which sits on top of schedulers and uh, it specifically deals with that.
Then we of course have a specific module for digital workspace endpoint. Uh, procurement is another area. It goes into the business realm, but we have a module there and software assurance for release validation.
So those are really the, uh, if you will, packaged products on this platform that we take to the market.