Cribl’s Nick Heudecker on the Evolution of AIOps in the Generative AI Era
Nick Heudecker, head of market strategy and corporate development for Cribl, explains how artificial intelligence for IT operations (AIOps) is evolving in the age of generative AI to address the needs of both IT professionals and AI agents alike.
Transcript
Hey guys, thanks for the throw. We're here with Nick Hoecker, who's head of market strategy and business development for cribble, and we're talking about AI ops and how it's evolving in this new age of gen ai. And well, there's a lot happening.
Hey, Nick, welcome to sha. Great to speak with you again. Thanks for having me.
We've been talking about AI ops for a while, and I think for the most part it was always in the context of machine learning algorithms and predictive stuff. And these things were gonna learn our environments and surface some interesting insights in the way we went. Now it feels like we're moving past Gen AI into two phases.
One was the copilot phase, and then there's the rise of the ag agent phase, which has a little more reasoning and is a little more autonomous as you kind look at this. Where are we on the journey for AIOps AIOps? Well, I, I think AI ops, it really turned into, you know, you, you're seeing that morph from AI ops into, you know, what analyst firms are calling event intelligence solutions, right?
AIOps is kind of, it, it really turned into a zero billion dollar market. Um, there was a lot of, of, of hype there. There was a lot of ambiguity about what an AIOps solution was supposed to do.
Um, and I think, you know, just calling it AIOps, right? People really focused on the AI and they got into that hype and they didn't focus on the value. And so now that we actually have some AI out there, um, people are realizing, you know, the AIOps story didn't quite deliver and now it's time for a bit of a rethink.
Gen AI is gonna be part of that, A agent is gonna be part of that. But, you know, let's foc let's call it something new, right? Let's focus on the value that this can provide, um, and move forward from there, right?
So I think you're seeing kind of a new, a new type of category emerge that hopefully allows IT leaders to focus on, you know, solving their actual problems, not chasing this kind of mythical hype beast of AI when trying to resolve their IT operations issues. And how tailored are these solutions gonna be? And I asked the question because every IT environment I've ever been in is just pretty much a snowflake and doesn't really have a lot of commonality with different things.
And how do I kind of take all that and expose it to something that looks like, you know, multiple kinds of AI models to get something interesting? I, I think you're, you're gonna have a lot of training periods that, that take a while depending on how diverse your IT environment is. Right?
At cribble, we work with really large organizations. You ask them what technologies they have in-house, they say yes 'cause they have everything. Um, two or three of the same things that, that are very identical.
Uh, some older, some newer. And so you're exactly right. Like there's gonna be a huge training phase, uh, as these agents learn the environment, um, the new systems that are coming out, right?
Or this new category of EIS, they're gonna have to figure out what the topology is. So they're gonna need really clean, accurate data that is going to delay. I think a lot of the advantages that, uh, companies hope they're going to get out of these new, you know, ag agentic systems, the companies that are going to deliver on the best solutions are the ones that already control that data.
Companies like, like IBM or ServiceNow, right? They, they know what's in your environment. They have access to the ticketing systems if they don't already run them.
Um, configuration management databases, all of these things are gonna, you're gonna need access to that. You're gonna need, um, really clean contextualized data to make these things work. How will the role of IT people, DevOps people and all the folks that make up the communities that we have been around for years change as we kinda get to this new way of thinking about applying AI to IT operations, no matter what we call it.
I think that's gonna take a while to really flesh out, right? Like there's been a lot of conversations happening in, in social spaces around, you know, IT operations managers, IT operations leaders as well as like software engineers. The one theory is they're gonna be running networks of agents, right?
They'll be managing these agents that are doing this work for them. Maybe that feels a little optimistic. Um, but I think it's going to be, you know, a lot of augmentation, right?
Me, as an IT operations leader, I'll be relying on AI to tell me what's important. I will still make the final decision. 'cause I don't wanna outsource, you know, potential downtime or other operational costs or opex to an agent without me blessing that.
Uh, so I think you're gonna see like a mixed model for quite a while, right? Where you've still got very much human loop, uh, for the next few years at least. Um, and then over time that may get more and more automated.
As these agents get smarter, the models get better and frankly, the models get cheaper, right? One thing that people don't really consider is these models in, in an operational context have to process a ton of data in real time. The economics don't really support that yet.
So we're not quite there. Even if I wanted to go full autonomous, whether it's my security operation center, my IT operation center, the money's just not there, right? You're still incurring huge costs to operationalize these models.
So I think there's a lot of things that have to happen, but I think over the short run, it's going to be, you know, I'll augment my teams with some of these automated solutions. I, but I'm always gonna have a human in the loop. And over time it's anybody's guess, we'll have to wait and see what happens.
But ultimately it does sound like people will be supervising AI agents that are performing tasks on their behalf. And some of those AI agents may communicate with each other to complete some tasks, but ultimately they gotta report back to somebody. They do, they do.
Somebody's like a, a human is ultimately responsible, right? You can, you can delegate the authority, but you can't delegate the responsibility. And so you're always gonna have to have someone who's approving or, you know, validating what these models are doing in the enterprise.
And that's a, it's a really different kinda role than a lot of IT folks I think are used to. So they're gonna have to ramp up on, on what that really means. How will AI keep pace with their rapid pace of change that we see in our environments, right?
I could argue that one of the issues there with AIOps was the machine learning algorithms were supposed to learn your environment, but you were constantly changing the environment. So they were constantly learning and not actually doing something. Um, we can see that issue getting, starting to get even more complicated when we're using AI coding tools to create more code and update environments faster than ever.
How do I keep pace on the other side where the AI models are supposed to be trained and updated on a regular basis? Is that gonna become continuous or how do we think about that? I don't think enterprises can really adapt to kind of continuous, like, you know, you may, you may be deploying code, you know, 20, 30 times a day.
You understand what that process looks like, you know, kind of what the splash damage is going to be. But if AI is going to be running your entire environment, it can be difficult to really understand what the second and third order effects of this is going to be. So I think you're going to see like kind of staged rollouts of these updates and then, you know, AI will continually learn and understand that topology, whether you're in the cloud on-prem or some kind of hybrid solution.
So you're always, this is another kind of instance of this is where human in the loop makes the most sense because you still want, like, you, you don't want the AI to kind of make a mistake for something new that it maybe hasn't seen or hasn't trained on before. That then leads to downtime. One of the issues too, you hear people talking about is like a lot of the gen AI stuff especially is probabilistic and then, you know, it's guessing what, getting better at guessing, but a lot of the IT tasks or, um, well they're, they're basically need to be done the same way every time.
And I can't have an AI agent that does it differently every time. It has to be precise. And how do I marry those two things because they're not quite perfectly aligned.
Yeah, it's, it's the, the 95% use case where AI is really gonna be useful here, right? The things that it sees all the time, it understands those environments, it understands what the impact is going to be. It's gonna be that 5% where the AI calls for help and says, alright, I've never seen this before.
My models maybe aren't better than what a human would do. Or, you know, a, a linear regression. So, you know, I can, I can defer that to a human and that's where that human's gonna be stepping back in to make those decisions on behalf of the ai.
Maybe that becomes a pattern that the AI can then learn from and then that 95 turns into 96 over the course of years. Um, but it's gonna be, you know, the idea that I'm just gonna flip a switch and I've got HAL 9,000 running my data center, my cloud environment, that's still very much the realm of sci-fi. One of the things that, you know, as you described that that comes to mind is most of these AI are trained to be overly helpful, shall we say, and they kinda want to come up with an answer no matter what.
So are you sure they're gonna kinda raise their hand and say, I don't know. Uh, so it's interesting that you say that because in our own, um, AI capabilities within CRIS products, we, you have to instruct the model like do not hallucinate, right? If you don't know, say something, right?
Or, or don't, don't try and guess don't make something up. Uh, and we've seen, you know, even recently, right? Instances where, you know, citations don't exist in reports that were released.
Um, so I think you can, if, if you take that into account, right? Like, I don't want you just making things up on the fly, that will be, you know, something that these systems have to have as part of, you know, their overall operation as well as, you know, very good governance over who can instruct the AI as well as, you know, what else is it allowed to do versus not. So you definitely need boundaries here, right?
And I think that that's, that's an area that has not been discussed enough, right? People are looking at this as a magical future, but you still need guardrails and those guardrails, right? Why do you have brakes on a car so you can go faster, right?
You have to have governance in the AI domain so that you can really take advantage of these things. You have to know where the guardrails are. Well, isn't this ultimately becoming a massive data management challenge, especially with telemetry data that is, you know, unique in its own attributes?
And do we really have the mechanisms in place to kind of manage that? So You're hitting on my sweet spot here, right? Mm-hmm.
I'm a long time data management person, right? I, I view every problem as a data integration issue. Um, and so for our customers managing telemetry data, they don't think of themselves yet as data managers, but they are.
And so, you know, they are, they don't, maybe don't use the right terms, but they're thinking about metadata, they're thinking about data structure, they're thinking about data integration. Um, and yes, these AI systems are going to need huge amounts of data that is clean, that is contextualized, that is enriched. And so, yeah, this fundamentally comes down to a data management issue.
The algorithms, the models are not that interesting, right? What's interesting about them is the data that they train on and then operate on. So how do you manage petabytes of telemetry data a day?
Uh, you need very dedicated solutions for that, whether that's simply getting data in from a variety of sources and normalizing it, uh, or storing that for long-term retention or using that in a variety of operational scenarios. So yeah, I, I agree. This, this does fundamentally come down to a telemetry of data management problem, and there's a lot of elements to that.
Yeah, it's funny, I think if I went back to AIOps and originally there was a lot of skepticism and a lot of people didn't believe it. Now we seem to have pivoted over the other side of the world where there's a lot of hype around all things ai, but are we getting to the point now where maybe I just don't want to do this IT job without some help from AI because, well, it's getting too damn hard. Why not?
I mean, you know, it budgets are not increasing. So I'm not getting headcount in many cases. Um, maybe I'm being advised to consolidate tools or reduce the number of vendors I'm working with, right?
Doing all of that work, managing all this data. If I'm not getting people I need something. And so it's, it's easy to see why there's a lot of interest in ai, you know, both for SREs, DevOps, IT, operations, cybersecurity.
There's just too much to do. There's too many alerts, there's too many things that, you know, require human intervention. So yeah, AI seems like a, at least the present AI seems like a really interesting way to bridge some of these, uh, staffing gaps and, and capability gaps.
So yeah, whether we'll get there or not, um, remains to be seen, but certainly there's a lot of VC money, um, interested in finding out. So among your customers, what do you see them doing to get ready for this kind of transition into the AI era that you kinda wish everybody else would pay more attention to? Yeah, so what we're seeing, it's, it's really interesting 'cause you know, we see, um, you know, companies using that, companies that are, that have embraced AIOps, you know, before it became event intelligence solutions.
Uh, and even today they're using our solutions to normalize data, to add additional context to it so they have a better idea of like, where this data came from, where is it going, uh, what other additional payload elements can I interrogate before I land this data? Either in, um, some kind of event processing solution or long-term storage. So on, on the cripple stream side, we're seeing kind of the data preparation, uh, happening upfront.
And then on, on some of our other products, we're seeing companies create dedicated data lakes for data they're gonna use to train AI models. Uh, they're integrating lots of different data sources together from a variety of sources and linking destinations together in order to get a better idea of like, alright, here's what I can train on. Here's where I can turn the data scientists loose.
Here's where I can, uh, let the models go crazy, uh, and, and train those. So what we're seeing is like people preparing, right? They're starting to think about data as a, as a, as a strategic asset, not just as something that like, eh, I gotta, I gotta store this for seven years.
But they're starting to think more broadly. They're starting to think in value, not just cost, which is really interesting to see. And we're doing a lot of advisory work with our customers to help them get there.
Hey folks, you heard it here. No matter what era it is, guess what it always was about the data and always will be. Hey Nick, thanks for being on the show.
Thanks a lot. All right, and back to you guys in the studio.