OpsRamp’s Varma Kunaparaju on How AIOps Is Transforming Observability in Cloud-Native Environments
Varma Kunaparaju, Senior Vice President/GM, Cloud Platform & OpsRamp at HPE, explains how artificial intelligence for IT operations (AIOps) is transforming how observability is applied across complex cloud-native computing environments.
Transcript
Hey guys. Thanks for the throw. We're here with Barma Kuna.
Paju, who's the CEO of ops ramp, an arm of Hewlett Packard Enterprise. And we're having a little chat about AIOps and observability and cloud native application environments because, well, it's complicated. Pharma.
Welcome Michelle. Thank you. Thank you for having me.
Observability has been tough, and I think in a lot of cases, all we ever really managed to achieve was some basic monitoring in the first place, but in cloud native environments since even harder. 'cause there's always microservices. They get spun up, they get spun down.
Nobody seems to know which way is going. And, um, by the time you take a look at it and start to sort it out, it makes your head hurt. So as we kind of move into the age of AI ops, is this gonna get better?
And how does AI ops kind of change the way we need to think about observing these environments? Well, Greg, uh, great point. You know, first of all, the, the complexity of cloud native and AI native applications increasing, right?
You know, more and more distributed microservices are playing a bigger role in making it to deliver business services, you know, to the lines of business. So if you look at that and look at what the AIOps era, if you looked at original AIOps era that started pre transformer days where, you know, the LLMs and, and the transformers were not there, the original AIOps journey was all about machine learning algorithms, trying to kind of find, you know, those patterns and analyze those patterns using historical data to figure out what is a deviation from a normal behavior. And the second application of that, uh, AA ops in the pre transformer days are all about, you know, potentially looking for, you know, a, a large volume of alerts and, you know, ingest all the, all those alerts as signal and, and figure out which is signal and which is noise.
And, and really try trying to kind of troubleshoot in terms of figuring out, extracting a signal from a bunch of, uh, noise, right? So those are the pre early days of pre transformer ai, um, techniques that are user for AI ops. But with LLMs and with the transformers, the AI applications and ai, the way that the AI is used in the observability is massively shifted, you know, not just at the large language models, but foundational models that are domain specific that can be fully fine tuned for being able to kind of do that probable root cause.
So where I see the AI and agent AIOps in the, in the context of post transformer and post LLM and you know, the current generation is all about how do you kind of really get the probable root cause or a, a real close to the root cause by using the, the foundational models and by using more natural language queries to kind of make sure the human operator is getting the assistance from context since two and domain specific intelligence around the observability data. Mm-hmm. Am I still having to instrument those microservices to collect that data or am I just pulling the telemetry data in its kinda raw format and dumping it into something that the algorithms then make sense of?
Yeah, good question. You know, you know, in the, with open telemetry, you know, the collection of the data is really making it, you know, much more seamless, you know, as opposed to in the old days where, you know, you have application specific instrumentation with agents deployed to, to really get that observability data to now, um, with Open Telemetry and EBPF, the advances in EBPF really makes, you know, in certain cases auto instrumentation to get a, a really good set of telemetry data that you could ingest for your models and your, uh, foundational models to absorb if you wanted to apply some customized foundational models for observability or even otherwise, you know, regular observability data, the instrumentation, um, both with open telemetry and the advancements in EBP of, um, you know, the instrumentation is becoming much, much more uniform and open. Yeah.
Early on, AI ops makes use of a lot of machine learning algorithms, and it took a while for those machine learning algorithms to learn the environment. And in the age of cloud native, the environment's more dynamic than ever. So how quickly do, do those ML algorithms first learn the environment and then how do they keep track of all the changes?
Yeah, so the original machine learning algorithms are all based on pattern recognition, right? So you have to kind of really do instrumentation, instrumentation absorbed data to feed along with all the historical information to really recognize the patterns. A lot of those machine learning algorithms in the first generation are all based on, you know, learning from the past data, very little chance to do zero shot, uh, being able to kind of, uh, for the first time encountering some things.
But with the new, um, transformer based models, you have a very good shot at being able to kind of go after Jira shot, um, predictions and, and predictions that you have not determined that you have observed before because the models are trained and the models can really come back and, and, and, and, uh, significantly, um, pinpoint saying that this is potentially the, the, the reason why this is happening, even though that pattern did not account encounter before. So that's the big difference that we see with, uh, the last two years of advances in ai. Mm-hmm.
As I understand it with AI too, there's generative and there's predictive and causal algorithms, and am I gonna be using a mix of these things and an AIOps platform and which type do I use for why when? Yeah, no. Um, if you look at, um, the, the entire instrumentation and observability data, you know, I think there is room for playing a number of these, um, AI techniques, right?
You know, you are not replacing the machine learning completely with LLMs and you know, you're not replacing full-blown transformers. So if you look at the, the, uh, advances in the last two years, you made the human interaction with the observability data a lot more natural language specific. You can do prompts and prompt based techniques to kind of interact where a virtual operator from, you know, um, agent or ai, um, you know, LLM driven AI models to GU that in human interface is more prompt driven.
That's number one. But it doesn't mean that, you know, you are completely replaced with the old EML technique. So it is complimentary to what took place in the past.
And, you know, leveraging those, uh, algorithms and moving the data to, in some cases frequency domain and analyzing that data in the frequency domain to kind of freely understand those, those ultimately is the needle in the haystack to determine the po potential root cause still exists. So, to Sean answer is, it's a combination. You know, Will we still need traditional monitoring tools?
We've had 'em for decades. They kind of track a bunch of predefined metrics, but observability in my mind was always about, you can dive in and analyze stuff and tease out root cause issues. Uh, are these things converging or are they always gonna be somewhat, uh, orthogonal to each other?
How do you see this evolving? Yeah, I think, you know, most often the industry says, you know, monitoring is all about, you know, uh, finding a problem, right? By, by putting some thresholds and determining, and that's how the monitoring definition came in the early days, more and more the observability is all about how do I reason, how do I find won't cause that actual problem?
That because you're now providing more context, Jan, than saying that something failed because of the, the monitoring alert that showed. So to answer your question, in my mind, observability is a more superset, you know, and, you know, and observability came in the, in the cloud, cloud native stacks with logs, metrics and traces all coming together to give, but it also, you know, doesn't preclude for making the same observability extend to your network, extend that to the edge where the actual consumer of those applications are really sitting and being able to kind of really find why an application performance or an outcome that the business user is expecting is not getting delivered. So, in my view, this is a, a super encompassing thing.
When you kind of really take cloud cloud native or AI native stacks with Edge and the edge consumer, um, of those applications and, and the data that is needed for finding that root cause is complemented, then that's when you get the full, full view of the entire observability. So in some ways, monitoring is a subset of the overall observability that is happening today. Early on, the folks who built these cloud native applications on things like Kubernetes or small groups tucked in the corner somewhere full of specialists, but it seems like in the last year or two or so, these apps are now going mainstream, and they are now part of the general IT landscape.
Are we gonna see some unification of the way we manage legacy monolithic apps in these new cloud native apps in some way? Because otherwise we still got two separate teams just managing more stuff than ever. Yeah.
I think these silos and and existence of these silos happen over a period of time because more and more cloud and AI native applications are becoming relevant for new stacks, but traditional enterprise applications still sits in the enterprise. Now, how do you bring these two silos together with, uh, call it digital operations command center, you know, for lack of any other words, where you bring the traditional applications, the edge infrastructures, and the observability around those consumer, uh, of those applications and those edge infrastructures and the observability data along with modern cloud, cloud native applications, uh, instrumentation, bringing that together is where, in my view, the actual it, and it is proactiveness to make sure that the business users and the outcomes that they're expecting are really, you know, is getting delivered, is going to play. And how those two things come together in my mind is a, a, a, um, a concept of digital operations.
Our digital operations command center, where you're bringing all the entire, entire state of the IT ex across bottom of the infrastructure, all the way to the applications, and taking traditional and modern applications together into one single place. You know, AI ops in the original days was all about alert correlation to create that digital command center. You know, the so-called AIOps tools were originally designed around, you know, Hey, throw me all the alerts from your traditional applications and modern applications, and we will process and give you a, a qualified incident that you can pass it to your IT teams.
But the concept of DevOps, tech ops, IT, ops, SRE, ops, all of them coming together in the modern digital operations command center calls for a new way of looking at that command center. And that's where observability is leading to now. Mm-hmm.
Do we need to rethink therefore the way the IT organization is structured? Because historically we had all these silos, we had the networking people over here, the storage people, the servers, bunch of folks managing applications over there, and they kind of created their own little fiefdoms. Um, do we need to kind of just make a concerted effort to start breaking those walls down?
Yeah, no, I think, uh, um, you, you, you brought a very important point in the, in the context of pandemic and how cloud and cloud adoption accelerated those silos needs to be completely broken. Because at the end of the day, the business outcome and the business service impact is what it is expected to deliver to, you know, lines of business, almost like it acting like a service provider, right? So if you wanted to make it act like a service provider, you know, throwing 20 people on a bridge to determine what happened is not going to kind of help.
So one needs to kind of provide a, a qualified observability data, not just in the full stack observability at the application level, but all the way to the edge where the actual business user is consuming, you know, whether it is rum, whether it is application performance data, whether it is full stack observability data along with network that is going to play a major role in terms of consuming this application. All of that needs to come together to really find, you know, a, a unified IT that acts like a single service provider to, to give the business outcome, right? That's where I think the, the advances in AI advances in observability, the advances in machine learning, all ultimately resulting in bringing that probable root cause and making the outcome, uh, to the business application user is what needs to be broken.
And that breaking down is happening as we speak. Right. So ultimately, what's your best advice to organizations that are rapidly embracing cloud native computing, but they definitely have legacy applications.
Is there some smart way of getting started? I think a lot of them have thrown tech at the wall for many years now, but as we look at AI ops, sometimes it's intimidating and other people are fully embraced, but how should they think about getting started? Yeah, I think the best way to start, in my view is, uh, an approach of integrate to consolidate as opposed to rip and replace, right?
So you have this modern, new, more cloud, cloud native applications that you, IT organizations did full stack observability, let's say, and then they have the old, you know, the traditional applications which are happening in, in, in, in either their data centers are in their edge locations. More and more edge locations are becoming like a mini data center. If you go to a store, you know, behind a large, a grocery store or a, you know, or any other store, there is a, a small mini data center that is running for what needs to be run at the edge.
So those applications and those, uh, edge edge applications along with their cloud, cloud native applications, they need to be brought together with an integrate to consolidate approach where you bring this New year's, um, agent ai, um, slash operations management command center kind of a solutions and integrate some of these traditional old instrumentation that one would've done along with the new observability stacks, and then over a period of time modernize that entire stack. So that way you have one single, um, operations console for most IT organizations to be able to deliver business outcomes. All right, well, folks, you heard in here, cloud native is everywhere.
It's not just in the cloud, it's all the way at the edge. And while that can seem a little scary, it also creates an opportunity to rethink how we manage it altogether. Hey, thanks for being on the show.
Thank you. Thank you for having me. All right.
And back to you guys in the studio.