How Generative AI is Revolutionizing the Field of ITOps | DataOps Day
In this session, Blair Sibille, the field CTO at BigPanda, will discuss how generative AI is revolutionizing the field of ITOps and bridging the gap between humans and machines. Sibille will highlight some historical challenges of automation and AI in ITOps, which were often hindered by limitations in technology and financial constraints – and the ways in which the advent of large language models (LLMs) and generative AI bring newfound hope for fulfilling the promises of AI in ITOps.
Transcript
All right. Thank you for joining me today. My name is Blair Sibille.
I'm with Big Panda. I am, uh, currently the field c t o there. Uh, and today I'd like to go through a little bit of our research, really, and the results of, um, you know, what everybody's been asking lately, which is, what in the world does generative AI mean for the field of IT operations?
Right? Um, and we've been getting that question a lot at Big Panda, right? Because we're an AI ops company.
Um, I mean, it's there in the name AI and operations, right? Uh, we kind of view our, um, mission really as a company to bridge the gap between human beings and, um, artificial intelligence, um, and how they can apply it to their day-to-day, and, uh, how it can be trusted, um, how it can help, uh, facilitate, you know, whatever metric you're looking to drive inside of it. Ops, you know, M T T R, availability, sustainability.
Um, there's tons of different metrics that it is responsible for. And ultimately, at the end of the day, if you're gonna utilize AI in order to drive one of those metrics, it better be accurate, right? And it, uh, it better, um, do do its job, uh, so to speak, uh, just like you would, um, expect another engineer or another pair of eyes, right, that you would hire within your organization.
So, at Big Panda, just real quick to, to go through who we are, um, we're an AIOps, uh, company. We, uh, serve, we serve a AIOps platform referred to as Big Panda. Um, we've been at this for quite some time, about a decade now.
Uh, and the reason why I'm kind of showing you the brag slide here is because when we talk about artificial intelligence, especially generative ai, you know, if you go and, and ask a data scientist, um, you know, what's the hardest part about AI and applying it, uh, in any applicable way, right? To operations, they'll say it's, it's not the algorithms. Those things are easy, right?
In fact, look at generative ai. It's an a p I call, it's a keyboard stroke, you know, in order to get, uh, output from it. But it's the source data that is the hardest part.
It's organizing that source data. So on the right hand side there for Big Panda, you can see the source data, uh, you know, examples that I have, um, and the customers that are running us, um, across all IT ops, you know, no one uses, uh, an AI ops platform just for network stack or the database stack, or the application stack. Um, most if not all of our customers we're deployed enterprise wide.
So, you know, we're, um, inside of trading platforms, uh, clouds, uh, anything that you could think of, branch networks, uh, for environments, uh, you know, even, uh, microservice, uh, even driven architectures, uh, and really, really cutting edge, um, edge compute, uh, for a lot of organizations. So it's, it's a culmination of all of those different, uh, endpoints, right into a platform, uh, which is big panda. So what does it do?
You know, and this is ultimately how we utilize that data, that source data for generative ai and where generative AI comes to, uh, action inside of IT ops, right? Uh, our platform ingest, you know, different alerts, changes, topology, historical information, and it's modularized in a way that we apply intelligence at different, uh, stage gates, right? Intelligence on the alerts and the events and the information coming in.
And then ultimately the end goal, right? Action, which is usually an incident, or it is a page or a chat, or a runbook or a piece of automation that someone wants to run. And that's really the functional architecture of our platform.
So where does AI kind of fit into this? Um, we have been building, first and foremost for the last decade, a modern event management platform, right? Um, the ability to, if you ask me, you know, is big Pan and AIOps company, sure, we're an ops company, but the first thing that we have to do well is organize data, right?
Um, and again, with that source data problem, and if you ask, um, if you go look at reports from the analyst, uh, community and, and everywhere, the, the number one issue and the barrier to entry when it comes to leveraging generative ai or AI in general, is bad data. You know, the, the old, uh, moniker of garbage in, garbage out, and that's what we've been solving for, for quite some time, right? And you can see that, you know, up until 2021, um, you know, us building an enterprise platform that, uh, was interoperable agnostic, could ingest any source of IT information or data into it, um, so that we could layer all these great things.
And you see the amount of velocity, um, that we've had over the last couple of years, um, with things like incident similarity, automated incident intelligence, which is actually what you're about to see here. Um, and looking forward into things like conversational interfaces that will be driven by generative AI and large language models, uh, in, in different facets. But we've been dealing with AI and ML in different ways, uh, and applying that technology to, uh, it as a whole, utilizing different ways, right?
And the, the fact of the matter is, is that generative AI is a Ferrari. There's only, what, four or five companies on the planet right now that can host a large language model itself in its entirety, right? Trained on the entire internet, um, with all of that information.
So they're good for certain use cases, right? Especially inside of it. Um, and they're not so good for others, right?
There's a lot of security concerns with them as well. Um, but what is the barrier to entry for them? And it all kind of starts with that source data, right?
Having the ability to have a malleable, um, back plane or, um, data lake, whatever you want to call it, right? Your source data, the information, the IT ops information, be it incidents or events or alarms or logs or traces. Um, 'cause ultimately, at the end of the day, you want to ask questions about that data.
You want to be able to get accurate results, um, and you wanna be able to use it in operations, right? So how do we bring a large language model and all of its amazing capabilities and allow it to interface with our IT operations teams? Um, we asked that question at Big Panda probably about seven or eight months ago, and, you know, we published a lot of the findings, uh, online.
And I'm not gonna bore you here with what happened, uh, in, in every single finite detail. And, uh, I'd love for everyone if, if you want to learn more about, um, our research, what we did inside of our labs, what we did with some of our customers, our alpha, uh, testing, um, and how we pit pitted basically OpenAI, uh, Google, Bard, and a w s is, um, bedrock, which runs, uh, a 21 labs in the backend against each other and kind of let them fight. Um, because at the end of the day, you know, we here at Big Panda, we really sympathize with, uh, outage outages and, and the knock and IT operations.
Uh, I myself, uh, from a background perspective, I got my start as a third shift l uh, one network engineer. Um, so I was always to blame whenever things went down. Um, uh, so I, I'm used to those scenarios, uh, and, and, uh, the stress involved, uh, with an outage, right?
Um, so as we were going through this, you know, we, we did things like send, um, unenriched even data to these LLMs, right? Non-correlated, all the things that AIOps does, right? We cross correlate across different observability platforms.
We integrate with the myriad of tools within organizations in order to bring intelligence into incidents, right? And point to root causes of outages, be them changes, um, so on and so forth. So we're doing a lot of that data organization.
That's what we do at Vic Panda. So we wanted to see what would happen if, you know, you, for example, watching this right now, uh, took a bunch of your events, hand correlated events together, and sent them to one of the large language models and started asking questions about the data. Um, spoiler alert, uh, it wasn't accurate at all.
Um, in fact, it was scary, uh, with some of the answers that it gave. It, it became very speculative. It was just like a human being, right?
Um, when you're on an outage and you are in a hot seat, you don't want to make a call unless you have all of the information in front of you, right? So, as you can surmise, we tested it with our data, with sending it, you know, observability, incident alerts, run books, and knowledge articles, change data was a change related to this particular incident that was created. C M D B information, uh, topology map information, as well as service maps, trace topology.
Um, and that's really what we do, is we combine all of those things today into, into an intelligent incident. But I was sending now this intelligent incident payload to the different generative AI models and asking it three questions. Um, and this was the original use case.
It was bearing in mind that every single time an outage occurs, or any time that an incident, not just an outage, but an incident occurs, um, it operations, the, the, the, the, um, the engineer that's picking up that incident, um, he does three main things. Uh, he tries to understand the impact. You know, he reads through everything.
He tries to speculate, uh, as to what is impacted and what the summary would be. Change the ticket subject to something that's, uh, humanly readable, because nine times out of 10, they're cryptic because they're being sourced from an event management platform. Um, and then speculate on root cause.
What are the next steps? What are you going to do to fix this? Right?
And we wanted to augment that approach, um, not replace it, but augment it. Um, because trust is paramount when it comes to outages or incidents or IT operations, you know, the age old moniker of what is a level three or a higher level engineer inside of an organization do when he gets a ticket or an incident escalated to him from an L one or an L two, he does all the work that the L one and L two did again. Now, is that a trust issue?
Kind of. It's more about ownership. Because if I own a particular piece of my company's architecture or infrastructure or application ecosystem or business services, and there's an outage occurring every single minute that goes by for some of the organizations that you saw on the, uh, the first page where I was talking about, uh, some of our source data, our customers, uh, within minutes, that's my salary in a year, maybe even multiple years.
So it's a sense of ownership, right? It's your environment. Um, and that's really universal across operations.
You know, we care, um, and we care deeply about, you know, the uptime and availability of what we're ultimately responsible for. And we won't make a call unless we have all of the data present. Um, and if I am going to start utilizing generative AI to start suggesting what the root cause or change the subject line, or give a short, short summary of impact, being accurate is paramount to gaining the trust of operations.
And again, that's what big Panda needs to do, right? That's our, that's our driving force, is to unify and train, um, different types of machine learning and artificial intelligence models, um, on you, and make it tailor made and purpose built for you so that you can trust it and work a different way, work with another pair of eyes, um, like automated incident analysis. And that's really what it is.
And this is what we found, um, that 95% of the time in our alpha testing, we tested with over six customers over the period of, around four months, and found that 95% of the time when it's speculated on root cause, it was accurate, it got it right the first time immediately. Um, which is kind of unheard of. Find me a human being that speaks database application, architecture, software architecture, network, uh, storage, and can glance at an incident and within seconds come up with a root cause, um, and be accurate 95% of the time.
And that's what this is. You know, the, the generative AI model that we're leveraging, uh, that, that ultimately won was OpenAI, just f y i, spoiler alert, if you're reading the blog. And we quickly realized that when we start sending these intelligent incidents that we're building inside of Big Panda for our customers to this generative AI model, we are leveraging, you know, ultimately when you think about it, OpenAI's original purpose, when they built DaVinci, which is the backend of chat, GPT, the large language model itself, it was built first in order to assist software developers with coding.
Uh, it's a predictive next word model. It's a generative AI model, right? So it predicts what the next line of code should be, or the next word it should say, next sentence based upon the context that you give it, right?
The question that you ask. So if you give it a rich incident with a lot of context, all that data that I showed you before, and it understands the topology in question, it understands the customer when it, when it knows an application is down, of course, it knows software architecture, infrastructure architecture. It's been trained on the entirety of the internet.
It is a full stack DevOps engineer. It is a SS r e. Um, it knows.
But the amazing thing that it knows as well is that applications are important for organizations. Applications have a purpose. If an application is down for a customer, people are angry.
Uh, it is the bottom dollar, half the time for a lot of organizations and their digital experiences that they provide. It knows these things as well, and it communicates in that way, communicates in the same way that someone that cares and someone that understands that it's not just a technical problem. There's also an impact here.
There's also an urgency involved with the fix. So we asked it three questions, and you can see the ultimate output here. Uh, this example right here is a real world example from one of our customers where they had a bunch of ATMs, um, out at a location that went down, um, logically they couldn't see them anymore through the network.
All the V P N tunnels to these ATMs went down, um, and they could no longer process transactions over them. And the root cause of that was a needle in a haystack. There is, there was hundreds of events associated with it.
And just one little alarm that said that an SS SS L cert had expired. That's it. Just that one little alarm in that, that rat's nest, right?
Of all these events, the sea of red, um, to harken back to my NOC days. And what's so surprising about this generative AI model, you know, if AIOps is a polyglot, right? Someone that speaks multiple languages and is able to, to intelligently group those things together and convey that in a ticket, that's really what AIOps does.
It can talk all these different types of, in this case, it's network alerts, application alarms, synthetic transaction alarms, um, as well as infrastructure, um, you know, OSS level alarms from SolarWinds logic monitor net, IM and SEC 24, you know, correlating across all of these different telemetry products. That's what AIOps is doing. But what the generative AI is doing is synthesizing all of that information into natural language, into a voice, into purpose built, um, insight, right?
Right away. You can see right there the summary, um, and the title that it's giving it, instead of it being cryptic that 19 hosts are impacted. It's a storage failure database latency and web timeouts.
The root cause analysis picked the needle in the haystack out saying that the root cause appears to be the invalid S SS L certificate on that host, which ultimately led to all these tunnels going down, right? And the reasoning as well, not just, here's the root cause, but why A S S L certificate alert associated with everything would plausibly cause this particular outage. And in fact, one of the great things it does as well is there's never one right answer, uh, when you're on an outage, right?
It might be multiple different root causes. That's why we call this speculative root cause there might be more than one. Um, and it'll list them out in order of, um, uh, of its, uh, of how well it thinks, uh, without a lack of a better term, it thinks in order of plausibility and probability, which one is the most likely one.
Um, but it's, it needs to be speculative, right? But again, the more data it has, the more source information, the better organized that source information is, the more accurate it is. Um, and the better it is at communicating what that, uh, root cause is, what that subject line should be, what is the impact of this being down?
And that's, um, that's basically what we did inside of our labs, really. That was the first thing that we could come up with, was how do we bridge that gap and bring generative AI in a way that makes sense for operations on every incident, right? Um, as opposed to from an admin or an analytics perspective, there's great applications in that way as well, and we're thinking about them too at Big Panda.
But we want to end outages that are stressful, right? Outages are going to happen. But like I said, you know, being in that hot seat and doing triage every single time, um, it can be stressful.
I mean, mean, think about when there's a major outage and you don't speak Oracle databases and you pinging your Oracle database team and they leave you on red, right? Um, and it's just not important to them. You know, this, that coupled with, you know, people with C'S and D's in their title coming and wrapping you on the shoulders and asking, what's the impact?
What are we doing to fix this? You know, um, it's, I I, I, I've found my, myself and operations teams always saying those phrases, we need another pair of eyes. Give us a minute.
Um, we're looking into it, right? Um, so on and so forth. And that's what we wanna do.
We wanna augment, um, operations with immediate access to a DevOps engineer, full stack engineer, a software developer, um, application developer, um, that can assist them with this typical triage, right? Because on the left hand side, you can see what AIOps does, right? Taking in those millions of events and alerts and siphoning them and bringing them and compressing and correlating them down into only 27,000 incidents, intelligent incidents, those are the incidents that make their way to generative ai.
You tried to send those millions in good luck with the type of information you would get out of them, but those are still going to human beings, right? Of course, AI ops is doing everything it can to build an intelligent incident, but regardless, when a human being picks up an incident, this is what they do. They read the entire incident, they update the title and the short description of what's going on.
They search through change records because the first or second thing that everyone says on an outage or when they look at an incident is, who's the cowboy? What changed? You know, 80%, it's, it's, it's no why, and this is true across our customers too.
80% of all incidents are because of some type of change within an a, um, an organization. And then they have to update that information, everything that they've gleaned with the summary, the impact, and then ultimately speculate on what the root cause is. What are the next steps?
Who do I need to pin? Who do I need to involve in this? But when we introduce generative ai, this happens in seconds, right?
Um, the human is still in the loop. He's still there. Um, but now he has another pair of bot.
He has someone that is already surfacing up that information, and all he needs to do is verify, you know, is, are all the signs that this is pointing to? True? Great.
I have a plan. I have a mo way forward. And truth be told, you know, when I, when I think back to, you know, 10, 15 years ago, um, from a background perspective myself, I, I used to be, uh, the director of enterprise architecture for about a decade at an automation company.
Um, and one of the things that I did as, as their, uh, enterprise architect and, um, basically their, their seal, their C T o was talk to customers and figure out what they thought automation was going to do for them. And the number one and number two, use case for automation. Now, this is IT based automation.
I think we know today what I'm about to say, automation is near impossible to get to do this for you. But they all wanted automation to communicate the business impact immediately whenever something went wrong and not fix things immediately, because they didn't quite trust it yet. But just tell me what you want to do.
What are the next steps? What is the root cause and what automations can we kick off? Or what should we look at in order to fix it?
Automation has been mis-sold for years, and that promise was just never met, never met, but we're doing it now. And that's really what generative AI is capable of doing, is delivering on those promises that automation platforms. And what we thought at that time, and were promised that automation would do for us, um, that we all now know, you know, uh, finite, state-based, uh, infrastructure automation or event-driven automation, it's near impossible to get it that customized, right?
Uh, given the current, uh, uh, state of things and things are getting even more complicated with event-driven architectures and self-healing environments, gray failures, just because something is green doesn't mean that it's not read right? Um, which just overcomplicates things. That's really what generative AI can do for us.
Um, but first we need to organize our data. We need to have our data, um, in a way that can be consumed and speculated upon and actioned upon via these large language models and generative ai, um, and build trust with them, just like we would onboarding a new person inside of our organizations, funny enough. But I'm happy to report that.
Um, if what I was just talking about and, you know, where generative AI is going, interests you, um, you know, please check us out, uh, at Big Panda. Um, we released what you just saw, uh, to general availability on, on seven 11. So we were the first IT operations, um, platform, uh, software vendor, quite frankly, to release a generative AI capability to the market.
Um, a lot of other capabilities are coming out now. You're seeing them, um, they're launching betas and wait lists and so on and so forth. Um, as they try and figure out, all right, how does this work, um, the approach that we are taking is we're agnostic.
We, we don't care, right? We wanna bring in all of your data from all of those different vendors and, and things that you're using to deliver, um, it as a whole and introduce generative AI to it, um, as opposed to generative AI for your observability tool, generative AI for your ticketing platform. This is for all sources of data.
So you can see some nice things here that people have set about us so far. Um, but it's doing exactly what we said, right? An automated summary, automated title, uh, suggesting the root cause.
Uh, and this is just the start. Um, from here we're going and looking into things like incident similarity, um, vectorized databases, so on and so forth. Uh, contextual based, uh, things, and of course conversational.
You know, I'd love to be able to have a conversation with my incident and ask how and what are the steps and point me to knowledge articles to fix things. So, um, thank you so much everyone for hearing me out today. I wanted to, you know, instead of this being speculative in nature and what could generative AI do, where could we go?
I, I, I just wanted to show you something that's in place today that's practical today. A way for you to answer that question that probably, um, you are asking and some of your organiza, your organization's definitely asking is, alright, this generative AI thing, uh, do we get on the train or, or what? Like, where are we using it?
Where can we use it within our organization? Um, you know, I invite everybody to, um, our website. Uh, we have a demo, the demo that I just showed you to screenshot.
Um, you can actually go play with this, um, and, and go check it out. Uh, have a self-guided tour. Um, and yeah, feel free to reach out to myself or Big Panda.
Um, if there are any other questions. And I appreciate you hearing me out today. Um, thank you so much.





