Artificial Intelligence and IT Operations – Phil Tee, Moogsoft
Moogsoft CEO Phil Tee sets some realistic expectations for the impact artificial intelligence (AI) will have on IT operations teams as advances continue to be rapidly made.
Transcript
This is Techstrong tv. Hey guys, thanks for the throw. We're here with Phil T, who's the c e O from Moogsoft, and we're talking about AI and IT operations, otherwise known as AI ops, and just how far down the path are we these days?
Hey, Phil, welcome to the show. Yeah, thanks. Thank you for having me.
We've been talking about AI and IT operations for some time now, but I'm not quite a hundred percent clear that everybody's on board yet. There's still some skepticism, and of course there's a lot of new technologies coming down the pike. But from your perspective, where are we on this journey?
You know, uh, I think it was Bill Gates who once said that, uh, you know, technology's always overestimated. Its first two years and underestimated its impact over 10. And we're probably, um, somewhere in that sort of, you know, five or six years on from people recognizing AIOps as a market.
So it's kind of playing out. Uh, you know, we we're in that sort of mid game, I suppose, you know, moof, we've been doing this, uh, for 11 years. We were probably the first in market with a genuinely AI and ML based approach to IT operations.
Um, but, uh, there's still a lot to be done. Uh, I will say, uh, you know, one of the things that is of course helping, uh, adoption is just general attitudes to AI have changed quite dramatically over the course of the last decade. You know, when we started out, people thought it was very mysterious.
Now of course, everybody has it on their, you know, uh, smart devices. Their Apple watches, their televisions, uh, it's all around us. To your point, there was a lot of initial skepticism.
Are we rapidly approaching some point where maybe IT folks will say, you know what? I don't wanna work for an organization that doesn't have some AI capability because it'll just be too hard. I think we are there.
Uh, if, if I'm honest, uh, you know, and, and, and of course, you know, chat gpt and all the hypes surrounding that has, has, has really accelerated people's desire for looking to look for AI as an assistance in their role. And just to kind of pair it back, I mean, if you think about the, okay, so what are the alternatives? What does life look like if you don't have AI in IT ops?
Well, you know, you're back to the rules-based systems and the kind of the manual slot, the toil, uh, in a more sort of modern terminology around s r e of having to, you know, chase down every alert that you get, um, you know, rely upon tribal knowledge, work it out for yourself from the beginnings, you know, it, it's very resource consumptive, and a lot of that resource that it consumes is much more valuable doing high level tasks, being dragged into the kind of the menial, uh, sort of job of sort of, you know, literally sometimes referred to as clicking the knock, right? I mean, going through every alert and manually assessing whether or not you need to worry To that point, what exactly is AI these days? Because there are a lot of folks who have rules-based systems, and it's not quite clear to me that they're learning anything.
And other systems may learn some things, but not a lot of things. So just what is the definition of AIOps these days? So, uh, it's, it's almost the inverse.
So rules-based system, um, is not learning anything you, you define beforehand the behavior of the system, and you have to encode into it Every response to every possible situation that it can find. A AI doesn't require you to do the upfront bit. So, so we, um, tune our algorithms, uh, to give it a hint as to when to assume two bits of data that it may never have seen before are correlated in an incident that needs working.
So the, if you like, the working set of definitions or configuration that a true AI system works from is tiny by comparison to rules-based system. You may be going from tens of thousands of rules and more to a handful of, um, of definitions for your correlation engines, which is kind of how that looks with, with Moogsoft. Of course.
You mentioned the, uh, a hundred thousand dollars question phrase these days, generative ai. Yeah. Where does that fit alongside machine learning and what should people be expecting from here?
Well, I, I'll be honest, we're still digesting, uh, you know, the impact of, of, of generative ai. Um, you know, obviously chatt and OpenAI being the, the one that's on everybody's mind at the moment. Um, but I think it will have some quite important impacts, uh, on the operational use case going forward.
I in particular, I, if you think about what it does, it's essentially a completion engine, how it works when you go to chat G p T and you give it a prompt, you give it a question. You, you know, before we got on, you know, I asked about you. I said, what do you know about Mike Vizard?
And, uh, you know, what it wants to do is it wants to complete the task. It is looking through all of its trained data to work out what the best next thing to say or the follow on, uh, from the prompt that has been given. And if you think about that in terms of, um, our industry, I see two really obvious places you could go with this.
Um, one is a correlation use case and anticipating the what is next, um, using whether it's a, you know, uh, chat gp, PT itself, probably not chat, G p t, probably a standalone transformer trained upon a large corpus of operational data to anticipate the likely next set of alerts, likely next set of things to occur, um, from the given state in your infrastructure. That I think is, um, you know, super interesting and we're definitely experimenting with that at moof. And the second piece is more around the explainability of, um, of, of your system.
So when you have an incident and you know, the data that you have about that incident is really pretty sparse. It might be some alerts, it may be some, uh, contextual data from A C M D B, it may be some user, uh, commentary that you have from that. Explain to me what's going on, give me context, give me relevant assistance, uh, in the problem resolution phase.
Uh, and it's ironic cause I think Stack Overflow is kind of banned the user chat G P T to provide answers back to it, but in a sense, chat, BT looks like a very clever version of Stack Overflow, or very easy to interrogate version of Stack Overflow in that context. So do you think the meantime to remediation will be better as we go forward with generative ai? We may use machine learning algorithms to discover the issue and then figure out how to fix it using generative ai, or how smart are the machine learning algorithms?
I guess I'm asking who does what these days among the algorithm? Yeah, so I mean, it's a, it's a bit of a, um, hat and mouse game in a sense that, uh, you know, if, if the infrastructure stood still, I think the categorical answer is yes. Uh, use of generative AI will shrink, uh, M T T R, uh, down the line.
Of course. Um, systems keep getting more complicated. Uh, the, the world keeps getting more complicated.
Everything keeps getting more complicated. So, you know, the demands upon the intelligence of the, um, AI that you using will become bigger as you go forward. My guess is, um, is it will be a bit better than the old promise of the pc, or if you remember this, Mike, it's like, you know, the next generation of PC comes out, you know, windows is finally gonna run fast, and then you upgrade windows and it runs slow than it was before.
So I, I think we'll do better than that. Um, but I don't think it's gonna be a, um, a silver bullet that's gonna solve everything overnight. There's, there's still a lot to be done.
There's still a lot of, um, like I say, we're experimenting. We're trying to understand the best use of some of this technology, and a lot of it is fitting the capability to the use case, um, and the domain knowledge, um, that's required to do that in a way that maximizes, um, upside. I think part of that subtlety is that the large language model may not be trained or updated in real time.
The machine learning algorithms are learning in real time, so they may have an advantage as the environment changes. And yeah, the longer you go between updates of LLMs, the more likely it is that the recommendation may not be optimal. Yeah, I, I completely agree.
And of course, um, the, the chat G P T that we all know in a sense is a generically trained model. I mean, it's built on a, um, a soup of, you know, uh, the, the general web crawl Wikipedia books, um, I mean, it's all out there in the public domain about how they do that. So in a sense, it's a very clever generalist.
And what we may need, uh, in the IT ops focus use case is a very clever specialist, which sort of says to you, you want to, you want to take the technology and change the corpus that it's trained on or the corporate that it's trained on, um, to produce a more targeted response. And of course, if you're shrinking that corpus down, you know, yes, that will, um, negatively impact the, uh, the, the accuracy of the model, but it will also speed training times. So there's probably a little bit of a trade off, um, in terms of thinking about how you sort of bend the model to be more specific to the IT use case.
As we think about all this going forward, I mean, the DevOps community has been committed to quote unquote ruthless automation since time began. So are they be at the forefront of this AI ops revolution, or who's got the most enthusiasm in the IT side of the equation? Yeah, it's an interesting one that, um, okay, so here's another promise of, uh, of generative ai.
Um, the, you know, coding becomes a menial task. I mean, the, you know, one, one of the first things one of our, uh, team did was, you know, go, go ask, uh, chat G p T to, you know, produce a hound chart and you know, which is your configuration for Kubernetes. And, and you know, it's bits one out that's not bad.
You can get it to, um, you know, to produce code snippets that are, you know, reasonably useful and maybe just maybe, um, that, you know, supercharges the, um, the, the productivity of, uh, people at the coalface, um, in, in that regard. What it's not gonna do, and, and this again, is super emergent at the moment, is it's not gonna replace the human, uh, involved in this because it just shifts the job from, I'm gonna sit down and write some python, or I'm gonna sit down and write some, uh, you know, shas scripts or whatever it is you're, you're doing to, I'm gonna sit down and design some prompts. You know, they're very famous prompts engineering, uh, that one does with these types of, um, uh, tools.
But you've still got to the, the, the novelty that is coming from the human being is couching the problem that you're solving and deciding to solve that problem. Um, and yes, there was a lot of hype and auto G P T where you sort of couple the output of one, um, G P T to the input of another and get it to produce lists of tasks. But I think that's a way away from, um, being real world.
How much data will we need to create some of these LLMs if they're domain specific, it seems like maybe know we're not gonna be waiting to boil the ocean. We just need to kind of get the right data into the right L l m and with the right set of recommendations and not overthink it. So, um, the interesting thing there is, uh, o of course, these, these models, they get better with more data.
Uh, that that's absolutely the case. But even in the context of how chat G B T works, you know, there is a, you know, a a core training corporate, um, which is used, and then there's, you know, various levels of, of reinforcement learning and supervised, fine tuning that is done, uh, you know, to improve the outcomes. I could see a shared, um, it specific, um, partially trained model being out there at some point, um, which, uh, you know, independent software vendors, end users, uh, may provide that last part of the pipeline to tailor it to their particular requirements.
Um, but you're still gonna need, uh, a large amount of data. I mean, there are, you know, the, the, the, the, the publicly available chat, G P T is trained on a vast, vast corpus of data. I mean, it is the entire internet and then some, uh, you know, so, you know, that's pretty big.
Um, so my guess is, is that it's probably beyond the scope of one customer to generate their own corpus to start off with. It's gonna need some collaboration, uh, you know, to get it up to a, a point where it may be useful. All right.
A lot of wags are talking about this notion of AI winner. What's your take? Are we gonna see some, uh, perhaps, uh, tro of disillusionment or where are we going next?
Oh, for sure. I mean, I think that, um, you know, between here and there, uh, there will be, you know, the gradual cool down of the hype. We're still in the hype regenerative ai, um, people are still running around, um, uh, you know, pulling their hair out.
I kind of, I kind of watch with a rise smile about the kind of the fear pr that's going on. You know, the open letter saying, we need regulation. The robots are gonna kill us.
It's the end of the human race, which is absolute and total nonsense. Um, but it's done for a very specific reason because if you are the owner of this kind of technology, you want regulation cuz it stops startups challenging you. Um, so, um, I I would say that probably there's gonna be, um, a year or two of people sat around going, well, where are the applications of this?
Um, and it'll be maybe, maybe five or six years down the line that you start to really see generative AI playing to the operational space, just, which is kind of the cycle time of a couple of generations of startups, um, to, to get you there, which is where this is gonna come from. So I think, um, I would look to back off of this decade where this starts to have a significant impact, but the impact will be profound. Um, there'll be a lot of improvement to mtt r We may have a bead on self heating networks for the first time, um, regenerative ai.
And certainly I think that we are gonna automate away a ton of what are considered to be frankly, quite high value, um, but low, um, you know, low complexity tasks around automation. All right, folks, you heard it here. There's still a long way from here to there, but as one wag once said, it's one thing to be wrong, it's an another thing to be wrong at scale.
So be careful. Phil, thanks for being on the show. Absolutely.
Back to you guys in the.