AIOps Evolution with ScienceLogic’s Dave Link
ScienceLogic CEO Dave Link dives into the next phase of artificial intelligence (AI) for IT operations following the release of his book, “Innovation: Journey and Outcomes for the AIOps Revolution.”
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Dave Link, who's CEO for Science Logic, and we're talking about a new book that they have out, it's called, uh, innovation Journey and Outcomes for AI Ops Revolution.
That's a mouthful right there, and we're gonna dive into what all that means. But, um, Dave, what was the thought process behind writing a book in this day and age? And it's like, I haven't seen a good book in a long time, so, Well, that's exactly why we wrote it, Mike, because we knew that you needed dump bedside reading material to keep you inspired for the future.
You know, we've been working on this journey helping customers really revolutionize their operational processes and the way they run, operate it. And if I've told the story of starting the company, why we started it, what was the light bulb moment once I've told it a thousand times and I thought, you know what, it is time to memorialize that story because it's one of those stories that's interesting. We started the company in our garage in a basement.
We didn't have any money. We had five credit cards. I took out a mortgage on the house.
This was a company started and inspired by a belief that we knew a better way. And from having worked in operations for global service providers, we felt this was a time to say why we started and what's happened during the journey of the company, how technology has evolved, and how we have had to really reinvent the company two or three times over to stay aligned with the future. So we wanted to kind of capture that in not just a, a story about innovation, but a story about a company's journey from start to an ongoing global, a customer state that we now serve that really inspires us to do even more going forward.
Initially, I think when people started talking about AIOps, there was, you know, a massive wave of interest and then it kind of petered out a little bit and people were skeptical. They felt like the machine learning algorithms weren't learning their unique IT environments, and then we've come a long way since then. So what's your sense of the current state of things?
Where are we? There seems to be lots of different types of AI models now. I mean, are we there yet?
There is still a gap. There's a gap from the market architecture and the storylines and essentially the desire of a customer to have a system, self-learn and self-manage environments. In a perfect world, the system would understand historical pat patterns, realize the change to a historical pattern, an anomaly, and then rush ahead to figure out why and fix it, adjudicate the problem, resolve the issue before you have any service disruption.
So we've really been on that mission for quite some time in it and we've all made progress there with scripts, with event-based automations, with, um, a variety of techniques with machine reasoning over performance data and some fault data. But I don't think the industry really got there. This past year was really a year where gen generative AI took over and we're starting to find some very interesting storylines with lots of different data sources beyond telemetry coming from kind of the digital exhaust fume of devices to other data sources like power outages, fiber cuts, other things that create service outages, data sources, fusing them together from lots of different, not just the technologies themselves, but other, uh, data sources, weather data sources for some of our satellite customers, merging them together to get to a better understanding real time of what might be causing a problem so that we can get to a faster meantime to innocence or mean time to repair, which is everybody in it is goal objectives.
I kind of feel at the risk of a gross oversimplification that gen AI is becoming like the front end for predictive ai and that all I really wanna do is walk in the office and say, tell me the three things that are most likely to get me fired and in language that I can understand. Well said. I think what we're re what we're realizing is that the, uh, the gen AI technology is really good for a certain class of data sets, the dataset that are human generated data sets.
So to your point, chatbots, the, the notes on issue resolution, kind of the, the workflow of how do we resolve issues that are known issues in our help desk, the ticket information for resolve tickets, all of those human captured elements are really good for conversational ai. To your point, marrying that up with the, the raw performance fall configuration data, which we've been using machine reasoning and we've been using algorithms since the beginning of the company to really look at data in a different lens. So we're capturing trillions of pieces of data objects literally every week across our customer estate, but then we have to post-process that to really make sense of it.
We're now fusing that with other human generated sources to actually bring the best of insights to both data sets to deliver a different experience and outcome for the end user. We think we have built what we'll believe, um, the industry will see as a leapfrog interface where I think if there are any, anything that we've learned like over the last few years is we've got this immense trove of data performance, fault configuration data, but we're not using it to the best of our ability and we need to fuse it with other data sets real time, in some cases in a federated model where we see an event and we query, is there a power outage, is there a fiber cut? Do I have a knowledge base article about this that gives me clear line of sight to what recommendation I want to have for a, a root cause of this issue?
Um, those are the things that as you pull these technologies together as one overarching system, they're getting smarter to your point, to give you really the recommendation engine when you walk in the door. Here's what you need to worry about this morning, What ultimately will be the downstream impact on the IT profession as we know, because it's made up of administrators on one hand and specialists on the other, whether they're DevOps folks or networking folks, whoever they might be. Is the bar for managing all this stuff starting to get lower even though the i IT environment itself is incredibly complex, but are we getting to the point where maybe mere mortals can do this?
So Mike, the future of jobs in it, I think they will change, they will evolve. Specifically there's one segment of, of jobs that we think we can materially improve and perhaps make a lot more efficient. What we're seeing is there's a tier of level one, level two jobs in it.
The first triage of solving a problem that comes into a help desk, giving that person really a set of tools to have an expert copilot by their side where they become exponentially more efficient at figuring out what a level three engineer previously because it was 10, 5, 10 years, 15 years of pattern matching. He's going to immediately know, oh, that's, here's what we wanna do for that. We can run all the playbooks, we can write all the playbooks, run books possible.
Somebody can search through those. So we've done a lot of that good work for that. Level one, level two, what we haven't given them is an, an ability to query something and get to a root cause recommendation across these data sets that I mentioned earlier that I think will be exponentially different over the next two to three years, maybe five years.
We'll start to see level one, level two people become much more capable to do level three work. Now what that means to our business is when we have an escalation, let's say somebody's having trouble running our product or they run into a technical problem with our product, with our platform, they call the help desk. If it's a tricky problem that gets booted to level three, that's a software developer, that's a serious engineer who's gonna look at the code, who's gonna really look at that problem with, uh, a really deep set of analysis and tools, really sometimes at a level of tools down to very, very low levels of logging, we're gonna be able to automate a lot of that so that engineer can really stay put on his deliverables rather than getting blindsided by issue escalation deger.
So I think it will make the whole organization more effective, more efficient and more streamlined as they do their work. So working that through for a minute, let's just say every IT person has their own co-pilot. Will these co-pilots need to interact with each other?
I mean, are we gonna have a scenario where, you know, my co-pilot will call your co-pilot and will resolve an issue? Or how do you think that's gonna play out? I think these co-pilots will be interconnected, but to the core essence of every software company is building gen AI capabilities into their platforms so that those platforms can be enriched to deliver better outcomes, better business solutions to the customer, better outcomes to the end user interfaces that are more intuitive to the end user.
Um, if you can imagine asking a co-pilot a series of questions, you know, that you might have to go through five or six dashboards to navigate around a tool to figure out the state of the union, but saying, I'd like to know of all the devices, which devices have the most CP utilization in this region, um, that have this process running that might exhibit this behavior because of this, um, we'll call it security issue that, that we've just become aware of. And putting that in a natural language search and then getting the result, that's a very different user experience. So you have a much more sophisticated ability to kind of use the, this immense trove of data that we have in a really intelligent way.
Almost like as your brain would think about the problem, how would you think about characterizing the information you need? And instead of having to kind of hunt and peck for it across, you know, a number of different locations, ask the question and get a response back. That I think is where we're going in terms of do the copilots need to talk to one another perhaps, but not in the use case description that I just gave, Have we reached a point where it, as it currently stands, is just too complex for humans to manage without the help of AI and machines?
'cause as I look around, there are monoliths microservices, serverless computing frameworks, event driven applications, and everything's more distributed than ever. It seems like it's unsustainable in its current form. IGI agree completely and that's why I love this business because whether it's the DevOps teams or the IT ops teams or the network engineering teams or the application developers, there is such a diversity and heterogeneity of technologies that come together to deliver a service outcome to you.
In some cases, we've seen applications that use a sliver of almost all of the architectures you've described as one integrated service. Sometimes those services have nested services underneath them that do special things to deliver an outcome for an aggregate service. These are very complex enterprise applications, but for the IT team to try to manage across it all is really hard.
And that's one of the things that the team at ScienceLogic has inspired to work on each and every day. We've really been really crossed the chasm during our history now working on observability metrics, traces and logs on the microservices applications, the traditional highly virtualized applications to the mainframe applications, to the serverless applications. We have a number of service services applications we built to run and operate our business.
So we, we actually kind of see across it all now and what we're finding is customers usually have five or six tools that they're using to try to do that. And that's really hard to manage across Mike. So that's one of the problems that we've been, we set out to solve that problem.
Um, really code day one for the company, but the next technology, the next technology keeps, you know, it's like waves crashing against the beats, the waves don't stop, they're quite relentless. So, um, we've really built an open extensible architecture to solve that problem so that we can adapt, adapt the, the platform to the next wave that's coming. Do we need to maybe have a little more patience with ai?
And I asked the question because it takes a while for this algorithms to learn the environment and the environment is constantly changing and I think maybe we expect magic out of the box, but at the end of the day it's math and science and it just takes a little time, but you get to where you want to go. I would say yes, yes, for certain application profiles, um, algorithms are really good. There's a whole series of algorithms we use just on performance analytics.
Um, interestingly just looking at performance counters, you can't use one algorithm. We'll actually test the algorithm for a period of two weeks to make sure that the algorithm we've selected automatically is still relevant. You can have a CPU that's pegged at 92 to 93% and that's actually normal behavior if you use the wrong algorithm, that'll create a series of notifications.
If you use another algorithm that knows that tight tolerance is is where you want to be. 2, you're gonna get an event. If you use another algorithm that looks for the performance counter to go from 93 to 50 to 60 to 30 to to 50, and you have much more variability, that's a different algorithm that we use to look for an anomaly with that kind of a performance counter profile.
So yes, we have to be patient, we also have to realize that things change dynamically in it and you can't just rely on one counter and one algorithm to behave the same all the time. You've gotta, you know, we now have had to over time learn that we've gotta retest, essentially retest our algorithm to make sure it's still relevant to what's going on all over again every couple weeks. The other really pertinent thing that we're finding, especially for log analytics is that you really, the the hardest log message and log set of errors to find are the ones we've never seen before because we don't have a regular expression match.
We don't have a rule, we're not looking for it. It's not a known security pattern that we would be looking for proactively with a signature saying, if I see that, that's a problem. 'cause we don't know it, we've never seen it before.
So we've now built a self-learning algorithm to look for very rare events that we've never seen before that are likely correlated to a root cause. So that's another kind of algorithm solving a different part of the problem in it. And, and I think to your earliest point, it is really complex and you can't just manage it all in one pathway, in one direction with one set of assumptions.
That is a clear pathway to a relevance and it needs a much smarter way to handle the variability that we face each and every day. So organizations that are getting value outta AIOps, what is it that they do differently that others don't? That is a great question and you're, you may not love this answer, but people who are super organized, super organized in the way that they think about running and operating and a methodical workflow process for kind of leveraging data sets to the maximum potential.
So a lot of companies when they get started in it, when I'm in, when I talk about organizational naming convention seems somewhat like, you gotta be kidding. That's a 25-year-old concept. What are you talking about?
You would be shocked how many companies don't use tagging appropriately. Don't use community strings in a, in a, or naming conventions in a really methodical way. We've seen companies that are super organized and super good at maintaining a consistency standard of operating procedures and policies for onboarding a new device.
It has these key attributes every time, no matter what, you'd be surprised how little things actually prevent great outcomes with AIOps, but a consistency in standardization and how you name things and how you describe things and how we then set up a service view around those. That is where we see some exceptional outcome. For AIOps.
What we're seeing is you really can't monitor things anymore just on a set of performance KPIs. You can, but that won't provide the outcomes that we talked about at the very beginning of the session, which is we should walk in the office and it should tell me the three things that matter. Well, to do that autonomically so that it just works so that it just happens.
You really have to, standards we know are a really good thing because they help us create a very consistent set of outcomes. And, um, and so the companies that are most organized and are most fastidious about naming conventions and standards and how they describe services within the technologies that are part of the service, uh, we can then discover that, create a service view, create the service context. And I think our belief now is companies are really, they care about the user experience, the user specifically the user experience to a service outcome.
And the service is made up of many different heterogeneous technologies that all have to work together to deliver this perfect experience for the, for the user that I think we see that the companies that get that right, the service view, the consistency and standards, and then have the right tools in place to leverage that organizational structure do really well. All right folks, you heard it here. Our IT environments in reality are pretty much a reflection of our own natural intelligence.
So maybe we need a little AI to bring some, uh, order to the chaos and it's probably cheaper than the therapy session anyway, so onward and upward. Hey Dave, thanks for being on the show. Thank you, Mike.
Great to be here. All right, and back to you guys in the studio.