AI’s Impact on Operational Intelligence – Kannan Kothandaraman, Selector AI
Selector AI CEO Kannan Kothandaraman explains how artificial intelligence (AI) will drive the next era of operational intelligence in the enterprise.
Transcript
This is Techstrong tv. Hey guys, thanks for the throw. ai, and we're talking about how AIOps is gonna be applied to operational intelligence.
The world is changing as we know it. Con, welcome the show. Thank you.
Thank you. There are a lot of folks talking about AIOps over the years, and some have been more successful than others, and of course, we've been doing various forms of operational intelligence with any number of analytics platforms. How has this whole space evolving from your perspective?
How are things changing at the moment? Thanks, Mike. Yeah, so the, what is selector?
I'll just do a brief introduction. You know, we provide observability and aiop solutions for our customers and focused on mission critical networks and application infrastructure. Uh, right.
You know, the team here is from, you know, Juniper, Cisco, VMware, and Nutanix. And we've built products that run some of the largest networks and application infrastructure in the world. So what we are doing here is bringing, you know, ML machine learning analytics into the operational world for, you know, networks and applications.
Uh, so why is this required? I, very briefly, there's been massive architectural changes, uh, in how infrastructure is built and delivered and what it is used for, right? The, the business transformation, right?
Uh, applications are no longer monolithic delivered in a single data center. You know, they are definitely, you know, stating obvious lot of movement to the cloud as well as edge enterprises are taking a hybrid approach. All of this, of course, is in the interest of optimization, but it also creates big churn and new complexity, right?
So new solutions for observability are required to tackle this complexity and change in how infrastructure is supporting business requirements. We've been saying for years that the network is the one source of truth in the IT environment, but we always seem to have a hard time getting visibility into not just the network, but what's happening in the stack of software above it. So how did you guys approach that issue?
That's a, it's a good point. Uh, if you look at networks even, yeah, from five years ago, it was primarily about connectivity. Uh, right?
You know, like, you know, it evolved into secure connectivity, but, hey, how do I get from A to B? And, you know, the SLAs were based on that. That's no longer the case.
Uh, when whenever I'm building networks, or I'm selling network services, you know, how my applications are performing, my user experience has become, you know, more important than just connectivity from A to B. And that basically means that, you know, networks have to be automated and observability has to combine both underlying network infrastructure, application, and user experience, right? So these are inherently very different domains that, you know, existing tools we're not looking at, you know, holistically, right?
And, and of course, the data and the volume that you need to look at is also increasing. So this is where the promise of, you know, applying machine learning, but still having that context. You know, you can't just bring in machine learning that a Google or a Facebook has used to recognize our photos or, you know, recognize, uh, you know, Karen dog pictures into networking.
You need to have domain context. How do you tie underlying infrastructure to applications? And then ML is just one, another one of the tools that can enable the, these various domains, uh, to coexist, right?
And in the end, of course, to deliver these sort of services, you need to have automation, but you can't automate what you can't see. Uh, right? So we approach this in a, in a very layered manner, right?
You know, you start with understanding the network. How do you automate visibility, you know, from the network in various aspects. If you look at a network itself, there is data center, there is van, there is, you know, SD van, there is, you know, network applications.
And then how do you tie that to the application, uh, layer? It requires, you know, working closely with, you know, our, our, our customers to understand their business processes and, and their business workflows so that those tie-ins can be made, right? It's not going to be something that, you know, is automatic and comes out of the box, right?
You have to work, uh, with, you know, subject matter experts and practitioners who understand, you know, the, the timesin between the two, right? And then the ML is finally, it's, it's an automation tool, right? So once you understand these, the, the, the, the various workflows and the connections ML basically automates, you know, rather than a human tying the, the dots together.
Mm-hmm. Today, it seems like we have a lot of AIOps platforms, but maybe they're optimized for particular silos. So do we need to kinda a converge that?
And then b, is this a step towards, um, today if I look around the world, there's NetOps and IT ops and DevOps, and everybody's got an ops for something, um, is all that gonna come back together in some cohesive way? Uh, I mean, great point, actually, you know, first your earlier point on yes, everybody, you know, everybody is adding, you know, some ML capabilities to their tool. Uh, right.
You know, so that's why if you look at it, you know, if you are, uh, a, you know, an infrastructure vendor, you have to have AIOps, you know, in your, uh, moniker, so to speak, right? But the key, what, you know, end practitioners operations teams need is, you know, I cannot have an AI ops from vendor one, another AI ops from vendor two, because you, you, you are just basically now back to dealing with multiple vendors, multiple domains, uh, right. You know, I, I might have, you know, as I mentioned before, multiple aspects of infrastructure each with its own AIOps that each of those individual tools are bringing.
I might have a wifi vendor, yes, I have AIOps, but you know, is the problem really in wifi? Yeah. Then maybe the AIOps there, you know, could help to troubleshoot, uh, wifi.
But what is my problem is somewhere else, you know, my d n s is broken, right? So who is going to bring these various domains together, right? So that is what, you know, uh, a solution that, you know, we are bringing in helps, right?
You know, is your problem, you know, in L one layer of the stack, or is it basically, did somebody make a configuration change? Uh, right. You brought in, you know, DevOps, right?
CI/CD, uh, you know, constant application changes are happening because developers are encouraged to push changes on an ongoing basis. CI/CD promise. Like, okay, hey, I, I make a change.
I, I tested using a CI/CD pipeline and I push it out to production, it might break something. How do you now know that, okay, you know, there is no problem in the underlying infrastructure, but this change that somebody made is the reason that, you know, my applications are, uh, failing. Bringing that in, you know, to the overall observability stack, uh, right.
Those are the use cases that, you know, selector is focused on is, you know, we are agnostic to the data and the domains. Uh, right. You know, the more data, the better.
The, the overall solution is how do you then, you know, tie them together, right? How do you do the correlation and also build the intelligence, you know, before the correlation can happen, right? And that data can be anywhere.
And because for operations teams dealing with issues, the data really comes from anywhere. So they cannot say, oh, you know, I, I, you know, this is a tool that I don't know about, so I'm gonna ignore it. Who's gonna fix the problem?
Right? So that's the true premise for what I call multi-domain AIOps, right? So, you know, there are, there is, there are single vendor, single domain AIOps great, uh, you know, that is needed.
They might be able to give us better intelligence for their specific, you know, device or domain, but who is gonna stitch it all together? And that's basically where a solution like select AI comes in. It also seems that the systems themselves are getting smarter in this regard.
Originally, we gave people an analytics tool and a query tool, and we told 'em it was an observability platform, and they could query that thing to go find out, you know, whatever the root cause of an issue might have been. The challenge is that nobody knew what question to ask. So, um, you know, are we getting better where the algorithms are gonna just tell us what the issues are, and we don't necessarily need to know what questions to ask as much as we used to.
Definitely. Right? I think this is where, you know, the promise of, you know, generative AI or, you know, large language models, you know, start to come in, but you know, it's you, you get what you feed it, uh, right.
You know, the, it's, yeah, it's, it's go, it's, uh, if you feed, uh, you know, just random data to it, you know, it's going to provide random feedback to you, right? So, uh, if you, if like to, if, uh, if you look at, you know, where large language models are evolving, there is this notion of a context layer that, you know, that is needed. Every enterprise, you have raw data, but the LLMs should deal with that context layer, which, which is basically, you know, to provide, you know, semantic understanding of what the underlying data and business processes is, and then applying LLMs on top of that can, you know, ha holds the promise of, Hey, I can ask questions, you know, as we would to each other, rather than, okay, I need, you know, c l I from vendor X C L I from vendor Y, right?
And I, and I have all of these, uh, you know, uh, tribal knowledge that I need to apply. That context layer is, is built from, you know, the intelligence that the underlying data is providing, right? Uh, you know, somebody might ask, okay, hey, now, uh, you know, do I have a problem with my, you know, SD van, uh, circuit for my application X, right?
Automatically, L LLM is not going to understand that because, you know, there's a lot of, you know, domain sense, uh, or, or, you know, uh, business context that needs to be built that, you know, layer in between and, uh, you know, is crucial, which then means that, you know, users can interact, right? And, and in, you know, typically interacting with a particular device or a particular, you know, vendor, you know, that is okay, but when you have 25, 30 different, you know, vendors and you know, you are, look at the amount of change that is happening in the infrastructure, I'm no longer dealing with, you know, just, you know, Cisco or Juniper or Microsoft, right? I now have cloud.
I, you know, I might be going through, uh, a, a, a cloud connectivity provider. How do I suddenly start to bring all of that in and, you know, do I ask questions, you know, uh, of each other, uh, or, or specific questions of each of those parts of the, the infrastructure each with its own, you know, context. Uh, right now somebody has to normalize it so that I can ask questions like, Hey, you know, is my external partner down and why are they down?
Right? So that could be where I'm going through a wifi network to an sd a network to a cloud connectivity all the way up to assess, and then I might have, you know, uh, you know, cloud networking or infrastructure to deal with. How do you build that context layer end to end, right?
So that's when these things will become truly useful. Mm-hmm. And it sounds like we're gonna have to build a lot of, uh, large language models that have domain expertise in a particular area, and we just can't rely on a general purpose AI engine that was built for, you know, data scraped from everywhere and ended on a certain date.
Yeah. It's a great starting point, but very quickly, users will see frustration because you know, it's going to provide generic answers. Uh, right.
And I like to joke. I mean, you know, uh, without a lot of parameter tuning, you know, and pro tuning, uh, what LLMs or chat GPT like tools will return is, you know, what a cio, uh, or a, sorry, A C E O might say in a earnings conference call, right? It, it's, it's, it's so generic, right?
They're saying a lot of, you know, words, but, you know, it's, it's sort of very generic, uh, right? Whereas what operations teams and operators need is very specific outputs that allows them to solve the problems that they have in front of them, right? So that does require a lot of context building here.
Do you think as we go along that maybe we're gonna have to reorganize the way it teams are structured? Cuz it seems like we've allowed a lot of silos to be built up over the years, and maybe it's time to kind of rethink that. I'm not suggesting that various specialties might go away, but the way we think about how folks invoke various capabilities might be, uh, more accessible because networking's gonna get a little easier to deal with.
Uh, you know, I think if, uh, I'm not an expert on that, I'll, I'll, I'll, I'll admit, and I'm, I'm always being on the product side of the house. Uh, right. I, I, you know, I am not, you know, I'm not convinced that will happen because each of the, you know, domains are getting more complex, uh, right.
You know, you do need to have expertise, uh, and it's not, nothing is getting simpler, uh, right. To where, okay, hey, I can have one common layer on top, but providing tools that make it easy. Uh, right.
The, the starting point is, you know, can we have ops engineering teams not drawn in the complexity today? Which they are. I mean, you know, they, they, their lives are, you know, really, really hard and complex.
Can we make it simple and at least to, you know, understand where the issues are. Right. You know, we focus on, you know, the three crucial things we focus on first, you know, meantime to detect, uh, meantime to innocence, uh, right.
You know, and meantime to repair if you, if you get there, you know, if we can, you know, be effective here, right. It'll free them up to do, you know, more valuable work right now, you know, because there is so much heterogeneity, uh, and, and, and, and tribalism in, you know, tools and domains. They are drowning just in, in, in these three aspects.
Forget about mttr, just meantime to detect and meantime to innocence. Right? And, uh, that alone will save them a lot of cycles where then they can start to do things that are truly, you know, outcome and value driven.
Mm-hmm. So what is your best advice to IT leaders? Cuz many of them are feeling a certain sense of paralysis of analysis.
They're like, there's so much coming down the pike and there's so much change ahead, they're not quite sure what to do next. Yeah. Uh, you know, uh, starting with, you know, the basics on, you know, the, uh, like what is the goal, right?
Uh, I think everybody is clear. Their goal is automation, right? You know, a cloud is just one method to get there, right?
How do I automate as much as possible so that I can reduce the burden on my teams and, and, and, and save costs and, you know, improve overall, you know, customer satisfaction. Uh, and you can't automate what you don't know. Uh, you can't automate what is not visible.
So it definitely starts with, uh, understanding the data. Uh, right. You know, you cannot apply a lot of those new technologies on ML till your data is, is, is, is, I'm not even gonna say clean, but your data is usable.
Uh, right. So, you know, starting with those basics, uh, right. You know, do I have, you know, good practices on data, uh, you know, am I able to get, you know, basic intelligence and insights out of that?
And then use that to, you know, drive transformations and automation, right? So that process, you can't jump straight away to optimization and, and optimization till you do that, you know, uh, the, the, the basic workforce, right? All right, folks here heard it here.
It starts with data, it starts with science, and there's never really not that much magic involved. Hey, Conan, thanks for being on the show. Thank you.
Thank you, Mike. All back to you guys in the studio.