AIOps Service Assurance – Chris Menier, Vitria Technology
Chris Menier, president VIA AIOps at Vitria Technology, discusses how observability and VIA AIOps enable a new service assurance operating model across applications, networks, and in-house, hybrid and cloud infrastructure. This new operational model accelerates the resolution of service-impacting events through process automation and substitutes intelligent automation for high levels of human expertise in the event monitoring and incident management process.
Transcript
This is texturing TV. The great pleasure being joined by Chris meniere. Chris is president of viops at Via aiops.
Excuse me. Welcome Chris. Thanks Mitch really happy to be here.
Happy to have you here. It's a great a great topic great area for us to be discussing when he first give you a chance to introduce yourself and tell us a little bit about about the company. I absolutely thank you.
Yes, so, my name is Chris meniere. I'm the president of via aiops here at vitrea. I've been leading this effort for about the last five and a half years and prior to that really grown up in the analytics space have about 25 years and in these industries working for service providers equipment providers and analytics companies as well.
So was was early in the Big Data space between before Big Data was was really a term in the analytic acceleration of data warehousing at a company called natiza, which ended up being acquired by IBM. After that, I led product in marketing and then eventually the group CTO the company called guavas, which was another acquisition by a large conglomerate Talis who was looking at Network analytics across various data centers. And you know, I I dabble in it from a technology perspective as much as I can and you know have some patents behind some of the work that we've done as well.
So I really like to get into the space. It's sort of my hobby and my profession at the same time pen test. Well tell us a little bit more about vitrea and Via AI Ops specifically.
Yeah, absolutely so victory has been around for a bit. It's about 25 year old plus company and we grew up always supporting this model driven Computing Paradigm and sort of pointing that at the problem of the day. So this when the company was founded it was really back on.
How do I get my legacy infrastructure on the internet effectively. So if you think about You know just Legacy companies that wanted to get into e-commerce and things was a lot of business process Automation. And so we had a very successful product offering called business where which is still actually deployed in A bunch of companies around the world the this excess of that let us to get very good at just handling data.
So whether it was a big data Technologies high volume data messy data that came in out of order and all that other good stuff. We just got really good at bringing data in normalizing it processing it for workflow. And so when I joined the company about five years ago or so there was a an emerging space in this AI op space and there was a problem to be solved there.
And the problem that we wanted to solve was in these large complex Network environments where the service delivery is dynamic and and you know, some of the problems we had there were not just the basic data volume problems, but you had false with performance data with a lot of change that was all coming together that caused service disruption. And so we've pointed our technology at that for the last five or six years and we've had some great success in some of the largest complex Service delivery environments in the world. You know it.
I have always had respect for network operations operations organizations and but it's it's grown. So even so much more now, it feels like you have to be a fighter pilot to be an Ops because you're either watch everything that's going on in the cockpit or you look at the heads up display to tell you, you know what to do next and where you're headed but you've got to have all that other sensory data available to you to then kind of dig into triage respond to which of course is a data challenge. It's an automation challenge process challenge so you can you can both respond human and wise and automated and then we can apply things like AI.
To it. Love to love to hear your philosophy on this because given your background wasn't in just in network operations, you're bringing the data the automation, of course AI in the platform and approach to it as well. How do you think about solving this problem?
So the first thing is we try to keep the service experience in mind and more importantly the customer in mind at all times. So we see the land we see things through the lens of the customer the consumer of the application or the consumer of the service. Yes.
I understand we can get a little bit esoteric and and a pedantic about what what the consumer is. Sometimes it's another service I get that but you know, let's take some simple ones if I'm a large Enterprise my end user might be my employee and that employee needs to have VPN service so they can sort of do their job and work from home. If I'm a communication service provider.
Obviously, it's the person on the other side of the smartphone who is the consumer of this first and foremost. What we want to do is eliminate the customer from being the canary. We need to get help our customers get ahead of those service disruption issues.
The customers aren't calling into customer care saying I can't get on such and such. I can't execute such and such. So our job if we think about it really simply is to get ahead of that.
And so first things first and I'll expand on your your pilot analogy your cockpit analogy a little bit here and to say that to be a little confrontational. We are the anti-single pain of glass companies as well because the problems that we see today are outpacing human scale and you mentioned that on the data side of it. So it no longer works that we have these fixed Service delivery ecosystems.
They're dynamic. It no longer works to have everything on a single screen and think well if I just take all of my element Management in my Telemetry systems, my apm's the tools and I put all their red dots on a single screen. I is a human I'm gonna be able to consume and triage that it just doesn't work anymore.
And so what we think about is really taking in that variety across the total ecosystem of fault and performance and change data we enrich it with valuable information such as the inventory basic information like the make and model of devices topology, which is essential in core into this we can learn the topological relationships through the data itself. And then finally the service dependencies and then when we have that rich data now we've taken a small amount of data and are a large amount of data and made it a ridiculously large amount of data which complicates the problems even further. So that's where the AI starts to come into play of pulling it all together in a logical.
Doing that correlation for you that a human just can't do well if you think about it, you know, my cockpit analogy was less about single plane of glass and more about you can only look at so many things everything else around you has to be automated in some form. There you go. Okay now listens.
Yeah, so not not to be right about everything but that was my why was my failure analogy so I can make my confrontational single pain in class. It worked it worked, you know, what's interesting. So, you know the the rich the mated data the additional information, you know often times we call that context right?
It's often times whatever Network or operations person in observability tools looking for is I don't know what that is. I don't know what that means when that happens or that happens in sequence or in relation to other events and the time it would take a person to correlate that connect the dots and realize that it means X and that actually is well, it's probably already happened not for telling a problem, right? That point the insurance happened the customers unhappy but we're in such a of an environment where the volumes are so high but also the environment is changing, you know, it isn't a rack of equipment all the time anymore.
That's right. It's software talking to itself doing it. That's right in the seven layer stack or inside the application stack.
Yeah, and so we think about it as you have an application or a service and if we if we just want to normalize it a little bit. There's a client in a host and that client and host have to communicate and they communicate over a very complex Network. Sometimes the public internet or you know, at least private Network and that application or service is residing on some virtual infrastructure and that virtual infrastructure is underpinned by a physical infrastructure.
And as we know it's dynamically changing, where is my microservice running in what VM? My container was ejected and brought back to life. Now my whole service dependency and service topology has changed so our job is to understand that in real time.
So we can do those correlations and get down to understanding the the impact of an incident itself and I'll give you a very simple example in the network and and not to keep coming back to the network side. I mean, we play across Enterprises but these are sometimes just easy examples to for us to consume in a network one of the most frequent things that happens is link failures, right? You have an interface go down on some router or a switch or something like that one of the most important things to know and it seems so simple is what's on the other side what's on the other side of that?
And is that going to be service impacting so our ability to learn and understand and correlate immediately. What's on the other side? Then allows us to understand is this service impacting?
What's the priority and then what are some automation routines we can do to restore the service? Yeah, a very very clear example of something where yeah, what is the impact and how would I know that without? Either years of experience of knowing my topology which I was all changed now, which is all changed.
Yeah. Well and I like your comment about microservices because I think security for Network that that's brought in a whole new dimension of You know the application isn't just surrounded by you know, a bunch of software hardened and then a few interfaces in and out microservices. It's apis talking to itself all over insightly containers are organized by orchestration like kubernetes or whatever.
It's sort of its own little Network either within or across wherever that all the pieces of that application Live that's a hard thing to get your head around because you know kubernetes is spinning things up and that's growing spinning up new microservices instances of the same service. They may not be there by the by the time you look at the incident that's occurred humans can't follow it that quickly. It's just not possible.
No. No, it's not and there is value obviously in institutional knowledge of the the services that you're operating. So a couple things on on your comment there is one we understand that that instit knowledge is important but it only goes so far.
So we like the concept of AI plus hi artificial intelligence plus human intelligence is going to be greater than artificial intelligence by itself. Right? So we allow that institutional knowledge to basically be injected into our process flow and our analytics to apply some best practices and some policies and some labeling if you will of what's going on.
I sort of laugh at at our industry because we we subscribe to Magic like we love magical thinking whenever there's AI in there, you know people have this magical thinking and so when I see companies who are out promoting no configuration required no policy required. No rules. No anything.
It just works. I sort of chuckle and think okay. There's that magical thinking in our industry again because we're we're really missing an opportunity if we're not applying that institutional knowledge.
But to your point it doesn't scale. And one of the reasons it doesn't scale is because of change. So we at Via think about change as a first class object or a first class event in terms of operational troubleshooting Operational Support.
And so 37 or some odd percent of service impacting issues on one study our resulting of as a result of change, right and as we see devops and like you said you have kubernetes that's doing things and orchestrating these in sort of real time. That number is probably a little bit low and so our job is to identify that change and then immediately identify any impact of that change sometimes change is good. We see kpis increase we see fall to go down but a lot of times change will actually cause issues and you know, one of the first things that happens in operations.
Well, what did you change what's changed? Yeah. No, I didn't change anything.
I didn't change anything. Here, right, but that script remember that thing or that orchestrator, it changed a lot of stuff and it did that because you wanted it to but again as we start to introduce more devops more and more cicd life cycles more and more just rapid constant change. We put our so we open ourselves up to that change potentially negatively impacting the service and not always just positively impacting it it could well just bad at a time before we need to leave.
I'm curious just you're you're ideas about where AI is playing in your kind of product strategy looking forward, right? We're here a lot about generative Ai and how many jobs that's gonna take, right? Yeah.
I'm not prescribing to that. But I'm curious for our Ai and ML and your world. What what's what increasing or expanding role do you see it taking?
Yeah, what I what I love about our product in in how we use Ai and ml is we really use in this pipeline effect. So every step along the way is that data is coming in being enriched pulling out key information from you know, log files learning the topologies and dependencies of that classifying and categorizing them. We're able to do this in real time at massive scale across a pipeline.
So what we're gonna see in our industry is not the evolution of some magic new algorithm. It's really going to be how that is being applied so it can be some some basic AI that we see today, but it can be the application of it at scale through a pipeline which I think will continue to evolve and from that you're gonna get the results like our customers are getting where they're cutting their mttr down by 40% for outages in 80% for impairments. We have one of our scale cable companies who's pulled out over a hundred thousand technician visits out of the business that they're directly attributing and that's not from some magic algorithm.
That's from applying that Ai and machine learning across that pipeline in a scalable way. Buckaro is a real money to Telecom and there you go in there. Yeah.
Absolutely. We're can't working folks. Go to find out more check check out the technology and what you all do.
com. com, you'll be able to sign up for a demo with real data real use cases. Not just not just smoke and mirrors and actually pull down some real case studies and white papers that I think will help you on your your AI Ops Journey.
Yeah. There's a great demo on the victory on the Via. Yeah.
Yeah Ops page definitely check that out. Talk about optimizing Services Assurance. So well thanks.
Chris has been a pleasure having you on. Hope you come back again and tell us more about what's happening. So we'll try to get into the same city.
We're on the same Coast this time. There you go home connect another way. There's been a pleasure talking to you Chris Manya who is president of via AI Ops at vitrea technology.
Thanks Chris. Thank you. Thanks Mitch.