Why AIOps is Critical for Networks – Andrew Colby, Vitria Technology
Monitoring network, service performance and fault identification increasingly require intelligence and automation to meet the needs of new and significantly more complex 5G and telco networks. Vitria Technology’s Andrew Colby, vice president VIA AIOps, discusses why AIOps, artificial intelligence and machine learning (AI/ML) are key capabilities for organizations to continue providing the level of performance and reliability customers expect.
Transcript
This is texturing TV. The great pleasure being joined by Andrew Colby Andrew is VP of AI Ops Adventure. Welcome, Andrew.
Good afternoon mention. Thank you. It's a great topic.
I'm excited to talk with you about it. You know, we we could go down the share War Stories in Telco experience, which could be about 10 episodes of a different show. But the you know the today in the Telco environment, we're just in the business environment in general, right the economic conditions competitive pressures looking for areas where we can you get more for Less.
There are a lot of different parameters that have shifted or changed or maybe tightened and we're currently working with them. I'd love to get your perspective on that. Certainly, and and thank you.
Yeah, I'd say we we see. Cautious optimism obviously, I'm based in in the US and the DC Metro Area Maryland and the US the government entities and quasi governmental entities have been tightening the economic structure in order to tame inflation. Fortunately that has not driven our economy and and had the potential recessionary effect that was feared but people are still cautious businesses are still cautious that said it's hard to hire people and it's really hard to hire technical people.
So a lot of companies are continuing to look towards how to leverage Technologies and automation to build efficiency so that they can do more with either the same number of people or retask their people to hire. Value purposes and let the technology do some of the more menial and mundane tasks especially and we can explore this a little bit especially in these new complex Service delivery and network environments. It's very difficult for me to imagine how an engineer who's gone through anywhere from two to eight years of college education is going to really be happy going and spending their days collecting a lot of data across Network container management VM and other infrastructure systems to figure out what's going on.
I mean really that's that's where a lot of the automation provides a significant amount of value to let the engineers to the smart difficult things that we want humans to do. You know in a lot of pressures around meantime to recovery even looking at resiliency, right? How how do we stand up under a test full situation?
Whether it be a security attack that might be going on or some unobserved condition that our systems and networks have never been under. Oh, there's so much of that so much is changing. It's not just a person like you or me behind a smartphone that can actually report that there's a problem but it's sensors and equipment that won't necessarily report right away.
So it needs to be detected. So that's a whole nother additional Dimension that service providers large Enterprise it organizations are under which is to be able to have this kind of real-time awareness of what's going on. Whether the service is real time like the video conference that that we're on or or not.
They're really is a desire and expectation to have real awareness of the service delivery to be able to detect what's going on react to it address it before the user whoever that is the customer the employee detects a problem identifies a problem and actually reports it that's really kind of the last line of defense that you want to have now that said when your customers are reporting problems I think it's really important to have a way to hear that and that's another thing that we are able to do with via AA option particular. We have customers that use that care event data, which is it's kind of messy data. It's human generated.
So it's late in the in the timeline, but when the users are telling you something's going on. You need a way to be able to hear that so that you can definitively react. I mean, I know for most people it's there's nothing more frustrating than having an interaction with your technical support teams to tell you.
Well, there's really no problem. There's nothing wrong or I don't see a problem. That's when customers feel like they know more about it than you do right.
I am experiencing this problem. Help me, please absolutely. Well say something I do want to hear some more about the AI Ops.
You can't just about can't stumble out the door without tripping over AI in some form maybe generative AI at the moment. So I'd love to hear some more about your approach around AI Ops Certainly. I mean it really starts with.
Ingesting wide variety of information structured semi-structured unstructured information enriching it in some cases classifying it and and providing additional context around it. So we take this tremendous amount of data the Telemetry The Faults and others and by enriching it we actually grow it even larger and then need to be able to To process it in a way that provides meaning for the business now historically you'd have expert Engineers who would say? Oh, well, this value should be here and above that or below that would be where we care about it.
There's just a there's two things that are happening simultaneously. There's too many measures. I was just working with a team turning up a pretty simple system on a container management system a CMS and they said well, there's like 2400 metrics just in that one system haven't even gotten really to the application measures yet.
So tell me what's important. So that's one of the challenges is just you can't set the values for everyone anymore. And then the other challenge is to know.
Which of them are important again, it depends a little bit upon how the service is consumed. Is it one way or bidirectional is it real time? Is it latency sensitive is a packet loss sensitive.
What are the characteristics of it? And depending on those is going to be what's important and then what's normal? I mean every service has some some sort of pattern throughout the day and throughout the days of the week.
The machine learning young supervised learning is able to learn what those normals are and then identify when those measures and values deviate from that normal and alert on it and more important or equally important to just alerting on it is to be able to bring together all of the related information and that's where the enrichment becomes important. It's not just enough to say. Oh, well this measure changed but what is it related to what are the services that run over that infrastructure?
What are the other measures that? Changed among those same infrastructure at about the same time. What were the plans changes or or manual changes that went in in the prior 15 minutes because maybe one of them had an impact maybe it didn't but that's really important and useful information to know.
So there's all this bringing together of the end of the enriched information in order to get a full picture of what's going on. And and that's that's a lot of the value of the Automation and the AI Ops brings together. You don't need to have your highly skilled expensive Engineers do a lot of that manual data Gathering let them let the machine intelligence bring that together for you and let the engineers do the the value that humans can do effectively which is interpret that identify new patterns.
See what's new especially if you're rolling out a new service or you have different behaviors going on. Now there's a wide variety of AI and ml that goes into this you mentioned generative AI which is a very interesting topic and has a lot of both Technical and and popular coverage and attention right now and there are places where that's effective to be used for things like the data extraction and ingestion, which is learning what the data is and understanding how to interpret it as well as potentially developing a suggested fix or likely cause or likely fix that could be generated from all of the different inputs and individual root issues that are identified. But there's also a variety of other traditional machine learning and AI techniques that really provide a lot of value and are a part of via AI option and we leverage very extensively.
You know where I feel like networking security. They're very broad topics, right? You could you know, a few words you're talking about many many different Technologies.
The same is true. I think also for AI there's there's so many fields of it. It's a bit overwhelming to folks.
I know you prescribe kind of take an incremental approach, you know, don't try to eat the whole often at one if you will that old adage. But what is that incremental approach? How do you work with customers to start to figure out that path?
They should go down. It's great question. I'd say there's a couple of of steps to that one is To focus on the business value.
What is it that you want to achieve and how can you tell that you've achieved it? It's not just enough to collect a lot of data and produce something that looks a little interesting. Aha.
Can you do something about it? Does it have an impact on the service? What is the measurable impact on the service?
So being focused on business value and even directly measuring it is one important piece the other is because this is all about Network and service operations. It's it's generally something that's done across an entire service or entire company. It can't it's not so easy to to try on maybe a departmental level.
So it's it's a big decision for companies. Because of the way VIA aiops is built and structured we enable it to be delivered or we enabled delivery of what we call incremental transformation, which is the ability to augment the existing or augment the machine intelligence and AI with the institutional knowledge. So to be able to specify policies, for example for things that are to differentiate what's important for the business from what's important in just a measure to leverage the value of existing Investments.
Nobody's starting this from scratch. So there's always investments in application performance monitoring or other network and service monitoring tools. They're not bad, but they just may be sort of siled or they may have a real key piece of information in one area and we want to leverage that across the entire Service delivery.
So that's another way to to provide that incremental transformation and Leverage. And overall improving the efficiency of the operation staff and being able to deliver these as really as individual use cases incrementally to continue to provide business value over time. And you mentioned automation before?
You talk about metrics and and things to be what should you should be observing it all starts with getting the data. And as you mentioned bringing, it's not just correlation anymore. Is that correlation of events?
Yes, we know do that. We need to continue to do that. But if you have You're a ten factors that are actually part of what's going on.
That's usually bigger than what someone looking at a monitor or screen up in an OP Center is really going to be able to put up together and this seems like that complexity or maybe the speed of that happening is also a big driver for where you might consider AI agree disagreeing. Absolutely. I mean the kind of highest level metrics are things like the meantime to understand mttu and mean time to restore mttr.
Those are top level metrics and you want to build down and drive down from those which is what what goes into them. What does it take to understand a problem in one of our customers the challenge was not only to understand the problem but identify which part of the network the problem was in so they can get to the right fixation quickly because sometimes you create in an incident management system a ticket, but if it goes to the wrong team, they have to triage evaluate it and they say oh and not us it has to go to maybe goes to the firewall team instead of the network instead of the land team. So that's that'll take a lot of time.
So getting that. the likely root issue and likely fix to be in the right area is really important to be able to do that. So it's areas like water what what are the components that drive that mttu?
How do you measure it? We have a variety of customers and they take a variety of different approaches. In the end what they want to achieve is a higher percentage of incidents that are handled through automation.
You can do that by decreasing the overall number of incidents as well as by increasing the number that are handled through Automation and and we're able to take both approaches. Because back to what you originally talking about not being able to hire enough people with the right people. They may not exist.
Right? So people would like to have some purple unicorns, but that's the people don't always have those skills we're looking for so it isn't always just about downsizing people. Sometimes it's bringing that curve of what you need to hire.
Right sizing that just to get so you can handle the workload with the staff that you have. Right and and some of that is to be able to to free up the staff from things like monitoring screens or systems that are just telling you red and green or up or down because that's that doesn't have enough context, but you need to understand and and often it's not just simply red or green. It's gray.
Or purple which is well it's working but it's not working to the level that we want or need or expect in order to provide the level of service that our customers expect or that our service requires underlying service requires. So it's being able to provide all that nuance. As well as that level of detail and insight all that enriched information we go back to that again so that the the right action can be taken eventually once the right action gets taken 10 or 50 or 100 times my expectation.
Is that the The trust will have been built in the system. So that that audit that action can be now taking in an automated fashion. And again, that's an opportunity to accelerate that time to free up an engineer from doing something that they've done 98 times before and be able to more quickly allow the action to be taken.
Obviously you'd still maybe want to have some post analysis review to see well did we take the right action is the right is that same action being taken every day? Maybe there's some other problem that really we need to address but still as long as they can deliver the service in a way that meets the customers and expectation and the service expectation. They're achieving their goal and they can they can enable engineers and others network operations to kind of free up to think at the higher level about what they need to do and can do to continue to achieve that effectively Yeah, we live in a world where it's very easy for customers to say your service taking too long.
I'll jump to the next app and the next side or whatever, you know, it's it's there is loyalty but there's also patients that tries that absolutely trust and loyalty. I love to have you read a few minutes left left to have you talk some more about the incremental approach. You know, I think all the buzz about generative Ai and you hear everything from we won't need programmers.
We won't need these kind of people. We don't need those kind of people, you know, we have Robotics and factories, but we still need people in factories, right? I'm not a subscriber to the the pendulum doesn't Slam against the other side all the sudden this kind of changes don't happen very often.
So it's probably somewhere in the middle where end up so how do you knowing that if you take that assumption, how do you Get aggressive enough that you're getting some value from aiops. And I mean so conservative that you kind of missing the opportunity. I see again, it goes back to Aligning your actions with the business values you want to achieve?
I'm working with a a customer who says they want to achieve a value of running their operations with significantly fewer people like maybe half or more less in a than a traditional Network Operation Center would have in in a circuit switch world. Let's say That's the goal from the top down the people that are responsible from the from the bottom up. They're like low slow down.
Hold on don't don't do all of your automation yet. We want to look at everything first. We want to see because that's how they've been used to dealing with and that's some of the place where there's tension in being able to do this when the tests are done and you measure response time from issue currents to issue detection to issue resolution.
There's a lot of human think time in there and it's not wasted. It's it's people doing their due diligence. Is it really down?
What can I restore it myself? Do I need to take an action with an outside team, but those all need to be aligned? And that's how you can achieve a result and it's hard in a large organization and in all these cases were we're talking about large complex organizations with large complex Service delivery environments.
And you mentioned earlier the trust that's built in that process to right? We don't just throw AI into the mix and say it's in charge now, let's walk away and I hope it does. All right.
No, we need to know they need but it's gonna do is the right thing. Right? Right.
We need to build that trust and and that's actually one of the reasons that again we go back to this incremental transformation. We think that's really important to be able to do it in steps so that you can see. All right.
This is what the system shows and we think the action from that system should be so we can give you a button to say all right when you believe it's be click on B. And again, you do that 50 times after a while. You're gonna get tired of just staying well every time it's be why don't you just do that for me automatically and that's what we want to get to and supervised learning right for yeah.
Very much. Well Andrew, it's been fascinating talk with you. I hope you get a chance to come back and chat some more and really appreciate you sharing sort of on the ground experiences with customers as you're working as in in a large setting is one thing to adopt AI in a startup is another in a large Toko environment that generally doesn't make those kind of shifts, you know on an hourly bad daily basis that takes a bit to get that flywheel turning in the new direction that you're headed.
So working folks find out more about via apps. Thanks Mitch. com on our homepage and on the resource tab, you'll find a suite of information about how you can actually realize this business value.
And the types of capabilities vitrea and aiops can deliver. Great a lot of great resources under that resource app. So be sure and check it out.
Thanks for joining me. Again. Andrew Colby who is a VP of AI apps at vitrea.
Thank you soon. Look forward to