Transforming IT Operations: The Role of AI and Observability with Travis Greene at OpenText World 2024
This discussion focuses on the evolution of IT operations management through observability and AIOps. The integration of these technologies reshapes ITSM and DevOps workflows, emphasizing collaboration and the potential merging of IT roles. AI agents are crucial for managing complexity, while predictive AI aids in anomaly detection. CIOs prioritize cost reduction, leveraging automation to enhance efficiency.
Transcript
This is Textron tv. All right folks. And we're back at Las Vegas at the Open Text World Conference, and we're talking about how IT operations management is gonna evolve with my good friend Travis here.
Welcome to the show. Thanks, Mike. Hi, how are you?
You know, I'm doing this probably too long, but I feel like it, operations management has never been seeing more change than it is right now. If I look at observability and AI ops and the melding of ITSM and DevOps workflows, all this stuff seems to be converging at this point. Why?
Well, there's a lot of change, but it's still foundational to how operations covers everything that you see here. I look around at all the booths and people are talking about content management, and they're talking about security. And none of that matters.
If the, the infrastructure, the, the, the databases, the networks, if those things aren't carrying that load sufficiently, then it's, it's critically important. And so I think the change that you see is being driven by the need to keep pace with the demand that's coming from all of these consumers of the things that IT operations has traditionally provided. One of the things that I hear from those folks all the time is you mention the word observability and they nod their head, but in the back of their head, they think it still means like basic monitoring.
And I'm looking at a bunch of predefined metrics. From your perspective, what is it about observability that's different than what we might have thought about with just traditional monitoring? Well, observability, uh, I think there's a group of people who think of observability as dollar signs because it is enormously expensive in the ways that people are consuming it today from some very popular vendors who are out there.
So one of the challenges that we wanna address from an OPEX perspective is how can we reduce that cost of observability so that it does come back to more of what we traditionally have seen from a monitoring perspective. But I think that most people consider observability today to be something that's focused on the applications, because that's where it came out of. I Google coined the phrase, developed this concept of site reliability engineering and has now provided it as a means of collecting those metrics, logs, traces.
And we would, we would extend that to events. We actually call it melt, uh, metrics, events, logs and traces for being able to instrument the applications and make sure that they're performing and available. One of the things that I saw at the show is, uh, OpenText extended its observability into the realm of networking.
I think before that you add applications and infrastructure. Yeah. Are all these things gonna converge into like one job function, or is it just more like we're trying to make these different silos be able to more easily collaborate with each other?
I don't know that there will be one job silo. I still think there's a need for network engineers and, and network operations teams. There's a, a, a specialized skill set for applications and databases and infrastructure cloud.
All of those things still need specialized approaches, but what we are trying to do is bring together infrastructure application and network observability, which, you know, we, we haven't thought of, again, observability has been thought of more as an application role, but how can some of the disciplines around observability, like, uh, having error budgets, like having service level objectives that we are measured against, how can we bring some of those disciplines from the observability side of the house into the way that we manage our infrastructure and networks as well to fully support the needs of the business? I'm asking that question a little bit. 'cause we have seen the rise of AIOps and it seems like if I, to make AI work, I have to be able to pull all this data to train the models.
And if I to do that, I need some kind of platform like yours to drive that. So if ai, if we rely more on, uh, AI agents to help us manage IT operations, is the future of IT a mixture of AI agents and people working hand in glove? How do you see that all coming together?
Well, You probably heard our CEOs say this a bunch of times, but let the machines do the work. I mean, we're talking complexity and scale that has never been as great as it is, and it will never be this simple again, we we're, we are in a world where complexity is continuing to expand and we're having more things thrown at us. So there's no way to approach any of this without some level of AI and automation.
And so there is a, sometimes a debate about the difference between observability and AI ops and traditional monitoring. Uh, we see AI ops as the place where all of those things come together. So, and, and of course, applying AI to sort through all those needles in the haystacks, or as we heard the keynote this morning, the needle that's been broken up into four pieces and stuck in the four haystacks and now trying to, to make sense out of all of that.
So yeah, AI is gonna play a huge role in that, but it certainly doesn't displace the need for the people to be able to, to deliver, to be more facilitators rather than knowledge holders. Because what we see is there's a lot of heroes out there. There's a lot of people in it who want to be and, and are take pride, and they should take pride and being the people who know how to repair things when they go down.
But the scale and the complexity just doesn't allow that to, to be a way forward for the, the environments that we're living in today. A wise man once told me that if a process requires a hero, it just means it's broken. Yeah.
That's another good way to put it. Um, do you think that, uh, as this whole space continues to evolve, uh, what do I need to know to be an IT expert? Right.
I used to have to get down into the weeds networking. People had their little CLI and they're like, I'm holding onto that forever. Aren't we evolving into supervisors rather than people who perform toss?
There's still gonna be a need for people who have sort of the low level understanding, you know, from an architecture perspective and, and how, how these things work together to create this, this beautiful set of, of solutions that people interact with every day and, and make it so that IT operations is the, the unsung hero in, in the background. Maybe to take the hero analogy a little different way, is we, we work best in IT operations when no one knows that we're there when, uh, other than our teams, of course, but when, uh, the rest of the organization is just their, their apps are, are humming along and producing revenue or supporting the, the needs of the users. So I think that there's a, a sense of how do we align ourselves better to those sets of metrics rather than looking at the traditional metrics of, you know, uh, how many changes did I implement today?
Or how many tickets did I close? That sort of thing. We have to realign ourselves to what, what are the, the actual top needs of the business and are they performing at a level that we, that the, that the organization wants for?
I don't know if you've ever heard the old joke about, you know, what's the one thing that an IT administrator and an application developer can agree on? Not working guy's fault. Yeah.
Oh, of course dogs. And I'm mad telling you that joke because so much of the history of it involves a lot of finger pointing, otherwise known as more rooms and mean time, the innocence. Can we get past all That?
I think that, uh, one of the big challenges in a war room is that everybody brings their own tool to the table and uses it to justify why they are not responsible for the issue that's happening. And yeah, there is the old joke that, you know, it's always the network's fault and, you know, we ought to just start there. Um, I, the, the war rooms are gonna evolve.
I, I, I still think that when, when a major outage occurs, uh, sev one tier one sort of outage occurs that inevitably people, especially executives, are gonna ask what's going on. But rather than interrupting the people who are actually trying to resolve the problem, it would be better if the executives could go to some sort of agent and this becomes this agentic AI that we've been talking about and get their questions answered. Do we, do we know what the root cause of the problem is yet?
What's our estimated time to recovery? Is there a plan to prevent recurrence? These are all things that AI can take that load off of the teams that are working to resolve things.
And then when it comes to preventing recurrence, it's about what can we automate to, to make that happen? And so the, the war room process becomes less of a finger porting and more of a let's bring all the information together using AI to synthesize that and help us to get to root cause faster and then make it the, we we could, we can borrow this from the DevOps teams, right? The, the blameless postmortems that are really all about finding ways to automate things so that they don't happen again.
Yeah. That blameless stuff is important. 'cause otherwise people would just, you know, they'll hide.
Yeah. Right. Whenever there is an incident though, like the first thing anybody goes to look at is the CMDB, right?
They're like, how are these configurations? What's, what went wrong here from, you know, when when's the last time it worked to what changed to now it's broken? Yeah.
What is the relationship between those C MDBs gonna be? And as we automate incidents and we add more ai, is, is the, the way we think about C MDBs, does that need to change? Well, of course when something breaks, the first thing everybody asks is what changed?
And the CMDB should be that place that we go to, to find that information. The problem historically has been that discovery tools are out there, uh, doing an update maybe once a day, if we're lucky. And there's a lot of workload that's put onto the network load, put onto the, the systems and the applications to do that level of discovery on a regular basis.
But what we've done at OpenText with our universal discovery and CM DB technology is to make our, our change, uh, or rather to make our configuration updates every five minutes. And the way that we do that is through differential discovery. So rather than having to repo the entire environment every time we want to do an update, we just look at what's changed across the various systems and applications that, that we're supporting the networks and putting that all into one place.
And I know there's this argument around should we rely on a single source of truth? Should it be federated? That sort of thing.
And we don't really have a, a perspective on that. You cannot, you can do either. We can, you can create the single source of truth or you can rely on multiples with, with federation to bring that information together.
But yes, having a sense of what's changed can dramatically improve your ability to respond to known errors. And the, the, I I would say it's a, a renaissance of CMVB that we're seeing right now because the discovery has gotten so much better. It's never been a problem with where do we store the data?
And, and the CMDB, it's always been about how do we improve discovery so that it's more accurate, more real time, and then yes, the reporting side of it, how do we make it easier to pull that information out? And AI's gonna be a great way to do that too. If I can just use the AI and interface to ask the question, what changed for these configuration items that make up this service?
Uh, then, then I have a much faster way of pulling that information out. Do you think we might get to the point where we can eliminate all these gremlins and gremlin in my world is like, this is an anomaly. It happens intermittently, it's like every two weeks, but I can't tell you exactly when it's gonna happen, but it ha impacts performance.
I should be able to get to the point now where, you know, if you talk to some IT people, they're like, it takes me three weeks to find something that takes me two minutes to fix. Yeah, Yeah. Well, and I think what you're referring to there is what we call predictive ai.
So, uh, the topic of AI has really reemerged as a result of generative ai. And that's been part of what we've been talking about here, is that natural language means of of, of finding out information. And, and I think that's important, but a lot of people are skeptical of that right now.
We had a customer advisory board yesterday, um, and there was a lot of skepticism about what generative AI value can actually bring versus the cost. Uh, so I think as vendors, we're really in a prove it mode right now around that. But predictive AI and causal AI being the third piece of that has actually been around for quite some time.
So when we talk about AI ops, we're really talking about can I find the root cause of something using ai? And then when it comes to predictive ai, it's, it's exactly the scenario that you described. Something happens every two weeks.
And because it's such a a time window, we, we don't normally see those spikes on our, on our screen display. But if, if our predictive AI can notice those sort of repetitive things or notice things that are, are outside of the normal things that, because, 'cause we've, we've been in, in operations, we've been focused on sitting thresholds for so long and maintaining those, and that's its own challenge. But if we could just alert on things that are abnormal, then it would be a another way to check to make sure that we're, we're staying ahead of the problems.
Yeah. It seems like to me, Jenna, I adds value for summarizations and things like that. But to your point, um, gen AI is probabilistic, right?
So it's, it is most likely to give you a good answer, but it is a very deterministic thing. It's gotta be right a hundred percent of the time and done the same way each time. So is that where the gap is between, uh, the awesome capabilities of gen AI and the reality of our day-to-day things we need to manage?
Uh, it's, that's a good way to look at it. Uh, I, I think that probably part of the, part of the holdback, or part of the reluctance of it to adopt these things is that we, we saw this with cloud, right? A lot of it.
And, uh, it operations teams have been resistant traditionally to taking on the latest and greatest. They wanna see somebody else succeed with it first before they're gonna be willing to take a stab at it. Uh, the challenge is, is that I don't know that the complexity of the environments that we live in today and the pace of change really gives us that luxury much anymore.
We're going to have to start taking advantage of these technologies if we, if we hope to do what the CIOs are telling me, which is cut costs. Yeah. If I think about it a little bit though, I agree.
I get your point. Not everybody wants to be the canary in the coal mine. Yeah.
But, um, more folks I talk to are coming around on AI to this point of view. You're kind of like going, I do a lot of stuff every day that is kinda low level tedious and, and monotonous. And after a while it burns them out.
So are we getting to a point soon where maybe a lot of IT folks are saying, well, I don't think this is gonna kill my job, but I don't think I want to do this job without AI because it's kind of tedious. Yeah. Well, I mean, that goes back to the point I was saying earlier is we, we treasure being knowledge holders, but at in IT operations, we have to become knowledge facilitators.
We have to figure out a way to leverage the best of AI and automation that's available to today to offload that tedious work. And maybe even some things that we would like to hold onto because that's where the organization is driving us towards. We, in our, I mentioned our customer advisory board yesterday, uh, one of our customers told me that they had built as a hundred thousand.
They're, they're at a hundred thousand automation, uh, automated processes that they use today. And that's everything from responding to known errors to the way that they apply patches to, uh, and they have this great concept of automate and iterate. So they have a team of people who are out looking for mining for opportunities to automate.
And as we think about AI and then using that to automatically mine for more opportunities to automate, it just becomes a, a cyclical iterative thing that builds upon itself. And that's why they've been able to get to this place of a hundred thousand automated processes. When does security fit in this conversation?
Is that gonna get melted in? I talked to some folks and there's not enough security people. So one of their strategies is move SecOps over to the IT team Yeah.
And see if that was a place to automate. Yeah. Well, security's always been about policy first and understanding, you know, what are the mandates of facing my organization, whether that's from regulations or from, uh, maybe internal policies, things of that nature.
And then they tend to throw that at IT ops to go and execute. Now, uh, they may be dissatisfied or frustrated with the pace that IT ops is doing things. Let's, let's look at vulnerability management.
Uh, 'cause this is probably the, the strongest connection between IT and securities. Um, you know, the, the, the, uh, security guys have great vulnerability scanners. They come up, you know, they're, they're constantly scanning.
They add tickets for every, uh, patch that needs to be applied to close that vulnerability out. And so the ticket queue just gets longer and longer. And IT ops is left with, you know, how, how I can't meet if the policy says I have to correct a critical vulnerability within two weeks, um, and I'm con consistently not hitting that.
That's a pretty good indication that some automation is needed there, but automation in security, it's trust but verify. So yes, I'm going to trust that the IT ops team is going to get that done. But then having a, a feedback loop of verification and being able to look, it's always about identifying the top level in, in this case vulnerabilities, but apply any sort of policy across security.
Being able to prioritize the things and making sure that those priori priorities get rolled out to the IT operations teams and then that they're actually executed upon. And that builds trust between ops and security. That's sorely needed on A lot of the IT ops guys I talk to.
They don't necessarily wanna be the one to apply the patch 'cause they're gonna get yelled at by some developer if it breaks. Right? Well that's, so now we go back to classic change management is who needs to be involved with these things?
And we've all wanted to get out of the, you know, the, the cab meetings that take two hours and 500 people on a phone call, just basically rubber stamping things that, that, that never helped anybody. And you know, we, we, we should be outta that business, especially in the world of cloud and SaaS and the pace of change. Uh, so there's a, and and there was always this sense of, uh, that there's standard changes.
So there's things that people can do that don't need approvals, but yes, we should know that where there's critical applications that can't go down, then who are the people that need to be involved in understanding when we're gonna make that change to make that happen? And so the coordination of that activity, again, is another place that's ripe for automation. And, you know, maybe a whole board doesn't have to be involved in that decision.
Maybe it's, we just do a checklist of, you know, have, have these teams approved the, the window for making this change and do we have a rollback plan? And if so, then the people who need to get the work done can just go ahead and get the work done. Alright, come in full circle.
What's your best advice to folks who are running these IT shops today? You know, what's that one thing you see the most successful IT leaders doing? Well, I, I started to mention this earlier, but CIOs that I talk to are all about costs.
They, they, they see IT ops, and this is not a new story, right? It's, it's CIOs, uh, want to take cost out of operations and ship it to the new things that are gonna make the business more profitable. Uh, so if, if you are in IT operations and you are not thinking about ways to reduce cost, then there's a good chance it's going to come to you as a request to, to reduce cost.
So, uh, and this, this goes back to where AI and automation are gonna have a, a huge role to play. So when we think about whether it's patch automation that we've been talking about, easier ways to find the root cause of problems, um, making sure that we are keeping pace with the changes in the complexity that's happening in the environment. Um, be proactive about how you are finding ways to reduce costs.
We, we don't like to talk about how these things impact our team members, but sometimes it does involve having to reduce people. If you, if you've put on your CIO hats, they're thinking that way. Uh, the ones that I talk to.
So, uh, and, and it doesn't always have to be that obviously we can reassign people to things that are higher priority, but, um, it's gotta be an all of the above strategy as, as we look for how do we take costs out of IT operations. Alright, folks, you heard in here the total cost of it is once again front and center. Hey, Travis, thanks for being on the show.
Great to be here. All right. And we'll be back in a minute folks.