Why Reliability Is Now the #1 Decision Criterion for Enterprise AI Infrastructure
Mitch Ashley (Futurum Research, The Futurum Group) and Scott Robohn (Solutional) present key findings from a recent market study report by Futurum Research, in partnership with Nokia, based on a global survey of 100 enterprise IT infrastructure leaders. The discussion highlights a clear and consistent result: reliability has surpassed all other factors as the top decision criterion for modern data center networking.
They unpack why reliability now anchors infrastructure design, operational practices, and business outcomes. Survey data shows that 86% of respondents ranked reliability as a top decision factor, with outages directly tied to service degradation and revenue loss. The conversation explores the primary contributors to incidents, especially human error, and why traditional approaches focused solely on training and process controls fall short. They also examine the current state of automation and AIOps adoption, revealing a gap between interest and operational maturity. The takeaway is clear: resilience is no longer optional, and enterprises must rethink how reliability is engineered and sustained in AI-driven data center environments.
To learn more, please read the full Futurum report here.
Transcript
I am Mitch Ashley of the Futurum Group. I'm Scott Roon with al. We're here today to give you an overview of the Nokia Data Center Fabric Reliability study, a survey that addresses some key issues in modern data center networking.
We ran a future and research survey of a hundred IT infrastructure leaders from large enterprise IT organizations with the goal of understanding how data center network reliability is decided that's delivered and measured both today and into the future. Mitch and I wanna cover three main takeaways in this video. First, reliability is the number one decision criterion.
Second, operational challenges, especially human error, still drive incidents and third teams claim meaningful automation, AI ops, adoption. And we wanna unpack that a little bit. Yeah, three important messages.
Number one though, reliability is not a nice to have. It anchors the de design, the design, the operations, and ultimately in the business outcomes. Resilience is the end game.
Yeah. Not a huge surprise, right? That reliability was the top priority.
Um, there's some interesting supporting stats that go around that. Mitch, can you talk us through 'em? I think first of all, 86% of the respondents ranked reliability as a top decision criterion.
So it wasn't just, uh, just above the midpoint. It was well almost, you know, get, you don't get 86% in, in responses for very many questions. And the things that, that it was sat on top of were the ease of integration operations.
We know those are also challenges. So why does this matter? A single hour, hour downtime is a widely expected to hit service levels and also revenue.
So 47% foresaw a major service disruption risk, a 68 expected direct revenue loss. So it's a big deal. 74% of organizations said they had greater than one incident of an outage in the past 12 months.
So it's not a rare occurrence. Um, when we see that many organizations saying they're having at least one out one outage a year, and that can be due to hardware failures, human error. Those are common top causes.
So now regarding human error, let's talk a little bit about that. We saw that amongst, um, multiple operational challenges, um, that drive those, uh, the drive, the outages and incidents that we're seeing. What, um, more is underneath those statistics, Mitch?
It's a significant factor. I mean, the, the respondents rated it 80% said that human error impacts service 17 point half percent, called it a frequent top cause things like that. So it, it's certainly just more than a factor.
It's an important aspect of when there is an outage, but it's also more than that. Um, there are often other failures that come alongside with human error at some point in that process. That can be things like hardware or software failures.
So how do we address this? When we asked the respondents, 35% said that they emphasize strict process and training. 25% said focus on resilience and recovery.
Recovery and only 12% said they aim to eliminate errors via automation. Meaning we know that errors are gonna happen, but we have to be able to handle those, respond to those we wanna resilient architecture, implementation, and also as well as the implementation or the automation that we're doing. So, you know, teams are struggling to meet the evolving needs of the business because we know those are under constant change and also limit the scope or run extra planning cycles.
Those are things that, that they're struggling with. Oftentimes they'll even postpone important tasks due to confidence levels, uh, when they're not sure if that's something they're ready to implement or if this is the right timing to do that. Last but not least, of course, skills always come up, but it's a significant gap.
54% said that that was skill gap was an issue and several in incited cited that state versus desired monitoring limits were a factor as well. So on the implementation and use of automation in AIOps and the actual adoption, um, versus interest in automation and AIOps adoption, what did you find in the, in that bucket of responses? Well, they, they said that here's what they're using today.
Uh, a 67% said that they're using automated monitoring. 50%, actually 58% said they're using infrastructure's code. I've particularly found that interesting.
And of course that's, uh, you know, followed by things like ticketing, auto failure over, but ai, ML based incident prediction was pretty significant at 54 4%. So I think this says that we're investing in A IML as part of the, the solution set, but also I think we know that, you know, tooling does not necessarily equal positive outcomes. Only 36% reported dedicated AIOps tooling as of now and many are advanced practices are still in the maturing stages.
So, you know, adoption is both planned and underway, but separating tooling use from operational reality gains is still key. Yeah, that separation, uh, and you know, that fine, fine grain understanding of are we just interested in automation and AIOps versus we're really, you know, going full force. We're gonna see that journey continuing I think for years with many enterprises really just getting started in earnest.
Definitely tracks agree with you So much that you covered in, uh, in this survey. We're only touching the tops of the trees here. Where can people go to get the full report and, and plow through this and understand the fuller picture?
com and download the report from there. There's a section for analyst reports and latest analysis that we've done. We'll also include a link with the video to make it easy to go right to the report.
It's free, download it, you've got in seconds. You'll be looking at some really compelling and interesting information. Definitely agree, Mitch.
Thanks for the pointer and for the readout. You bet, Scott. Thank you.