Modernizing a Global Data Center Network Without Downtime: Nokia IT’s Automation-First Journey
Scott Robohn (Solutional) and Tom Hollingsworth (Tech Field Day, The Futurum Group) introduce a multi-part video and blog series detailing how Nokia’s IT team modernized a global, brownfield data center network while keeping mission-critical applications online. They set the stage for the series by outlining the operational pain points that accumulated over years of organic growth: inconsistent designs, heavy manual work, limited rollback, and disconnected tools.
Scott and Tom describe how Nokia IT defined clear, outcome-driven requirements and adopted a NetOps operating model built on Nokia SR Linux and Nokia Event-Driven Automation (EDA). Key architectural decisions included a leaf-spine fabric, API-first operations, infrastructure as code, CI/CD pipelines, and the use of a digital twin to model, test, and validate changes before deployment. The video previews the live migration strategy used to transition workloads without disruption and highlights early results, including an 80% reduction in incidents during the migration pilot. Beyond technology, the discussion emphasizes the importance of data quality, communication, and operational discipline in delivering resilient, automation-driven outcomes.
Watch Nokia’s presentation at Tech Field Day’s Networking Field Day 39 here.
This video is number 1 in a series of 6. To see the other posts, visit: https://techstrong.tv/videos/modernizing-the-data-center-nokia-its-netops-playbook
Transcript
Have you ever been responsible for modernizing a global data center network while keeping critical apps online? Nokia is IT team, to just that? They performed a Brownfield migration from a mixed legacy fabric set up to an automated fabric in multiple data centers composed with Nokia's Sr.
Linux and their event driven automation management system. Ida. I'm Scott Robohn and in this video, Tom Hollingsworth and I will give you an overview becoming blog and video series that breaks all this down step by step.
I'm Tom Hollingsworth Scott and I interviewed the Nokia IT team behind the project and dug into their planning and migration materials. What we're sharing today is how they turn pain points into an automation first operating model and what you can take away from their journey. You'll also hear about the people side of things.
Why data quality communications and ops discipline matter just as much as the tech choices. So let's set the scene. You know, over time, Nokia's network and data center environment grew organically, different pods, different tech stacks, different operational patterns, adding stuff here and there over a long period of time.
And with that came the usual friction, non-uniform designs, too much manual work, limited traceability or rollback and tools that just didn't talk to each other. This all had impacts on the operations of the business. If you lost application heartbeats for just a couple of seconds, you'd have a database go down, taking two hours or more to recover.
And if you disrupted factory operations, you could easily cause a 500 K or million dollar loss per incident. The biggest challenges were with the infrastructure. People became afraid to do the simplest things like adding A-V-L-A-N.
They needed to move to a NetOps deployment model with better tooling and observability. This wasn't a buy some switches situation. They laid out specific requirements that had specific outcomes.
It covered hardware, software services and migration execution across multiple dual data centers. The production fabric requirements included API first operations with zero touch provisioning, programmatic overlays, robust routing protocols, jumbo frames, multicast QOS and dual IPV four IPV six Stack operations On the management side, small failure domains, programmatic VLANs, strong aaa, and tight integration with ticketing monitoring and logging. Okay, so how do you architect for that?
Well, the team leaned into leaf spine clove physical network architecture with a layer three underlay and a programmable overlay using vxlan. They also specified a digital twin requirement to model and test future deployments and pre validate changes and NetOps operations with CI/CD and using that digital twin for dev and test environments. Sr.
Linux and IDA are the heart of this new tool set. They opted for EA via SaaS to keep the infrastructure up and running no matter what happens in the environment. IDA is Kubernetes native.
You treat your network constructs like resources, you keep the network in a desired state and then you extend it with custom apps, think connectivity, diagnostics or related alarms and logs With proactive monitoring and forecast, The team made all these choices to drive programmatic access to the network with a shift to infrastructure as code CI/CD pipelines and superior observability. So now let's talk about the migrations themselves. They used a live migration method with a handoffs from the legacy infrastructure to the FMO SR Linux and the IDA Fabric.
They rehearsed everything in the EA digital twin and executed changes as code. The process went a little something like this. One, build layer two VLAN n extensions between Legacy and the new SR Linux fabric.
Two, make sure that critical loads are dual homed and then swing redundant links in batches. Three, activate the host on SR. Linux, deactivated on the legacy 'cause your gateways are still on.
Legacy. Simulate that whole thing and verify that it all works. Four, move the servers and frames one by one and lastly, cut the gateways and fabric exits over to the new fabric.
And you have quick rollback baked in. Every step in this process had pre-checks, approvals and a clean rollback path. Post migration is where the winds really show up up faster automated implementations, fewer inconsistencies, fewer outages, and measurable cost and time savings.
You want a concrete example? The team saw an 80% reduction in incidents during the initial phase of the migration pilot. That's, that's striking.
80% reduction, that's a big deal. And the human factors investing in high quality network data, keeping communications with the team and motivating strong operations. Operational responsibility does not vanish.
You still own the outcome. All those team human factors came into play. So here's what you'll learn about all this.
The more detail in the coming series. With more detailed videos and posts on interviews with the Nokia IT team, we'll start with the team's pain points and their desired state. We'll dive into their specific requirements.
We'll take a closer look at the target architecture with SR Linux and EA. We'll walk through the live migration method and then we'll wrap up with, uh, speaking to the long-term day two ops and desired outcomes. Be on the lookout for posts on Techstrong.
We're gonna look forward to going through all of this with you. I'm Scott Robohn, I'm Tom Hollingsworth. Thanks for watching.


