Industrial DevOps – Build Better Systems Faster | DevOps Experience 2023
In today’s rapidly evolving landscape, organizations face the ongoing challenge of building better systems faster to remain competitive. Traditional development and operations (DevOps) practices have proven effective in improving collaboration for software and operations. How can we take the benefits that software has achieved to the system level? The approach requires beginning with tests, leveraging multiple horizons of planning, managing flow, utilizing the entire digital engineering value stream and, most importantly, a focus on people and culture. By adopting Industrial DevOps, organizations can learn faster, accelerate the development of CPS, enhance system reliability, and enable rapid adaptation to changing requirements.
What Will Be Covered:
1. Industrial DevOps Principles
2. Cyber-Physical System Challenges
3. Applying principles to the challenges
Transcript
Hi, I am Robin Yeaman, and I am currently the Executive Space domain lead at Carnegie Mellon, the Software Engineering Institute. Um, I have spent the last 30 years of my life working in large system delivery. So I spent 26 years at Lockheed Martin building everything from submarines to satellites.
And I have a lot of passion for being able to deliver large scale safety critical systems at what we would call the speed of relevance. So I'm gonna tell you a little bit about industrial DevOps. So lemme give you a little background.
Uh, the first thing is, you know, 21st Century development approaches things like Agile and DevOps. They've been around for a bit. Um, however, they're traditionally applied to software and they've hugely benefited these small teams, uh, with, you know, faster schedules, reduced cost, higher quality, and even product or employee morale.
However, what would happen if we took these practices and we applied them to cyber physical systems? What do I mean by that? I mean, what happens when we apply them to a satellite or a submarine or a, uh, aircraft?
Well, my hypothesis is you would get those benefits even larger. And so what I'm gonna tell you about is my experience and an approach to apply Agile and DevOps. And for some of you, uh, you know, if you don't consider security within DevOps, I usually do, uh, we can call it DevSecOps, but how do I get and apply those practices to cyber physical safety critical systems?
Now, the first thing we have to do is begin with why. I'm an Agile analyst and, uh, super fan of DevOps, so I know why, but it's not good enough to just want to transition to Agile or DevOps. We wanna understand what's the so what behind it, and the big, so what is this for?
Simple systems or systems that we've built multiple times and they're very predictable, waterfall could probably do the job. However, where Agile and DevOps thrives is in complicated and complex systems. Now, one could argue that satellites are complicated and complex systems in a highly volatile domain.
So agile would thrive in this arena, and that would be one of the benefits. What does that mean? It means that I have the ability to, to make changes throughout the lifecycle, which is really important because a lot of these systems have a lifecycle, even a development lifecycle of multiple years.
The motivation for agile is typically vuca, right? So volatile, uncertain, complex, ambiguous environments where in the next industrial age, and inherently pretty much everything we build comes into this space. I had the opportunity to attend a AI hackathon, and as we were walking through some of the algorithms, one of the, uh, developers mentioned, you know, an algorithm they used and they said, well, that was, that was 30 days ago.
Um, that's been deprecated. Well, 30 days, that's no time at all, right? How, how volatile is this?
Another example is when I talked to a general recently, and he said that he expects that we're gonna see more change space in the next five years than we have in the last 45. Well, that's huge. So all of these reasons point to why.
So recently published a book that will be coming out in October, and we referred to it as industrial DevOps. And based on experience, we have decided that there's approximately eight, and we could argue even more, but eight principles that help us deliver industrial DevOps, organize around the flow of value, apply multiple horizons of planning design, or, you know, decide based on empirical data we want to architect and for change and speed, and then we really need to make sure we're managing flow. All that being said, we wanna begin with the end in mind, which is going to be a shift left approach, integrate early and often and apply a growth mindset.
I've said a lot of words, uh, let's see what they mean. So I'm going to deep dive into each one of these principles and maybe give you some examples. First thing, organizing around the flow of value.
Now in the, uh, you know, age where we were doing, you know, a lot of new manufacturing cars, et cetera, you know, our goal was to, to uh, really optimize that delivery, but we were in much more of a predictable environment. And so what happens is a lot of these large organizations have very functional stovepipes, right? They've got the program managers together and they've got the systems engineers and they've got the hardware engineers.
Um, however, if your goal is to deliver at the speed of relevance, then we wanna change the way we organize and really focus around that value stream, right? What is the outcome of, let's say, building a satellite? Well, one of the things we have to do is, is bring these teams together, and that's not easy.
One of the, the problems that we see pretty regularly with these large organizations is program managers often talk about things like P M I, they talk about schedule, um, they talk about lean, cool. If you talk to the systems engineers, you're going to hear a lot of things like systems thinking or design thinking. If you talk to hardware, you're going to hear things like rapid prototyping.
And in the software arena, we do hear quite a bit about Agile and DevOps. When we talk to test, they're gonna talk to us about things like shift left and in operations, you're going to hear a lot of things like it, t i l or IT infrastructure library. Now, the key is all of these approaches are designed to optimize the flow of delivery.
However, we're talking past each other, right? Communication follows the org structure. And so the program managers can't quite talk to the systems engineers quite can't talk to the hardware engineers, right?
We're talking past each other. So we wanna organize around valuable outcomes. What's a valuable outcome for a satellite?
Well, one could be guidance, navigation, and control, right? Most things have to perform guidance, navigation, and control. And in order to implement G N C or guidance, navigation and control, I need some program management.
I need some systems engineering. I need potentially some hardware because I've gotta work with the interfaces. I've got definitely quite a bit of software and a lot of tests.
So I'm gonna bring these groups together so that they can actually communicate and deliver those outcomes for the satellite. Now, does that mean that I don't need, uh, education and specialization in these areas? Now, we may have something like a homeroom or we may have, you know, communities of excellence where, you know, I can learn the best things about software engineering, but my day-to-day interactions in order to deliver faster, really has to be around that value stream as opposed to that functional flow.
The next thing we're gonna talk about is applying multiple horizons of planning. You guys might be saying, Hey, what's, what's the difference between that and an integrated master schedule? Well, an integrated master schedule will frequently tell you, um, what I am doing the last week of October on a Thursday in 2027.
And we know that that's probably not true. And so we have to look at that now for a large program, so something like fleet ballistic missile or an aircraft, or even a car, we are talking about a multi-year project. Now, what's the difference is that I want to have regular intervals while I'm getting feedback.
Now, here's an example that I'm showing you, and you can see that my multi-year roadmap decomposes down into an annual roadmap, right? So I'm gonna review things like that, and I've got feedback loops and demonstrations. That annual plan that decomposes into a quarterly plan.
Again, feedback loops, demonstrations, and then at the lowest level, if I'm following a scrum approach or even just a time boxed approach, I may have one or two week iterations. They're also providing feedback and demonstrations. Now, if I am looking at my sprint plan, and for easy math, I complete eight out of 10 things, right?
If I do that a couple of sprints in a row, because in typically there's about six sprints in a quarter, then it's likely that I'm only gonna be 80% complete. And if I have lead times and other suppliers that I'm working with, that may not be good enough. So what I have to do is at the close of each of those sprints, I need to evaluate my plan and further inform the next horizon, right?
So maybe after one sprint, I wouldn't change, let's say my quarterly plan, but certainly after two, I'm gonna start investigating and communicating with people that actually have to integrate with me to tell them, Hey, we're these things aren't gonna be done. Um, or I have to make a, an adjustment. If my quarterly plan is off again, just like my sprint plan, then it's likely at the end of the year I'm gonna to be off.
And when I'm doing things like I have to order hardware to come to a particular facility at a particular time, that could have a huge impact. So I don't wanna just plant plan sprint to sprint. I don't want an integrated master schedule that's got so much detail that it tells you what I'm gonna do in Thursday in the last week of October in 2027.
But I do want a multiple horizon plan where in each of these horizons, I'm evaluating empirical data and I'm using that to communicate and further inform my next planning horizon. Next thing we wanna do is decide based on empirical evidence, decide based on data. And many of you may consider that we do this already, but what we typically see, and what I typically see is a phase gate approach.
Phase gate is something and, and a waterfall, um, approach is one of the phase gates. A phase gate approach says, okay, I completed all of my designs and I'm done, and then I've completed all of my, you know, software. I've completed all of my tests, I do one at a time.
The problem with doing things like this is I'm not actually validating any of that work, which means I could be creating a lot of rework because I could be wrong. Um, I don't know about you, but frequently that system doesn't just magically come together at the end. So how would I change that?
Well, let's decompose the system into smaller components, smaller areas, and then get real feedback validated, observable feedback on capabilities that have already been integrated. Now, for here, you can see that I have an example from Lockheed Martin, fleet ballistic missile. And these guys, um, have quite a bit of hardware associated with their, so they have an initial breadboard, and they have been able to validate the work that they've done so far to be able to ensure that, you know, they, they meet the criteria, they meet the non-functional requirements, they meet the goals and objectives.
It's different than validating paper because I can actually touch it and feel it, which buys down risk, especially for these large systems. So while I'm not gonna deliver a plane every two weeks, I do want to buy down my risk exposure and minimize rework every two weeks. So that's what this does for you.
Next thing we have to do is really look at the architecture. This is a big deal. Most legacy systems that I, uh, have had the chance to see, and I've seen quite a few, um, are highly coupled.
There's a lot of dependencies. They're, they're what we would refer to as monoliths, right? We've created these great big systems.
Um, however, to make a change to that system, I have to make a change to the entire system. I have to retest the entire system. If I make a change to, let's say an aircraft, um, that is a monolith and tightly coupled, I have to validate every single thing to get flight worthiness certification again, why?
Because everything touches everything. If I decompose it and I create modularity, and I'm not saying everything has to be a microservice, right? Context matters.
We have to understand that as we decompose things, we may be adding a level of complexity too. But if I create a modular architecture with standardized interfaces, for example, in this autonomous vehicle, I may be able to make an update to the camera lens without having to test the entire vehicle, right? That module come in, module will come out, it's a lot faster to update, it's a lot faster to test, it's a lot faster to, to get into operations.
Um, it does require a little bit of a different architecture approach on the front end, right? We have to think about it a little bit more. Um, however, you know, over and over again, this has proved to be very valuable.
Uh, and it's not new. If you were talking to some of the hardware engineers, they'd be like, oh yeah, I've done, I've done product line engineering. So not all of these concepts are new.
We're basically taking the best tools in our toolbox and applying them with Agile and DevOps to actually instantiate a better system faster. We want to look at the flow of delivery. And again, this is not new concept.
We are talking about the theory of constraints, and we're talking about systems thinking. Now, as much as we know this, it's still a problem we still have, and it's because we have very functional organizations, we have very large batch sizes. Most people do not limit work in progress because we're pretty sure we can do five things at one time and we do not have a common cadence.
Why is that a problem? The bigger the batch size, the larger amount of rework, the bigger the batch size, the longer it takes to go through the system. Um, if I don't have a common cadence, I can have parts of the system get complete, but I can't validate the entire system.
The other thing is, as Don Reson will tell you, when we have these large systems that are in or development, we want to exploit good variability and reduce bad variability. Well, if we have a lot of variability in the system, we have a lot of noise, it's hard to do that in manufacturing. We want all variability gone, but product development, we don't want all variability gone because there's optimizations, there's things we can get that are better, but if I have everybody on a different time cadence, if I'm developing in a different speed, if I'm not actually doing integration, I have a lot of noise in the system, which makes it very difficult for me to find those opportunities that that good variants could, could buy.
So really looking at the flow of delivery is huge. A lot of software engineering teams have have heard about shift left, right? Test-driven development test-driven development's typically done at the the unit test level.
This is a little bit different. Let's begin with the end in mind for the entire system. Let's start with how do I validate the system at the beginning?
Now, you might think that this happens already, but given that a lot of large system developers follow a waterfall approach, many times they don't even purchase the test equipment until later in the lifecycle, which means we have a whole lot of work in progress that's never been validated. Now, in the past, that's been very expensive to make the case on why you should shift left, why you should buy your hardware in the loop earlier, things like that. And maybe the most just expensive exquisite systems would have that.
However, we've seen that a lot of these costs have gone down and we can take that physical system and move it into cyberspace. So what I'm showing here is what I would refer to a digital twin. One thing that you have to remember about a digital twin, um, is there are multiple, a whole range of fidelity, right?
There's not one type of visual twin. So if I have a system that doesn't have a lot of safety criticality, maybe I've just got portions of that. Um, if I have a system that does impact human safety, I have much higher fidelity.
Now I know I have a digital twin. If I'm taking that physical system and I'm connecting it to the cyber system, so I'm using data from the sensors and I'm feeding that into the digital twin, which allows me to test a whole variety of behavior for large safety critical cyber physical systems. We often will see emergent behavior.
Emergent behavior is something that, you know, we need to understand as soon as possible so that we can remove that, right? We, we've got safety problems if that's the case. So leveraging a digital twin in the beginning, um, and what we would refer to that is a either a digital shadow or a digital prototype, right?
Again, it may not be connected to the end system yet, or maybe it's just connected to a test layer. And then over time, growing that fidelity. And every time I make a change to the system, validating that in cyberspace, is it the same as doing airworthiness flight test?
No, but I bought down risk. Again, I'm reducing the amount of times I've gotta go backwards in the value stream, which allows me to go faster, and it allows me to increase quality. This is another problem.
Another area we have to integrate early and often, and this is difficult because if I'm in my functional stove pipe, I'm gonna tell you that I go a lot faster if I don't have to integrate with anybody. And it's true for that one particular slice of time. You may be able to get through something faster, but the system doesn't get through it faster.
All right? So where you're saving your ti time and schedule, is that, that big bang integration at the end, right? So if I purchase the hardware, if I have hardware in the loop software, in the loop modeling and digital twins, I can start integrating very early in the life cycle, right?
So here you can see again, I created kind of value stream based teams instead of functional teams. So you don't see that I've got a software team or a systems team, but I do have an avionics team that's focused mostly on the, the avionics hardware and software. I've got an environmental control team, uh, thermal and maybe orbital structures.
Obviously you can see here I'm, I'm looking at a space launch vehicle and I have different integration points. It does not mean that every team can integrate at every integration point, but the more frequently they can integrate, the more likely I can get that system out and available, um, in a much shorter time. Traditionally, at least within what we would call traditional space, it would take 90 months to build, deliver, and deploy a satellite.
And that's just too long. Now, a lot of our, our new innovators, let's say SpaceX, relativity, uh, planet Labs, they, they've debunked that and, and so really forced people that were in traditional space to, to look at how we build systems and do it faster, but integrating early and often is critical to make that happen. The last principle I'm gonna talk about is applying a growth mindset.
And you've probably heard about this before, but what does that mean? Well, it means that I'm gonna continuously learn what I knew before. I may not know, you know, coming back to that example I told you in the beginning, which was, Hey, you know, I was, I was working with the, the AI hackathon and being told that the algorithm that was only 30 days old was already deprecated.
That was beyond my mental model. That's not how I think. I'm like, well, it's just gonna take me 30 days to even learn it.
Well, apparently that's not gonna work. Um, another good example is for most of my career, I have always been telling people we have to have feature-based teams. And right around 2018, team topologies came out and they talked about cognitive load.
And maybe there's more than just one type of team, and even me who is, you know, head over heels, you know, a fan of DevOps and agile approaches was like, whoa, whoa, whoa. I know what I know. And then I had to remember that I knew what I knew, right?
But that we keep learning and we keep getting better ways. So it's kind of fun to do the impossible, huge fan of Disney. All right?
So we want to apply a growth mindset, and sometimes we wanna look at ourselves and say, am am I, am I thinking about this? Am I thinking about it the way it was last year? My previous mental model?
Has something changed? Do I have a new idea that I can bring to the table? Couple of examples of companies who have done this.
Fantastic. Here you're looking at Alton, and you can see, you know, you can, you can pick whichever cycle you want, but you can see they're on three week iterations and they're doing software, mission engineering and electrical engineering, and they're showing empirical data and validation at each one of these. So they've got their plywood prototype and they've done some hydraulics.
Now I've got my near, you know, flight ready and verification, huge benefits, and they're able to get capability out much quicker than your, your classic, you know, uh, manufacturing company, Bosch, A lot of people think of Bosch when they think of innovation. I always think of Bosch. Uh, Bosch does a lot of that.
Uh, virtual prototyping, digital prototyping, digital twins to make changes and updates to a whole array of their systems. And this, again, allows them to validate that work early and try things out. So another one of my favorites, so Planet Labs, they're more of an earth observation company.
Um, but they launch satellites every three to four months, right? Not the 90 months, I was just telling you. Um, they make changes to their delivery process all of the time.
Uh, they are definitely modular, even to the point of their instantiation. And so Planet Labs has multiple cube SATs in space, you know, hundreds. And they're even able to, you know, allow a commodity based satellite, right?
So commodity based parks not custom and, and maybe let one of those satellites die and still have the same outcome and images they had because they've got basically a whole array of satellites. So a lot of redundancy. Um, so there's, you know, huge, huge benefits with things like these approaches.
But I wanna tell you about Agile and DevOps, and especially industrial DevOps is I want you to use all the tools in your toolbox. Don't think of just the software or just cyber or just requirements or, or just anything. It begins with a need and delivers value to the end user.
I'm typically within the government space. So I've got the war fighter here, but it could be any user. And you know, I've gotta look at optimization through planning.
How do I do requirements, right? Um, how do I integrate that with the models? How do make, make sure that those models inform my digital twin?
How do I ensure that I've got intuitive interfaces? How am I leveraging artificial intelligence and machine learning to find patterns faster than I ever could? How am I pushing the bounds of edge computing so that I can get faster processing power?
And how am I ensuring that I've built cyber into the whole system, right? So I would refer to this as the digital engineering value stream. And I think that we have to look at systems across all of these areas.
And one of the things that we do have a tendency to do is specialize and specialize and create stovepipes. You know, DevOps is a fantastic example. We were supposed to be bringing the development teams and operations together, but how many of you have heard about the DevOps team?
We created yet another silo, which slows us down. It doesn't make us go faster. So the goal is to learn faster and to do that through validated delivery of value to our stakeholders.
The last thing I wanna offer you is an excerpt to the book. All right? So this is Industrial DevOps.
IT Revolution is publishing this book. It will be coming out October 11th. Um, however, if you are at does or the DevOps Enterprise Summit in Vegas, we will be providing a number of these books, um, for free ahead of time.
But if you wanna just take a look at the first chapter, here you go. Now you're not here to ask questions, but I will tell you that you can certainly reach out anytime I'm on LinkedIn and a variety of other places and ask me questions. I'll respond.
You can ask anybody. I have a lot of passion for this topic.





