Modern Governance and Operationalization – Techstrong Con 2023
In this session, John Willis will explore this new concept of modern governance and how operationalization, specifically operational definitions, applies.
Transcript
Hello everybody. Yes, John Willis also known as bakugaloop. Modest presentations called modern governance and operas rationalization just on the right there.
The operationalization part is some work. I've been doing on researching Dr. Deming for about 10 years now which turned into a book so it's gonna be I'm told by my publisher it Revolution that will be available for pre-order in about a week or two on Amazon.
Stay tuned follow me. I'm about to gloop and learn more. It's pretty much something.
I've been focused on for 10 years. So I want to combine sort of what we're talking about with the things we talk about in sort of cloud native risk and SecOps and and tie that to some sort of hundred year old Concepts. So as I said, I'm about glue.
Most of you probably here have seen me or we've talked or you know, who I am. But I've written a number of books about 12 books over the last 40 years. I've done about 11 or 12 startups.
Over the years a couple of things that have note like I was the co-authorax handbook the two books in the middle really describe modern governance. I've done a lot of presentations where I cover a lot more about the concept of modern governance, but those are two books. Once the reference architecture actually one's a novel based on some sort of a phoenix project and then there's my damning book currently.
I work for a company called costly Oslo and but I've done I've sold company to Dell darker. I was wondering earliest people involved in Chef. I spent a couple years and executive team at red hat so So what is modern governance just at a high level?
Like I said, I've got other presentations where it covers a lot more detail. But it basically we like to think of it as a higher level system design, you know, if we think about like the way we think about design and Cloud native and sort of modern infrastructure and all those things that we talk about like post devops post devops. This is really sort of fits as an overlay in that world.
It's autonomous. We really relying on moving the human where possible we never completely removed you it probably the primary goal of this sort of working groups that that have been involved in this is to reduce tutorial and increase the efficacy of how we mitigate and monitor and reduce risk. And so therefore we want to reduce and mitigate real real risk not basically service now record or so internal audit where everybody's handshaking but we don't have real data.
So really removing humans from the gating process. Gaps, you know make that automated and removing from this sort of evidence. You have to stational data.
So this like I said all comes with a bunch of working groups primarily from Gene Kim's it Revolution starting in 2015, 2018. We wrote a deer auditor letter. We did a reference architecture which then turned out to be a best-selling book right now about it paying that fails an audit And we gotta let you either find the book read it or because he's my other presentations, but I really wanted to focus on this idea of operationalization as it ties to what we're doing today with devops risk.
So the guy over a hundred years ago called Percy Williams Bridgeman and he wrote a book called the logic of modern physics. So he's a physicist. He was very frustrating about at that time in the 20s 1920s about how we did measurements and and I'll talk a little more about that.
But here's this quote. We do not know the meaning of a concept unless we have a method of measuring meant for it. So hold that back ahead.
And so Percy Bridgman was a Harvard Professor. He wanted a mobile price for physics in high pressures. The thing probably got a most interesting in rethinking this concept of operationalization or how we do measurement was he was basically working on creating synthetic diamonds and his gauges get breaking just the pressures was so high whatever Gage he had didn't work and he was fascinated intrigued by Albert Einstein appear his at the time special Theory relativity where Einstein used measurement used in time and space.
So basically he was doing measurement in sort of a non-conventional means and and that get Richmond thinking about like for example in this book. He talks about the length. He doesn't use a street but like the length of a street might be in feet or yards or me.
There's or versus the length of a planet which might be light years, right? So you can start seeing that we've all show you have this. already ambiguity in very simple examples that children learn, you know in elementary school.
And so Deming who was a you know, haven't admirer and researcher of Williams work Deming was a physicist himself turned into sort of a management scholar and theorist and he said that and so out of the operationalization. Definitions came this idea about operational definition and so Deming, it's been a lot of work describing it in his two major books one his last book called new economics. He said an operational definition is a procedure agreed upon for translation of a concept into measurement of some kind right.
So basically they're saying that you can't just say there's a measurement you have to have some operational definition of what the measurement is. And this is the one that like probably where people's head start hurting a little bit Deming also goes on to say in his new economics. There is no true value of any characteristic state or condition that is defined in terms of measurement or observation.
Now that one probably hurts a little bit right because what the heck is he talking about? Well, let's walk through some exercise that Deming had shared in his book. So the first question I'll ask you is Count me count the number of horses elephants and pigs from these animal crackers.
I'll give you a second. Yeah, the answer is you can't right because we don't there's a bunch of like what you might find in any animal cracker package or box is a bunch of broken crackers. So we can't tell if that's the foot of the elephant if it's a you know, so we don't we don't have a clear definition of the criteria what we're asking I got better example.
This should be simpler right count the number of people in a restaurant sure piece of case now, I'm not gonna make you do that, but let me ask you some questions about that question. Do we count this woman here? Is she leaving the restaurant?
Is she just cutting through the mall to get out, you know, kind of come around the corner. How about the woman behind her she just looking into the restaurant from maybe the mall or the street to see if it's a restaurant that they might she might want to. Come into today, maybe tomorrow maybe next week.
How about the waiter or waitresses? Do we count them? Do we count the kitchen staff?
These real question. Why are we counting in the first place? Right.
So this gives you some insight to why operational definitions are important right? Because if we're counting for Logistic restaurant Logistics, like food ordering food drinks and a logistics sense. Then we're counting a certain we're counting.
Basically the people are sitting at tables if we're counting for. For fire code then we probably want to count everybody including the woman that's working on the side there the kitchen staff if it's a restaurant on a cruise ship we might be counting for weight distribution, right? So the point is the operational definition is very important when we're asking questions about measurement.
So this is a long-winded thing. But this also Percy Williams goes on to say that the final length of an object. We have to perform certain physical operations.
The concept of a length is therefore fixed. When the operation by which the length is measured or fixed, sorry. So therefore the concept of length involves as much.
As and nothing more than a set of operations by which link that's fine. It's a tough one. But the point is it is Clarity that it is really the operational definition that declares what we're trying to Define.
So in other words, so in portugaloop terms, you know in other words words matter. And and a lot of times I think when we're thinking about the things we commonly do in it and devops SecOps and I would say in what we'll call in modern governance. We have to be a little more cautious in terms of how we use a word.
We use words like lead time mean time to repair. So we use this word time quite clear often and then I'll give you some examples of like what do we really mean by time? You know root, you know if you filed any devops or SecOps conversations like root cause analysis.
Root cause what is a root Root Root is from a plant humans are not plants right zero defects. Or really on zero trust right now. What is zero?
Is there you know, is there really a such thing as zero in it? Her you know, you'll hear some like SLSA. We'll talk about hermetically.
So, you know as an attestation for hermetically sealed deployments. I mean, I mean I can be a little silly here. But what is a hermetically sealed infrastructure like that means nobody can log in there's no access.
There's no jump boxes. There is no like so again, I think I'm not saying that there's anything wrong with root trust or I mean, I'm sorry, you know, well root cause analysis I might say to something wrong with but but zero trust, you know, I think so trust is a great sort of domain way of thinking the question is what do we mean by zero? And that's what operational definitions about and then service, you know service level agreements.
What is the service? Hopefully you get the point by now. So Deming had introduced in his last book this idea of a system profound knowledge.
I cover a list in quite and shortcut by sopk. And basically it's scientific method. It basically is scientific thinking and goes back all the way into the enlightenment.
For instance bacon. We can see it through epistemology. And and then, you know not to get too heady.
But this is in my book. A lot of what Deming picked up was from a guy named Walter schuet Walter shooting Deming were big fans of the pragmatism movement the the philosophy of pregnancies and which was really a sort of first American creative philosophy the and so so a lot of this goes all the way back to this and like I said, I spend a fair amount of time going into the the craters of pragmatism. But anyway coming back to the the point here most people know this this sort of model is playing do study act.
There's all sorts of variance of this was, you know, people talk about utilities cybernetics all those things. It's really sort of the same thing. It's it's a sort of a cycle of learning right?
And so when we tie this to operational definitions Deming said that remember we talked about early, I'll go back here a few slides. Sorry. I should just put the site in here.
There's no true value of any character state or condition that is defined in terms of measurement or observation. So what he's saying here is you really have to come up with. with first and foremost a criteria So if you're going to measure something you need a standard against what the evaluate the test it provides a human judgment.
So this would be the basically what I would call the plan portion of of the pdsa and and then for the theory of for system were found knowledge, which is made up of sort of Four quadrants basically first the theory of knowledge. How do we know what we think we know so this is an element of profound Mouse the second criteria for for creating an operational definition is the test. And this would align very much with the do phase you plan.
Then you do. Well in this case. You know, how is Clarity determined who performs the test?
How is the test performed? This allows it to do like I said, it also aligns with a second. Piece of the profound knowledge lens which is called Theory variation and I'll cover like a little bit statistical control charts or statistical analytical statistics a little bit.
I know this seems like a lot but I hopefully if you sort of think about it and maybe watch this again and slow down or certainly reach out to me and ask questions. And the third piece that Deming would Define for operational definition is the decision. Right, so you have the criteria you have the test and then you have decision and this aligns with the study and act portion of the plan do study Act.
So whether or not object or material met the condition. Right so that the test results are used determine if the characteristics meets the criteria so your criteria in matches the characteristics of the test, if you will again scientific method and this aligns very well with the the study and act as I said, so so the point being that that there's no true measurement unless you clearly can Define the operational definition Deming says that in order to find an operational definition. You need to criteria the test the decision And so if I take this now to things like the commonly used metrics that we Define from like Dora like lead time and time to restore right?
So when we talk about time here, what is the criteria? What is the test? What is the decision?
If we're just going to use these as industry standards like lead time and you know people who do this lead time less than a day or high performers or or medium performers or low performers, right? What what do we mean by time and like you there's the point that Deming would not object to Dora like you'd be fascinated by Dora data and you know data was a very important of his, you know, his thought process but he'd be very critical of like, how are we defining? And so if we look at lead time or TTR, you know, what the first question even back to the right like you using the restaurant analogy.
You know, why are we asking you know, what is good or bad? What is the standard? What are we judging again?
It's just an average the average of all the times per day per week per month. Do we worry about the floor of averages? Right, the f a l a w floor of averages are our teams aligned on what the actual criteria is is, you know is the you know is the teams within an organization or the divisions with an organization.
So if we're aggregating up the lead time for the organization or a division or the whole organization are we actually using the same? Definition for the time aspect of for a toner store. So on the test, right?
What's the what is the measurement? So like for lead time are we measuring the delivery or deployment because continuous delivery is different in case deployment right? We know that are we measuring from when the commit so if we are so like the so most people tell me well John the canonical definition of lead times from commit to production like, okay.
All right. So is that when it hits the branch the master Branch or is it on the pull request? What is production does is it on a dark launch is it when the future flag is when the future Flags turned on and then again, I'll go back to are we measuring mean the mode the standard deviation?
You know, what? Are we? You know, what are we asking what's good or bad and the decision?
It didn't meet the criteria. What actions do we take? Are we continuously improving the pdsa the other thing when we talk about mttr, right?
That's fraught. With examples of poor operational definition I mean lead time is at least people can say from commit to to production. I would challenge that that could be if one team thinks it's You know for any time it hits the first commit or other team thinks it's a pull request.
You're gonna have a skew and what time matters and in but if we get into mtcr boy that we can that's clearly for it like what first off Mean Time right like just throw that out the window because you know, like the average of an outage that is five minutes versus when it's three days the you know, the floor of averages right? But the the bigger question is when does it start in most organizations? I've been around it's sort of napkin based right?
Like somebody says, you know, I think you started 305. Yeah. Okay.
When do we fix it? When was it fixed? How are you actually measuring fix like some of these problems and complex systems manifest for days before it actually happens the the fixing it.
You know, we don't need you know, there's some examples of like massive outages that are sort of partially fixed throttling. Anyway, I can go on and on the whole presentation on that so If I didn't so the other thing I think is really important is there's important discussion we have about how you interpret the data the decision and one of the things Deming was a big fan of this idea of analytical statistics mostly portrayed in something called statistical process control. Now, I've done a lot of work on this very simply it's a model where it will take the data in sequence.
and then generally map the standard deviation across so it sort of mitigates in a certain sense the the floor of averages because instead of just doing meantime, it's using Standard deviation, but really what it does is it measures the data in between a Six Sigma sixth standard deviations three above the mean three below. I notice getting a little heady, but there's a ton of ton of really cool data and stuff about this can give you far more accuracy of measuring standards and what data is measured against the key point is so it's a couple key points one key point is this These type of tools have been around for 100 years. At least a hundred years.
And more importantly they're used for creating anywhere from toasters to nuclear power plants. Right, you know and and we don't use this technology or these type of statistical analytical statistics very rarely. So I see is in it very powerful stuff.
So I have a little bit time. I can show you some examples if we take this idea of system profound knowledge for risk. I took three risk controls that you might measure when you're actually managing risk.
So let the the first is the container security scanning and so like basically the the city at the station that we might be checking is all dependency to build satisfy anyone vulnerabilities. It's a pass or fail or we might look at unit tests. Like that a percentage of coverage for unit tests.
and then last maybe secret leakage, so let's take the risk control for security, so So one of the things that Deming talked a lot about is this notion of numerated versus analytical statistics. So if you look at the left what it this is is a 17 week run of of basically container scans. So container skin failures, right?
Okay, and so in the left, there's a histogram right? Which okay gives us some information tells us a frequency chart, but in right we're in a statistical house controls right now. I'll tell you if you know how to read this apart control and if you follow my blogs and you you want to reach out to me and learn more about this that actually shows that that's in a process data.
So right off the bat you would notice if you didn't know sensible process control, but the chart on the right which is an analytical statistic versus the chart on the left, which is enumerated. The turn right is actually telling me might not tell you until you learn what's took apart, but the randomness, you know above the mean tells me that there's basically what they call common cause variation. But here's the thing.
So if I ran it for the next eight weeks and I look at the histogram on the Left Right the histam really again doesn't really tell me so this is a numerated statistics on the Left Right the frequency. I noticed that the frequency in week 24 is very high. And but but if I look to the right, it's telling me a whole lot of information.
It's telling me right about at week 17 we have this trend in fact. It seems that we're in process where like the upper control limit UCL over there is basically the three standard deviations above. And but the point is we're going up so something's not right the Federal in fact.
If we got to week 22 would be like, hey, actually, there's no anomaly. There's no red. 7 percentile, right?
And that's really bad like that. And so the other thing really cool about animal statistics is Demi would say that the statistics job is not to solve the problem. It's to give the subject matter expert the data to go find the problem.
So in this exact example what happened they added a new team to to that that literally wasn't following the procedures for pulling in container repo images. So we had a lot more scan failures until they were able to correct but like nothing on the left would tell you that but it still Park control chart would tell you just and and there's a lot more there's a whole Opera 100 years operational science behind this and if we look real quick as I wind down here if I want to look at sort of unit unit test coverage. So this is another control chart of 25 weeks of basically unit test and and so one of the things we see here is there's no anomalies everything.
In fact, it's pretty tight variance between sort of 52 and 40 like say 47 if you will and that's great. But here's the other thing just because it looks normal or what you call is common cause or in process variation. Doesn't mean that's exactly what you want.
And let's say in this example like everything looks normal, but we don't we think we should have better test coverage in general So we then create this sort of pdsa this Theory basically it's you know, predictive analysis. I'm going to predict that if I do a couple of things differently, I might get a better results. So actually what we in this particular case, the organization Wanted to sort of come up with this.
Like let's do something in this case. They decided to do two sprints. So if you see starting in week, You know basically about week 17 to week 19, they basically did sort of non-functional Sprints where they focused on training.
Maybe they brought in in this case. They brought in a training a tdd, you know some you know, they just sort of wanted to revamp all the developers to basically, you know, maybe what they learned about tdd was old maybe some people never learned the classic training and so they bought in some Trainers for two weeks. They really focused on tdd and what we see there is the plan cycle was yes, you know, we made this plan here and actually it dip for a couple weeks because you know, basically we expected that we're you know, we're not gonna focus on we're gonna focus on sort of non-functional in this case learning, you know learning how to do tdd better.
So, but then what we find is seems to be increasing Which is maybe we're getting the results are hypothesis is that if we if we basically focus a couple, you know a Sprint. Let's say two weeks on improving how we learn and do test driven development. Possibly will increase and so if we look at sort of the first 17 or 18 weeks and we can go back and we see that that the average is from about 47 to 53.
Which is good but we really want to increase the successes that the in the Texas control gate of maybe we set the control grade at 80% test coverage and we're finding that like 50 on average 50 percent actually average is 50% is that we're sort of where it is the I'm not 50% I'm sorry 50 is the average the mean that we're getting. what if we could try to hypothesize that if we did a couple of things this can improve maybe the average goes up and so we look and look what happened in the next 25 weeks. Basically, the average is now at about 60 65.
And now with some Randomness. Anyway, I know this is a lot but I wanted to at least open up in give the opportunity to understand a couple of things that there's a new way to think about governance. We're calling modern governance is a couple of books written on this that there's some hundred year old.
methodologies from you know from from different parts of statistical analysis from physicists from Deming from shuart that we can use to improve the way we work. And then there's the act and then I'm looking a lot at the state of devops report and data. As a mechanism for doing this and there's a lot tons of data there.
I'm working some other organizations of collecting data and really try to grind it through these hundred year old processes of you know, operational definitions statistical control analysis. Just bringing back the old Pareto reports fishbone really good stuff. So anyway, thank you so much again.
I'm John Willis. com. com pots glue and then these two these two these are two blog posts that really go into detail about devops and operationalization and sort of more detail about enumerated and analytical statistics.
So I'm Anyway, I want to thank you so much.





