3 Metrics to Gamify DevSecOps Transformation | DevOps Connect: DevSecOps 2023
You may be using the percentage of applications being scanned or the mean time to remediate to track the effectiveness of your application security program, but both of those metrics drive unintended behaviors. Maybe you’ve dabbled with the DORA metrics. This talk will explain why those are bad and provide two alternatives that avoid the pitfalls of those metrics as well as introduce a third metric that is even more important for the success of your DevSecOps cultural transformation.
Transcript
Hi, I'm Larry Mat and I'm here today to talk to you about three metrics you can use to gamify a cultural transformation at your organization towards DevSecOps. And, and the very next slide will explain what I mean by DevSecOps, um, which is, which is, uh, aligned pretty well with, with most of the other speakers today. But it, I, I sort of take a slightly different flavor of it.
So the flavor that none of the speakers today are going to, to, to talk to you about, but, uh, but the industry has sort of drifted towards, unfortunately, is this idea that you can slap DevOps lipstick on a traditional security pig and call it DevSecOps. And, and, and that doesn't work. And that's not what I mean.
What I mean by DevSecOps is really about empowering engineering teams to take ownership of the security of the products that they're, they're building. And so that's sort of the top line of it. The way you do it in the middle here is also critical.
And, and that is you have to use the three ways of DevOps, or you're not doing DevSecOps. And then basically my point is, is that you're not doing DevSecOps if you aren't also doing DevOps. As a matter of fact, you could even drop the SEC because the way DevOps is defined, if you're doing DevOps right, you're doing the SEC part.
Um, of course we're here to talk about the security aspect of it. So Des DevSecOps is, is appropriate. And then the bottom line, which is, which is, which is something you can never forget, is that you, you have to never forget that the reason we do software development, that we invest in developers, that we, we, we, we, we do the process of, of actually producing software is to deliver value to customers or users in the form of working software.
And so that's paramount. That is the underlying, most important thing. That's why I don't like SEC DevOps, um, because you can't really, you know, do security.
It's not worth it to do security if you have no product to secure. And so, so the development really does need to come first. So anyway, so that's my definition.
Uh, a little bit about myself before we get going. Uh, this is gonna be a number of logos that are gonna pop up on the screen here. I'm not gonna talk to each of them, but hopefully some of them sort of connect with you.
There's two things I think you, you really should know about me for this talk. One is that I was the head of application security at Comcast for five years, and over the course of that period, I transitioned the entire program from a traditional ABSEC program to a developer-centric ABSEC program. Oh, and by the way, I mean, my definition here of DevSecOps, uh, developer-centric is sort of the, the sort of the way I like to refer to it, cuz I think it, it doesn't get confused and doesn't have this baggage.
But I mean, DevSecOps or shift left or now the term that, that my company now I'm working for contrast is, is shift smart. And so, so what all of that means roughly the same thing to me. Um, so at Comcast, I led the AppSec program, transitioned it to this developer-centric shift, smart way of doing application security over the course of five years, 600 different development teams, 10,000 developers involved.
So I did it at scale and I used metrics from day one to essentially drive the entire program. And so, so that's, that's the sort of the connection to, to today. The other thing, the second thing I I think you should know about me for this talk is that I'm an active developer.
I write code every day. Um, one of my open source projects that I, I manage, I'm the, the primary contributor of, and I lead the development team. All of the other contributors, uh, gets a million downloads a month.
So I know what I'm talking about. And I'm not just preaching, I'm practicing what I'm preaching. And that's the way we run the whole program for all those, those, uh, the, all those open source projects, the dozen or so that I am the primary author of.
So there are three primary reasons for using metrics. Uh, and, and this, this is, uh, I'm gonna talk about it in terms of DevSecOps, but I really do think these three reasons, at least in some flavor work, no matter when you're using metrics, the first is to gamify improvement, you know, to basically motivate improvement. And that's the title today.
And that's probably gonna be what we're gonna focus the most on. But I I, I would be remiss if I didn't tell you these two other purposes. The, the second purpose is, is to use metrics and the visualization of data in particular.
And I'm gonna talk a lot about how you visualize the metrics, and I think that's as mu as important, if not more important than the actual metric you choose to build sponsorship. And, and, and the, the, the idea here is that you can be completely convinced that shift left or shift smart or DevSecOps or applic developer-centric application query is the way to go. But unless you can convince other folks, you're not likely to get the funding to actually pull it all off.
And, and I had to basically show the metrics every year to get the funding increased to expand the program the next year when I was at Comcast. And if I didn't have these metrics, I would and, and be able to visualize 'em in a way that in a few seconds makes it clear the impact I wouldn't have been able to to to, to get that sponsorship. And then the third way is this feedback for the program itself.
I mean, a DevSecOps program is essentially a list of practices you want development teams to adopt. And how do you know the practices you're pushing in, in the particular flavor of the practice you're pushing is effective? Well, unless you had feedback that sh that showed that the adoption of that practice highly correlated with good outcomes, like risk reduction in production is the only outcome that really matters, it, then you're gonna have a hard time, um, actually deciding.
And, and so my talk at the R S A conference in 2020 was the impact of DevSecOps quantified, where I showed that that, you know, two out of the 10 things that we were telling development teams to do were actually net negative value for the company. And if we didn't have the data, we wouldn't have been able to do that, show that. And then of course we changed the program based on that feedback based on that correlation.
So those are the three main purposes. Um, so let's dig right into the three metrics. Now, the first metric I is practice maturity.
And so I said that DevSecOps is essentially a set of practices and you adopt these practices to the, to the level of culture to, in order for it to be fully adopted. And I'll describe in a little bit what I mean by to the level of culture, but the other thing you, you have to be careful about what the list of practices are. And so I host a workshop at, when I launched this program at other companies and I helped them pick the dozen or so practices.
There are 13 on the screen here, I think, right? Seven plus four, 11, yeah, 13 on the screen here. Um, and, and these 13 sort of show up two thirds of the time when I do these workshops.
And, and, and, um, a couple of things to sort of note about the way these practices are defined. And this was learning from that feedback, uh, that third purpose that I talked about just a minute ago. The, there's a couple of things that sort of, sort of lead to that.
And, and then I needed to add to this list in order for it to be effective. And, and I didn't originally have the, the, the practices defined this way when I launched the program. The first is that every practice is weighted.
And, and this is important because merely doing good things is a failure in your duty. You, your duty is not just to do good things, you have to do more than that. You have to pick the few good things that are gonna produce the most value because honestly, you can't get to all of it.
And, and that is, that is, that is a sad truth. So you have to pick the ones that are the most valuable and you have to make a decision this is twice or three times or 10 times more valuable than this. And, and that's tough.
And, and, but I, I can accomplish this, this waiting at least the first draft of it with an organization in about a half a day workshop, including defining the, the practices themselves in that half day workshop. The waiting is important and the waiting becomes the points you gamify with. So if the team gets to culture for team working agreements, that's five points.
If they get to culture for, um, critical clean for third party code, code, you import s c a sbam, if they get to critical, that's 12 points. And so you actually sort of motivate the team by saying, okay, we're gonna have a leaderboard. And the folks who have gotten the most new points, gained the most points, their score has gone up the most in the last 90 days are gonna be at the top of the leaderboard.
And we use this to gamify improvement. And, and, and without that weight, they don't know sort of how much it, you know, how much effort to really put into it, how valuable it really is to, to the risk, uh, reduction, uh, for the entire company. The other thing to note is, and and I I already sort of mentioned it a little bit here, is this, I don't say get a third party code, code, you import S C a, I don't say start using s e a tools.
That's not the way the practice is defined. You have to get to clean for all critical findings from that tool to get any points for, for this. You get no credit if you just run the scans, but you don't resolve them because of course that provides no risk reduction value to, to your organization.
Oh, there's some value in knowing, yeah, but it's down in the, in the minuscule, you know, fractions of a percent value compared to actually resolving the things that are reported. So we focus teams on this gradual improvement where every step, um, provides value and critical clean for S C A findings is 12 points critical clean for the code. You write like s SQL L injection and command injection findings at SAS or IAS findings is also 12% in this, in this example here.
And so the team can pick either one to start with, and it doesn't really matter according to this weighting, but you know, you could weight them differently and emphasize one, um, over the other. And then only when you're critical clean do you wanna focus on the highs because this is only worth seven points. And so get to high clean for those.
And then medium clean would be an even lower point value here. So, um, didn't make the list of the top 13, it was so lower value. So that's important.
Um, here is the maturity model, so to speak, that I use thoughts to words, to actions to culture. The acronym is thwack. So you've been thwacked by the, you know, by the time you get there, you know, whacked in the head to see the light, think of that, um, thoughts means you've not put much time thinking about it at this point.
Maybe you're just starting to think about it. Words means you're talking about what it would take to adopt this. You're making plans to adopt it.
Actions is you're, you're in the process of adopting it to the level of culture and, and actions really means you have to be like, you can't just be running the scans, you don't get, you don't get credit for that for actions. But if you have five apps and you're critical clean for three of 'em, that's actions for the critical clean for that particular tool category. Um, if you are doing three of the five, and when you get to all five that it makes sense to do, then you get to the culture level, that's meaning you wouldn't think about doing this anyway, any other way for a new project and for every older project where it's reasonable to do it, you're doing it.
So that's, that's culture. Um, and you, you notice how it's not red, amber, green. This is not an audit, this is not gatekeeping and this is really important.
It's just shades of green. You just haven't gotten to it yet. And that is very also important for mo motivating.
Um, the way we use this is, is, is we, we host this coaching workshop to define the list and to put the weights on it. Um, and then, um, we sit down with each development team one at a time and we walk through the list. And when I say walk through, it's not a, it's not a verbal survey.
It's a real conversation. Are you running any s e a tools? Yes.
Okay, great. Can you pull up, pull the, pull the data up for that? Let's take a look at that.
Oh, well, well here's how, how come there's a lot of open, um, criticals. Oh, well we're running the tool, but we're not actually resolving them yet. Great.
So let's, let's, let's put that on the list as a candidate to adopt. And then when you get to three or five things that the team could do, and you're gonna start with the things that are at the top of the list that are the most points or the things that are for requisites, for the things that are the most points. Um, when you get to three or five, you stop talking about the full list of 13.
You basically go into planning mode for the rest of the 90 minute workshop with the team. You say, what would it take to actually adopt this? And you have the product owner in the room and or the business person or whatever it is, who makes the decision about what the roadmap is gonna be for the next 90 days.
And you ask them, do you support the team doing this? And, and, and unless they confirm that they do, you, you, you, you keep, you keep evolving the plan until the product owner and everyone else in the room supports the commitment to adopt this practice to the level of culture in the next 90 days. And that's how you use this, this thing.
So this is about metrics. We haven't talked about metrics yet. We're just talking about sort of how you gather the metrics and sort of what the data that goes into it.
And this is the visualization for the metrics. And basically the metrics are the weighted average of the things that are this darkest green in this visualization. So this visualization is for a 10 team organization.
And so each one of these arcs represents a single team's adoption of this practice. Um, and so it's intentionally small cuz I don't really want you to use the list here cuz it's an older list. Uh, this example is this visualization here, but the, the idea is the same.
You basically count up the dark greens and weight them by the weight of the points for this pie slice. And that's the score for this, for this. And so there's the metric for you.
Um, this is what it looks like when you do a single assessment for a team. And you've got 40 practices like we did after five years at Comcast. So we kept declaring victory on some and adding others to keep essentially improving the whole organization's capability.
And in the end we had about 40 practices that we, we would selectively pick some that only applied in certain conditions and anything was marked in this dark gray was not applicable for this situation. So there was the option for that. And, and that's really important as a nuance later on.
So that's metric number one essentially is, um, practice adoption metrics. Um, w done very qualitatively. Um, so I'm working on, uh, I, I wrote a tool or I had a team that helped me write a tool at Comcast to, to guide this coaching process and generate these visualizations at Comcast, but Comcast owns that.
So I'm working on developing it as an open source project. dev and sign up for the beta when it's available. Um, I do have a number of companies using it in Alpha now I only let companies use it in alpha when I've done the coaching to the, to launch the initial workshop so you could get access to it even quicker.
Um, in order to make it a beta, I have to make it sort of self-service and there's a lot more work I have to do to sort of make it really good as a self-service. But if I'm doing the coaching, I can sort of help you and get you through it using the tool. Um, okay, so is this effective metrics to get sponsorship?
So these are the summary metrics that I use to get sponsorship, um, at Comcast. And, um, the teams that had gone through my program even halfway through my program where they'd adopted half of the, of the essential 11 practices, um, five, actually less than half o o of the essential 11 practices, they had one six as many vulnerabilities found in production as teams that didn't. And they also had one six as many after they got through the program than they did before.
We did what's called within subjects analysis to verify and control for other variables. And that showed roughly the same reduction in vulnerabilities. Vulnerabilities found in production were detected by sort of port scanner tools, likenesses and quez.
They were incident, uh, records like an attack that actually occurred. Um, and they were pen tests that were done in production. So teams that adopted this approach had one sixth as many of those.
And, and the cost of coaching and tool smithing to run the program this way as opposed to the old way of gatekeeping and doing a lot of manual sort of triage of tool output, et cetera. Um, this costs one fourth as much. So it's really a 24 x improvement in, in value.
So this 24 x ROI is what got me the budget to expand all 600 teams over the course of five years at Comcast. Okay, metric number two is median time to resolve. And M t MTTR is the acronym for what I just said.
Most people, um, define m t t R as, um, the mean not median, um, time to remediate, not resolve. And I think that's really important. We'll go through why here.
Um, so, uh, time, uh, open, uh, for findings resolved in the last, uh, end days. So how long take the, the 90 minute, 90 day sliding window, look at all the findings that were resolved in those last 90 days and, and then take the median of how long it took to resolve each of those things. Um, this is the best metric for retrospective analysis, annual reporting, um, uh, it's, it, it it's really a, a, a, a great metric for for that and everyone gets it.
And everyone's familiar with M ttr. Um, even though I redefined it a little, I still just gloss over that when I, when I don't have the time to explain it. Um, and everyone just sort of accepts it, it's close enough.
Um, the problem with this metric though is that if you're trying to drive improvement with behavior, if you resolve a really old finding, it wasn't in the 90 minute sliding win 90 day sliding window before, but now it is, but it was a 200 day, uh, to resolve, well that's gonna actually make the score worse. So you're, you're, you're, you're sort of dinging the team for doing the right thing and resolving this really old finding. Um, and so, so it's not great for driving behavior.
Um, so I, we're gonna get to metric number three is the one that I use to directly drive behavior here in a little bit. Um, but if you do insist upon using M T T R to drive behavior, there is a slight fix for M T T R that I think sort of does make it now suitable for driving behavior. And that is, instead of just looking at the vulnerabilities that were closed in the last 90 days, you also include the vulnerabilities that are currently open and how long they've been open in the average.
If you use mean, don't use mean use median in the median calculation for that. Um, actually, you know, uh, uh, mean median is the 50th percentile, the 75th percentile and the 95th percentile can actually be a better, uh, driver of behavior in these cases. I use the 75th percentile for mediums and highs, and I use the 95th percentile for criticals.
Meaning what's, what's basically the worst case of how long it takes to fix a critical finding from a tool like this. Um, uh, resolve not remediate. So resolve can mean you marked the tool finding as a false positive.
Validly marked it as a false PO positive. You didn't actually remediate it, you didn't modify the code to fix it, you just marked it as false positive meaning the tool was wrong, not your code was wrong. Um, and so that's a resolution.
And so that counts as assuming you, you marked it correctly that way. Um, and then the median, I I talked about that. And the reason for median over mean is that these are highly asymmetric curves, distributions, and the mean is not a good representation of the middle of that curve.
The median is when it's highly asymmetric. You should use the mean when it's easier to calculate the mean than the median, though you should use the mean when it's a normally distributed curve. These are not normally distributed, they're highly asymmetric and so, and they have a fat tail.
So the mean is sort of really a bad, uh, representation of the, of the overall center of the curve. Um, even better, as I said, it was the P 75 or the P 95. Um, and you measure it separately for each, each severity.
So criticals use the P 95 and measure that, and that's a number you track for mediums and medium severity and, and high severity findings use the P 75 and that's the, the number If, if, if you, if, if you, if you're sort of uncomfortable with percentiles, you could use the median for all of them and, and it would be pretty good. Um, this, this P 75, P 95, 50 50, 70 fifth percentile, 95th percentile is sort of a tweak of the, of the, of the um, P 50 median, um, uh, thing of course when you go to P 75 or P 95 and you can no longer officially call it M T T R, you can, you can, if you use the 50th percentile, cuz that's the median third metric, the one I said um, was the better one for driving behavior short term, um, is called cumulative days open. And then you normalize it, that's where the slash n comes from.
So it's basically you focus on not the ones you've closed in the last 90 days, but you've focused the team on the ones that are still currently open and that's why it's better for driving behavior. And you basically just sum up how many days those ones that are still open have been open and you, you add that up. So if one's been open for one day and one's been open for a hundred days and those are the only two that are open, then you cumulative days open is the 101.
And then you normalize it by the total count of findings that you've, you've had open in the last, um, year. Uh, and so that's one way of normalizing. You can normalize it by the size of the code base, the size of the team, a number of other ways you could normalize it.
But the, the one that's in the same data set that you're using to calculate the numerator is, is the, the, the quantity of findings, um, that you've gotten from that particular tool category for that particular severity in the last year. So that's why I, I use, uh, that for the, the denominator for that. Um, again, measure for each severity.
So CDO for, for, uh, critical versus CDO for high versus CDO for medium severity is important because you care differently about those. Um, so here's some metrics you shouldn't use. Um, percentage of applications scanned.
You get ne negative value out of scanning until you resolve them. So why metric that? Why count that, why focus on that?
Um, useful if near zero, but C D O is better, you know, the number of open vulnerabilities, um, uh, the number of closed mo vulnerabilities is essentially a vanity metric and you definitely don't want to do that. Um, so cuz what if you closed a hundred but there are thousands still left, you know, and you added 200 in the same time period that you close the hundred. That's not a good performance, but you, but if you just brag, we close the hundred vulnerabilities, that's, that's, that's rewarding bad behavior.
And so you don't wanna do that. Um, so there's this, this idea of, of sort of the theory of constraints, which is, which is that you should try to, to make your improvements, um, not in the links that are not the weakest link of your process, but the only make your improvements at the weakest link. And so this concept of accumulative photo diagram was invented to visualize and sort of figure out where the, the, the problem is in your process.
Um, developers, uh, engineering teams are very used to using CFDs now because the lean agile movement sort of introduced it. And every agile tool out there, um, every Jira, you know, has a cumulative flow way of visualizing their work. Not the vulnerabilities, not just the defects, but all of their work moves through process steps.
And so you basically visualize it with each different color band in this stacked area chart is, is a different state that the vulnerability in this case can, can move through. And the way you read this is the top line goes up when new vulnerabilities are discovered, the green line catches up with them when they're resolved. And then the blue line is, is, uh, false positives noise.
And that that's sort of as simplified. Your tools may have more states than this, but these three, every tool, sort of all of their states mapped to these three, these three sort of canonical states. And so when we did this to Comcast, we had 12 tools feeding data into this system.
And so we could map every state for any of those individual tools to these three states. And so this is what we use to drive behavior. Um, notice if you just sort of rewarded, uh, vulnerabilities resolved this, this looks pretty good.
They resolve half of them basically, um, you know, a thousand vulnerabilities, this is a, an org a a business unit, a whole business unit. Um, this is a different business unit, so that's why the y axis is different here. Um, and, and, and so, but that's not really good because this orange band is not getting smaller.
They're not actually less risky now than they were back here. So you, you don't wanna reward that you, this is an unhealthy chart, this is a healthy chart. Um, a a as things were found, it took 'em a while to get moving, but after a couple months they got to rapidly resolving them.
Um, and then they stopped injecting them. This, this, this line essentially levels off. And, and as soon as a new finding does happen, um, it, they basically resolve it.
And so this is a healthy visualization here. Um, also useful on the cumulative flow diagram is the vulnerability arrival rate. I mentioned that this curve leveled off.
That's the vulnerability arrival rate. So there's a metric you can actually visually extract, um, or you can calculate it, um, as well. Um, you can also focus on the false positive rates and, and do some, uh, you know, response to a team that marks a lot of things as false positives.
You can look at the current open trend. Um, and so if the area under this orange, this orange area is not getting smaller, then that's bad. And you know, this is not getting smaller, that's bad.
Um, uh, so anyway, I mentioned my talk at the R S A conference in 2020. Here's a link to that. And here's the thing that showed that, that, um, the, the one six, uh, as fewer, fewer risks, um, found in production.
Um, if, if you wanna watch the video, you'll understand sort of how you extract that one six from, from this, from this table. Um, that talk also lists the correlation of practice adoption with risk reduction in production. And this is sort of the payer wise correlation, not the multi correlation.
Um, so that these numbers are exaggerated from reality a little bit. So it's not controlling for the other variables of the team's adoption. Um, but w but the talk sort of, sort of drills down into that a little bit.
Um, and that's all I have for you here today and that's all the time I have. dev, connect with me on LinkedIn, ask any questions, ask for the slides, set up a time to chat. Love to, to keep the conversation going.
Thank you.





