Why SCA Isn’t Enough – From The Source EP6
In a time when software supply chains are under constant threat, relying solely on Software Composition Analysis (SCA) is no longer sufficient. Join Sonatype CTOs Brian Fox and Ilkka Turunen as they delve into the value of a more comprehensive approach to security. This webinar will explore the strategic advantages of combining tools to go beyond basic SCA. Learn how these solutions work together to provide continuous monitoring, proactive vulnerability detection, and enhanced policy management, giving your organization the edge it needs to secure your software supply chain effectively.
Transcript
Hey everyone. Welcome to, um, a, a very festive, uh, episode of From the Source with your host, uh, Nan. And I am Brian Fox.
And Brian, um, we've got a very festively, cheery, uh, topic that we wanted to touch on today, but Should we, should we do one of these, uh, or work? Oh, yeah, yeah. Get, oh, yeah.
That's a great idea. Actually. Let's get distracted, uh, just for a moment.
Ho, ho, ho. Nope. Can, can't get it to work.
Well, um, so Brian, we got a bit of a spicy topic, uh, that, uh, I wanted to touch up on today. Um, and in fact, the synopsis that we've got going on in here, oh, look at that confetti. Love it, love it.
You're getting in the, in the spirit of things. Well, uh, just in time for the holidays, uh, a little bit of something for the, um, uh, for the, um, holiday table discussion why SCA isn't enough. So, you know, this is something you and I talk about constantly.
It's been, you know, something that I feel like we've been saying for years in one way or the other, but, you know, software supply chains are under a constant threat. We're, uh, we are seeing, you know, huge amount of new threat vector, uh, entering into their amount of risks from just technical negligence managing it. We're seeing, uh, you know, attacks, uh, being executed, leveraging it, and from one paint place to another.
And because of all of these things, people often see SCS sort of the silver bullet. And, you know, the reality is SCA totally is not a silver bullet and solely relying on it, uh, is not sufficient. So the topic that we wanted to touch on today was really to do a little bit of a deep dive on, um, what that means and why that means.
And I'm, I'm gonna start us off, Brian, with a little bit of a controversial statement. I think most people kind of use SCA as a bit of a, a pressure valve and, and a bit of a sort of thing to, uh, to say, Hey, I am doing something without actually doing much at all, because the thing is alerting me and therefore I'm safe, even though I don't address any of those alerts. But, uh, look, Brian, you've been thinking about this a long, long time.
What's your thoughts there? Yeah, I mean, I'm a little bit on the fence. I mean, I feel, I feel like a big party industry isn't even doing s ca as evidenced by the fact that so many people still use, uh, known vulnerable versions like the Log four J that we've talked about in the past, that they keep using the vulnerable versions.
And I don't think that's an intentional choice. So, um, there, there are shades of people are doing it, not doing it, doing it a little bit and ignoring it. Um, you know, and, and I think you, you, you almost have to explore that a little bit.
I mean, clearly, uh, the angle we've always taken is, uh, just doing the analysis is a little bit passive. Producing reports that are summarily ignored by developers doesn't do any good. And, and that's why we've always tried to, you know, be deeply integrated into the pipeline so that you can actually define guardrails for your developers, uh, break bills, block releases when you have to.
Hopefully that's a last resort and not the first move, right? So, you know, I think, uh, traditional SCA if there is a such a thing, you know, um, commodity SCA is really just producing reports. So clearly an element of you need to do more, can be thought of as, you know, you need to think about it as a supply chain, get integrated, be able to provide the, the information to the developers when they're making these decisions upfront, not, not scan and scold as we like to call it later, right?
That's definitely an area. Um, you know, and then there's, there's going beyond that in terms of being able to address malicious components, right? So those are, those are kinds of the things going on, uh, in my head as I think through this, this Topic.
Yeah, I mean, it's a tough nut to crack, but, you know, let's take a look at some facts. Um, you know, one of the reports that we've been publishing, we've just released a sort of malware report where we do a deep dive, uh, into sort of the malware distribution. We'll, we'll touch on that a little bit later, but, but, um, it is the state of the software supply chain report.
We did, we've done a number of deep dives on, on the show, uh, about that. And, uh, you know, one of the things, if you just look back 10 years, it, that kind of stands out, is really, uh, the fact that there's actually been not so much changing behavior in terms of vulnerability c consumption. So, like you said, people still download vulnerable versions of Log four J period, vulnerable versions of Log four J to log four, shell to be specific.
And that is literally the most published security part of CEO o of all, that's like literally no excuse to be downloading it at all in the first place. And, um, and now when we, when we kinda look at sort of the technical reasons as to why that's happening, part of is ignorance as in not doing, uh, SEA at all, any, any kind of activity. And part of it is also sort of, Hey, we will accept the risk and we'll do that.
I think part of it is also technical, you know, a lot of sort of what you just referred there to as commodity SCA, really the only thing that it does is compares, um, in dependency manifests, you know, Palm XMLs package files, uh, package logs, things like this, uh, to a public database of security vulnerabilities that, you know, may or may not be accurate mm-hmm. To the level of information that's needed to get an accurate reading. And, you know, kind of early days of SCA, the biggest hurdles that we people always had to overcome was the fact that it was super noisy.
It was super full of false positives. And I think to this day, if you use something like, uh, something like depend upon or other, other tools like that, you're still inundated with sort of false positives because it feels like it ought to be a very simple problem to solve. I've got a list of components, I've got some database left joint, that's my things to fix.
And, and, and that's literally it. But it turns out there are no good databases. The databases have to be built by vendors such as ourselves.
Uh, there are no good ways of actually matching dependencies, uh, from just manifest. You have to look at the files, the installations, the environmentals, you know, and it's different for every single programming language. And that leads to a situation where I think a lot of people, even when they are doing S-C-A-S-C-A through some tooling or another, aren't actually really doing anything that's impactful, aren't really reaping the benefits because they're not really truly covering all the risks that that's there.
And even worse, they're sort of drowning in things that aren't actually necessary things that you needed to address in the first place, which unfortunately us all being engineers in this industry leads us down this path of, well, I'll find a way of filtering those things out. I'll find a way of looking at my call path, or like, I kind of going down the execution flow of the application to try and skim down that sort of poor result set that I have into something that I can at least action. So, so I think a lot of the, a lot of the issues that we see, the in effect that we sort of observe are actually sort of very technical in nature.
They're, there are things that, uh, that, uh, people do, but there's also sort of, I feel like a human element thing. Um, and I think it's sort of the ostrich problem, which is, Hey, if it isn't the most critical thing in front of me right now, I'm just gonna ignore it. 'cause hey, it, it ain't broken and it hasn't come to bite us yet.
And that, that's sort of, I think the same analogy as blood pressure. You know, it isn't a problem until it becomes a real problem. Yeah.
And by then the damage is done. Yeah. I mean, the, the thing that I see this happening a lot is, is of course, within the malicious open source components, right?
That for organizations that are doing some form of SCA, they, they often think that that's gonna handle the malicious component because they're thinking about, well, I don't want these bad things to be shipped downstream to my customer. But the problem is, that's not the goal. The goal is to attack your developers and the development infrastructure.
Um, and so if you're scanning these things after the fact, the damage has already been done, right? These components leverage, um, the post install scripts and things like that, and NPM and Python or, and what have you, and fire, as soon as a developer downloads and installs this. And, and so the, the attack happens at that moment, not after you build and release the software.
And so if your SCA process, even if it's deeply integrated into the infrastructure, the pipeline might not see it. And if it does, it's still after the attack, right? So that's part of the problem that, that I think, um, for, for many folks, they're conditioned to think about, about, okay, these are dependencies and I have a process for managing my dependencies.
It's like, yes, but this attack is happening far left in the process. And so you need different types of defenses to catch and stop those. The only way to prevent it is to intercept that request and make sure that it doesn't land on the developer's machine.
And that's a very different dynamic that I think, um, is getting lost in the noise of SCA. Yeah. I, I think you're absolutely right.
And, you know, one of the other sort of things about SCA is it is a useful tool for the job that it was originally designed to do, which is to enable you in flight, during an engineering workflow, make better decisions earlier, and then give you some sort of gating to make sure that you know, hey, if some decisions that weren't great were made, uh, you can, you can at least prevent them from going forward in the process because kaizen, you know, pull the Andon code, stop there, you know, and, and, and fix the issue, the, the problem or agile manufacturing for that matter. But the problem there also is it ignores a problem on the right hand side, you know, so we certify something today as good, you know, it's gone through a scan, uh, uh, et cetera. We, we ship that out.
Even if we continue to build and we continue to throughput things, that doesn't mean that the thing that was shipped out the door is still good. Um, because that in on itself might have different things. That's part of the reason why we've had this introduction of, of all the SBO laws as a sort of direct descendant of, of introducing SCA into our workflows, uh, you know, had drop an SBO monitor that om, uh, in perpetuity for every release, uh, in order to have there.
But again, I think in other places, the serological assumption is that s om is, is ubiquitous to SCA AI can just, you know, grab one, generate one, and then, and that'll be good. It's just a snapshot of time in, in fact, SBO and SBO is quite a living thing that, uh, now increasingly it's gonna have legal consequences if, if you don't manage it, right? Yeah.
I mean, for the most part today, without SEA, you can't create an accurate SBO m because it's, it's not just a case of pull together all the SBOs of my dependencies, add them all together and ship it. Part of the problem is most of those don't exist yet. So you don't have, you don't have the building blocks.
You have to do the analysis yourself. You know, over time, um, I think it will morph more into SCA is able to validate and augment SBOs, but right now it's really the way to produce the start of an SBO that gets shipped down downstream anyway. Right?
So the fact that, um, we're seeing so many people have to lean on SBOs kind of, kind of goes back to my first point that there's still still not enough people are able to do it, because if they don't know what's inside their software, um, they're probably not doing any of these things at all anyway. Yeah, Absolutely. And I, I think what's really interesting about that, uh, as well is that, um, is that, um, they're often being driven by very different personas, uh, even though they are actually very related.
All these three topics that we've kind of touched upon so far are very related. There's actually a fourth element to this, which is license analysis. Um, so many people do SCA just with a, just with this or security perspective, Hey, I wanna make sure that the software that I build is devoid of no more, the most critical security benefits is usually the version that you see.
Um, but you know, an interesting thing over here in Europe is, 'cause we have pretty strong copyright laws in some of the countries. Uh, we've always had a bit of a aversion to misusing licenses. So some of the SCA work that, you know, traditionally started here was all about license analysis.
0 license thing or a feral GPL license thing. Um, but increasingly we're actually seeing a lot of risk, especially with repackaged, uh, lms, especially something like wman that doesn't actually have an open source license. It has a no commercial usage license, uh, on it.
You know, when somebody repackage like fine tunes it, repackages it and publishes it with a different license, you may feel like you're actually, um, actually, um, you know, free of any, uh, licensing terms. 'cause it, that retrained model came out with an MIT license. But underlying it, you're actually also covered by the original terms of service, which might or might not come, come back and bite you.
So what I mean to say about all of this, you know, talking about this licensing element, is they're often driven by very different personas. You know, SCA came from DevOps wanting to do more DevSecOps, uh, malware analysis really truly came from a sort of security perspective. Hey, there's this new attack vector that targets developers.
How do we, how do we deal with that? S om is a compliance thing. It's a, again, a completely different persona that, uh, has to address in completely different things.
Um, and one of the sort of big challenges is, uh, or big inefficiencies, I feel like we're gonna see, you know, uh, this sort of, uh, next few years is people are gonna try and invent each and every one of those separately when in fact, to me, they're just facets of the one and the same problem. Yeah. Yeah.
I mean, I think when you think about going beyond, you know, just the SEA, the produ production of the inventory, you know, we, we, if you think about it as a part of a, a supply chain in the system, you need to be able to monitor those things for changes as well, right? And the, the analogy I use a lot is, you know, we all have cars and we're used to recalls and things like that, right? And imagine if the manufacturers never had to do a recall as long as they printed out the list of parts and put it in the glove box and they sold the Car.
Funny enough, I literally got alert today that my car is under a recall. So yeah, I feel that. Right?
Right. But, but it, it absent that continuous monitoring of that bill of materials, it would put it on you as the end user consumer to I know it, to basically do it yourself and how would you know what to look for, right? And so that's kind of the modern equivalent of the SBO m is just the start.
The SCA producing the SBO m is just the start. You have to be able to take that and use it for infor important information. A, to be able to, uh, follow up, see when things, uh, when bad things happen to these components.
You know, it's not that the component changed. You're not monitoring that the SBO m changed, you're monitoring that the state of understanding around the po the components changed, right? When, when cars are shipped, hopefully they don't know that the airbags are gonna be faulty, but later mistakes happen.
They figure it out. And, and so the information about that change, even though the part didn't generating the alert, the, the equivalent needs to happen on, on the software, right? Yeah.
Uh, for sure. And I think, I think that's, that's sort of, I think why, uh, why, um, you know, reducing the supply chain problem into just, hey, you know, do some scanning during build time is sort of right, sort of a very natural reaction. That's sort of the minimal possible thing that we're all used to doing.
Now it's a, you know, I've got a scanner of, it'll compare my list of things to other things when in fact, this sort of, the domain of the issue is, is far wider and leads to that logical conclusion of I'm selling out a product and that product may or may not have quality defects, and I need to be able to execute fixes. Now, historically, you know, it's, I think, you know, we, we spend an entire episode talking about, you know, the regulation. But, you know, one of the big changes, the CRA and the PLD are gonna lead to a place where if you don't do stuff like that, there are gonna be fines and, uh, sanctions that can be imposed as a, as a result of not following us or best practice.
So in, in that sort of vein of continuing on our predictions, uh, from our, uh, last time, um, I, it's actually, uh, it's actually not just, uh, not just, uh, a consideration that some of us have to do. I think it, it'll be, it'll be sort of something that we're all forced to really think hard and true about. But the good news is that foundational element of SEA certainly sets you up for a starting point.
So, agreed. Uh, man, you know, funnily enough, uh, funnily enough, you know, if you think about, uh, think about, uh, all of these things, um, all of these things just in general, you know, there's been, uh, actually a couple of, uh, couple of sort of supply chain incidents, uh, quite recently that we've seen, uh, where, um, you know, I think there was a, there was a JavaScript, uh, component that actually got taken over by malicious actors. There's also, uh, uh, something called web free js fairly recently, a very popular, uh, JavaScript library used in crypto and, uh, sort of, uh, web free, uh, type activities.
And that actually got compromised, uh, periodically or temporarily. And I think that's a really good example of a type of supply chain attack that's now being executed. You know, I, I saw on Hacker News and, and a bunch of other places, and, you know, our research team kind of picked up on it, uh, relatively quickly.
But that, um, I think it's a good example where, you know, having protection at the network or ingress level is really actually the only way to avoid the type of attack that it tries to execute, because it really tries to execute on the developer machine using developer tooling rather than poisoning your software and then carrying, uh, itself, although that too can be a risk. Um, and so having that sort of more realistic approach of, Hey, look, you need, you need something to filter things coming in. You need something to filter things as they're being built to minimize as sort the technical debts that you're accruing, and then you need something to work out for it once it's out the door.
Um, I think it's pretty important. Yeah, I mean, we're gonna, we're gonna continue to see new and novel attacks on these types of components, whether they're fake components, the type of squatting we've talked about before, whether it's account takeovers, like we've seen, um, just this week there was another one, uh, what within Python, where, where the pipeline itself was actually compromised in publishing some malicious components, right? So it's kind of all over the map, and that's why you have to have these defenses in place to be able to deal with It.
Indeed. And, you know, just to, um, uh, just to, uh, plug a little bit of something that we, uh, did this week as well as we released new reports, uh, that dives into some of the malware, uh, publication patterns that we are observing. Mm-hmm.
And, you know, sort of surprising, not so surprised, uh, I'll spoil you, uh, one of the findings in there. There's a lot of it in NPM, in fact, MPM accounts for, uh, a huge, huge amount, and there are a myriad of reasons mm-hmm. Why that is, mostly because it's very popular, you know, by itself.
Um, but regardless, uh, regardless, it just shows that that sort of risk landscape is following wherever. You know, we have popular, uh, development happening, you know, NBM and Python as an example. Uh, you know, they are constant targets for this type of activity.
Well, hey, uh, Brian, we're kind of coming up on time. So I guess it wasn't that controversial at all. Uh, you know, I think we, uh, uh, we have managed to cover, uh, quite a lot of it.
So, um, I would say, uh, I would say that, um, uh, if you're fancy, have a read, uh, if you need something for your festive table, you know, perhaps under the tree or a stocking filler, print out our report, uh, at a read and, uh, you know, see what that does for you. But, um, you know, I, I guess the takeaway for me today is as I've been doing most of the talking here, uh, is, um, uh, I don't think the SEA is, is the answer. It's a part of an answer, but you need to think a little bit bigger.
Yeah, for sure. You need to be protecting, you need to be thinking about what happens after you produce the inventory. You know, how do you defend your developers against, uh, these other types of attacks?
Right? It, it, it, there, there's challenges that go in both directions further to the left, further to the right in your, in your pipeline and your, your, uh, A DLP. Indeed.
Well, with that, uh, Brian, thanks so much, uh, again, uh, for a, uh, lovely discussion today. And to everybody listening, uh, thank you very much indeed, uh, for sticking with us. Enjoy your festive period.

