Enhancing Open Source Software Integrity with TestifySec’s Mikhail Swift at OSS Seattle 2024
TestifySec’s CTO Mikhail Swift explores the company’s focus on ensuring software integrity in open source projects, emphasizing the importance of provenance and open collaboration. TestifySec prioritizes open source, providing assurances about where and when software is built, while tackling challenges like key management and integrating into existing development pipelines. Looking ahead, they aim to enhance policy distribution and security, leveraging projects like TUF to ensure the right artifacts and policies are enforced securely.
Transcript
This is Textron tv. Hey, everybody. Mitch Ashley here at Open Source Summit in Seattle.
We had many great conversations, uh, with some technology business, thought leaders, uh, people doing, doing the legwork to make stuff happen with great open source projects and products that are built on that. And we're gonna be talking about some of that today. So I'm joined by Mikhail Swift, who is CTO with Testy sec.
Welcome. Nice to be here. Thank you for having me.
Good to be here. I mean, good to be here with you. Yes, exactly.
Um, let folks know, tell, tell folks a little bit about Testy sec, what you all do, and we're gonna get into open source and talk about that. Yeah. So the, you know, the core of our idea is we want to prove how, where and when your software is built, and give you assurances that the software you're running in production is actually the software that was checked in by your developers, built by your infrastructure, by your pipelines, and, and really try to solve that problem of, Hey, I, I pushed something from my developer laptop at midnight into production, so it probably shouldn't happen.
Right. And be able to catch those types of things as well as like, uh, ensuring artifact integrity. The code that was checked out was the code that actually got built and the code that got tested and scanned.
Mm-Hmm. As well as, uh, collecting all of the evidence to storing it in a one central location for querying and, and compliance and policy usage instead of the 62nd rundown, I hope of What we do. Mm-Hmm.
Yeah, the fancy word provenance. Right. Where's all this stuff come from, and how do you know that's where it really came from?
Right, Exactly. Yeah. We really, really wanna hone in on that problem.
It's a problem we hear about all the time, you know, the whole XZ attack that just happened. Who knows how, you know, when you have a a, a malicious maintainer like that, how he could've caught it, but the fact that files were changing between, you know, release and building, it should have been a little bit of a red flag that it's Kinda a variation of an insider threat. Right.
You know, and dis comfortable employee or someone who's, you know, kind of doing something for whatever malicious reason, but in this case, can happen to open source too. It's not widespread, but it, it has happened. Now we know it's very visible.
Right. Yeah. You know, we Allus had his, uh, keynote this morning and, and talked a little bit about, you know, how the, the stuff with the colonel, the university that did the experiments on malicious commits and everything.
So people have been thinking about this type of attack for a long time, but to see it actually executed in the wild was something different. It's one thing to theorize it, and, you know, usually those things eventually happen. Well talk about, so open source is a big part of your strategy, I know from talking to other folks at your company.
Um, and, you know, there's open source and then there's really op people who really get and live and breathe and are committed to open source. And I know from your strategist, much more the latter. Tell us about your strategy for open source.
Yeah, so we, we really embrace open source from day one of the company. It was important to us, Paul and I, when we started the company, that this software that we're riding this, the, the functionality that we wanna bring to people shouldn't be locked behind a paid gate or some proprietary software. Uh, we've, you know, Cole and I have benefited a lot from the open source community.
We've worked in the open source community for myself my entire career. It's been very near and dear to my heart. Um, just everything the ethos stands for of like collaborating and open ideas and big ideas.
The more voices you have, the better and open source is one of the best ways to do that. Um, so it was really important when we started the company to really embrace that and hold true to what we view as our, the most important thing we could do. Uh, so we started witnessing our visa, our two open source projects, uh, open source from day one.
They're the core technology of everything we do. Mm-Hmm. And to really hone in on that commitment to make sure that, you know, we're not gonna take our ball and go home at some point.
You know, we donated to the CNC athlete last year under the in toto umbrella, so they're all under the in toto projects. But yeah, we, we really think open source is the, the best way to have these conversations outta the open. Everyone can agree on the specs, how things will happen.
You get better ideas that way, better implementations and more secure software for it. You know, there's different variations. They're always evolving to a business models with open source and there's everything from, it's a hundred percent, we just operate it for you to Yeah.
We kind of, we, we, we make it open source, but we're really a prop proprietary software company. Obviously, you're not on that end of the spectrum. How, how much, what's your strategy in terms of how much goes into the open source?
Do you do an enterprise version? How, how, how do you, you have to monetize it in some way. How do you do that?
Yeah, it's definitely a delicate balance at times, but we default open first. If, if there's something that winds up, you know, being enterprise only. Like I said, the main, the main goal there is to find things that, like the open source, it's not really as applicable to them.
You know, a lot of, a lot of problems that the open source has aren't necessarily one-to-one with the enterprise have. So you can kind of find the line there a little bit and then figure out what makes sense and what doesn't. But the, the first strategy always for us is default open.
They'll go in the open source. I can't think of one feature we've gated behind it because there's been no real need for it. Mm-Hmm.
Yeah. We, we do have our, uh, enterprise platform. We're building around it for the more management piece, but the core technology has always been open first and foremost.
Well, and there's also using the information that you're creating about the provenance of all the sources of this, and that has to be packaged up and utilized and communicated in different forms, different ways, whether to a compliance officer or the folks that are worried about, uh, software decomposition analysis and all of those factors and Absolutely. Yeah. Yeah.
And we've really, even, they're tried to do a lot in the open source. One of the big problems of, uh, you know, we, we work a lot with the in toto community, which is the, the open source framework and specification for all these attestations. So, mm-hmm.
Finding that common language, how, you know, when a SLSA providence gets created out of your MPM package or, uh, your running witness and actually collecting attestations that way they can all talk to each other. You can use them. Um, but one of the problems was that discovery and the, the transmission of 'em.
So we, we created VIS to really help out with that problem and open source it as well. So put that out there. It's all under the, in Antonio umbrella.
And, and we again, yeah, we really don't want those, these problems to be gated by any one company. We really think it's better that we all work together on these problems. Very cool.
It is a big problem space, right. No one is solve it alone. You mentioned attest stations for folks that dunno what that is, you know, compliance offers officer will Yeah.
Oh yeah. That's bread and butter for me, but maybe not everyone writing software knows exactly what that looks like. What, what form does that take?
Yeah, so I mean, at the end of the day, the all an attestation is, is just some, in the case of a toto, it's a js ON object that's got some information about what happened, what went into something, a process, what came out of it. Mm-Hmm. And the, uh, the environment that it all happened in.
So, uh, functionally, what this looks like for us in witness is it collects information about, you know, the, if you're on a cloud machine, the AWS metadata server, so you can say, this happened on this AWS instance, or these were all the files that the process touched when it was doing something when you ran go build or whatever else. And at the end of the day, that's all an attestation really, is just that signed document of mm-Hmm. This is what happened.
Here's the modifications that happened on some process, and here's what came out. And you could build those chains of attestations to prove that your software wasn't tampered with. Like, the stuff that I checked out was the stuff that got built.
There was no weird thing that in there. Yep. And that, uh, signing that all helps make it, you know, validated, immutable, you know, it's valid information.
It's trustable. Absolutely. Yeah.
The signing's always the hard, I guess I should say signing sometimes is the easy part. Verification's the hard part, but getting the keys to the right keys to the right place could be difficult. And luckily there's a ton of great work happening with like six store and the open source as well to make signing even easier.
Key management is no small task. It's really not. Yeah.
Well talk. So talk about, um, one, one thing about DevOps and how we create software now is there's so many tools that we use that, you know, there's got a lot of data exhaust, right? And, you know, there's sort of the po there's a pony in there somewhere, strategy of I should be able to recreate what happened.
'cause I've got all this data through all these log messages or whatever data that I have. But that's, that's, that's a puzzle. I mean, that has a major task to go back and figure out what all those things mean.
Plus it's not, you know, really a testable to what it's just messages anybody could have created Mm-Hmm. Uh, or falsified if they wanted to. Why, how do you take approach where you know you're gonna have that information that you can pull that chain of events and, and validate here's what, what happened, what occurred, when, where, and what was in it when it happened?
Yeah. So it all kind of boils back down to like that key management piece of like getting the, the right keys to the right spot. So we use, uh, workload identity around, uh, the key distribution.
So making sure that the key only goes to, excuse me, a specific machine that's authorized to run that workload. Mm. And from there, you know, we can, we can do queries about when you write a policy over the attestations and defining that chain, defining what should have happened and the allowed parameters of everything, we can tie back to what we call those functionaries, which are those cryptographic identities and say, you, this is a little more trustworthy.
Maybe something that I just generated a RSA key on my machine real quick and tried to push an attestation and, and then really start establishing trust with the data, not have to trust like a centralized service to do it. Mm-Hmm. That's one thing we believed really heavily when we started Veta, is we don't want you to have to trust Avita.
It's just hint and silver, like what attestations exist, but then looking into those attestations and digging in and them to find out like, is this an actual trustworthy bit of data I'm being told about here is something wrong? Yes. Not, not oversimplify, but it's not just about, you know, our back kind of controls about access control.
It's really, I'm curious, you know, as developers we use keys and tokens and all those things all the time through APIs and many other things. Um, what are some of the challenges when you get into the domain of key management and applying it to software providence and attestation? Yeah.
It, it really hones in on the, the, just the, the fact that secret material has to exist somewhere at some point, but limiting that blast radius to make sure that only exists for the very specific amount of time it needs to exist to do its job and then doesn't live in a long form state somewhere. Mm-Hmm. Uh, you know, that's really one of the problems that we, we, again, calling back to the six store, what they're really working on in the open source is that that key is a one-time use short key that you just get tied to, you know, your, your GitHub workflow, JWT or whatever exists, it's one and done and it's gone.
Mm-Hmm. Um, and that's kind of the, the model we're really trying to coalesce around or these shortlived ps you know, spiffy spire is another great tool for that of Going Down to even like tpms if you have to, to really get some verification on that workload before you issue the key. And then once that key's there, it's there for 15 minutes and it's gone.
Mm-Hmm. Uh, because The world has changed, it's no longer that I need that five year certificate from my server. Right, Exactly.
Or checking in, you know, uh, or having an environment variable set to some token that you just hope, hope your, uh, provider handles appropriate, never leaks in a log somewhere That that private key that doesn't stay private Right. Isn't our source source code or some variable file somewhere. Um, I'm, I'm curious about, there's always the, it's better if you can set things up from the start to be able to do this.
What's it like when you're having existing developing a lot of software already and we want to really up our game to, for software supply chain security, and be able to produce this information. Maybe we're under stricter compliance now or under the gun of something might have happened. How do you apply this to an existing working development organization?
Yeah, so we have a few different things to ease the burden there of trying to integrate it into an existing pipeline. So we have a GitHub action that people can just pull in and use and it reduces the friction quite a bit, uh, for a witness run and actually start and generate these attestations. Uh, we're working on some other stuff with like some integrations with GitLab to make that easier as well.
We just published A-C-I-C-D component to their beta catalog, so you can, you can import that and start getting attestations. Those still involve having to wind up editing your, your pipeline definition. So we're really starting to experiment with ways we could even make things smoother and, and even more frictionless.
So whether that's a runner integration, so like the, anything that runner does automatically spits out attestations appropriately. So you don't even have to think about it in your pipeline. Your dev team doesn't have to go change.
Things are, are really kind of where our area of focus is right now and what makes sense there. So we have some proof of concept stuff, but nothing ready for the real World either continue to work and evolve it. Absolutely.
How far, kind of back up into the process, all the way to the developer setting up their environment and to start working on a software project, where do you need to kind of get in into that chain or that flow of work? Yeah, I mean, the closer, the closer you can get to the, the source of truth, the better. You know, in an ideal world, what we'd be able to do is, like the compilers themselves would be generating these attestations Mm-Hmm.
The compiler, you know, is at the end of the day for software supply, like actually building something, that's the thing that you're trusting the most. Mm-Hmm. So if that thing could give us better provenance and better information about what came about, and there's a lot of effort going in that regard, go and a release, uh, you know, a hand, I think one 20 or something added in a bunch of new build metadata that gets built into your binaries.
Um, but when you're starting fresh and, and looking at it, the closer you can get to the source of truth, the closer you can get to the actual process to it, the better. So it's, it's a tough problem when, you know, going back to your last question about, well, we already have these existing things, how do we do that without having to go touch everything? Um, but yeah, that, that's really the best way to tackle it early on is right at the source of truth rate it where the hatch will work is happening.
If you get better data that way, you get better guarantees and better trust. If it was easy, we'd already be doing it. Right.
Exactly. Yeah. You know, it's, uh, and then this is just a kind of a random aside, but you know, Dennis Richie wrote about like the, uh, what is, I can't remember the name of the article, but the, what happens when your compiler, you can't even trust your compiler Mm.
Like the compiler's mutating stuff. And, and that was back, I'm pretty sure in the seventies. Yeah.
Oh yeah. You're saying you're going back a ways here. Yeah.
If, if, if it were easy it would've been fixed back then. But if it's not, an easy problem Is there's trust all throughout the workflow, uh, through tools. And of course, you know, your job is to validate, you know, whether what happened actually did happen and where the source of those things came from.
What about, I mean, it's easy to think well about, relatively easy to think about ID and the developer writing software and somebody creating it. What about all the other sources, you know, whether it's RPM, package managers or whatever, you know, platform you using, um, sources of images, containers, you know, pre-configured virtual images, whatever it might be. That's all.
Plus we're talking to third party services too, right? Mm-Hmm. That we're exchanging data with.
How do you manage those sources of, of code? Absolutely. Yeah.
So the ingestion of, you know, third party software in your is always the hard part. 'cause Yeah. You, you either have to in inherently trust it, do some manual review currently, or we're starting to see, uh, things like NPM, um, actually start generating these SLSA attestations and other attestations at the level that actually tie it back to do, get commits that happen.
You know, previously, uh, you know, attack factor with some of these package managers was the, the software that's being pushed to 'EM doesn't necessarily have to match the, you know, be built from the get that they say they're being built from. Mm. Um, so we're starting to see a lot there, uh, happen as well that that is starting to ease that burden a little bit of bringing in external software and external packages.
A lot of Linux distribution now, like Debbie and Arch both have really great, really great efforts going on to make their builds completely reproducible. Yeah. Which is great.
It doesn't, you know, still solve that inherent trust. You have to trust something at the end of the day, but the fact that it's all reproducible you, it raises red flags earlier. Mm-Hmm.
You know, um, so it's a tough problem. We're really looking on, you know, a lot of enterprises have a very manual process of establishing who's trusting it, who's not. Mm-Hmm.
Building out stations around that to make automated compliance around is something we're really thinking hard about right now. Of like, okay, the open source office said this is good to go. Uh, so we have a little more trust in that than something that's just being pulled from public internet or somewhere.
It's, it's not an easy problem. Right. Whether you're gonna put it all in your own repository, build it yourself, that's a whole effort.
Yeah. And that, as you said, doesn't guarantee what happened before to, to create that code. It's such a, you know, the whole concept of open source, so we can have eyes on it, everyone sees what happens, but no one functionally or practically actually reviews every piece of code that goes in their supply chain.
It's just a, it's an untenable problem to solve. Like no, there's no not enough time in the day to do that. You see what you're doing as testify sec applying to open source software projects themselves, right.
So that we someday might have a source who we know we can go back and validate of what happened on that open source project. Yeah, absolutely. We really, really wanna make it easy for open source projects to adopt this and get up and running.
Uh, our, our opinion is, you know, the more ations that are out there, the better. It gives you better data, it gives you more insight into what's going on, and it gives you the audit log of like the retrospective of trying to figure out where something did go wrong, if, if it's happened to get through. So open source adoption to these things is something that I'm really excited about.
Uh, like I said, I mentioned NPM creating some of these, they use GitHub actions like mm-Hmm. That's great. Like getting those attestations for almost free is an amazing thing that, that I'm really excited about.
And we're starting to see more uptick in other open source, uh, package repositories and, and platforms as well. I would imagine, you know, con contributing as open source makes it even more tenable or attractive for someone to say, great. Right.
Just make this part of how we do our work. Yeah. If, if we were a forest closed source solution, no open source projects had of take that and do it just be Kind of tough.
Yeah. Yeah. No one would want to do that.
I've run open source projects, I wouldn't do it. Mm-Hmm Mm-Hmm. But so what's next?
I mean, you mentioned a lot of things that you're working on and, and, uh, kind of innovating, involving where the Cisco, what do you think the next kind of big challenge to, to focus on is? So one thing we're really focused on right now is the, the policy distribution and making sure like you're using, you have the latest and the right artifact. So now that we have attestations about these things, how do we make sure when I'm pulling it, there's not like some downgrade attack or even if we have the attestations, it's not the right version that we're using.
Yeah. Or for policy enforcement, making sure that the right policy is being enforced instead of having some old policy that maybe is still has trusted functionaries that we've now known aren't trusted any longer. Or maybe that policy changed since it Oh, yeah, exactly.
Yeah. Or contributed. So we're really focusing in on that.
And, and there's another really great open source project called T that really helps with that. It's called the update framework. Hmm.
So we're working on integrating that in ARCA Visa, which is our IS station store. So we can actually start distributing policy securely in the right policy at the right spots. Mm-Hmm.
So that's our, our really next focus for the short term, uh, for the company. We're also releasing our platform hopefully shortly, uh, for some early partners. So from the company side, you know, we're excited about that, but the open source is really coalescing around the, the tough idea right now.
And Ruby Gems and I think Pi Pi have both implemented tough and their own package repository big. So again, even more. Yeah.
Big. Yeah. We're starting to see it everywhere.
It's great. Yeah. It's those, one of those having a lot of things come into place really helps pulling it all together.
It does. Yeah. And like as the, like we were talking about at the very beginning, you know, just with opensource being such a large community, more people on it, more people looking at it to see all these individuals like come together, like saying, yeah, this is a great idea.
I should put this in and Ruby Gems or whatever else. It's just idea of, of validation at one point. And it's also just great to see people who care and passionate about it, you know, working on this stuff.
Well, it's clear you're passionate about it too, and your experience running open source projects, massive help. You know, you're, you live, live the dream and the challenge at the same time and can appreciate what some of those challenges really are in a practical sense. So it's gotta be super valuable for you.
Absolutely. Yeah. Well, this company wouldn't be where it is today without the open source communities that helped us, and we've hopefully helped foster it as well.
Mm-Hmm. But you always know by engagement and use, right. Uh, if you're adding value or helping people.
Absolutely. Yeah. Well, Mickey, it's great talking with you.
Keeps us up to date. I know you're rolling things out and, you know, open source, community and commercial parts of the offering. They look forward to hearing more about it.
That's great being here. Thank you for having me. Absolutely.
Great to be talking with Mikhail Swift, CTO with sec, another example. The fantastic people who are making our world easier, better, safer, secure. You know, it, it's a tough job to secure, uh, software supply chains and not just open source pride, our own code as well.
So thanks to and team at Testify sec from taking on that challenge with us. So we'll have more interviews for you coming up soon. Stay tuned.