Gremlin Expands Proactive Reliability with Disaster Recovery Testing
Kolton Andrus discusses Gremlin’s launch of Disaster Recovery Testing, a new capability designed to help organizations safely simulate large-scale cloud and infrastructure failures. The platform enables enterprises to test region, zone, and datacenter failovers to strengthen digital resilience, support compliance initiatives, and ensure business continuity during catastrophic events.
Transcript
Hey everyone. Welcome back here to Tex Trump tv. You know, I, I, if I don't have this gentleman on every two, three months, I start wondering what's going on.
So it's probably been about three months. Um, I want to introduce you too. Well, you may know him already, but if not, meet Colton.
Andrews Colton is the founder and one time CEO. And now CEO again, ed Gremlin. Hey, Colton, it's good to see you, my friend.
Happy New Year. Well, it's a little late, but I hope the New Year's off to a good start. How have you been?
Yeah, always a pleasure to chat with you, Alan. Thanks for having me on. Things have been going well.
Always excited to come share the latest and the greatest or dive into the details on the, you know, on the technical side. Very cool. Colt, and I see you're the founder.
You were CEO then you weren't CEO for a while, then you came back as CEO. Um, but beyond that, prior to Grambling, you were at Netflix and, um, you know, that this is when Netflix was not that they don't innovate anymore. They buy things now, obviously, huh.
But, um, but you know, they were really an innovative technology company. They kind of pioneered the whole thing around like chaos engineering and talking about scalability, right? They kinda wrote the book, a lot of cloud native kind of, uh, functionality and projects were spawned out of Netflix.
It, it was a great time to be there. Yeah, no, I'm, I'm really grateful. I mean, uh, for those that, dunno, I got to do a, a four year stint at Amazon focused on keeping the retail website up and available.
We built fault injection, chaos engineering tools there had a lot of success. But I remember being at a, a, a Velocity conference, I think in 20 12, 20 13, and, uh, hearing, uh, these companies talk about hearing Netflix talk about what they were doing on the resilience, reliability space and Chaos Monkey had just come out and, uh, I thought that was neat. I thought they could have done more there, but, but it was, it was a good first step.
But they did a great job promoting it and helping people understand why it was important and really driving that innovation. So I was excited to go there. I showed up, you know, there was room to pitch in and help out.
I helped us get another nine of availability, helped us build some grade tooling, and really that's what allowed me to have the opportunity to go found Gremlin. Um, I was giving a talk at a conference and I'm running into some VCs in the lobby, and we went back and forth and they said, and I was like, I'm gonna bootstrap, you know, I'll just wait it out. And they were like, hold on, you could start tomorrow.
Let's get going. Uh, and three months later we founded Gremlin. And coincidentally, uh, as of Sunday, it's been 10 years now, we've been out in market Another overnight success.
Yeah. Yeah. Not, not what I read on Twitter back in the day.
No. Well that, you know, that's the deep dark undervalue the whole thing. All these people who think, oh yeah, startup guy, you know, overnight success.
No, this 10 plus years, it's not unusual to have that kind of effort in there. And you, you know, and you, but you're still a startup and still building and still learning and, and, and doing all that. Um, now Gremlin's mission has, I don't know if you wanna say expanded, changed, evolved.
Chaos engineering is still obviously part of it, but chaos engineering now has, you know, it's, it's the life cycle of a product to a feature in tech, right? Uh, today's product become tomorrow's features as you, you know, March inex inextricably towards a platform and, and all of that good stuff. So what, what's rambling about today, Carlton?
Yeah, I mean, I think that's, it's such a great point. Chaos, engineering, cool idea, fun project, not really an enterprise discipline that drives reliability. And that's what we learned early days is a lot of folks wanted to like, you know, kind of have fun, we'll break some stuff, we'll see what happens, we'll order some pizzas, you know, we're doing good things.
And a lot of that didn't really result in the type of outcomes that they wanted because it was isolated. It was not repeatable. And to your point about things becoming commoditized, things becoming part of a platform, things really becoming just how we build software.
That's what we've seen a lot of growth in in the last few years, is people moving to, Hey, every team needs to do these basic tests. It's just, you know, it's unit and, and integration testing for our distributed systems. And when we do it, we find things on a regular basis.
We fix them before their issues and we're just building better high quality software. Absolutely. So what, what we've got coming out right now is just along the same lines, uh, you know, having grown up in a lot of enterprises, a thing that almost every enterprise company does is some form of disaster recovery testing.
Hey, what happens if we lose a data center? Hey, what happens if a WSU East U US East one goes down, or GCP loses a region or a zone? How does my software, you know, behave?
Can I keep operating? And a lot of companies do this in a pretty manual process. Hundreds of engineers, a whole bunch of prep.
They're running it once or twice a year, and they, they, you know, they've gotta, they've gotta invest a lot of time. Then some companies, this is a weekend where everybody's on call and you've gotta just show up and be on that call just in case something goes wrong with your software while we execute this large event. So we thought, you know, this is a place that we could do better.
Uh, and what we did is we took Gremlin, which is already good. Our customers are already using Gremlin for this purpose. Some of the largest banks are using Gremlin to go do this kind of dedicated disaster recovery testing.
But we said, well, let's make it easy to do the right thing. So we built, uh, built into the product the way to model this large scale event this way to have the right safe preconditions. Let's make sure everyone's run it on their own first.
Let's make sure everything looks good. Do we wanna run it, uh, one big event or do we wanna spread it out over a couple of weeks and let people, you know, prove their piece independently? Um, let's make sure we've got that halt button if things go wrong, you know, a way to clean it up and revert it so that we've got that safety net while we're running the experiment.
And I think part of how we've grown up is the last piece is if you can't measure it and you can't turn back and present to the business or the auditors or the compliance folks, the evidence that you've successfully completed this, then it's just not as valuable. And so we spent a lot of time building really all of the, the right reporting and detail so that you can run this event, you can run it more often, uh, because you don't need as many people involved and you can streamline it. But then you also have everything you need on the back end to go hand to the business to say, we've done our due diligence.
You know, Colton, I think back to my days, you know, running companies that were, well even here, ra right? You know, disaster recovery plan. We're, we're totally sass.
We, we have none of our own infrastructure per se. So disaster recovery you plan is basically getting on the phone and begging, right? Um, I'm being facetious, of course, but you know, in other companies that I was involved and we, we did have much more dr right?
We were running data centers and stuff like that. Um, I wish we had something like this back then. 'cause as you said, it was quote unquote a manual process.
And, you know, living here in South Florida, you would think everyone would have, what do you do in a hurricane as part of their DR testing? But you'd be surprised. You'd be surprised, you know, oh, I didn't have that on my bingo card.
Well, you're in south Florida, you didn't think a hurricane might knock you out. You might be flooded. You might not have power for a few days.
What do you like, you know, where was the thought process there? And it's, you know, so when we, Manuel is a, is a, is a good term, Colton, for covering up, just, you know, you can't cure stupid. And, uh, there was a lot of, I've seen a lot of stupid over the years down here with, with people, you know, who thought they had their DR plans in, in, in, uh, place until, until the stuff hit the fan.
Yeah. Um, Such a great analogy. But It is.
But you know, so having something like this, you know, this, this is like a, you know, sleeping under the blanket of security, right? That that kinda keeps you, you, you feel like you actually really do have it. It would also seem to me, Colton, that this might be something that maybe AI can help with in terms of, of, uh, you know, setting up, what, what are the test parameters?
What is your test coverage here? What should you be testing, what, you know, like we probably don't test a lot down here for snow, right? I, well, who knows, but ha are you using AI or anything here to, to help with that?
Or is this really based on real world experience that you're now, you know, being able to scale up to multiple or infinite amount of, of customers? Yeah, so part of what we have in just built into the Gremlin product is a set of recommendations where if things aren't going right, then we're gonna tell you how you can go fix them, how to go adjust them. Um, and one of the things we're always tuning and improving is just how do we make it easy for people to do the right thing?
And that's really tell 'em what they should be doing. I think this is one of the things I've had to learn in my career early as an engineer, come from Amazon and Netflix. Just assume everybody's done this a hundred times.
They know what they're doing, just give 'em the tool and get out of the way. But truthfully, there's a lot of folks that maybe haven't done this, or maybe they're, you know, if they've got some junior folks on their team and they really need a little bit of guidance on the structure, on how to set it up, on how to model it, on how to run it, on how to interpret it. Uh, so we've got some of that.
We're always improving that, uh, within the product. So I'm not gonna say it's AI driven. I don't want to be sensationalists there, but we've built in a lot of recommendations that tell people how to fix what goes wrong and what they should be doing.
Absolutely. Colton, um, you know, thi this became available on February 3rd. It's in ga, February A as of the third then, or Yeah.
Okay. Yeah. I told my team we've had it ready for a while and we've had it in beta with some of our large customers.
And one of my quality bars is have we run this ourselves in production? And the answer is yes, we have, we ran it and we learned some things. We passed the test, but we found a couple things we didn't like that we went and fixed, which is exactly why you run these exercises to uncover those things.
Absolutely. And go fix 'em. So that was why, that's why it's ready to launch.
'cause I said, engineering team, when you've run it and you feel good about it, that's when we're ready to go take it, run, then We're ready to go. You know, and that, and that, right there was the beauty of of Chaos Monkey too, right? Because it's like a live fire drill.
Mm-hmm. You know what I mean? Until you actually go through it, it's hard to anticipate everything.
You, you've gotta kind of have been there, done that kind of thing. And that's what Gremlin brings to this. So it's an excellent thing for people who wanna go check it out.
Colton, what's their best kind of on-ramp for this? com. We've got links to, you know, we'll have it, we have it up on the main page.
We can tell you details about it. We've got a free trial. You can go in and play with it yourself.
And of course you can reach out to me or my team. We love to not just show people the product, but guide people in how to go model these exercises. How to run them safely and effectively and partner with them to make sure they're building not just, uh, a chaos engineering experiment, but a reliability program for their company.
Absolutely. You know, too many people, Colton, treat VR as a checkbox, you know, cyber insurance company asked, do you have a DR plan in place? Yes.
Have you tested it? Yes. You know, and That's far as it goes.
How often is a tabletop exercise or, you know, maybe once Yeah, we did a year ago Or five. You know, there, there is some of that too. 'cause not everyone is in Amazon or a Netflix.
There's a lot of, you know, small medium businesses who have a, a data center or some colo rack space. Well, as you mentioned, if you're hosted in the cloud, as, as a lot of us are, uh, look, it was kind of a rough end of 2025 in the cloud. We saw Yeah.
Major outage with basically every cloud provider. So yeah, the truth is, this isn't like wishful thinking. This is something that happens regularly and you can turn it Into an event.
It's not if it's when Yeah, it's not. If it's when I get you, Colton, thanks for coming on here and telling us about this. What do you call it?
Is it called Gremlin Disaster recovery testing? Yeah, gremlin disaster recovery testing. We, we went back forth on a name, we had another internal name and we said, look, this is what people run, what people call it.
Let's descriptive, let's just call it what it is. I love it. I love it, man.
Colton, good luck with this. I think it's a great product. You know, I've, as I said, I've firsthand knowledge on this kind of stuff of where it comes back to bite you in the butt.
Man, that bites heart. You know, it'd be a fool not to do it. Um, keep up the great work.
Keep us posted. You know, the clock's running now, right? It's end of January or beginning of February, actually.
Um, we'll hopefully see you back here by April. Yeah, that's the plan. I love it.
All right. Hey Colton, be well. Colton Andrews founder Gremlin here on Tech Trunk tv.
We're gonna take a break. We'll be back.