Gremlin Launches Reliability Intelligence to Enhance System Resilience with AI
Kolton Andrus, CEO & Founder of Gremlin, announced the launch of Reliability Intelligence — an AI-driven solution for analyzing and remediating reliability concerns in modern, complex systems. Through a combination of automated fault injection experiments, continuous resilience analysis, and a Model Context Protocol (MCP) server for LLM integration, Gremlin’s Reliability Intelligence decreases downtime and improves performance for online businesses.
Transcript
Hey everyone. Welcome back here to Textron tv. You know, I'm really glad, I'm really glad to welcome my next guest back.
It's been a, it's been a minute since he's been on, but welcome back. Colton Andres Colton, of course, is the CEO and founder of Gremlin, actually past and present, CEO, founder of Gremlin. He was always the founder, but took a little hiatus from the CEO role.
We're going to hear about it, but Colton, welcome back to Text Drunk tv. We've missed you. Thank you very much.
I appreciate the kind words, and, uh, it's always, I was, I was excited to talk to you today. I'm always excited to talk to you, Alan, so thanks for taking the time. Uh, our pleasure, man.
So, let, let's get right to some nitty gritty. You were, you found Gremlin's been your baby, right? You were founded, you were the c longtime CEO, you took it up to kind of the top of the chaos engineering marketplace, and then you ditched me.
I didn't, I didn't hear from you all. What happened to you, Colton? Yeah, well, we, yeah, as you said, you know, we, we came out, we pioneered the, the category.
We did a lot of hard work. We, we grew the business and we hit a point where, uh, I needed to go focus on the product and engineering side. And so I brought in a CEO to help run the business.
Uh, he did a great job. He was around for a few years. He helped us, uh, run the go to market side.
And I really, I went to town on the product. I think what we learned in 2020 was chaos engineering. Great idea, hard to execute, hard to really get the value out of it.
And that hard part was really the people, the process, the accountability, not really the tooling. And so I had to go back into the lab as, as, uh, the team has said, and I've gone and, you know, had to go focus on really making the product amazing, you know, not just a chaos engineering tool, but a reliability platform. And we did that and we got it in a good spot.
We got engineering humming along, and it was time for me to, to take back the reins of CEO because people like you were missing me. And I need to get back out there and make sure, you know, we were, we're spreading the good word. Absolutely.
And, you know, all all joking aside, Colton, you, you as a founder, you realize it, right? I'm a multiple multi-time founder. I realize it sometimes, Sometimes it, you've gotta be the person who carries the flag, right?
One of the most important, and I don't get it 'cause I'm not a military guy, but one of the most important positions in the military is the guy carrying the colors, right? Because that's what the troops rally behind, and that's what the market kind of focuses on, and that's where people look for signal. And if you don't have someone waving that flag up high, who people look up to or, or are looking for it, it kind of, you know, the morale of the troops goes down, things disperse.
It's just not, it's just not the same. And, and, and it, and it's not a knock on anyone. I'm not knocking right what the CEO over the last couple years has done.
But not everyone carries the flag. Some people are administrators, some people are, you know, very efficient managers, but they're not necessarily a thought leader. Maybe plain and simple, I don't know, a better less sugar.
I'm not sugarcoating it, right. Uh, I don't know a better way, easier way to say it. Yeah.
Well, and I think, you know, SRE, uh, reliability, you know, dealing with outages, it's not the most glamorous work. It's not always front page news and less things have gone horribly bad. And so I'm honored that, you know, you look to me as a flag carrier and I can help carry that flag because I think there's a lot of good work that doesn't get recognized.
And I think there's a lot of good work that could be done that would make everybody's life better. And if we don't stand up and, you know, advocate for that, if we don't justify it, if we're not there making the case and helping the business understand, hey, this is what's best for you and what's best for us, then it just doesn't happen, as you've said. And it's, it's a bit of a shame sometimes to just see, see things lull or, you know, regress when you would hope we're always, you know, making forward progress.
Absolutely. And Colton, I don't wanna make you note, Colton wasn't here. So the whole chaos engineering and S-R-S-R-E space went to hell in a hand basket that, that's not necessarily what happened here.
But, you know, even when we look at the broader SRE market, I, I think, you know, we've seen an interesting confluence of, of factors that, that come into play here. First and foremost is ai, right? And we're gonna talk about this, right?
How ai, and whether you're talking about generative ai, agent, ai, ml ops, ai, you know, what, AI in all its many flavors and forms has, has really made a, a, a profound impact on the SRE and, and chaos space. But then also another thing called this, I think the rise of platform engineering, understanding that, right? We, we have to have a platform that this factory that where we build software and maintain and run software is running on and, and the SRE space and, and chaos engineering.
But you know, as part of it, they have to also play on this platform, if you will. And so that's the rise of platform engineering, I think has had a, a, a role in this, the continuing evolution of DevOps, uh, all, all the above and more. But that's my take.
What's your take? Yeah, I agree with you on platform engineering. I think, you know, I love the factory analogy.
Uh, the Phoenix project is what really helped me. Mm-hmm. Think about software like a factory, and by, by the way, a little peek into what I've been doing the last few years.
I was fixing my software factory and making it run efficiently so that we were not just building cool software, but we were building cool software on time, on budget, all of those things. So yeah, Definitely that people want to use. Yeah.
Yeah. So the factory analogy, the, the rise of the platform, I think SRE has always been in a tough spot because, you know, we, we wanted to embrace DevOps and we wanted to ask the engineers to do a lot of this work. And then what we ended up with is kind of like a rebranded CIS ops team in some regards.
And I think that, you know, it makes sense from an efficiency and expertise point of view, but I, I see a lot of companies struggle because they can't have, you know, a handful of ses fix every problem in the company or look over every engineer's shoulder. And, and by the way, with ai, you know, turning out more code and, and, you know, helping velocity go faster, that problem gets worse, not better. Yeah, Absolutely.
Well, I think it forces you to say, look, if AI's gonna drive twice as much code as we did before, we need AI to help us deal with twice as much code as we did before. Which kind of brings us to this whole AI driven reliability intelligence that, that gremlin is, is talking about. Now.
What do we mean by that, Colton? Yeah. So I think one of our goals at Gremlin has always been make it easy to do the right thing.
And, you know, we started with build a great platform that has all the bells and whistles people need that, that's doing, you know, enables them to do what they want. And I think what we learned five years ago is that that chaos engineering platform is great, but you also need to go into the organization and influence the right behaviors. Leadership needs visibility into the baseline of where your reliability's at, the improvements that have been made.
If you can't measure it, it, it doesn't happen. And one of the things that always tears me up is there's a lot of teams that do a lot of good work, but if they can't quantify the work, they can't quantify the value for the business. They don't get credit.
And sometimes that means budgets get cut and teams get cut. And sadly, I've seen that amongst my customers, you know, otherwise it means people get passed over from promotions and they're not really recognized, oh, the system, no outages last year, guess we don't need our, you know, reliability team instead of really rewarding those folks. So we spent a lot of time scoring risks tracking, it's a lot like security, let's really measure it, and let's give leadership visibility.
So I think that addresses some of the organizational problems, but it doesn't really address that. You know, I take this tool to an engineer and I say, great, go run some tests and fix your system. Well, the first question they have is, what test should I run?
And so we did a bunch of work. Let us tell you that the recommended set of tests, we've got a stock set, it's the same 10 things that go wrong on computers. Let us guide you through that.
So, okay, an engineer runs the test and they're looking at their screen and the test failed, and they don't know what to do. And so that's where we decided reliability and intelligence really comes into play. Hey, you ran a test and it failed.
Well, you know what? At Gremlin, we've run a million experiments over the last decade. We got a pretty good idea why your test failed.
And not only that, we got a pretty good idea how to go fix it. So that's, that's the, that's it in a nutshell is we're, we're gonna analyze your data. We know a lot about your system, we know a lot about the experiments you're running, we're tied into your observability, so we've got a sense for the health of your system and how it's operating.
And then we are going to analyze that. We're going to say, based on what happened, here's the analysis, here's what happened, and then here's what to go do about it. And part of that is being really focused on credible, actionable outcomes.
You know, not just, oh, hey, the system broke. You should go figure that out. But very concise, Hey, this dependency failed.
You don't know it's a critical dependency. It either it is a critical dependency or you need to go, you know, make it a non-critical dependency. So that, that's really our, our goal is just take the expertise we've learned and build it into the product.
Now, to me, that's a bit of a slippery slope, right? We talked a little bit before we went live because he, he, there's a line between tools like Gremlin helping the SRE, the engineer, whatever the DevOps, do their job better, do their job faster, do their job, higher quality. Versus in today's world where we have this vision of some autonomous, you know, agent or AI empowered agent actually taking the place of the human, whether it's an SREA DevOps or what have you, where does, where's your thinking and gremlin's thinking on that?
Yeah, so my, my personal opinion, there's this great quote from IBM I've been throwing around for the last year, and it's, it's, uh, here in, wait, I have it up. 'cause it comes up. A computer can never be held accountable, therefore, a computer must never make a management decision.
So I think this is an accountability problem. If you want to fully automate SRE, the core of what an SRE does is make hard judgment calls. com on the side of I five on my motorcycle.
I had to make judgment calls in the fly that if you get right, everything gets better. And if you get wrong, everything gets worse. So I think, uh, companies and teams willingness to trust AI to make those decisions, I think we're years off of that.
I think we need a lot more comfort there. And so that's, that's kind of my, my split, I think, look, if AI is augmenting you to do something that you're not great at, to do better at it. So if you, if you never have done reliability, if we can help make you 50% better at reliability, that's a net game.
But could we just do it for you? And we've talked about this, the idea, the original idea of Gremlin was we're gonna release autonomous agents into your system to go break stuff and find the failures. And let me tell you what, that didn't market well in the first couple years of the company, if I bring that up to somebody, they, they panic the blood drains from their face.
They're like, oh my gosh, please do not. So uhhuh, you know, and I, and I agree, you know, it's like, look, if it's safe, if it's, if it's, you know, safety safety's key here. You know, we, we're here, we're about preventing outages.
We never want to go unleash something that might accidentally cause an outage. I, I agree with you a hundred percent, Matt. Hey Colton, let me turn back to Gremlin.
So obviously in your role as CCTO, you have, I don't wanna say re-engineered, but let's say refreshed the whole tech stack of gremlin here. And in doing so, I don't wanna say it fundamentally changes sort of gremlin's mission, but let's say it brings Gremlin's mission up to today. Not when, you know, just the thought of chaos engineering.
You had people saying, oh, I wanna be like Netflix, right? And, and so yeah, if it, if it works for Netflix, it works for Mikey, you know, and, and so they were doing it just for that reason. But we need more than that today, right?
We, we need to show ROI, we need to show, like, like you mentioned before, insecurity, when nothing happens, you do your, you're doing your job, but it's very hard to go get more budget because nothing happened. It's, it's very, it's a very similar argument here. Tell us today's ground lit.
Yeah, I think it all, what Is it about, It's, it's all about that que you know, it's the, it's kind of the classic, like if the tree falls in the woods and no one hears it, does anything happen? Like, if you prevented an outage and it didn't occur, did anything happen? And, and again, I've struggled with that.
I've, I've seen my champions and, and my peers struggle with that at their companies. They're doing great work. Mm-hmm.
They're now, I think there's two categories. Let's be clear. There's, I see a lot of people that wanted to do chaos engineering.
They wanted to be like Netflix. And they went out and they kind of ha you know, they didn't put the full effort into it. They had one team do it.
They tried it on a couple applications. They didn't build a process. They didn't build any structure.
They didn't build any repeatability. And then when it was time, when it was the business said, great, how are you solving reliability for the company? They didn't have an answer.
'cause they were solving reliability for a pocket of it. I think my best customers, this is one of the things we've learned, this has to be part of how you build software. And, uh, and it really, it's hard for it to be optional.
I think that's the other analogy from security is like, you know what, at the end of the day, security isn't optional and neither is reliability. And so we need to, I I, I like this term field of dreams. DevOps is what we've had for the last 10 years.
If you build it, they will come. If we build cool mm-hmm. Tooling, the engineers will arrive.
Well, I'm here to tell you, I wish that were the case. And it's not. The engineers show up where their bosses tell 'em to, and they work on what their project managers tell 'em.
It is important. So if it's not important to the company, then the engineer isn't gonna be given the time to work on it. And so we need to make sure there's time, we need to make sure, and then there needs to be follow up.
Hey, the Ss the SREs team problem, we'll just let them figure it out. No, who's the vp? Who's the C-level that gets called when the, when the service is down, when the website's down, who has to go to the board and report?
We had an eight hour outage that ended up on the front time and the front page of, of the newspaper. Those are the people that need to care. They need to be bought in and they need to see the progress.
And I think in general, for lack of visibility, those people, I'm sure they wanted that visibility, but they didn't have it. And so it was a lot of trust and a lot of flying blind. So a lot of what we built was that visibility.
Let's have a reliability score. Let's define your services in Gremlin. Let's track that score over time.
Hey, you started at a 50, you know, a score of 50, that's okay, everyone starts somewhere, but where are you a month later, three months later? Are you making progress or are you standing still? The other part is build it into the company.
Pro pro, the company culture. This is what I loved about Netflix. This is what I loved about Amazon.
You know what, when you told people at Netflix we're gonna do some failure testing, nobody balked. They were like, yep, that's a thing we do here. Sounds good.
We understand why it's a good idea. So getting that cultural buy-in that yes, this is a good use of time and yes, we're gonna invest in this, but, but also educating the business. We're investing in this to save the business time and money outages are expensive.
Yeah. They cost engineering time. They lose revenue.
They impact our brand. If we can just prevent those outages, we've saved the company a whole bunch of time and money. Well, how do we do that?
Do we need to drop everything and spend all of our time on reliability? Nope. We could spend an hour a month, we could spend an hour a week and we could make substantial gains over the course.
So, so a lot of what we built scoring, uh, a set of detected risks, there's a whole set of things. We can detect our problems before you ever run a test. So let's give you a low risk way to calculate that.
Um, we built something called dependency discovery. So originally we thought, oh, we really need to know how this, these pieces fit together to test them correctly. And originally we thought, ah, tracing solves this problem.
We'll just integrate with tracing. We don't even have to think about this. Well, as you probably know, tracing is not ubiquitous.
It's, it's, it's a hodgepodge and it's different services. So we couldn't rely on that. So we went and built our own, you know, network, uh, traffic analyzer to understand how these services were communicating.
And that was one of those like, hidden values of, uh, hidden, hidden gems of value. We didn't, you know, we didn't think we, that was a means to an end for us, but a bunch of customers were like, oh my gosh, I didn't know I depended on this thing. That's a huge, that's a huge learning for me.
So again, giving people visibility into the system, helping Shining a light's happening into what would, what was dark, right? Shining a light on what we're kind of dark, just you don't realize they're there until, until something hits the fan, you know? And, and, uh, it's important, you know, Colton though, one of the things that we see here, and as I mentioned, you know, we're part of future.
So they, analysts do a lot of research on like, spending on, in, on the software industry and how people are allocating dollars. You know, they, they, they just came out with a report, uh, uh, software, uh, spending would I think hit 300, almost $350 billion last year. It might go up to as much as 600 billion in two years.
It's crazy. A lot of money. Yeah.
But it's not necessarily net net new dollars that are going towards building, this is the enterprise market, by the way. Mm-hmm. Building enterprise software.
It's dollars that are being reallocated from other things, from people, right? From people spend from other areas within IT and other areas outside of it. Um, so I, I think what you're saying resonates that, hey, if you're gonna ask people to put their hard earned budget dollars, you know that people lost jobs over into building this software and running this software, this software better work, it better it, it better do what you're saying it does or should do.
And we need reliability and we need, we need to be able to put our finger on that reliability right. And point to it and prove reliability. It it can't be that reverse.
Well, nothing happens, so it must be good, right? I don't think that works in today's, it's, it's a tighter budget world than we've lived in before. And it's hard to get lucky that long is the truth.
No. Right? If this like stick your head in the sand, it, that's the truth is like those, those failures exist.
Would you rather know about 'em or not? And mm-hmm. And that's the key.
And so yeah, but like to your point about, you know, being good stewards of, of our customers resources, you know, I think that's the other thing we saw a lot of people struggle with the chaos engineering tools to really get that value. And that's a problem for them. But that's a problem for me.
I want them to get value. I want them to see the same success I saw at Netflix that we see at Gremlin. By the way we do this at Gremlin, I'll, I'll tell you a fun quick story.
2019, we were chaos engineering experts. We had the best tooling in the world. We weren't using our own tooling consistently.
We were doing it spottily kind of like everyone else. And one of the things I did when I took over as CTO is I made it part of my on-call handoff. Every engineer gremlin is on call.
Every on-call runs all the reliability tests that week or the ones that are scheduled, they follow up on 'em and they fix them every on-call handoff. We look at our scores and if the score goes down, the team knows I'm gonna ask 'em what happened. So they're already on top of it, they're investigating it, they're paying attention to it.
The side effect of that is we very, very rarely have outages. We're in the five nines plus territory and gremlin over our lifetime. And it's been super solid for the last three years because we've gotten so much better at this.
And we feel comfortable. Our all, you know, our, our team sleeps well at night because they know if something bad, you know, most of the dumb stuff isn't gonna bring us down. And if something really bad happens, we're gonna know about it.
We're gonna be able to deal with it. 'cause everyone's had some, some reps at bat, some practice. So of course the question is, are you drinking your own champagne or eating your own dog food?
I love that. Jeff Bezos Square. I was at that Amazon all hands where Jeff gave us, who are you?
Uh, uh, champagne bottle. Uh, I keep my magic cards in that champagne bottle actually in that champagne bottle case, but yeah. Good for you.
No, I think we're drinking our own champagne and, and I feel good about it. And it's like the new reliability intelligence stuff. We just launched my principal engineer for a month has been coming to me.
Colton, guess what I found? Hey, guess, you know, this dumb thing that was happening. We found the problem because reliability intelligence zeroed in on it and told us what was wrong.
That's what gives me, you know, a lot of confidence to go out to market. When my, when my team of really sharp engineers is getting, getting educated, finding value, then I know, you know, the vast majority of people are gonna find value. I love it.
Hey Colt, we're over time. I gotta I gotta wrap up. I apologize we didn't mention.
People wanna find out more about this reliability intelligence from Gremlin. What's the website? com That hasn't changed?
Hey Colton, welcome back. It's a pleasure to have you back here. I expect to see you on here regularly now.
Right? I'd love to. And uh, we'll keep this conversation flowing.
It's a pleasure. Thanks for having me, Alan. Alright.
Colton Andres, CEO founder, CEO again, and founder of Gremlin here on Techstrong tv. We're gonna take a break. We'll be back.