Kit Merker on Nobl9’s Commitment to Software Reliability | AWS re:Invent 2023
At AWS reInvent 2023, Kit Merker, chief growth officer at Nobl9, discusses the company’s commitment to software reliability, helping organizations measure uptime, set service level objectives (SLOs), and navigate software releases. Amidst the bustling atmosphere of AWS re:Invent, Kit emphasizes their multifaceted engagement, from managing the booth and meeting customers, to connecting with partners, AWS personnel and the broader tech community.
Transcript
This is Textron tv. Hey, hey, hey everybody. We are back here at AWS Reinvent 2023 in Las Vegas with another great conversation.
This one with Kit Mercer. Mercer, sorry. Kit, uh, who is, uh, chief Growth Officer with Noble Nine.
Welcome to be talking. You good? Good to be here.
Good to see you. I'm glad you still have a voice. Yeah, it's Been, we're all fighting that when, Uh, the Carnival Barker down at the booth, you know, it's hard to keep up with, uh, all the noise down there in the expo.
It is, yeah. The noise level pretty, pretty high. Um, so for folks that they don't know, noble Nine, tell 'em a little bit about the company, what you do.
Yeah. Well, no, nine is, uh, focused on software reliability. And what we really, what we do is help you measure the, uh, uptime of your services, whether you're using Amazon or any cloud, uh, any observability tools.
Let you set service level objectives, which are essentially reliability goals. Um, so you can figure out which services need mission critical five nines, which ones can get away with two or two or three nines, and really get everybody on the same page about serving customers, uh, managing cost, managing software releases, um, get really, uh, more, uh, management of software reliability as you're adopting cloud or, or running, uh, different digital services. Yeah.
One thing that's cool about being AWS is it is kind of the whole ecosystem showing up at once, right? Yeah. So if you wanna talk to customers, partners, A-W-A-W-S, people, investors, whatever press, you know, media folks like ourselves, what do you spend most of your time doing?
Are you mostly in the booth, you mostly meeting with customers, being growth officer? I kind of can imagine what that might be, but tell us a little bit more about your engagement. Yeah, I, I, I have the privilege of doing a little of everything, and so it keeps me quite busy.
We have a booth here, it's an awesome booth, booth, 10 90 smooth operations is our theme. Um, we've got card games, we've got finger surfboards we're giving away. We're trying to be pretty unique about it.
Um, also great demos and opportunity to, to meet the team there, get some stickers, et cetera. So we're, we have a, a really a, a good team here, technical people. Our CTO is here as well.
So really hands on. I've been meeting with, uh, pretty much everybody, partners, people from AWS, um, press. I've been going to parties.
I've been talking to people from the open source community. I mean, it's really a great chance to connect with so many people. And, uh, it's nice to be back traveling again.
Although, you know, Las Vegas is not exactly my favorite place to go, but, you know, it works and to get to see everybody. And, uh, yeah. I think one of the big things I've, uh, seen this year for us is the maturity of people's thinking on reliability in general and on service level objectives.
You know, in, in years past, we had people that were a bit skeptical and said, oh, this is that Google thing, and is this really what we wanna do? Now I'm hearing, you know, wow, this seems like something that could really help us. Um, we've gotten a bunch of credibility from the ecosystem.
You know, over the last few weeks we've did a launch with Cisco. Uh, we're, we're, I'm staying here next week for Gartner. We've got Gartner cool vendor this year.
We're nominated for the third year on DevOps Dozen, which is congratulations. Incredible credibility. And, uh, hopefully, you know, people will, uh, continue to vote for us.
We won two years in a row. This is our three-peat year, so let's do it right. Um, but it's become now a thing where you're seeing different vendors, different open source projects that are actually trying to not just do observability and try to see what's going on, but really understand what is the goal for our service, and how are we gonna manage the trade-offs between keeping the service up and running, keeping our team happy, keeping our customers happy, and also the bottom line of how we run, uh, services.
So I've been really encouraged by that, um, and really kind of impressed with the level of sophistication we've seen from customers stopping by and asking questions. Yeah. Yeah.
We, we, we've been in this cloud native world for a while. Not that everybody's doing it on all applications, but we certainly have made a lot of progress. But I think at the same, that's kind of the forefront of architectures and cool things that we're doing, serverless, et cetera.
There's also kind of fundamental blocking and tackling, right? There's a big emphasis around reliability, uh, resilience. Mm-Hmm.
It's a good conversation. You know, you kind of came out of the chaos engineering ish world. I'm not sure that people really refer to things that way as much as we we did at one point.
Yeah. But my observation is it's not about one point of solving one problem. I've gotta have context of what's happening all over my systems, my, you know, the horizontal and vertical kind of integration.
Um, similar, similar kind of thought that you're seeing. That seems to be the conversations that we're having. I, I think complexity is not going away.
And I think, you know, having the, the full ecosystem right, the right tool for the job, and not, you know, kind of looking down on people for not having perfect architecture and perfect systems. We, it's a, it's a messy world we live in with software, and there's been a lot of wrong turns and dead ends and, you know, different, um, techniques and vendors and, you know, all the change that's happened. And so, when I think about the, you know, the challenge for an IT manager or somebody who's really thinking about that infrastructure for an enterprise, especially for a large enterprise, it's a vast, you know, complicated system with all kinds of reasons.
And so now what they really need to be able to do is to understand the progress they wanna make and not be bound to an all or nothing approach. They wanna be incremental microservices might make sense in some context, serverless in others, mainframes in others, right? And so, what, what I'm trying to do, and what our team is trying to do is really meet customers where they are say, look, you, you don't have to be, you know, 95% cloud native to get, you know, a a reliable system.
How can we help you manage what you have and improve it over time? Set realistic goals, prioritize it appropriately. And I think that's really resonating with people actually, because, um, I think there has been a lot of zealotry over the last few years, and, you know, I've been, look, I was part of it too.
I was a Kubernetes guy right? Early on. So we were pushing for different technology solutions.
As we've gotten a little bit older, a little bit wiser, have realized that, you know, people can do things their own way. Mm-Hmm. And that's okay.
Actually, I, it's a good thing. We don't need every company to be completely on the cutting edge. They should, where it matters, where it's a differentiator for them.
Absolutely. And now everybody's flocking to AI as they, you know, rightfully should. Now to your point about blocking and tackling, we gotta take the boring stuff off their plate, right?
Make that simple, easy out of the box, um, so they can focus on things that are gonna be differentiated for their, for their company. It seems obvious, but it's not the behavior you always see from organizations. And we're increasingly, I'm seeing that that is what people are trying to do.
They're really having to, in this, um, more difficult economic environment, focus their energy, right? The only way that they're getting more, um, more differentiated value is by deprioritizing or finding good enough solutions to solve the basics. And that's an opportunity for us.
We wanna help customers kind of tell the difference and build the right level of complexity into their system. Mm-Hmm. Yeah.
You know, in that timeframe too, I'm guessing about the same timeframe. You know, we've, I mentioned Case chaos engineering before, but platform engineering. Yeah.
SRE we've had a lot of discipline disciplines kind of evolve, or maybe there's specialties around that. I, I'm curious, how has that affected how you work with your customers? Because I remember when SLOs were kind of a new thing, right?
And, uh, you know, book just came out about it, right? And, uh, you know, is this gonna get adopted? People can figure out how to use it.
It sounds like, you know, you're getting customers that are applying that. I mean, you know, the, the SO thing is your colleague, uh, Mike Ard will remind me all the time that SLOs have been around forever. Um, and, you know, in the context of using SLOs for infrastructure and cloud computing, this is something that came outta Google and it's been in books, and it was adopted in these kind of cloud native ish companies, but applying it to the way that people run infrastructure in general actually works, works pretty well.
And once people get over that kind of almost psychological burden, you know, hurdle of, okay, like, I'm not Google, but I still need to, to think about this. You see a lot of, uh, adoption there. We've also, the tooling has just gotten so much easier.
It's, it's point and click. It's, you know, um, uh, you know, we have analysis tools, so people are, are not having the same difficulty that we saw, but we're seeing, you know, great adoption. One of our customers that I'm very proud of is Ticketmaster.
Um, you know, they, they have really embraced SLOs as an approach. They're partnering with us. We've all become Swifties, uh, at, uh, noble Nine because of, uh, how great it is to work with, uh, Ticketmaster and make sure that their service is always up for, for, uh, ticketing.
And those, you know, those kinds of companies that really, you know, they care a lot about reliability, but it's not a simple thing. It's not, is it up or not? It's, it's really about understanding the shape of their business.
They can't, you know, do all things. No one can do all things at all times, but they know what's important. And so we help them figure out where they should focus their energy.
That is, to me, the, the really the critical, um, part of this whole story. So, um, so yeah, it's, you know, the chaos engineering and these different techniques, the sre, the DevOps platform engineering. I think there's a lot of semantic discussions and I, you know, I'm happy to participate in those for SEO reasons, but for the reality of it, right?
When we talk to real people, you know, there's usually some folks that are very concerned with reliability, performance quality, right? Mm-Hmm. They have different roles in the organization at different levels.
Those people, that's our, that's our people, right? They, they understand inherently what we're trying to do. And they see this as a way to help bring the whole organization onto the same page about the priorities.
That's, that's the key, whatever you call it, that is the key of what we're seeing. Yeah. Focus on the end result.
That's right. Um, how about the adoption model? How has that evolved over time?
And obviously we want to get time to value to show what your product can do and what, what's that, uh, like with number one these Days? Well, yeah, great, great question. We actually, um, I was, uh, hesitating to mention one of our great partners, Microsoft.
'cause we're here at, uh, reinvent, but since they, They have been around. Yeah. So we, we did a launch with, uh, Microsoft at, at, uh, ignite a couple weeks ago in Seattle just before Thanksgiving.
Um, and you know, what we, what we launched there is actually an integration with the Azure Monitor, which is their, you know, um, observability solution. And it showcases the best, uh, integration that we've ever built with a partner. So you, you can take Noble nine point it, it Azure monitor all of your resources, show up in the Noble nine ui.
You pick the resource you're trying to monitor, we give you a 30 day look back with an SLO analysis that tells you what the SO should be. Mm. One more click create ss LO super fast time to value obvious no guesswork.
And, you know, that is the direction that we're trying to take this, because even sophisticated users would like to have an easy way to get started. You can always tune it later. You can always go back and sure, and, and, and question it.
But we've really focused on making that user experience of, okay, how fast can I get to an SO that's workable? Uh, what kind of examples can do, I'll, I'll mention, you know, we didn't launch it here, it's been around for a while now, but we did launch with Amazon, uh, a project called EKG. It's an open source project outta the Box, which stands for, uh, essential Kubernetes gauges.
Uh, but it's also like a heart monitor. Uh, and it's basically a set of SLOs you can run on EKS out of the box. Very simple, easy to get started, no guesswork.
That's really kind of the direction we're going. So from an adoption model, look, we've got a perpetual free tier. You can buy it on marketplace.
We have a teams edition that's very inexpensive to get started. The scaling, you know, economic units super clear, uh, the value's clear. So you, you get started quickly.
You see the value quickly, you scale with it and not get sticker shocked like some other people. You, you may have heard in the observability space, sometimes sticker shock shows up. We, we avoid that completely.
It's very important to us. Um, so yeah, adoption has gotten just a lot easier, I guess is the way to put it. Yeah.
Okay. Excellent. Um, I think you might've mentioned ai.
I'm just curious your, your, your take on the focus on it, maybe either now or how we're down the road, how you think ai, whether it's generative or, or not, how that might influence Yeah. Where you're headed. ai, and it's live, uh, on the internet right now.
So definitely go check that out. ai, um, partnered with Google on it. It's, uh, powered by, uh, vertex Palm two, but really gives you the ability to have a chat like interface and ask questions about reliability.
So that's an exciting project that is, is actually pretty cool. And, um, gotten some good, good feedback on it and good adoption on it already. It's a free tool.
You can just use it. Um, in terms of, you know, customers adopting ai, I've seen quite a few interesting use cases around one, one in particular, uh, that's come up recently is the cost of training. Mm-Hmm.
And I know that there's a concern in enterprises that that cost is gonna, you know, quickly spiral outta control. And so I know there's some people that are thinking about that as a, a metric to, to pay attention to, right? Cost of training.
Um, we've shown a demo in the past that basically would measure drift in a model as a SLO. And when that error gets too great, trigger a retraining as a way of optimizing that cost around, around the training. So that, that is a, an opportunity for sure.
You know, I had this, I think, discussion even maybe it was LA last night, maybe even talked to Alan Shimel about it, I don't know. But, um, about, you know, do we hand the keys to, uh, to the, uh, the AI let run our operations, right? Why do we need SREs when we could just have AI ops to solve it?
And the way I think about that is, you know, let, let's imagine that there is a future where, you know, it's maybe not a GI, but at least like, you know, a, uh, a GI, general intelligence more like, um, somebody you can treat as like an IT admin, right? But it's an AI as an agent of some sort, right? Let's imagine that that future exists.
Well, I would argue that to get to that future, one of the important steps is to give them a framework to learn, right? Them. I'm gonna anthrop ize it, right?
This, this agent, this AI agent in the future, you gotta explain to that agent what does good look like? What are the reliability goals? In fact, you need to do it much more precisely than you do with a human, right?
If I say, Hey, make sure the service is up and running, and I say, it should be always on, you know, I don't literally mean a hundred percent, you know, that you can take some liberties with it because users will refresh and things. Okay? But if I'm gonna do that for the robot, I gotta be precise.
And so I think the SLOs are a necessary, but insufficient step on the journey toward that possible future. There are people that are skeptics of this. I personally kind of lean toward being skeptical of it for a variety of reasons, but let's just assume that it's there.
I think this is a necessary step to get you there. So if that's where you're going, start with SLOs, right? Define the reliability targets you want today, use that for automation, basic automation or human response in the future, more and more, you know, automated AI runbooks that can take advantage of that data.
I think that's a valuable step along that journey. Yeah, I think many of us, I'm guessing you might also sort subscribe to the idea of it's about helping us be more productive versus replacing jobs. Yeah, sure.
I agree. And we move the cloud less people rack and stack, but oftentimes it's jobs change and they evolve. That's right.
And there's a lot of focus around coist and things like that. It seems in an operational capacity, there's a lot of opportunity for AI to, uh, work in assist capacity, not in a, I'll take it over and let you let you know when the world's distracted or whatever could happen. I, I think that's true.
I mean, I would, you know, one, one product I'm paying close attention to is Cody from sourcegraph. I think it's a really powerful tool, and I'm, I'm pretty impressed with it. Um, yeah, look, the, the, the copilots, the code whisperers all of these tools as if you think of them as almost like better spell check, right?
Better, better, uh, inspection and static analysis. I think that really is true. But, you know, expecting the AI to replace our, you know, our, our goal setting and desires and, and intuitions and other things, I think we're a long way away from that.
And, and I think there's also a responsibility issue, right? Like, I don't wanna let my car drive, uh, I don't wanna let my servers, you know, be responsible for their own patching. I, you know, I'm maybe old school in that way, but like that with, with greater automation comes a heavier reliance on human intervention in critical situations.
Mm-Hmm. This is the paradox of automation. Mm-Hmm.
Same thing applies in, uh, in ai, I think. And, and in these mission critical systems in particular, I think we gotta be cautious. And I, I think, um, it's, it is about not replacing jobs.
It's about changing, and we gotta get with the times. Like everybody's gotta adjust to this. We've seen it.
Look, it's a year ago that, that, uh, open AI launch, We just had the anniversary of, that's right, the Last few days. Chat, chat, GPTI don't think that they, uh, knew, I mean, based on what I've seen, I don't think they knew what was gonna happen, leaving aside the recent stuff. But like, even even this, the, the adoption, it's something that now people have embraced AI and, and asked all of these tough questions about what is it gonna mean?
And I think the it cloud computing industry should ask those questions, but we also know, um, as technologists, like the limits of these things and the, the time horizon. So I'm, I'm looking forward to learning more about where that, you know, what's possible, what's what works. But I'm, I'm interested in solutions that create real productivity and I'm less interested in science fiction type type.
Uh, That, that's for another app. Yeah. Those are fun too.
Yeah, that's right. Well, thanks for joining us. Uh, where can folks go see you at the show?
Where can they check out you online or in the marketplace? Yeah, absolutely. So at, if you're at AWS, come to the expo 10 90, we're right between PagerDuty, Grafana, and Splunk.
So interesting. The same place we sit in your stack. We are at the booth.
It's kind of perfect. Um, and, uh, we'd love to, you know, give you some giveaways, games, demo, anything fun. com is the best place to find the company.
Uh, we have a free tier, as I mentioned. We're also in the marketplace in AWS reinvent. You can get, get started there.
Uh, I'm on Twitter still, unlike anyone else. I know. I'm the last one on X Twitter.
I'm on LinkedIn. I'm easy to find, actually. So yeah, hit me up and I'd be happy to grab a coffee or go for a walk with you.
Yeah. Awesome. Great.
Thank Kit. Thanks for joining us. Great to see you and have a good rest of the show.
Hope your voice, uh, hangs in there with all of, well, all of us are hoping that. So we will, we with, I I don't know if this is our last interview. We might have one more.
So hang tight, uh, just in case we do. It's, uh, been a lot of fun bringing this to you from AWS reinvent and talking with folks like Kit from Noble Nine, lots of great things happening here and interesting conversations as well. So hang out, hang out.
We may have another one. I can't promise for sure. We'll see if it comes together.
But either way, we thank you for being with us and watching today. We'll see you soon.





