Meet Matt the AI-Enabled Reliability Engineer with Manish Gupta
Manish Gupta, CEO of Webb.ai delves into, Matt the first AI-enabled reliability engineer that automates troubleshooting to identify the root cause in less than 5mins. Matt troubleshoots alerts from Observability tools (e.g., Datadog), cloud providers (e.g., AWS), and infrastructure (e.g., K8s). Given developers spend more than 50% of their time troubleshooting and debugging, Matt enables developers to spend more time creating than troubleshooting.
Transcript
This is Techstrong tv. Hi everyone. Welcome back here to Techstrong tv.
You know, a as you've heard me say, I've been in tech a long time, 30 plus years, and along the way, I, I've had a chance to meet people and you know, some people you meet and you may never see 'em again. Other people, you know, like bad pennies, they keep coming back and they're involved in what you do and you run in parallel circles and parallel lives, so you get to know them. Over the years that's been the case with my next guest here.
It's my friend, man, Gupta, Manisha. And I go back, I'm going to guess it's 20 years that we go back. It's been a while.
And over that time, you know, we've both been in different roles, founding companies, co-founding companies. When I first met Manish, actually, he was an executive at, at McAfee when McAfee was these security company. Right?
Right. Pre Intel buying McAfee and everything, right? It's a different type.
Uh, since then, Manish is co-founded of several companies, one of which was called Shift Left, which we covered. I remember it was a early, early entrance in the DevSecOps and DevOps space. Orton, he's out with the new company he's gonna tell us all about and, and their first product.
ai. ai. Manishh, welcome.
It's great to have you on. Thank you, Alan. It's always great.
It's always great to get a chance to talk to you, excited about what we are gonna talk about today. Absolutely. Manishh, I don't think you have to give people your background, I did it for you, but let's hear about web ai.
What, what was driving this one? Yeah, well, so web ai, we are solving a very exciting, very complex problem. We are automating, troubleshooting, right?
So what does that mean? Uh, well, if you think about it, there are about 26 million developers around the world and, um, surveys, uh, show that a typical developer spends more than 50% of their time troubleshooting in deeper, which makes sense because you write a line of code once and you run it for the rest of its live. And so, um, if you do the math, that's like more than a trillion dollar problem.
Um, so leveraging the latest large language models, uh, we are automating, troubleshooting to find the root cause in less than five minutes. Wow. You know, and that jives, I've seen several surveys, manishh, that says, developers spend somewhere between 11 and 28% of their time coding developing.
Correct? Correct. They spend the rest of their time doing things like troubleshooting, debugging, scrum meeting, you know, slack channeling and, and everything else that gets in the way.
So anything that could let, and, and the funny thing is, developers, they just wanna develop, like girls wanna have fun as the song says developers want to develop. That's right. And uh, And also anything That facilitate that.
Yeah. And, and that is sort of where the human ability really shines, right? The ability for us to create, think about new cool things, features that we want to deliver to our customers, and therefore we wanna quote them into a product outcompete our competitors.
Uh, but, you know, over the last, let's say 10, 15 years, um, for all the right reasons, there's, there's been a lot of digital innovation, which has enabled us to develop and deploy faster, uh, DevOps and DevSecOps and cloud like AWS and microservices containers, Kubernetes. I mean, the list is endless. Uh, all of that stuff allows us to develop and deploy faster.
What that means is we are able to deliver our feature functionality much faster to our end customers, but when something goes wrong, everything comes to a screeching halt because now we gotta spend time to figure out what's wrong. Right? And that's the area where we thought, um, some innovation is needed in order to, uh, automate this as much as possible so that a, people are spending less time doing this.
Um, then they're, they can spend therefore more time, right. You know, doing more creative things. And B, this of course reduces, uh, downtime if you can find the root cause faster, uh, because in the sum of downtime, 90% of the time is actually finding the root cause only 10% is actually fixing it.
Um, and is downtime important? Of course. I mean, you look at Gartner, PON one, all of these companies have done studies where they say that a typical organization would cost them about 5,000 to $9,000 per minute of downtime cost.
I, I, I think that's, well, that's per minute. Okay. So it's 60,000 to 250 fat's about what I've seen too.
Yeah. Yep. I, I, I agree with you.
Agreed. The other thing, quite frankly about it though, Manish, is look, this is not a new problem per se, but you can't make wine before it's time. If you didn't have the AI available, you couldn't really solve this problem this way.
And that Absolutely right. You know, that's, it's all timing is everything in life, right? I I think we've hopeful learned that lesson.
Yeah, no, absolutely. Right. This has been a complex problem for the last 30 years.
Ever since computing has been around. There have been all kinds of folks who've done their PhD thesis on trying to figure out how to automate troubleshooting. And, uh, it has been a hard problem because at the end of the day, you need to reason like a human, um, yeah.
And that ability was very hard to quantify in hold. But with large language models, we now have the ability to reason. And, uh, you know, broadly speaking, there are lots of large language model or AI based applications that are training a model on large corpus of data so that when the next time after the training is complete, you ask a question, it is able to respond based on that training.
Uh, this is not a problem that we believe can be solved by training ba from, you know, training a large language model because such a data, such a large corpus of data on failures and the various, um, should I say, states of the system are absent. Um, and what a failure in your environment will be very, the same failure in my environment will be very, will have a very different set of metrics, uh, different set of states. So we can't train it.
And so the way we, we are trying to solve this problem is to actually mimic a human when troubleshooting. So, uh, you know, the first, for example, when, uh, an alert is received by a team of engineers, the one of the first questions they'll end up asking is, Hey, is this worth troubleshooting? Right?
Because people typically get lots and lots of alerts at a given day. Uh, and then once they've figured out, yes, this is important enough to troubleshoot, then the next question they land up asking is, Hey, what changed? Uh, and based on that, then they now have to look at multiple systems, GitHub for code changes, terraform for infrastructure changes, AWS for cloud changes and so on and so forth.
Uh, and all of this takes time, uh, and there's so much data, uh, that it becomes be, it be, it becomes beyond a human's ability to actually reason with all of that data very quickly, right? So that is what Mac does, Matt is the name of our product. Uh, we've given it a human name because it behaves like a human in order to troubleshoot.
Uh, it does the same set of things that I just talked about. Um, so yeah, and the results have been, uh, phenomenal. Uh, we've, uh, we took about 150 issues across the environments that we are deployed in today.
Um, and we manually analyzed each issue and found that we were able to find the root cause 88% of the times correctly out of these 150 issues that are infrastructure focused issues. Right. That's, and that's huge.
Sure. It certainly is. I wanna jump in you, so the, the, this first AI product you called Matt, correct?
MATT. So I'm assuming Matt stands for something? Yes.
Actually Matt doesn't, it's not an acronym. Uh, Matt is the name of a gentleman, um, who has been a design partner for us for the last 18 months. Uh, Matt is super smart, uh, and he's been instrumental.
Uh, so we, our goal has been to make our product, uh, mimic how that thinks and troubleshoots, so he, How cool is that? It's kinda like an AI version of Brent from the Phoenix project, right? Would it be cool to do that?
So I hope that is, is duly flattered by that. So let me ask you a question. It is interesting.
So you, you've tried to duplicate a human's thought processes and base of knowledge as an AI engineer, you know, persona. I don't think we've, you know, and I do a lot of these interviews, I haven't heard anyone doing quite that yet, like really basing it on a, a real life person's, you know, function. What, is there some kind of special sauce that goes into that you think?
I mean, how did you create an LLM of Matt? Yeah, so first of all, we haven't really created our own large language model. We are using OpenAI.
We could use, uh, Gemini from Google. Um, I think the, the, uh, what our innovation has been, um, like I was saying earlier, uh, one set of products actually trained the model and then leveraged the model to answer questions. Uh, in our case, we have used, you could call it chain of thought reasoning, um, which is, you know, exactly like I said, a human goes through a chain of thought process to say, I'm gonna do step one and then step two, and based on the output of step two, I'll either take step three or I'll take step four.
Um, and very similarly, this is really a combination of codifying, um, uh, using code, of course implementing this, but also leveraging large nine model. Because let's just take a very simple example. So we find, uh, there is an issue and, uh, based on our understanding of the changes, uh, that have been made just prior to this issue happening, because we automatically extract all the changes from a given environment.
Um, so we, we see that someone deployed a new version of this code just five minutes ago. Okay? So now we know that this change could be very relevant.
So this is all through code. This is not where large lag model comes in, but now we need to find whether this change in code could have resulted in this error. Now trying to quantify that in source code is extremely hard because you have to look at all kinds of, uh, possible situations.
So this is where the large lag model comes in. So we submit the a difference in code because we are aware of not only that the code changed, but exactly how did it change. So now we take that code and we leverage a large language model to reason with it to ask, Hey, can this change and code could have caused this issue?
Right? So, um, as you can see here is it's a, it's a marriage of, uh, large language model plus, uh, our I implementation in code or how are human thinks and troubleshoots. Yep.
Very cool. Very cool. Neesh.
ai, that's, where are you guys? Is this in GA now or you still better or? Yeah, so actually great timing, uh, for our conversation.
We go live on June 20th, which is Thursday, Which is today. Correct, correct. That is, that that is correct.
And the magic of tv, right? That's of tv. Yeah.
And uh, again, the great part is, you know, this is a product that is targeted developers targeted at SREs and DevOps. And so it's a self serve product. You can literally go to our website, sign up for free, try the product and find out for yourself how good is map at troubleshooting in your environment.
Excellent. I love it Manish, much success and, and you know, good times with web AI for you. I'm looking forward, we'll all be watching as this continues to develop.
Is there another person on your team that is for the second product here? Maybe a Joy someone, Lisa? That's right.
Maybe would over time, maybe we will come up with another product, but right now we've got a hands full trying to solve this problem. You bet. All right.
Manish Gupta, the CEO founder of Web Wbb do AI and their first product out. Matt, the first AI enabled reliability engineer that can automate troubleshooting and save valuable, valuable developer time. We're gonna take a break on Next Drug tv.
We'll be back in a little bit.