Webb.ai Comes Out of Stealth – Manish Gupta, Webb.ai
Webb.ai CEO Manish Gupta joins Mike Rothman to discuss his new company, which provides SRE and other DevOps personnel with a deeper view of the changes in the environment to accelerate troubleshooting and proactively optimize the environment. They discuss how Webb.ai is complimentary to existing Observability solutions, and what it takes to onboard to their system.
Transcript
This is Techstrong tv. Hi everybody. Mike Rothman here, head of Techstrong Research.
And I, I, today's an exciting day. Right? Exactly.
Can, no other way to say it. Right. Very excited for a couple reasons.
One is I'm joined by an old friend of mine, Manish Gupta, who, honestly, I'm not even gonna go into how long I've known Manish and how much mutually assured destruction we could have if we started telling stories and the like. Uh, but I've known Manish for a, a long, long time and, and, and we're here to announce, uh, his new gig, uh, the new company that he started up with, uh, with a partner or two. Uh, and Manish, again, longtime experienced entrepreneur, uh, has, has been there, done that.
ai. Manish, welcome to techstrong tv. How are you today?
I'm doing well, Mike. Uh, it's great to talk to you and, uh, thanks for that kind introduction. Yeah, man.
So, so, you know, you're, you're not for a long time, right? And we've been just, again, in and around each other's circles for, for many years, mostly in the security side, right? Yes.
So, um, you know, now web AI is targeted a little bit of a different, you know, kind of use case. So one, why don't we go through a little bit about, you know, what the general problem is that customers have, uh, you know, what web AI is really gonna help them do. Uh, and really then maybe we can dig into a little bit of the differences between, you know, when you, you get into a troubleshooting, uh, uh, an app ops type thing, an AI ops, you know, type thing, and security, right?
I mean, security can be pretty grumpy. So, we'll, we'll, we'll table that one for, for later on, but I would guess it's a little bit of a different community and constituency than we, we, we've been dealing with for, for many moons. So first, let's start, what, what's the problem and, and why, you know, web AI at this point in time?
Yeah, absolutely. So when I started this, uh, company, I was actually thinking about this mostly from a security point of view, Mike, uh, like you said, uh, Mo both our experiences have largely been in security, but as I started digging into the problem deeper, and it became clear that actually the problem was far bigger, and the problem in essence that we are solving now is if you look at DevOps, if you look at s e site reliability engineers, they have the TA task of making sure that the environment is up. I mean, for SaaS companies, what's more important, right?
Then uptime is revenue. It's lost customers when you have downtime. Now, if you focus on that problem, then what becomes very apparent is that the modern cloud platform, as we like to call it, um, which is sort of your distributed cloud architecture, is highly dynamic.
Uh, we of course, are all very aware of code changes, uh, that our developers make. Uh, you might have a two week sprint, you might be developing even and deploying even faster. Uh, so core changes is one, but then what is very, very fast pace of changes, infrastructure changes, especially with the wide adoption of Kubernetes and this concept in Kubernetes called Kubernetes Operators, where we have things like v p a, uh, vertical Port Accelerator, or h p a, uh, all of these tools are software that are making changes.
Yep. And then finally, you have config, uh, for things like LaunchDarkly or configuration for a particular service that is being used by one of the Kubernetes operators to scale up or down that service. So all of these changes make this environment very dynamic.
Yet the first question that the SREs, the DevOps want to ask when they're troubleshooting is what has changed? And today, there is no single product, there is no platform that is aware of all of the changes. So you have to start there.
Uh, but, uh, I'll pause, um, see if what other questions you might want to ask me, because, you know, clearly, like for example, uh, one of our design partners is making more than 5 million changes a month, that is more than 7,000 changes an hour. Yep. Yeah.
Right? And, and I think that is really the goal, right? I mean, you, you know, what we all wanted to do is get to this point where you have a developer, they do something right?
Make a poll request, submit some code, what have you, and it flows through the process to ultimately be deployed with optimally no human interaction, right? So all of those things are committing changes in some way, shape, or form. Sometimes it's the infrastructure, right?
Sometimes it's the application code. But when something goes wrong, that's where it starts to go haywire, right? Because you have just so many different changes happening all at once.
It's hard to, you know, regress. It's not like you have, uh, you know, kind of an opportunity to take, you know, kind of the whole thing down and go, all right, let's see, you know, what really happened here. And, and every time you make a change, you know, oh, when did that break?
And, and really to, to understand that. So it really becomes a totally different operational environment than a lot of organizations are used to. And again, the faster you wanna deploy, the more you have changes.
The more you have automation, the less human, you know, kind of touch that's in there, if, if any, again, the more important it is to be able to gather that telemetry and do the analysis so that you can, again, troubleshoot faster. That's really, you, you remember back, back in the day, right? You, you know, it was, we used that kind of have this whole thing about y you know, respond faster, right?
That was in the security lingo. It was always respond faster. And then my, my old partner, rich Mogul, uh, used to say, no, no, you know, faster's good.
But at the end of the day, you want to be able to respond better, right? You have better information, you have better, you know, kind of telemetry. You have better, I, uh, uh, a better idea about what actually happened, and you can make that more actionable.
So again, it's, and, and now you can get into what, what you guys are doing, but what we did is we took a lot of the processes that we were working on anyway, and we tuned them up to 11, right? You know, we just decided to go a lot faster than we had gone before. And that obviously creates all sorts of issues with, you know, kind of the existing operational motions that we had in place.
Yeah. And pro, you know, we want to sort of take this process out of it, and we want to automate, right? Because, you know, let's go back to the 5 million changes or 7,000 changes an hour.
Like if a system were to become aware somehow of all of these changes and just displayed this in front of an SRE needs garbage, right? You can't do anything about it. And so that is where we have innovated, where we establish what we call causation analysis, because we understand the infrastructure and the, all of these Kubernetes operators really, really, really well, so that we know that X happened because y happened, and then Y created Z.
And so this is causational analysis, not correlation, because as you can imagine, when you're seeing 7,000 quote unquote changes a particular hour trying to do correlation, even with a 1% false positive rate, you're gonna have like two many things that the SREs and DevOps will have to worry about. So the first step is to do the causation analysis, um, which establish this concrete relationship between changes. But then we also, uh, have built what we call the knowledge graph, because knowledge graph, think about it as a representation of your cloud platform.
How are services connected to the infrastructure and how the infrastructure connected to, to Kubernetes operators, and how are these all connected to sort of higher level constructs like changes and causality? Uh, because that therefore becomes the scaffolding that therefore becomes the, the, the knowledge graph, I suppose, that AI can now, uh, leverage to identify and conduct what we call continuous automated root cause analysis really said that is the end goal, right? As the world today consists of continuous integration, continuous deployment, continuous scaling, and up and down of the infrastructure because of customer load changes.
And in this environment, we can only start adopting these tools in mass volumes if we can actually do continuous root cause analysis. That's right. That's right.
So, so the idea is monitor a whole bunch of kind of sources, right? That happen up at the application layer all the way down into the infrastructure and hypervisor, you know, I guess that's an old school term, right? But the Kubernetes layer, lot of the orchestration, uh, idea, you know, kind of understand and and profile what those different environments should be looking like.
Look for situations where, you know, something is amiss, bring that forward, do some additional, uh, I guess analysis to, you know, really try to understand the extent of the issue and hopefully deliver something that is both contextual and actionable to, is it the s r e? Is it the application developer? Is it Yes, it could be, you know, both at any given time and, and really make, make them more effective at their job?
Yeah. No, our primary customer persona, primary user persona is s r e or DevOps. If an organization doesn't have s r e, but we do find, uh, uh, that eventually developers will want to use this product as well.
Why? Because when we talk to developers, they tell me that, Minish, I'm aware of the change that I made, but I'm in a team of 200 developers and I do not know what are the changes that all the 200 developers made, right? And when did they get deployed?
I mean, I can go into get and look at the core changes, but I can't really know what changes were deployed when, right? Um, and for example, I might actually start seeing bad metrics or metrics getting worse from my particular service, but that could have happened because a northbound dependency changed for me, right? Um, so really the two emerging use cases, as we talked to our design partners, and they're starting to deploy our technology that are emerging is number one, as you pointed out, is troubleshooting.
When something goes wrong, we want to be the destination of choice, uh, which provides the entire storyline from service layer to the infrastructure layer to the configuration, and ideally should be able to provide you things like, Hey, look, the VPA updater, uh, scaled, uh, scaled down three pods for this particular service, uh, for, because some configuration was changed, and then HPA scaled down another three pods for this particular service for whatever configuration it had, and suddenly there were six pods for this particular service that went away. And for that small period of time, which is a couple of minutes, let's say it started giving a high API error rate, right? Otherwise, how do you debug these things, right?
So that's one use case, which is the reactive monitoring. But the proactive monitoring use case also, uh, comes to the fore, which is, so imagine you're upgrading your developer and you're upgrading a service today, what you will do is you will go to an observability platform and you look at the metrics of that particular service, awesome, right? ai.
Yeah. So I think you just kind of talked a little bit about where I wanted to go next, right? And that's the existing observability companies, right?
They gather a whole bunch of telemetry, they're storing a whole bunch of stuff. So what are they not doing? Where's the gap?
And, and you started to allude to that a little bit, but let's really kind of hone in on it, right? Because every time you have a new company, you know, you're innovating in some white space that the, you know, bigger folks have left for you. Right?
So let's explain that to, to everybody. Yeah, absolutely. So first of all, we don't compete with observability tools.
We are complimentary. You do want to use an APM solution to look at a service level, especially because of instrumentation. You can create custom mess metrics, business metrics that really no other technology allows you to do.
Um, and so by service, by service visibility in an APM tool, yeah, it's still the best technology by extension. Therefore, if you have 200 services, it is still the best technology to monitor each service in and of its own. But now where, you know, especially with orchestration platforms like Kubernetes, um, there is an opportunity to build a system that is aware not only of the service layer, but the infrastructure layer and the config.
So you get to see the whole enchilada in one go, right? Um, and so we can now start to, because this is the other part that I should mention is observability tools because of what they do, are really good at giving you a lot of data, right? If you take a, if you think about a particular APM interface, there are lots of graphs on the dashboard, um, CPU, memory latency, right?
But a human has to understand what those metrics mean, right? On the other hand, I think now, especially because, you know, many APM companies started almost a decade ago now, there is an opportunity with leveraging AI and leveraging some of the newer graph technologies, um, that we don't have to provide our user persona with just data. We can now provide insights that I like to call across your infrastructure services and config, right?
Because, you know, as I've talked to more than a hundred, uh, customers in the last year, Mike, one of the common themes that I hear from customers is, Manish, when I'm debugging, when I've been paged at 2:00 AM I don't want to see data. I want to get insights. I almost wanna be given action items, try this, try this, try that.
Right? So that is what we are aiming to deliver. Okay?
So, so really reducing that window of exposure, right? You know, helping folks, you know, really remediate and address some of these issues on the operational front a lot faster than they could have, right? Is there a, is there a kind of a performance or a reporting or kind of a, a, a backend analysis type of thing?
I mean, obviously when you, you know, kind of the brown stuff is against the wall and you're trying to figure out what you know is actually going on, I, you know, clearly there's, there's, you know, a huge urgency there, but folks are always trying to optimize their environment. Folks are always trying to, uh, again, make it work better, especially as you start to see, you know, kind of increasing load and, and imminent scale, uh, issues. So is, is there something for web, web i web AI say that 10 times faster, right?
Web ai, uh, in that case where you're helping, you know, the SREs really understand what proactive change, I think you mentioned proactive before, but Yeah, Absolutely. Yeah. So I think tying perhaps the two questions, the last of your two questions together.
Yeah. For developers, APM products are the destination of choice, because as a developer, I am most interested in my service. Yeah.
Right? Um, now for s r e for DevOps and for engineering leadership perhaps, right? But definitely for s r E and DevOps, it's far more important to understand the entire landscape, right?
The so-called cloud platform, um, this could be across clusters, this could be across dev and product. This could be across multiple cloud providers. This could be across multiple applications cuz consisting of multiple services.
So yeah, proactively manage, uh, excuse me, proactively monitor this. Like the example that I gave you about a quote, uh, upgrade, um, you know, is a good one, or another one that actually just came up yesterday in one of our discussions with one of the customers was, so they, they created, um, a health check, but health check was created by the developer for the wrong port service was serving customers on port 80 80. Health check was being done on port 80 81.
How do we find issues like these? Right? Um, and so the, it is proactive monitoring and it is also reactive monitoring.
The idea is letting software understand all the changes, extract all of the changes first from the system, do causation analysis, build a knowledge graph, and then do continuous root cause analysis. That's the other thing that I would want to add very quickly, Mike, is historically we've thought about root cause analysis is, Hey man, I got a big failure, I got a lot of downtime and I need to do rca. Yeah, that is appropriate.
But with 7,000 changes happening every hour, why can't we left? Why can't we leverage software to do RCA for every change? Yeah.
Right? And if it is, if this change didn't result in anything fine, I don't need to bother anyway about it. But if this change did result in some bad, some business metric going haywire, well I gotta expose that to you for the letting the human decide what they want to do next.
Yep. Yep, yep, yep. So increasing coverage of, of all of this, you know, to use a developer term, right?
Alright, let, let's talk about what it takes to, to do this. How do I onboard stuff? Do I have to point, you know, kind of telemetry to you guys?
Do you have, you know, some kind of script that you use in order to, uh, onboard into the system? So how long does it usually take to, to get this up and running? Yeah, so it sh literally shouldn't take you more than 15 minutes to get you up and running.
Um, so all we give you is a helm chart which contains, uh, two of our agents. One we call the traffic collector, which leverages E B P F technology that gets deployed on each of your nodes. And one is what we call a resource collector, which is essentially a pod level agent that gets also deployed in your Kubernetes cluster.
And that is it. Uh, as soon as you deploy that, it collects all the data that we need, it starts streaming the data into our backend, which is SaaS. Um, and you start seeing all of the changes, you start seeing the causation analysis, we can, we, we start to see knowledge graph and you start to see insights.
Now coming back to that knowledge graph where that, you know, think about knowledge graph as an extension of the concept of service graph that we, that exists in, uh, APM solution service graph, meaning, meaning I have a software microservice and I have an upstream dependency and I have downstream dependency, right? So that view, but now if we are going to expand this notion to cover entire cloud platform, then this notion of service graph has to be extended to add all of the infrastructure elements. It has to be extended to add all the Kubernetes operators.
It has to be extended to add all the position analysis, the changes in the higher level, uh, knowledge that we land up creating because of all of this analysis. Uh, and you get to view it because it is apt entirely possible, especially early on that there was a failure and we were not able to do root cause analysis for it, right? In which case you get to see this graph visually and you get to see, ah, this particular note changed.
Ah, I see Y X Y or Z downstream, or a, B and C upstream could get impacted. Yep. Yep.
Cool. All right. So available today.
You're, you're launching the company today, so that's very exciting. ai, they can sign up for, is it a free trial? Is that how you, you know, end up doing stuff?
Or do they have to talk to you and go through a whole big process in order to do that? What does that look like? Yeah, it is indeed exciting.
Uh, Mike, as you, you know, we were talking about before we started the recording, right? This is like giving birth to a baby. Um, and, uh, yeah, it is, it is our baby.
We've been working on it for almost nine months now. Uh, yeah, please, uh, for interested folks who wanna learn more, web AI is the domain, is the website. Uh, we have not launched our product as generally available today.
It is available as what we call early access. Um, and yeah, absolutely, if you're interested, um, there'll be a form where you get to submit your first name, last name, email, so that we get in touch with you, um, to make sure, right, that both of our, um, you know, what you expect from this early version of a product, um, matches with what we are Going to deliver. You bet.
You bet. So good. Fantastic.
ai really kind of filling one of the gaps in terms of how, uh, s r e and, you know, kind of more traditional DevOps operational engineers really understand what's happening in their environment. A little bit of reactive analysis, a little bit of proactive analysis, uh, really just, you know, again, kind of illuminating one of the big issues out there with, uh, some of the, uh, challenges of, of scaling up, uh, a new infrastructure, a cloud native infrastructure type of environment. ai, uh, and as always, Manish, great to see you.
Congratulations on, uh, what we know will be a successful launch. Um, and we'll see you next time when you have something else to do to talk about, which I'm sure you've got all sorts of stuff cooking in the lab right now, and it won't be long before you're, you know, doing another drop of really exciting features. Truly appreciate it.
Mike. It's great to get in touch again. All right, be well.
Thank you. And we will send it back to the studio for our next interview.