Emergence from Stealth – Dev Nag, CtrlStack
Mike interviews Dev Nag, CEO of CtrlStack to discuss their emergence from stealth and how they are going beyond observability data to more effectively discover the root cause of application issues.
Transcript
This is Textron TV. Hi everybody, Mike Rothman here general manager text wrong research with another Textron TV interview. I am really excited today to be joined by devnag.
He is founder and CEO of control stack. They are recently introducing their company to the world coming out of stealth my goodness stealth for so long. It's it's hard right because you know, you start a company and you know, you want to tell everybody how exciting it is and what cool stuff you're up to but you're in style so you're not allowed to do that.
Unless you're a customer then they're happy to tell you to have welcome to the show man. Thank you. I'm doing really well.
Thank you for having me. That's yeah, you bet you bet. So what are you know, we don't make any assumptions anybody knows who you are man.
Oh, you have a long history and you know kind of development and devops and a lot of stuff. It was a little sense of who you are where you've been and and kind of a little bit about what you're planning to do with control stack and then we can kind of dig into it a little bit. Yeah.
So I was a pretty early software engineer back at PayPal, you know, 20 years ago Google 15 Years ago with the SRE kind of movement starting there was at eBay at VMware working with Enterprise companies. And then I started away front back in 2013, which was kind of the first really like, you know, kind of cloud native observability vendors. We had some great customers like Lyft and Snowflake and box and Groupon OCTA workday and so forth, but I saw a pattern towards the end of where fun before we got acquired by VMware, which is the same kind of gaps in the workflow that we had to kind of like, you know fill in with human effort and we've seen that across the industry.
We've seen like a lot of like, you know stats that showed that the kind of traditional observability stack of melt, you know, metrics and events logs and traces have not actually solve the problem that devops kind of set out to solve which is like in a delivery more customer value faster and easier and cheaper and so, you know looking at the data forward Like we did away front like, you know kind of only goes so far. Well, we do a control stack is look at the entire process backwards and work backwards to the data you need and all the other stuff to you all the ways to automate the process and investigation and deployment and understanding like what's actually happening inside your system from cause all the way to effect. So the general problem is that devops right and folks and it's one of the things that we're talking about, you know right now is we're doing more things for 2023 now show up at in our Trends and I'll show up at the the predict conference that we're doing and early January but you know our theme there is is really long with devops, right because everybody's talked about devops everybody, you know, and and then the last year it kind of felt like they were other things people were talking about other things and in my experience, right, you know, 30 plus years doing technology stuff.
It's when folks stop talking about it is when folks actually start doing it, right and what it's not, you know got all hyped up and everything like that. You know, that's really when we start to see, you know, kind of the knee and the the real adoption curve there. So so we're expecting, you know, kind of some pretty significant and continue to adoption for devops.
But the point that you're making in the problem that control stack is there salt is that we only have partial Ability and even when we talk about observability, right and traces and metrics and a lot of other stuff. Well, there's still a piece of it that's missing relative to how long our developers are spending trying to isolate root cause right trying to understand, you know, kind of where the issues are and that's costing, you know us a bunch of money. I know you guys did a studying you relied on some other, you know, folks study to to Really quantify how much that's costing developers and and what kind of productivity hits that are, you know, can you share some of those numbers?
Yeah. So we saw like, you know, what you you understand from like the surveys is that the more devops technique you actually adopt the higher the per capita tax on your team. So what we saw is like, you know folks who we're not doing devops at all.
We're very early in the curve kind of like, you know, just starting out very traditional Enterprises had a certain like cost per developer right of like, you know, help the production stuff for it, but as they started to do microservices or kubernetes or like you continue to appointment right those folks had an increasing tax on each developer and at the And either for the folks deploying at least once a day, we saw that you know about 72% of them said at least half their injuring time spent doing troubleshooting or debugging like you're higher Engineers. They're costing a ton of money. They're something like, you know hardest folks to get in the world and then you bring them in and the very first day they're chopped in half right productivity wise they're spend all this time, like working on kind of maintaining the path that's supposed to build in the future and that's like, you know, obviously like a very unfortunate like kind of outcome but also the fact that it's like so like, you know kind of increasing over time as you like adopt more devops is a bad sign.
We're gonna hit a wall here, right? You can't have a hundred percent tax. You have no time for future velocity here.
So yeah, we have to fix it up on volume, right? You can't make that up. The match just doesn't work there.
Right? So we have to actually solve this problem like right now. It's like, you know, very attractable.
It'll become impossible very soon. So what we're saying is like, you know, the the demand on these off teams is actually like, you know Rising yet comes like a machine scales kind of exponential going as as fast as application complexities, but the supply the resources of people and traditional ability cannot keep up, right the Gap is getting bigger. And so you see this tax you get like, you know kind of more agile and dynamic and complex applications.
That's right. And you know it to be honestly different. It's a little counterintuitive, you know, because we've been told by everybody that you know devops it helps you move faster if there's a whole bunch of automation into the process it it reduces and removes some of the silos that we've had between a lot of these different teams right there.
And then that's kind of the story right? That's the devops story maybe the mythology on that. But what we're saying here is is that it actually, you know, again depending on how you structure and really kind of the visibility that you have within all aspects of your pipeline.
It's not exactly that right, you know, you get to a point where there are enough built in inefficiencies that this tax really does start to both slow down the process go up the works right wrench in the wheel how well whatever metaphor you want to use but it's one of these things that again they don't tell you and devops right? They don't they don't you know, say oh Is all these things and do your drink into your getting lamb or your you know, kind of get home, you know type of thing and and all this stuff goes away and it's Kumbaya Everybody by the Fireside right? But but again in reality, you know kind of there are kind of impediments right there are you know kind of headwinds that you have to deal with and and what control Stacks trying to do is is take a little bit of a different perspective on that data collection a little bit different perspective on the analysis.
That really does isolate right what the core issues are in the pipeline and really help you bring one that process did I get that kind of right or yeah, I think really I think like, you know devops the aspirations are great. We still have the same aspirations. We did like, you know, five 10 years ago when the term kind of got bigger, but we haven't lived up to it yet.
There's a lot of gaps in the process and The data hasn't helped as much as it should have right so we kind of focus on data. I did this myself away from like I was a fantasy to have a very like, you know large other building company, but we focus on the data with like blinders on we didn't look at the process that it can fit into if you look at actual devops engineer, they have a whole sequence of events in a tools and like workflows that they go through they're looking at data for sure. They're looking at all the Melt stuff metrics events logs and traces but they're also taking actions.
They're like, you know doing terraform and going to the animals console and do guys stage and they're socializing they're taking their knowledge and there's kind of sharing with other people. Like what are you seeing? What am I seeing?
What are you doing next coordinating to make sure they're not stop each other. And so the whole workflow is actually much more complex than just the data alone. That's like the windshield but they actually are driving a car every day.
Right? Look at the whole car the whole workflow the whole process work backwards and I think that's what's missing. That's why I think the aspirations have not been lived up too is because we're not seeing the whole process and optimizing across the whole thing.
And something's a devops. If you do a little bit of it, it actually is much harder much worse than doing none of it or doing all of it. Right?
And so like that's I think the lesson that we've learned last year Purgatory right never build a campaign around that and that's so let's get into a little bit right? So, how do you on board into control sack? How do you know?
Where do you get the Telemetry from you get kind of how does it really work at, you know kind of that tangible level that folks can get a little bit of a better fence for what you're doing. That's different. Yeah.
So, you know one thing we saw like, you know during away from that we saw I think the last 10 years is that as new kind of data types came along people can't Proclaim. This is the new day type that eats everything right like metrics lead everything. No, no structured events lead everything on a trace Elite everything and actually we found that wasn't true.
We need all that data. We need all those different data types. There's no like one size fits all as far as it goes with data.
They all have different like in a pros and cons. So we don't actually replace metrics and events and logs face. Those are great.
They're necessary but they're not sufficient, right? We need more than that. And so what we do our two things so we looked at like, you know how people actually run You know applications how they actually troubleshoot in debug.
This huge tax is like, you know, troubleshooting debugging like what people actually do there and they do two things that are really interesting. So one thing they do. If you look at the win right when something happens they look at what just happened before that.
Right? What we're changed is that were made. They like scour slack chat messages.
They look at like logs. They look at like, you know other people's like kind of like, you know change systems. They look at ticketing systems.
They try silly what could have been the different input that caused a different output right? So that's kind of the one side of things just kind of like, you know, what happened nearby time the other part they look at is the where right if you see a symptom like what happened nearby like what other components upstream or Downstream like just a couple hops away might have had like some more symptoms that could be the root cause so if you see like a weird like, you know, drop in transaction rates, maybe you go like a hop over to like another microservice to another service like down to like maybe a third party service or to a data storage and you're like, oh I see the cash is actually having problems like much more misses than usual the traffic data shape change, right? So you're like actually Traverse is like almost a graph in your mind.
Right? Like here. I think it could be let's look at like one hop to kind of find in this like localized area.
Like what's happening? That's the way of things if you think about the wedding the wear observerability doesn't actually capture either of those things. They don't capture all the changes happening.
That's like Other system those aren't usually typically like, you know stored as metrics or events or logs and like the typical like, you know off stack and when it comes to the where they don't make graphs outside of traces at all, right, the traces are great for request traffic but almost all dependencies are actually outside of that. Right? So like most things like if you have a process sorting container that's not a request relationship is actually a different type of relationship.
And so we have to store it in a different way. And so what we do a control stack is actually we reflect the when and the where in data like on top of the Melt data itself and say Here's how we tie the data together, right? So like you have a metric here and event here in a log here, how do they connect together?
What's that like traversally you're doing your mental model? And how do we put it there in the the software itself? Because once you have it in software in the data itself, you can actually write code to Traverse the graph for you, right?
You don't have to have like someone who's an expert understanding how the system fits together. You can actually have code walk the graph for you and find the root cause much faster. And so we're not sitting here.
You're not trying to replace. Data dog or Dino trays exactly those or wave front. Yeah, you kind of have a lower weight, right, you know and and you you layer on top of that really, you know asking a different kind of question and doing different kind of analytics using graph analysis to be able to isolate root cause much faster exactly.
There's a whole like, you know, multi decade long like, you know kind of history of graph algorithms that are fantastic for this kind of like, you know root cause troubleshooting and Analysis and it has really brought to bear. So there was an idea for a while called dependency mapping but they were really like independent. They're kind of style it off and they're often mainly created like their diagons of like, you know, this component affects this component, but they were hard to keep up today hard to create hard attention data.
So we do is once you enter controls like takes like, you know, five minutes you integrate like AWS and get lab and you know have like agents for kubernetes. We actually construct that graph every few seconds, you know in entirety. So we like bring together all the processes services like this is like, you know kind of third party Services into a giant graph.
We tied together events all the things that you're doing to change your application and your system. You know when a future flag has changed or when code is deployed. Those are events that have a when and aware we tie them to the graph.
And so something weird happens like a metric system like a spike or a change in Baseline. We can actually walk the graph back or say here's the most likely root causes because we can see it's right and nearby in time or nearby and space right here isolate when it happened and then look at all the other activities that happen right around that exactly and all we're doing is mimicking what people already do was what we've done for like, you know decades just like what what happened you're buying time and you're buying space. We just haven't had the data to do that automatic encode and now we actually do so speaking of which how quickly does it take to, you know, kind of instrument it so that you know, you're feeding data to control stack so you can do some of this.
Yeah. So once you know AWS and you know your pipeline for like we get lab or GitHub and you have like the host agents or kubernetes agents that takes about five minutes within 30 seconds. You can see like those first snapshots coming into the graph you can watch changes flowing through the system.
It's exactly what they do. So we actually have two flows here. We have one foot which is like the the effect back the Cause right?
You have a weird effect a weird. I'm like, what caused that? What's I I Smell Smoke.
Where's the fire? We'll show you the fire and we'll show you like where the first match was. Right, but the other side is caused to affect right you make a change you're checking a code or deploying an image or whatever it is or you change your future flag what happened next right?
We can show you like from that change what happened throughout the system like what did like, you know touch what did like do as far as metrics or the impacted metrics and show you how the typology itself changed over time? We can play back that whole movie and see exactly like how maybe connections were broken or created? That do you see in some of your early kind of customer environments or folks doing it in fraud or they doing it mostly in Devon test to get a sense of hey, what are these impacts when I do make this change is a little bit of both, right?
So what what tends to be the use case that that drives folks the first started out there's a really interesting divide. So there's like two different like, you know placeops you can go so staging production. So staging actually benefits the most from like, you know, git commits, right?
So you check in code it gets out to staging right away. Right? What happened there for QA folks?
They can see like right away like, you know, here's like a weird change in the staging system. I can track it back to like, you know a commit or a small set of commitments very quickly and see what the problem there was. Now a production most people still don't do continuous deployment.
They're really like, you know batching up like, you know deployments like maybe every couple days every couple weeks. That's like, you know, that's still kind of I'd say the the typical practicing mystery for those folks. They don't care about the perfect stuff so much they care about the deployments a lot right when an image gets deployed from like, you know ECR into like, you know your services what happened next right?
A future flag is Switched what happened next right all these and careers events like, you know scaling pots up and down. Those are things that really touch production so tracking what happened after that and going backwards and forth between cause effect is the key value proposition there. Yeah.
Yeah good. Well listen Dad. Thanks so much.
You're saying you give a couple customers now you launching with when when you do that. So some folks saying good stuff about you. Yeah, we have a bunch of data about that and some some really fun like, you know other directions where we're kind of like exploring here as well.
So, you know, we had like this kind of like, you know, very strong team of folks. We were x-way front you came in been like, you know doing obserability for a decade for both sides of it. We also brought in a team that did machine learning at VMware.
So like, you know, we started this kind of like, you know, really Flagship machine learning product and VMware called you realize that cloud which really optimized just one small part of the VMware stack of v stand the storage array virtualization, but the approach was so generally like, you know, God dozens of paths for that technology. It can be applied to much much more and we actually had a bunch of breakthroughs after I left field right now that like you see across all Twitter, right? So like you see, you know, gp3 and stable diffusion these kind of like, you know generative AI approaches they're able to take like, you know, very high level, you know, natural language intent and create like amazingly complex artifacts right there create like, you know, long text paragraphs or they can create like images create videos in some cases that same approach like, you know can be applied very very easily to develops as well.
So we have a whole team work on those those purchase as well. That's cool. All right, Dad.
How do people find you guys out? I think to learn more about the company and the technology and obviously to you know, start playing around with yes. com.
com etrl get that ttrl. That's right. com.
That's fantastic. Well after congratulations, I know starting a company and getting to the point where you're kind of take the raps off and come out of stealth is is a huge milestone. So congratulations on that.
We're actually doing webcast together. That'll be on December fit. I believe so sign up for that.
com. So I'm excited for that. But yeah, thanks for showing up on text on TV.
We're really excited for what control stack is doing and look forward to hearing from you quite a bit in the future. Thank you Mike. And with that we will send it back to the studio.