Signal Intelligence Helps IT Teams Move Beyond Dashboards
Signal Intelligence Changes the Observability Model
Mike Vizard talks with Ronak Desai, CEO of Ciroos, about why observability is moving toward signal intelligence. Desai says traditional dashboards have created too much cognitive load for SRE, DevOps and operations teams. Teams receive thousands of alerts, but they can only investigate a small percentage of them. Signal intelligence takes a different approach. It correlates logs, metrics, traces and alerts so teams can focus on the issues that actually need action.
The conversation explains why observability did not always deliver on its promise to reduce mean time to repair. Many organizations invested in dashboards, but those dashboards still required humans to connect the dots. That made incident response slow and inconsistent. Signal intelligence is designed to reduce that burden by using AI agents to reason across data sources. The goal is to surface the right problems, explain what is happening and provide evidence for each recommendation.
AI Agents Reduce Alert and Investigation Fatigue
Desai explains that operations teams can move from alert fatigue to investigation fatigue if AI only produces more analysis for humans to review. Ciroos is focused on taking that next step. The system can group related alerts, investigate causes and recommend actions. In some cases, teams can also give agents permission to take approved remediation steps.
This shift is especially important as software environments become more complex. Modern applications span cloud platforms, Kubernetes, databases, security tools and distributed services. Human teams cannot watch every signal around the clock. AI agents can monitor those environments continuously and highlight the few actions that matter most.
Better Support for SRE and Operations Teams
The episode also looks at how signal intelligence can help L1 and L2 support teams become more effective. Desai says AI-powered analysis can help those teams ask better questions, narrow down issues and act with more confidence. That reduces the need to pull senior SREs into every incident.
For IT leaders, the value is not just faster troubleshooting. Signal intelligence can also reduce burnout, improve resilience and free experts to focus on bigger problems. As AI becomes part of IT operations, the next phase of observability may be less about building more dashboards and more about delivering trusted, evidence-backed actions.
Transcript
Hey guys, thanks for the throw. We're here with Ronak Desai, who's the CEO of Cyros, and we're talking about how, well, observability is about to give way to signal intelligence. Ronak, welcome to the show.
Thank you very much, Mike. Second time. Last time we spoke, we had just come out of stealth, so great to see you again.
We made lots of progress and just to set the stage up here, gone are the days where you want humans to be looking at dashboards, because dashboards were too painful, too slow for our teams, SRE, DevOps, our operations teams. And in the era of agent tech, you really want agents to be looking at your logs, metrics, and traces to really figure it out what's going on, rather than just clicking through a bunch of dashboards. And so that's basically what we've been focused on at here in terms of how do we reduce that alert fatigue our teams have it.
Because today, if you think about it, our operations team see thousands and thousands of alerts coming their way. They can barely take care of fraction or percentage of those. With system like what we've built at Cyros, we're looking at all of that as a cohesive correlating without you configuring anything from a human perspective and figuring out what really needs to be looked at it.
Clearly, this applies across your observability stacks, across all of your data sources. So that's basically what we mean by signal intelligence. Rather than focusing on just the alerts, let's look at collectively all the inputs we are getting it and really help customers get to those outcomes which they are looking at it to say, "How do I resolve these issues?
What are the recommendations? " So that's basically where we've seen a great traction with customers. To your point, we've been talking about observability for a long time, and I'm not quite sure it gained the level of traction that it was intended, largely because there's too many dashboards and no one could correlate anything, so they didn't know what they were observing.
And when they did observe something, they didn't know what questions to ask. So on the first end of this thing, are we finally getting the point now where we can just correlate all the signals into some sort of common bucket that AI can sort for us, and therefore I'm not dependent as much as humans trying to make sense of a bunch of signals that may be conflicting? Yeah.
No, I think it's great question. So if you just look at it how the observability, and to your point, observability had made big promises. In fact, I used to run full stack observability and app dynamics at Cisco before I started Cyros.
And the promise was we'll help you reduce the mean time to repair. And by building all of these dashboards. In reality, what happened was we built millions of dashboards, expected humans to sort of correlate and create that cognitive load.
And so to your point, the real benefits or the promises of observability never materialized because we had so much siloed way of looking at this information. With a system like what we built it, what we call a distributed reasoning layer. So I'll set up couple of things which I made the observation in my prior life, which is there is no way you can bring all the data from all the components for your application stack in one tool.
This is not possible. And it's couple of reasons. And then just to make it concrete, you wouldn't be able to bring all of those in Datadog or Splunk or pick your favorite tool because of two reasons.
One is when you're bringing so much data, all of those vendors really focus on ingest-based pricing and building the dashboards. So it becomes super expensive for our end customers. But then there are piece of data which will never be in your observability stacks.
Most of the large enterprises never want to put everything in one basket. But more importantly, your Git data, your CI/CD data, your Slack channel information, your Confluence pages, those are never going to be inside the data lake, like what Splunk or Datadog, any of those tools will build. So what do we do?
The way we've tried solving this problem is to really build that distributor reasoning layer. A layer which sits across all of those tool, like the way humans sort of extracts the information out of different tools and make sense of it. That's exactly what we have built with Cyros AI Teammate.
It is makes sure that you don't have to copy the data. So it's a zero data copy solution. It surgically extract whatever insight you needed based on the alerts or group of alerts which we saw, and really figures it out what should be the action to resolve.
We build observability tools to really help get to the outcome and resolution of the incident. And this is where we are really focused on how do we get that average time, which is order of 200, 300 minutes in terms of resolution time to less than 10 minutes. By building that distributor reasoning layer, we are able to deliver that for a large Fortune 100 companies.
And does that reduce the cognitive load on the IT folks? And I ask this question because more than once, I've been told, "This observability stuff sounds great, but I have no idea what question to ask. " Yeah.
I'll pull the threads in many different ways. So let's start. By building AI system, you clearly reduce the cognitive load because you're not asking people to look at these gigs and gigs of data or hundreds of dashboard.
" This is the thesis by I'm concluding that this is really the root cause, and that's what our AI team does. So it definitely helped reduce that toil of gathering this information from very distributed data sources. So that's one.
The other thing which we've seen at, Mike, is clearly you have lots of alerts coming into the systems. And if you think about the load on our operations teams with what's happening with AI coding tools where things are getting deployed at a much faster clip. You're going into production at a much faster clip.
Think about MITRE, you're patching your software at a much faster clip. So when you're pushing changes, either config changes or new feature into the production, clearly, your systems are much more vulnerable at that point in terms of outages. So that's where we wanted to build the system.
But now if you think about limited capacity of human and the alert volume which is coming in because of all these deployments, if you were investigating every alert and giving those results to the end users for them to analyze it, you just went from alert fatigue to investigation fatigue. You solve that problem of alert overload. Because you correlated them by correlation logic or causation, you investigated group of those alerts, and then you give the investigation output.
" So where we've done a lot of work is we look at all of those investigations. There are places where we can take automated action, and this is what I call it, not only correlate or signal intelligence, investigate the issues, but auto-remediate. There are things which you feel comfortable and you build this trust on the system.
So think about for some of us engineer, you give the white list of permission saying, "These are the task. " So by doing that, again, you're cutting down on humans need of getting involved into those things. And that's basically how we are sort of helping reduce the toil.
And there are multiple steps we've done it to really cut it down because the expectation from CIOs and the leadership is we want to deploy, and if vulnerabilities are released by vendors, we will patch it overnight. Not within 30 days, not within a month, not within couple of months. They will patch it overnight because of the exposure they have.
So you mentioned MITRE and all the things that are going on with security. Essentially, in my mind, that's us revisiting a conversation about technical debt that's probably long overdue. But is observability the key to understanding what my technical debt is?
Because it's overwhelming as it is. It might be getting bigger as we go along with new AI coding, and observability is going to be the mechanism by which I gather, to your point, the intelligence that I need to actually decide what to do next. Yeah.
See, I think the fundamental issue I have, or fundamental issue as a industry we've had on observability is when you speak about observability, people immediately go to dashboards. And as much as I love those pretty dashboards, those are good for humans. Those are not useful when you want to move at a rocket speed.
So to your point, yes, observability can help from a technical depth perspective, but what you really want is your autonomous agents or AI agents to really surfacing those five things you should be focusing on it in terms of what are those concrete action you need to take to reduce the technical depth. For example, if you patch a certain vulnerability, you're going to remove whole bunch of alerts which are coming in from your security centers. Can I bubble that up rather than you manually going and verifying it?
And what I call it is our agents are 24 by 7 by 365 watching your environment. So it's able to bubble up things which, as a human, is just not possible to do it. One of the customer scenario, we were doing the investigation for an issue they were seeing outage on.
In fact, in that process, we figured it out that there was a command injection exploit going on by some third party. So we are able to look at the systems holistically, your applications, and actually figure it out the latent issues, the issues which are sitting in your environment, and you're able to bubble that up. And so yes, it is helpful from observability perspective, from a technical depth perspective, but what really we need to do is to really pinpoint those four or five things which human needs to do to get the maximum ROI.
Let agent do the heavy lifting of validating this. There is no solution for AI slop. If you're going to use AI coding tools and not review the code, and you're going to push it, you're going to deal with consequences which are going to be very painful.
So there is no solution for that. The only way to make sure is you have a good engineering discipline before you push things into production. And agents, automated way to look at all of those things is helpful, but you need that engineering discipline for sure.
And if the signals are a mess, no amount of reasoning is going to make a difference. Ultimately, I need some sort of reliable source to reason against. Yeah.
So the great thing with modern systems are, if human can look at that piece of data and can reason through it, our systems can absolutely do it. " To an extent it is true, but in reality, if your humans are able to look at that piece of information and able to make progress and conclude, your AI systems should be able to absolutely do that. If humans can do it, AI system will not be able to do it.
That's given. AI systems are excellent in terms of looking at the large amount of data and reason through it. Humans get cognitive capability limitations because we always take abstractions and you lose a lot of details.
AI doesn't have that problem. So that's where I see a benefit with something like what we've been doing it in the industry at this point. When people hear the phrase signal intelligence, they start to think about, well, that sounds like we're going to collect a massive amount of telemetry data, and we're already choking on the telemetry data that we have.
So how do we do this in a way that doesn't result in massive amounts of data being stored in some sort of data lake somewhere that I can't afford to maintain? Mike, you're bringing up a point which I hear from all my enterprise customers. There's no lack of telemetry.
There's no lack of data in all of those enterprises. " Absolutely not. What they are struggling is able to interpret that data.
So when we think about signal intelligence, what we are looking at it is, let's look at all the data you already done. You've done a heavy lifting, Mr. Customer, in terms of bringing all of this data, instrumenting all of these applications and infrastructure.
Let's look at that and make use of that. Because the output of that is today, millions of alerts coming your way. Because we really have a lot of data and a lot of telemetry.
The place where we are really looking at from funnel perspective is not from application or telemetry or data collection perspective. It's post that. So we sit next to a lot of those observability tooling, which are spewing out those alerts.
Those are anomaly detections which are happening. " We really need to collapse that into a correlatable alert. Because typically when incident fires, let's say, for example.
Let's take a good example. Your DNS port is flapping on your physical infrastructure. " So if you have 100 applications, you're going to see hundreds of alerts just because of that one incident or event happening.
Now, when you're seeing those hundreds of alerts, if you have built that knowledge graph, you can quickly correlate that based on the dependencies and saying, "Okay, actually, these 100 alerts is actually one investigation because they're all related happening in the narrow succession of time. " Instead of 100 investigation, we're going to do one investigation. That's how we reduce the toil on the customers.
But then after investigations, we correlate it again. So there's an alert correlation, there's an investigation correlation. The goal is, can I really get to actionable things which you can do rather than just giving you tons of data for you to understand and inspect.
How automated will all this get? Because if I trust the output of the signals and the analysis, can I just hand that to some sort of automation platform to go fix it? Or am I going to go and want to review that, or will it depend on the circumstance?
Yeah, no, I think that's a very interesting point, and which is where I was talking about saying, look, the way... And I'll give you this Tesla example. In 2014, '15, when Tesla came up with autopilot, if you think about it, it was follow the front car and lane change.
Now you see Waymo driverless or FSDs. Still human is in control. " There are places where you would still want humans to be looking at it to make sure those actions are correct, and you can remediate that.
So it's question of building that trust. It's not a technology problem. At this point, we have to build the trust on the system.
And the other thing which happens with a system like Cyros is, as you use the system, as you give the feedback, it continues to improve and understands your environment better and better. So your confidence score continues to go, your analysis accuracy continues to go up. And so we are early, as in from an enterprise perspective, where they feel like, let these agents take automated actions.
We are not there yet. They are saying, okay, these are the set of things, like for example, opening the tickets, attaching the analysis, and if the issue got resolved, close the ticket. Do not need to get humans involved.
There are pod restarts, needs to, pod needs to be restarted, take care of that. Port needs to be flapped, take care of that. So there are set of action which they will give us, say, okay, these things are less risky, and it can remediate the issue.
We'll let the agents do it. But the industry is moving really fast. I think last year, if you asked me this question, I would have said, no enterprise wants me to take action.
So we started with read-only access to all the systems. " So Industry is also maturing very fast and people are building more trust on the systems. Right.
Within IT organizations, we have site reliability engineers, and a lot of times they are considered the gods of IT. They know how to do the analysis, and they dive deep, and they can sort out all the relationships. And then we have what I call mere mortals, IT administrators, who are the bulk of the crew and do most of the day-to-day work.
Are we getting to the point now where the IT administrators can benefit more from something that looks like signal intelligence, and they can do more, and they're not always dependent upon an SRE to go fix something? Hundred percent, Mike. You nailed that one.
So, there is a clear ability for your, what is referred to as L1, L2 in the support organizations, to be a lot more confident in taking the action, rather than pulling the experts, saying, "Hey, you know what? This is what is happening. " They could look at the analysis.
They've got this buddy, the Cyros buddy, and it can ask questions and get it to narrow down the problem and take those action much more confidently. So, you're right. Need for SRE to get involved, and this is another big benefit which you see by system like Cyros is because there are very few experts who understand the full gamut of your applications and deployment.
And they are typically the ones who have the most amount of burnout because they are the ones who are getting pulled from all directions. Now, empowering your L1, L2. You're freeing up the bandwidth of those senior people to not get involved into smaller issues.
Unless until there is a big outage, you can spare that bandwidth. And that other thing which is happening, and then this has been something which I deeply care about, which is we need to make sure that we create the next generation of leaders who understand these issues well. Because 10, 15, 20 years from now, if we don't have people who understand all of these applications, we're going to be in trouble.
So, how do you educate that L1, L2 people to become better and better, and grow their skill set, and also improve their capability so that they can do a higher layer task? The solution like this not only gives them the recommendation, but explains how it arrived at that. So now you're almost teaching them live on the job that this is how you troubleshoot, this is how you do it, and they've got a dedicated partner with them who's going to answer every question they have.
" Because otherwise, we're going to have trouble. Because there's not going to be lack of applications or less applications. There's probably going to be more applications.
We need probably more people to be aware about, to see what's going on in their environment. All right. Well, folks, you heard it here.
I think the one thing that everybody in IT can agree on is that signal-to-noise ratio is way too high. The good news is help is on the way. Ronak, thanks for being on the show.
Thank you so much, Mike. It's a pleasure being here. All right, and back to you guys in the studio.