AI-First Observability: Tomer Levy on Redefining Platform Design
Logz.io CEO Tomer Levy explains why there is a need for an observability platform that has been optimized to first be used by artificial intelligence (AI) agents rather than software engineers.
Transcript
Hey guys, thanks for the throw. io. And we're talking about giving observability capabilities to, well, AI agents.
Turns out that they're gonna have to check out our infrastructure, our applications, and all kinds of things that maybe we don't wanna do anymore ourselves. Hey, Tomer, welcome Michelle. Thank you.
Thank you for having me. All right, so you guys just rolled out this new platform and it's designed for AI agents rather than people. What's the thought process that went into that and and what is different about observability being provided to AI agents versus humans?
Yeah, for sure. It's been, it's been quite a journey. You know, it's been a little bit more, probably a year and a half almost, you know, since we started with LLMs and like everyone else, we said, oh yeah, we have an observability platform.
Let's duct tape, uh, a chat bot so people can ask a question. You see it on every product you go to, right? Whether it's Kayak or whatever product you use.
And we did the same. It was pretty cool. It was not very accurate.
And then, you know, we, we, we realized a potential in it. And, um, late last year we moved into AI agents. So you, you know, you send, ask a question, you send it to do research, it wakes up at night and does handle alerts for you.
It's pretty amazing. But then we realized we're still duct taping. We're taking, you know, a UI that is designed for humans and we're trying to duct tape an AI agent to start and imitate what a human does.
And, and, um, late last year, we took a decision that we're gonna look at AI agent as the first tenant in this future platform, and we're replatforming the entire, uh, product to be AI agent first. And I can talk more about what it means that this was the, the, the notion behind it. Well, let's do that.
And what does it mean to re-architect the platform? What's required? So you think about whatever tool you use as a user.
Uh, you know, our users come to log over, they use to having something, and it's designed for someone with a mouse and a click or a keyboard to ask question. And then I'll give you a simple example. If you were to take the UI that we have, the dashboarding technology and try to feed it to the ai, it'll be very hard for the AI to feed it, because it's a lot of data, a lot of numbers.
The context window is very small. So what we did, we've done is said, okay, let's redo the, the dashboarding technology first. Let's make a great one.
That's always good. But, but beyond that, let's make sure it's designed to be compact. To design that to be API first, it's designed to be created by AI versus created by a user and designed to be read by ai.
So now when an AI wakes up at night and handles an alert, it'll go to a dashboard. It's very easy for the AI agent to read the dashboard like a human read the dashboard. It's very easy for you to generate a new panel when it does investigation.
So we're, we're kind of announcing a big part of that, this new dashboarding technology to be, to be AI first. AI first doesn't mean human exclusive though. 'cause if I understood you correctly, it is human readable, whatever the output is.
So how will these AI agents and humans kind of collaborate in the age of observability? Yeah, a hundred percent. People pe the people are not going anywhere anytime soon.
Uh, they still need to, you know, do all the work, the dashboarding, the workflow. So it's absolutely human readable. It's just designed from the, from the, from scratch to be AI first and then, and then we, we think they're gonna collaborate a lot.
And you know, if you see in every product we look, Hey, AI is gonna have a magic wand, it's gonna be amazing. Gonna solve all your problems, it's gonna write your code, it's gonna order your flight. You know, the US engineers who are a little bit, you know, skeptic and, and, and we, we take a very much a crawl, walk, run approach.
So I'll give an example how we collaborate. So we tell customers, Hey, one day this can wake up at night and take care of your alerts and do automation, and it's gonna be awesome. But today, let's start with a simple model.
Let's take the alerts that you're just very familiar with. Hey, this issue always happens. I know what to do.
Let's have the AI agent wake up at night, take care of this alert, but not do anything. Just tell the operator, the user, the junior engineer, Hey, this is what usually you do. Here is the playbook.
I've already looked and investigate. And I think that's the issue. And it's a recommendation.
And once you feel comfortable with recommendation, I take the most mundane one and automate them. And only then we go step by step and get confidence and we improve the AI and make it more accurate and you improve. So it's kind of a journey and a lot of collaboration ahead of us.
How will these agents get orchestrated or managed or supervised somebody, a human? I assume when, who knows, maybe it's another AI agent is gonna need to interact with these things. So how do I kind of, um, ensure that they are automating some sort of process that makes sense?
Yeah, it's all today done by humans. You'll have to say, Hey, if I get too many error, like, let's say there was a deployment and AI wakes up after every deployment and it looks, Hey, was this a good deployment of code? Did I break anything?
Do I have too many exceptions? Now? Do I have too many error messages?
If the AI thinks you have to, uh, do something, it'll tell the user. So me as the user, I can say, Hey, I want an AI agent to do deployment validation. I can turn it off.
I can turn it on. It's fully controllable. It's not gonna conquer the world just yet.
So it's, it's good. Are these AI agents ones that you have trained specifically on the platform, or can they be any AI agent that just wants to invoke the observability platform for some reason? So, so we use cloud as, as our infrastructure.
And, and we calibrated in that lot of work on cloud to make sure it fits what we're trying to do. So yes, we will have the ability in the future to, so customers can bring their own AI agent, it can work on others specifically with cloud. We have worked a lot to optimize it and to do all the work to, and, and you know, one of the biggest issue in observability, and it's, it's exists in other markets.
There's a lot of data like terabyte and terabyte of data every hour, every two hours or every day. AI agent has a very narrow window of context. So how do you take all the world and all this telemetry and, and push it into an AI agent?
And that's a lot of where our unique technology, so the work is actually done not in the agent as much as it is done in how you prepare the data to make sure the AI agent has the full view of the system. And this is where we, we feel like we, we made the most impact. We also hear a lot of excitement about things like, uh, the model context, protocol, MCP, and then there's talk of these agent to agent protocols that might sit on top of that.
Um, how will we achieve the level of interoperability we're looking for and, and just how, uh, integrated will everything get? It's a great question. Uh, we are actually, um, have a, have a session in, uh, in, in a week where we and another company, PagerDuty, we're gonna show how agent to agent works together, how our agent talks to their agent, how will it all works together.
It's it's pretty good. It's very innovative. It's really the cutting edge when it comes to MCP.
Yes, all the protocol exists and we're gonna launch our MCP soon. That, that, the real issue is that you don't see many of them actually interacting. And it's still very, very early in the market.
And I think they can integrate. But there is a, there is a race in the market, in our market, and I think in every market, who is gonna be the agent to rule them all? And who's the gonna be the one serving the information?
Everyone wants to be the agent to, you know, to look at all the other pieces and tell you, Hey, this is where the problem is. And they don't want to be the one just serving the data. So there is always this, this, uh, this, uh, question is where is the right place in observability to really match and correlate all the information to find, you know, what you're looking for.
So I, my my, this was a sh a long answer to say, you know, yet to be seen how they're all gonna integrate. It's in theory they're all going to, but let, we'll have to wait and see Will we be thinking about observability in a larger context? 'cause most of the conversations we've had so far, very pretty much, you know, software development, DevOps specific, but I need to observe a lot of things including security business processes.
So as this evolves, we'll, observability platforms kind of expand their number of use cases. Uh, uh, a hundred percent. We already see that.
Literally today I was talking to one of our customers and they said, Hey, our business people, because they can look at the BI that will be ready in the week, and they don't have the ski, the BI takes time, but they go to log Zion now they ask a question and they get an answer, Hey, how many people from, you know, went to our app from, you know, uh, this country using an iPhone versus an an, uh, you know, um, an Android. So that's a question that you can use log Z to ask today. But if you are a business user, you don't know how to run the queries and it's complicated.
So suddenly it exposes other things to, uh, exposes the platform to other, uh, users. And what I think also will happen in the future, in, in, if you kind of go another level up, what is observability? You take a bunch of data, you integrate into all these system infrastructure.
You, you bring the data into one place, you organize it, and then you have all these UI and alerts and dashboards that can be used, as you said, to cyber, to business intelligence, to finops, to various use cases. And I think in the future, these are all going to converge. And the the strength will be how much can you integrate, bring all the data, can you do it in a very cost effective way?
And how good are your AI agents to be able to, you know, solve and answer the question and do the automation I need to do as a user. So I can see a lot of, um, convergence. It's maybe five or 10 years from now, but I'm, I'm sure it'll happen.
You talked about being AI native. Um, one of the things that has struck me about AI in general is the amount of telemetry data being kicked off is phenomenal. And so do we have the ability to actually process all of that and make sense of it and store it?
'cause well, storage still isn't free. Last time I checked, so, so, you know, how will we handle all this high cardinality data that we're gonna have on our hands real soon? I think the data is only going to grow.
We've been saying it for years. Everyone has been saying it, and they're right. I think the, the, the strength of the, of a good platform would be the ability to process a lot of high cardinality data at a very, very low cost.
Because AI needs a lot of data. It needs the right data. So, so I think one, you need good platform, then yeah, in the market there are 20 other platform, they're all good, but I think they need to be very, very cost efficient because AI is gonna demand more data to be able to make thoughtful decisions.
So I think, uh, platforms that are very expensive are gonna suffer because people cannot get the value of ai. That's number one. Number two, the second largest use case for our AI is actually to solve this problem.
And people use our AI and they ask it, Hey, can you do some research and tell me which data I actually don't need? I never search for, I'm not using a dashboard or alert. And the AI is actually doing fantastic job in scanning and telling you, Hey, you actually, 40% of the data you're shipping us is actually, you don't need to ship it.
Just put it some in a some S3 bucket so you won't be charged for it. And that helps them cut the cost. So I think on one, one edge, you want to increase the value with AI agent.
On the other hand, you wanna reduce the spend on these is, you said storage not for free. So it's, it's a, it's a common use case. Ultimately.
How will our user experience evolve? 'cause we're so used to kinda looking at dashboards all day long. Are the dashboards gonna go away?
Are they gonna be replaced by AI agents and or, or what will the actual user experience be like? I honestly am not sure. I'll tell you what I don't think will happen three or two or five years from today, you are gonna see these big screens and you know, the knock people on the DevOps or platform engineering, whatever, they're gonna look at all these dashboards and try say, oh, I see the spike here, and that spike there.
And that is not going to happen anymore. No one knows how the future interaction will happen. Um, my gut, and again, I'm not sure, is the dashboard will go away completely the way AI agents will work.
They will do the correlation analysis and they will generate a dashboard for you for this specific incident right now, Hey, I wanna show you something right now. Have a look at that. Our AI agents right now, and again, we're early in the journey, this is only getting started, is already you ask it the question, Hey, there was some too many error 500 or unable to log into system error messages, do some root cause analysis and it'll go and to start to create panels and dashboards for you to, as part of the reasoning process to show you what it went through.
So it's, it's pretty powerful. And I think the time of like the day that you, we have customers that do 10,000 dashboards, that this is gonna go away. As you kinda wonder about all this new AI capabilities and experience that we're gonna have, um, how autonomous will these AI agents actually be?
Because I can envision a world where I have a AI agent for observability that is connected to, uh, some ITSM platform that then the two of them go execute something. But am I gonna trust them to do that? Do I wanna have that be permission or over time will I just increasingly rely on them?
It's a good question. I think I still think even in 5G years from today, a lot of the things we still have ma like human supervision in the most critical systems, you still want to have some change management before you make changes and before you automate things completely. But I would focus on the mundane stuff, the 80 20 rule, the 80% of the alerts, the 80% of the incident, the 80% of the noise.
I think we are gonna get to a place of autonomously and it's gonna take some time. Um, we wrote a, uh, we wrote a, um, um, a story about autonomous observability and it's a little bit like, you know, cars, you know, getting the, the first level and the second level where it helps you drive and maybe it does the, the FSD and, but the, the the last type of full autonomously is such a big leap. But I think we are gonna get up the scale of autonomously and assisted and, you know, only supervision very, very, I think faster than we think probably in the next two to three years.
Well folks, you heard it here, observability, AI agents, they're gonna be pervasive. The only question is now is are we ever gonna be surprised again by these complex IT environments or will everything be known that today is all too often unknown. Hey, Tomer, thanks for being on the show, Juan.
Thank you for having me. Appreciate it. All right, and back to you guys in the studio.