AI Raises the Stakes for Observability Engineering
Observability Engineering Enters the AI Era
Alan Shimel talks with Liz Fong-Jones, Technical Fellow at Honeycomb, about the evolving role of observability engineering as AI changes how software is built, shipped and validated. Fong-Jones explains that AI is giving teams more ability to create software quickly. That speed also creates a new bottleneck: understanding whether those systems work as intended.
The conversation highlights why observability engineering is becoming more important as AI-assisted development grows. Developers may be able to generate code faster, but teams still need to understand behavior in production. They also need enough context to validate changes, investigate issues and reduce risk when systems become more complex.
Telemetry Is Not the Same as Observability
Fong-Jones explains the difference between telemetry and observability. Telemetry includes logs, metrics and traces that describe what is happening inside a system. Observability is the broader ability to use that data, along with human knowledge and process, to understand the system and make better decisions.
She describes observability as a living property of software, similar to testability or accessibility. A team is never “done” with it. Instead, the goal is to keep improving the system so engineers can move faster with more confidence. In that sense, observability engineering is both technical and organizational.
AI Changes the Observability Loop
The discussion also covers the second edition of Observability Engineering, the O’Reilly book co-authored by Fong-Jones and others from Honeycomb. The new edition reflects how much the field has changed since the first version. Teams no longer need to spend as much time debating whether observability is different from monitoring. They now need to understand how it supports AI-driven software delivery.
Fong-Jones says AI can help reduce toil by generating instrumentation code and helping teams sift through large volumes of telemetry. Agent-assisted workflows can also help engineers ask questions, surface insights and move through investigations faster. That does not remove the need for human judgment. It changes where people spend their time.
Cloud Native Becomes the Default Platform
The episode also explores the relationship between cloud native systems, OpenTelemetry and AI applications. Fong-Jones notes that cloud native has become so successful that many organizations now treat it as the default platform. Containers, Kubernetes and related technologies are now part of the expected substrate for modern software.
For Honeycomb, the business value of observability engineering is helping teams innovate faster while reducing risk. Fong-Jones says Honeycomb uses its own practices internally and then brings those lessons to customers. The result is a focus on helping engineering teams understand production systems, validate AI-driven change and move with more confidence.
Transcript
Hey everyone. Welcome back here to Techstrong TV. My next guest, always a pleasure to have on here on Techstrong TV.
They literally wrote the book on observability, which we're going to talk about in a second. But let me introduce you to Liz Fong-Jones. Liz is a technical fellow, and we're going to explore what that is with Honeycomb.
But Liz, it's great to have you here again. How's everything with you? Everything is a little bit chaotic.
I think that describes this era of AI that we're going through. It's possible to do so much, and the possibilities fill your head. m.
last night coding, which is so weird to be like, I am excited to do this, and also, I have this partner that never gets tired, and it's like I should probably at some point go to sleep. So yeah. You know what?
I'm excited and terrified- I do the same thing ... and that's the best feeling ever. Absolutely.
I have so many friends who frankly haven't coded in decades, 10, 20 years, who are now back to coding or at least prompting the AI to code, and then looking at it and improving and doing this. But even I'm up late every night, too, because that's exactly how I feel. I finally have someone I can't wear out.
With all the ideas and the stuff that just comes pouring out of me into it, and it's great. The challenge is what survives the following morning. But, you know.
Well, that's my rule. I do it all, and it's funny. My wife will be going to bed, and I'm just sitting there, and sometimes I talk to it.
And then I look at it the next morning to see how much of it was Alan on some sort of adrenaline ride- ... and how much of it was worthwhile. But a lot of it is worthwhile.
I've written a whole book that I'm publishing now called "The Indispensability Trap," and I did it arguing with my AI. Every idea and notion, we went back and forth on it and really thought it through and refined it. And I love it.
Anyway, but Liz, people want to know, what's a technical fellow? How did you become a technical fellow? Give them your story.
Yeah. So, I am the first technical fellow at Honeycomb, so there's not really a blueprint for it, but I think in terms of career progression, it's very similar to being a distinguished or principal engineer, right? It's someone who takes a holistic view of the industry and figures out where's the industry going, how can Honeycomb and its customers be as successful as possible in that new world, and how do we help get them there?
So, it's technical, but also there is a very, very strong focus on the business and on the customers. And how did I get there? Twenty-plus years of being a site reliability engineer and developer advocate.
Basically spending time initially working on running systems at a game company and then running systems at Google, and then over time working with Google's customers and then working with Honeycomb's customers to really get a broad perspective. Similar to what you do, right, in terms of talking to a lot of people, getting exposure to a lot of different organizations, kind of figuring out what are the patterns, right? What is it that we should be doing to level up as an industry together?
I love it. You're right. That's very insightful of you because a lot of people don't realize, and that's why I do this.
I love it. I love talking to all these people. Right.
Exactly. It's being a generalist rather than a specialist, right? Mm-hmm.
" For me, I really love going broad, right? I really love kind of pulling out the patterns of what are other organizations doing? How can we help make connections between people that really should be talking to each other?
You're talking to my soul because I just thought it was my ADD. Today I'm talking about DevOps, tomorrow I'm talking about security, the next day I'm talking about AI or cloud-native or platform engineering, or it's whatever my, where my brain takes me on a given moment. But you know what?
Thank you. Honeycomb. As I said before we started, we've been covering Honeycomb for as long as I'm doing this, right?
And I'm doing it a while. Yeah. And Honeycomb started 10 years ago, so there's kind of been this very rich history of kind of- Yeah ...
us both evolving with the industry, honestly. Absolutely. Now, observability 10 years ago was not the term, right?
It was out there. It was a term that some people at Twitter used, but- Right ... I think it's fair to say, right, Honeycomb did popularize observability as a term.
And then we kind of lost control of it, right? And I think that's fine. I think that the term took on a life of its own.
And now we have- It's kind of like being a parent. At some point, you've got to let them fly the nest, and Jonathan Livingston Seagull, if they come back, they were yours. If not, they never were.
That's probably way before your time, though. I apologize. But it was something I remember from reading as a, I was probably in high school, not even college.
Anyway, Honeycomb. So yes, observability is bigger than Honeycomb in and of itself today. It has become the word.
For a lot of people, observability has become very much entwined with open source, with the CNCF, the OTEL, the OpenTelemetry product, Prometheus. Some of the biggest products, or the biggest projects, excuse me- The biggest projects in open source are observability related. I think that illustrates how much demand there is for observability and for improving observability.
The one thing that I would caution people is that observability and telemetry are not synonymous. You mentioned OpenTelemetry earlier, and it is the third largest Linux Foundation project, second largest CNCF project. We're only behind Linux and Kubernetes itself.
But there's a lot of appetite for it, and also what you do with the telemetry that you generate matters a lot in addition to having and collecting the high-quality telemetry. So you brought it up, I got to ask you now to close the loop. In your mind, as I said, you wrote the book, we see the book up here.
What is the relationship between telemetry and observability? Right. So observability is like testability and accessibility and understandability.
It is a living property of your software systems. There is not a, we are done. We are 100% observable.
That's never the case. " We're 100% testable. So observability is this property of the system that we strive to improve because it gives us the safety to move faster.
And when we talk about telemetry, telemetry is the ability to produce and gather these signals that help us improve our observability. So when you think about signal types like tracing or logging or metrics, these are all mechanisms of gathering data about the events that are happening in your system under the hood so that you can feed them into a system for analyzing and actually giving you that observability at the end of the day. I can feed all of my telemetry data to DevNull.
I could have all the perfect telemetry in the data, but if I'm just feeding it to DevNull, I don't actually have any observability. So it's this property of the data and the people together. It is a sociotechnical property.
Got it. That was excellent. Thank you.
I feel, well, we should say off the bat, the current edition of the book is not the first edition of the book. No. In fact, it says second edition right here.
It says second edition on the spine. Yeah. And you'll notice that it is a lot wider than the- It's grown a bit, huh?
It has. Look, we all have. We all have.
But in any event, talk to me about why. What's changed? How so?
For the better? I don't know. What do you think?
What we saw was that there is a need to revise the book in light of two things. First of all, when we first published the first edition in 2022, something like that, very much we had to beat people over the head with observability is not monitoring, observability is not telemetry. People got it, and I think that we really wanted to move beyond the basics and into some of the more advanced stuff.
And secondly, we could see the need for observability growing as the AI movement grew. That people were writing and shipping a lot more software. That the writing of the software, paws on keyboard, was not necessarily the bottleneck anymore.
The understanding and validation was becoming the bottleneck. So we really wanted to help people understand how observability fits into that. What's the business case for really investing in observability as a priority to accelerate your AI efforts?
And also, yes, AI has changed how you do observability. And I think it was also important to harp on that a little bit. To say, you don't need to manually write all of this instrumentation anymore.
The bots are perfectly good at turning out boilerplate instrumentation code. We believe as site reliability engineers in not having excessive amounts of toil. Writing telemetry payloads is a little bit of toil.
How can you automate that process? And also, how can you automate the process of sifting through this telemetry data so that you surface insights to the humans that need it? That was a subject that we didn't really explore that much in depth in the first edition.
In the second edition, we talk a lot about the agent-assisted observability loop. How do you actually ask questions, get answers in a agentic fashion rather than in a purely human-driven fashion? Agreed.
I love it. So Liz, the thing about, though, I recently wrote a white paper for our Cloud Native Now site that the cloud native stack is the AI stack. Right?
And in my mind, when I talk about it, it's not just Kubernetes I'm talking about. It is observability. Can you really do AI applications or AI-empowered applications and get the most out of them without observability?
Thoughts. Yeah. One of the darlings that I had to kill was we had an entire chapter in the first edition about how observability relates to the cloud native concept and to the cloud native efforts, and why people who care about the CNCF decided to adopt OpenTelemetry into the project.
And why we graduated, although it hadn't happened yet. And as we were writing the second edition, we realized we were not hearing the word cloud native from our ecosystem anymore. It was just a given.
Really? No one was talking about- It was a given ... let's try to become more cloud native.
No, that's what it is. It's the platform, period, end. Yes.
And even some of the people who said, "Oh, we're going to repatriate our clouds back to on-prem," those people are still using the same foundational scheduling technologies. Yeah. On bare metal there, but they're still running the same stack a lot of them.
They're still running containers, right? So yes, I think basically cloud native was so successful that it became the default, and now we no longer need to talk about cloud native. And also now, right, with the advent of AI-written applications, AI-assisted systems design, right, it presumes this degree of being able to deploy to production fluidly, and then we get to the thing that we were talking about earlier with the validation crisis of how do you actually validate all of these components work together.
But the debate about the substrate and how you deploy it, that's gone, right? No one is debating- It's done been won and done ... should you use containers or not.
There are still haters. There's always haters, Liz. There's always haters out there, right?
I think- Right. There's the adopter curve, right? There's the laggards- Yeah ...
there's the early adopters, right? There are always going to be laggards, and that's fine. But my concern is primarily how do we help organizations that want to move towards the forefront?
That's really fundamentally one of the things that Honeycomb excels at is how do you speed up an engineering organization that wants to change? Agreed. Hey, we're almost up on our 15 minutes, but I got one more area.
So at some point now, right, Honeycomb pays the bills. Honeycomb is a for-profit enterprise. Where does all of this rubber meet the road when it comes to Honeycomb, right?
All the learning, all the knowledge, how does that find its way into Honeycomb? So Honeycomb really, I think we do two things really, really well. I think we really, really dog food extensively, and I think that teaches us a lot about the ways that systems ought to be built, right?
So we have a team of about 300 employees, right? We've got about 100 people who are in our engineering and product organization, right? When you compare that to some of our competitors, and our competitors have 10 times, 100 times the headcount, right?
So really our engineering methodology and being at the forefront of all of this is how we deliver results that are outsized in proportion to our engineering organization size. And then the second thing that we then do really well is to help translate that methodology into helping our customers achieve those outsized results, right? So observability is the how, but I think the why for us is helping people innovate faster and move faster and not have to be so concerned about risk because we help mitigate that risk.
So I think, right, that's fundamentally how we make money, right, is that our clients come to us because they have observability challenges that they think that we can help them solve. I love it. Eloquent.
Hey, Liz, we're about out of time. Did we mention the website? io?
Yep. io. Or just if you search Honeycomb Observability, you'll find us there too.
Okay. And then believe it or not, I'm already making plans for Salt Lake City and KubeCon. Honeycomb will be there?
Honeycomb will be, including some of our OpenTelemetry team members. It'll be great to see some of the Honeycomb people at KubeCon. And as part of the OTEL, they always have a whole section over there because it is such a big piece of it.
And also typically O'Reilly has a book, and they might be distributing signed copies of our book. You never know. I would imagine, yeah, they always do good bookstore stuff and promos at KubeCon.
But Liz, thank you for coming on here. Thank you for not only you, Charity, Christine, everyone, the whole Honeycomb. There's more than that.
There's a lot of good people at Honeycomb. Honeycomb friends of mine from over the years. Thank you all for the work you do around observability, and as you said, you're competing against giants with 100x the funds and the resources that you do, but yet seem to hold your own.
So that says something. Come back and visit us on Techstrong TV anytime, okay? Don't be a stranger.
Thank you. It's an honor as always. Thank you.
Liz Fong-Jones, Technical Fellow at Honeycomb. Checking out-- Hey, can you hold the book up for us one more time? The new edition of "Observability Engineering," an O'Reilly book.
Check it out at your local wherever you buy books. I was going to say bookstore, but a lot of people buy them online. Anyway, we're going to take a break here on Techstrong TV.
We'll be back with more soon.