Learning OpenTelemetry with Austin Parker
Mitch Ashley and Austin Parker, author and OpenTelemetry (OTel) open source contributor, discuss the launch of Parker’s newest book, Learning OpenTelemetry. This book helps expand the approachability of OTel to new users and the broader user community. Parker also discusses how OTel is more than just an IT or security, as he sees broader use cases such as manufacturing, business metrics and other operational processes. You can find Learning OpenTelemetry at your favorite bookstore, including https://www.amazon.com/Learning-OpenTelemetry-Setting-Operating-Observability/dp/1098147189.
Transcript
This is Textron tv. Well, the great pleasure of talking again with Austin Parker, Austin, who I've talked to many times on panels and shows and, and at shows. And we just both came back from, uh, coup Ka in Paris recently.
And, uh, you're director of Open Source. Open Source. Okay.
At, um, honeycomb Welcome. Yes. Honeycomb.
Um, yeah. It's great. Been there about, Ooh, nine months now.
Yeah, I remember when you first moved over. Yeah. That's all fantastic.
Hope it is all going well there. I hope It's great. Yeah, no, it's super, it's super fun.
Super exciting place. Um, we're doing a lot of really great stuff over at Honeycomb right now. Ah, There's some fantastic people there, including, oh yeah.
Love, love the people. Love. They're really great.
Love, Love charity. You love charity and, uh, Liz and everybody. Oh, Liz iss great.
Yeah. No, it's, it's so fun to work with a bunch of people that I've, um, you know, throughout my career, sort of in observability, I've interacted with so many of these people in the community, and now just getting to work beside them, um, and really help, you know, continue to innovate in the space is really fun. It is.
It's a very, um, highly respected and highly collaborative group of people. I mean, this, it's really good to see the great things that have come from all your work, one of which is you recently came out with the new book, and that's what we're gonna talk about. Now.
You're not a first time author, you've authored other books, but I think this one is a little bit different approach and it tells us the book pile. Yeah. And tell us about kind of what, what it's about.
So the book is called Learning Open Telemetry. Um, I wrote it with my friend Ted Young, who's also, uh, who I worked with at LightStep for five years. And we were both, uh, initial contributors, kind of founding members of OP of Open Telemetry as a project.
You know, one thing that people might not know at this point is a long time ago, or maybe not that long ago, but certainly five, six years ago, there was a project called Open Tracing that was part of the CNCF. And there was a project called Open Census that, uh, Google and Microsoft were backing. And we both, both projects had very similar aims, um, to really help push forward observability.
Um, mostly through making distributed tracing more accessible to people, open Census, um, had slightly different ideas about how to do it, but it was causing a lot of confusion that there were these two projects that more or less did the same thing out there in the world. And one day, you know, we all kind of sat down. It's like, eh, it's silly that there's these two separate things, you know, why not just have one thing?
And that's what led to Open Telemetry. So that was announced back in 2019. Um, we recently kind of celebrated our five year, you know, anniversary.
Wow. Um, and it's come just an amazing, uh, it's honestly staggering, you know, how far this project has come over the past five years. But with that in mind, you know, as it's gotten more popular, uh, open Telemetry has been for three or four years now, you know, the second highest philosophy project in the CNCF, um, at the end of 2024 or 2023, sorry, we had something around 2000, you know, almost 2000 contributors, um, over 300 companies contributing to the project, you know, which is a huge number.
And as time goes on in any com large complex project, it becomes harder and harder to kind of understand what it's about, right? Like as it expands and grows, there's just so much knowledge that kind of doesn't get lost, but just sort of gets, gets subsumed. It goes below the tide.
And we found ourselves kind of answering a lot of the same questions over and over about like, why, why were things in Open Telemetry a certain way? Mm-Hmm. And we have a lot of great documentation about it.
Um, and people have written other books about open telemetry, right? Good friend of mine, um, Alex Boon wrote a great book about, you know, using Open Telemetry with Python. But the thing that a lot of these books had in common is they're very kind of focused on like a, you know, on one language, right?
And how to do this stuff in Python, um, or in Go or in Java or whatever. And nobody really took a step back and said, well, hey, what, what is the point of all of this? Because when someone comes into Open Telemetry from, you know, existing sort of monitoring observability background, you know, they know things like Prometheus, they know Elastic Search or Open Search, they know maybe Zipkin or Yeager, or they know specific commercial tools, you know, they know Datadog, which obviously is a great observability company.
Splunk, there's so many things Mm-Hmm. And fundamentally, like, yes, that's valuable knowledge coming in, but Open Telemetry is trying to do something that is a little different. It's, and it's kind of a subtle difference sometimes.
Um, but we're really trying to push this space forward by talking about how all this telemetry, all of this data from your application is really part of not three pillars, but of a single interconnected braid of data. You know, a stream, a very rich stream of data that has things like inherent correlation, um, hard and soft context between different signals and between different events that happen in your application, in your infrastructure, um, conventions for how to represent this data so that tools and end users, you know, can understand it more easily. And there wasn't really a great resource for people coming in that didn't have that context, didn't have that background.
Mm-Hmm. And so we sat down and said, you know what? We're gonna, we're gonna fix this.
And we went out and wrote, wrote a book on it. And that is a learning open telemetry. It is, you know, the comprehensive guide, I think on, not necessarily like the API or the SDK or how you do certain specific actions, but we give you all of the knowledge you need to be able to interpret and understand like, what is happening in Open Telemetry?
How does Open Telemetry function, how do all these parts fit together? And we give you the knowledge that you need to, to go on and learn more, right? Because there is great documentation.
You don't need us to regurgitate the documentation to you in a book. Um, we have great docs on our website, but you do need to understand what all of these different things mean in order to really understand what the documentation is telling you. And so that is what Learning Open Telemetry does.
You know, it, it's really great. 'cause I, I've often struggled with this, the simplicity of logs, alerts, and traces sounds good. But it's a much more complex than that of putting that all together with context and what's happening and in, in systems and multiple systems and, you know, transactions that are flowing and resources that are involved in it.
Mm-Hmm. You know, it, it is, it is a, a braid, if you will, of weaving all of that information together as you described, to really understand what's happening. So just listing those three, um, sort of understates what's really involved, it makes it simple to understand.
But, you know, you could put all those in a database and have that, that's not what, you know, this is really about. Right. Uh, well, well tell us, I I, I, one of the things I, I really appreciate about the book, and it's great that you, you took approach, an approach like this.
Uh, let's just set the context of understanding kind of the whole system of view of a system of tele open telemetry and how that fits into an observability kind of model. And you talk about the difference of telemetry and versus analysis, right? Mm-Hmm.
Those are kind of getting that data and then doing analysis on it, on it are not necessarily in the same tool. And that's where most of us live, right? That's part of what Open Telemetry does, is separate that the data gathering, um, from, you know, being tied into individual commercial or open source tools for that matter, um, kind of giving you that one common plane of how you get all that information together and then access it in a way that you can do analysis across all those systems, on all those kinds of information.
Yeah. I think there's one, one thing that I've been sitting down with a lot recently is this idea that, you know, uh, again, to kind of do a quick history lesson. Let, let's think five years ago before Open Telemetry, um, when observability was sort of this hot new word.
Mm-Hmm. And if you go and you look at, you know, again, you look at these commercial tools like Splunk or Datadog or whatever, these were considered monitoring tools, right? You didn't have an observability tool.
You had a monitoring platform. And that monitoring platform will let you do things like take in metrics instead alerts on them, or search through logs, or do what was called, you know, what we call a PM, right? Application performance management, real user monitoring.
What are all these things, you know, and what do all these things have in common, right? At the end of the day, these are all taking events from your system, um, translating them in a certain way, storing that data in a certain way, and letting you do math over those events, right? Letting you do things like when this number gets bigger than that number for more than five minutes and be an alert, you know, ping, ping me, um, PagerDuty or something.
And very quickly, as observability became a word that people kind of cared about, um, all these monitoring tools just, you know, flipped the sign around and said, Hey, we're observability tools now. And I don't think that's wrong, right? Like, I think that some, some people, That happens all the time, every time you come up with a new turt, everybody rushes to become that Every, everyone rushes to say they're part of it.
And I, but I don't think that's actually necessarily wrong of them to do, because I do think that, you know, most modern analysis tools do let you do observability stuff. The distinction that I like to draw now is that a lot of people, you know, and this goes back to what I said earlier, if you're coming in from sort of the monitoring background or professional observability background, you're not really, you're, you're thinking of this stuff as still workflows. You're thinking of it as analysis.
So you don't think about logging as anything other than the input to sort of a log search system. You don't think about traces as anything that's necessarily other than the input to a a PM tool. And now with Modern tools, yes, you can have some correlation between those, right?
You can jump from a log to a trace or from a, you know, metric exemplar into a trace or from a real user monitoring, you know, from a web session into specific logs that make correlate with it. But the dis the difference, and I think Open Telemetry is trying to get at this, is that the actual data itself is more or lesser relevant to the analysis in the sense that right now, you know, the analysis drives the data of it. What should actually happen is the data should, the data itself should drive the data.
A trace is a particular, you know, in open telemetry, all of these things are just events, right? When I, when something happens, I'm creating an event and I'm telling open telemetry, okay, I want you to interpret this event as a particular type of instrument or signal. I want you to say that this is a metric or, um, a counter or a measurement in a histogram or a log statement, or a span or a span event or all this stuff, right?
There's a lot of, you know, varying details here. And one of the things about the book is we break all this stuff down for you and say like, okay, here's the difference between a span event and a metric and a blog. But ultimately, all of these things are just ways for you as a developer, or you as a library author, or you as some, you know, someone that works with software to make a statement about like, Hey, there are events that happen in my software.
How should you as a user interpret those events? And they give you ways to model the behavior of your software and system. And then the people that are using your software system can take those events with their semantic information about what is going on and what is happening and why it's happening.
Put those into analysis tools and use that to understand the emergent behavior of the system in production rather than, I have a bunch of logs, I dunno what these mean. Open telemetry gives you the tools you need to actually have a semantic understanding of like, this is what this log means, and this is why, even why this log is a log, right? It could be anything else.
It could be a metric, it could be a trace, but it's a log for a particular reason. And as tools kind of catch up with where open telemetry is, I think we'll see more being done to help end users understand like, oh, all these semantics, all of these, um, these affordances that open gives you, actually helps you become more efficient at finding problems in your system, right? Find understanding like why an incident is occurring, and even understanding how the system is really supposed to be put together.
I think we talk a little bit about this in the book, um, like you said, this difference between transactions and resources, right? A transaction being something that happens in the system, a resource being something that is consumed by transactions. And if you're a developer, understanding that is actually super, super critical because you can't affect resources necessarily.
You can give, you can put more in, and you can always give more ram to your pod or, um, more hard drive space, but you can't really modify them on the fly. You can only really affect transactions. And understanding that emerging behavior that comes out of hundreds or thousands or millions of transactions and resources, code co-mingling at the same time is where I would say pretty much all interesting problems, performance problems happen.
It makes a lot of sense. Yeah. So, and one of the things that I hear you saying is, in part, this is for end users to kind of understand open telemetry is a system and the elements in it, so you know what it's made up of and how it works, so you can better implement it.
But it also sounds like, you know, there, there's a great book for prologue of how to, how to develop and using hotel, uh, libraries and such. Um, but it sounds like it's good, good for developers to read this because you also will understand the whole, the system as a whole. And now, you know, when you're omitting transactions, here's how that interacts with other transactions of different types, that kind of thing.
Yeah. Fair to say. Uh, yeah.
One, one thing that, there's a little call later in the book that I really love, and it's that we used, you know, I don't know if I've ever told you this, but like, when I got started in software, I got, I was a software developer and test, right? So I was in qa, and it used to be that we had people whose job it was to understand, you know, how the system fit together and the interactions between different parts of the software. And we called them qa.
mm-Hmm. And so you would write code, you would make sure it builds in your machine, you would make sure it passes, you know, unit test, and then you would throw it over the wall to a QA engineer that would go through and do acceptance testing, right? They would go through and they would make sure that it works in the way that we think it should work, that there aren't unexpected behaviors, or that when I do something weird, you know, everything keeps happening the way it should.
And over the years, you know, QA has been, I mean, I think it's really kind of been a dying breed, right? Like, we don't have, uh, as much manual qa, I feel like as we certainly used to as an industry, but in a lot of ways, telemetry can really help us with this, right? Because telemetry is a way for you as a developer to actually say, this is what the system is supposed to do.
You know, you can define the relationships between things. You can say like, Hey, when I'm writing a, you know, when I'm creating a trace of my system, or I'm creating a span, I'm actually saying what's important, right? Like, how is, um, what do I, how do I expect this to work kind of in isolation?
And then you can record that by taking those spans and sending 'em somewhere. And you can have a nice like copy of like, Hey, this is how things are supposed to work. It works this way on my machine.
I can then take that and compare it to how it's working in production, right? That's a really powerful ability. And if you start thinking of like, Hey, telemetry is part of, you know, this is kinda like the modern QA process.
This is how we're actually making sure that things that we're building work properly. Um, but it's also much more powerful because it's something that runs all the time, right? Like this telemetry is just there, it's below the surface.
You can go look at it when you need to, to validate your assumptions. Um, and then when things do break, because things are always changing in a complex system, you can go and actually understand what's happening to my system in production. Why are things different than I expect them to be?
And again, you have that history, right? You have like, well, it works this way here and now it's working that way there. Um, and now I have all the data I need to really understand that.
And also I can take that data, put it into analysis systems and do things like get alerted on it, right? Or see, you know, track those changes over time, um, through metrics or get that kind of like snapshot of a single end-to-end transaction with traces. You know, that, that, that kind of explains, at least in part, one of the really interesting things I saw in the book too.
Um, you know, 'cause we talked about observability and development and test. Now it's not just an operational, um, uh, tool, if you will, if that's the right term, but, um, there was a call out in the book that I remember saying, sort of, don't stop at telemetry and analysis, right? Getting the data and then doing the kinds of math on, on that data to tell you what's happening.
This is more like a DevOps kind of adoption, right? It really, it's something that affects the whole organizations. And as you're talking about how this could help qa, right?
Mm-Hmm. Up informed key way of what the system is doing, what's happening and it does, is doing the same thing in test Mm-Hmm. Talk about how that, how observability is something, uh, that an organization adopts, not just a thing we use.
io, uh, inspired by some other people writing in the community about sort of the thing that I said before, right? Which is how every monitoring company has flipped the, you know, flipped their sign over and said, we're an observability company. But I think fundamentally like that actually kind of hides a deeper truth, which is we spend a lot of time talking about, you know, we, we've kind of spent so long as people out on the bleeding edge talking about observability as just better monitoring that we've missed a real opportunity to make it relevant to the rest of the organization, right?
'cause if you, a really simple example is this, we spend all this money on like these very expensive, you know, observability platforms. Um, but if you think about like a, a business in general, what drives a business? It's data.
Businesses collect so much data every single day, and we have these massive data lakes, data warehousing, you know, we business analytics. Like what is that? Well, it's just telemetry, isn't it?
Right? Like, aren't customer analytics, you know, aren't like tracking clicks in a webpage or how many things someone ordered or a customer profile aren't All these things actually just also telemetry, there's telemetry about a different thing. Why do, why aren't we kind of taking that next step and saying, well, look, perform, you know, performance telemetry from sort of your engineering is actually so much more related to that business intelligence telemetry, then you really think, because if I'm, if I have an e-commerce site, um, or I have any software that deals with, you know, any kind of software that people interact with, and I generate revenue off of that software in some way, the performance of the software and the actual business performance are incredibly important to correlate.
So why don't we, you know, and right now we actually can't really fundamentally ask those questions because we have this, this wall between sort of the engineering side and the, the business side. But wouldn't it be interesting if you could ask stuff like, okay, show me, you know, I have this conversion funnel of people that are on the site, and I wanna know, um, how many people are falling outta the funnel? Well, I really also would like to correlate that with, you know, performance, right?
Like, for people that are falling outta the funnel, what percentage of them had a page load time that was over 500 milliseconds or over one second, right? What percentage of them were in a different geo? Uh, break it down by device type.
Break it down by like any number of factors that are, again, blending analytics, you know, business intelligence, data with performance, uh, telemetry data. Now, open telemetry is really just about standards for that, you know, performance data, right? It's not trying to be a schema for bi data, but there's no reason that you couldn't extend it or that you can't use the fact that it is a standard to kind of correlate it with other standards.
I actually had a really interesting chat a couple weeks ago with a, um, someone that works, uh, involved in large injection molding operations for plastic bricks. Mm-Hmm, hmm. I can probably guess who that might be.
Oh. Um, and they're using, and they, so they obviously have, you know, very complex, uh, and very finely detailed standards for, you know, production and a lot of robotics, a lot of automation, a lot of manufacturing stuff. And it's very important to them.
They've extended open telemetry to get data about the performance of their injection molding process. Right? Interesting.
And you know, how the robots work in, um, and all of this stuff is obviously, you know, uh, like this is a huge thing for them because this is, this is literally the bottom line, right? Like, if you're physically making stuff, you wanna make sure that you're not, you know, you want to control waste, right? You wanna control like, how much do I have to throw away of a given batch because of x, y, and Z reasons?
So they're taking open telemetry and they're really extending it into sort of their business operations in a particular, in, in a one specific way. But it really got me thinking, it's like, oh gosh, like software doesn't do, software really doesn't do this that much, right? Like, you tend to see this, again, this silo of, and even in places that understand this, they're like, well, I would rather, you know, it's like, oh, don't, don't co-mingle these, right?
Like, take the, the customer analytics and use a different SDK and throw those into the, the BI data lake so that our analysts can look at them. Mm-Hmm. I'm, I, I feel like that's a missed opportunity for observability as a a practice.
And we can extend it even further, right? We can start bringing in code quality, we can start bringing in team health. We can start associating all sorts of other things with performance and ask interesting questions about like, well, and I mean, I don't wanna like sound like a jerk, but hey, maybe it would be nice to know how many, like, can we identify which team is improving their code the most in prod?
Can we see which team is maybe responsible for more? Um, failed pushes, failed mi, failed migrations, right? And not in the sense that we should punish them, but in the sense that that is useful data for our decision making process.
'cause maybe that tells us that people are getting burned out. Maybe it tells us we need to invest more in training, right? We can take this data, we can put it into, we can turn it into SLOs into service level objectives about like, how good should our code quality be?
How many times should we pushing? Where do we need to invest and move that up and up and up until it's sitting, you know, until that's what they're looking at in the boardroom. And at that point, I think in a lot of ways, like soft, especially now, you know, the economy is changing.
Money isn't free anymore. The biggest challenge that every single person I talk to, especially executives, it's like, we need to show value. We can't just throw money away anymore.
We have to, you know, we have to connect everything back to how is this making money for the company? And I think even in places where software is their job, r and d is seen as a cost center, right? It's not seen as something that produce, you know, it's, it's like we have to just keep giving the engineers more and more and more money.
And then sometimes they come back and give us something that we make money off of, but we don't really like that. Right? But it's like, well, why not?
You know, why can't we change that? And I think there's a lot of reasons why it's actually a much, we, we don't have time to talk about all the details right now, but I think one way that we can start to adjust this is by bridging, you know, this world of performance telemetry and how the application is running and what the system is doing to what the business is doing, right? Like, how is the business operating?
Where are we making our money from? Where, you know, what is important to us? Let's, let's tear down this wall and bring these things together.
0, right? Something that is more than what we have today. Something that is more valuable and more essential.
Um, and I believe Open Telemetry is a huge part of that. And I think by reading this book, you will kind of come to see like some of this, you know, where, where this starts at on the performance telemetry side is certainly Super fascinating. I can totally see the talk you're putting together, um, probably in your head and in a few slides, I would imagine.
So what's what's fascinating too, is that one of the things that really intrigued me about Open Telemetry initially was that we could start to measure things that, um, like end user experience, we could start to put business metrics in here. And what you're really doing is saying that that's really a wide open area. Um, we could do quality of our products of quality and a manufacturing process, quality improvement efforts, performance, right?
What are, we're more efficient lines or, or manufacturing process that might just to use that example, um, might help our business, uh, perform better or, um, respond to the customer faster on, because software's running all that stuff, right? So we're already there writing code or using software to operate those things. Why not collect the telemetry information that can help us measure the value we're creating or how to improve it, or Mm-Hmm.
What's working well and what's not that just in our system, but in the things that our software is helping produce. Yeah. I think that's a, it's a fascinating idea.
It's, it's definitely, I think it's where this is all, you know, it's, it's where this is all going, right? I don't even wanna say, I think 'cause like I, I fundamentally know, um, these worlds are getting closer and closer together. And it's not just because of, you know, I think some of it is definitely the macroeconomic stuff I was talking about, but a lot of it is just every company is a software company, right?
Like, you can't escape that fact. Um, and with the rise of things like generative AI with the likelihood that more and more and more of our lives are going to be impacted by software, you know, it's not enough to just have, you know, the, these sort of slow feedback loops in terms of how is stuff performing for the customer, because people have more options than they ever have before for permission, anything imaginable. And if you're not kind of putting customer, you know, that sort of end user experience and customer focus is the first thing, then someone else that is, is gonna come and take you lunch.
Right? Exactly. So, And it's too late to learn how to do that at that point.
Right? Right. Like, by the time you see by, and that's, I think that's what has been the story of so many industries and so many companies, you know, over the past two decades or so, certainly since the early two thousands.
But by the time you realize there's a problem, then you, the moment has passed. You can't be reactive anymore. And this isn't just, you know, about, like this is your business.
Yes. But this is also about your software system. By the time, if some, if you were hearing about a problem, 'cause the customer's reporting it, then you have already lost like tens of thousands, hundreds of thousands, maybe even millions of dollars because that one person that reported it, you know, people don't have this, you know, people don't give you a break anymore, right?
Like people, if, if somebody doesn't work, they're just gonna be like, oh, heck with this, and they're gonna move on to the next thing, or they're gonna find another way to do it. Um, this is true for so many things. You know, if you aren't proactive, if you don't know what's happening in production to your customers, like it's not, oh, we were slow and we caught it late.
It's like you have lost people that will probably never come back at that point. Yeah. You only get, you know, and that's just the way of the world these days.
It is, it's pretty cutthroat out there. I feel like, um, you people really have to love you to give you a break. You know, Everything's digital.
That means there's a lot of options, right? Yeah. Because we're working the same and, And with AI and stuff like that, there are, it, it reduces the barrier even more to people coming up with new options.
Mm. Um, Good, really good way to think of it. Well, hey, you know, whether you're wanting just to learn more about Open Telemetry, I really just wanna understand it better.
Or maybe I'm new to it or you, you want to understand it to, well, well enough to do more than just operationalize it. Really leverage Open Telemetry in some unique and valuable ways. It sounds like this is a fantastic book to help you set the stage.
Yeah, I think it's a great resource move that Way. So Yeah, we have practical guidance on how to, you know, under, you know, we have foundational knowledge. We have some practical guidance from actual users that we've kind of collected from our, from our open symmetry end using end user group.
So stories about different trade offs, stories about like how to start using it, how to roll it out, um, deployment scenarios. There's a, there's a lot of like practical stuff in the book and there's a lot of foundational stuff I would strongly recommend. You know, I said if this is a topic that interests you, if you're interested in observability, open to telemetry, um, either because you're adopting it or you're thinking about adopting it, you know, I really do think this is an invaluable resource and we, we tried to write it as something that, you know, is gonna be useful for the next five, 10 years, right?
Like Open Symmetry is gonna be here for a while. I think this is probably a book that should be on everyone's bookshelf that is working with systems, software systems. Sounds like it to me as well.
So Learning Open Telemetry, um, where would you suggest people go check it out? com. Um, we have links to Amazon to find a local bookseller or to, uh, where you can read it digitally on, um, o' rally atlas.
And we have a couple of, you know, you can find out some of our events. We're gonna actually have a, uh, book release party this summer in Portland, Oregon. Um, and as we travel and do book signings and stuff like that, um, we'll try to post them on there.
com. Well, I hope we'll get another chance to explore this some more. 'cause uh, you brought up about a hundred things I'd like to ask you more about, but we're obviously can't do that today, but appreciate you, um, well writing the book, you know, seeing the need for this and then, and filling that need, creating value through that.
So congrats on the book. Um, yeah. To you and Ted, you said Ted Ko?
Yeah. Ted Young, yeah. To great.
Another great person in the, another Wonder, another wonderful human being. Yeah, he is. Just super appreciate it much.
And folks, uh, definitely check it out. com and, uh, to get your copy in whatever form and through whatever source you choose. Austin, it's been a great pleasure.
Always good talking with you and love getting online. Love seeing you in person too. So I don't know if you'd probably be it, some of the same events here again, we'll catch up with you.
Yeah, No, definitely. See, see you around next, next time we pass, cross. Okay.
Take care. Thanks again. Bye Al.