Morgan McLean, Splunk | KubeCon + CloudNativeCon North America 2023
Morgan McLean, director of product management at Splunk, discusses the state of OpenTelemetry and why organizations are turning to it to ensure their data is fully versatile and transferable to various tools. He walks through major changes OpenTelemetry has seen recently, Splunk’s investment in the project through contributions, related projects and more.
Transcript
This is Textron tv And we're back at Kon 2023 here in Chicago of the great conversation. Then we're gonna take a little bit of a different track here. I'm joined by Morgan McLean, who's Director of Product Management with Splunk.
Yes. Um, and we're, but we're talking about some Splunk, but really talking about one of the roles that you have, um, with the CNCF, with Open Telemetry. Yeah.
So I'm one of the co-founders of Open Telemetry. I've been of the project, I mean, by definition then, since its beginning, since 2019 when we announced it, uh, open Telemetry is formed by the merger of two pre-existing open source projects, open census, open tracing. Those are both deprecated now.
Uh, and so really, I've been working on open source solutions this space since around 20 16, 20 17. Yeah, It's surpri, it's, it's amazing. It's been that long.
'cause I remember reading about, oh, open Tracings now part of Open Telemetry, You know, and to, to make, uh, announcements one thing to get to the point where it's been deprecated. Yes, Yes. That's a huge, huge effort.
Massive. And, and it's funny, at the time we thought Open Census and open tracing were relatively big and successful, and, and they were in their own way. But Open Telemetry so rapidly dwarfed them in terms of, uh, market penetration and success and features and everything else that, you know, the, the fact that they were deprecated relatively recently is, is, you know, interesting in, in the context of 2019 and the announcement, you know, fulfillment of original vision, and yet in some ways today kind of a small deal because Open Telemetry is so ubiquitous and so successful.
So, uh, to kind of rolling back history a little bit, what, what did you think was possible? What, what was you hoping for with Open Telemetry when you're first starting out? Yeah, I imagine it wasn't like, oh, we're gonna have, you know, all these technology vendors contributing.
If you had Told me then that within a year of announcing Open Telemetry, it'd be the second most active project in the C ncf f So the second highest number of monthly developers, I would've assumed you were lying. Yeah. Like, like that, that a project just on focused on data extraction could become sort of on the same order of magnitude as big in terms of contributions and code as something like Kubernetes, this compute platform that powers billions of dollars of industry every month.
It, it's quite incredible. Uh, and it's certainly been really exciting to be part of it. It's, I think it's, it's a model, at least one model for open source success of when the contributors will also say, okay, we know that, that we do that Yeah.
Too, on our own stuff, but this is Yeah. More valuable for all of us to do. Same.
Exactly. And, and if you go back in time to that era, before you, before this, before Open Telemetry got big, you had these established vendors in the space, particularly in APM because for APM you require integrations of actual applications. So to contrast of say, like logging or even system metrics where you generally have integrations of like Windows and Linux and maybe a handful of third party apps that people run to do distributed tracing.
For APM you need integrations with every language runtime and every single client library that people use on, on those run times, there's hundreds of thousands of these integration points that people need. And so no one company was ever gonna be able to provide that, let alone maintain it, even if they built it in the first place. And so that's where I think like Open Telemetry, the success is being very clear and that not only for end users, the people who want deep visibility in their own services, where they want to use Open Telemetry, export the data, but even the vendors who had existing sort of effectively offerings or, or components of their offerings and their agents have very rapidly adopted Open Telemetry because it's just, it's almost impossible to go it alone, right.
net, but not maybe Python or something else. And so I think Open Telemetry has really shown everybody the way, and it's been lovely in that the project has been so collaborative, both between the end users and all the vendors in this space. I think because it's scope just on data extraction, data collection from infrastructure and services, no one's commercial interests have really threatened by it.
And it really just has helped everyone. It's, it's been a very, like, honestly positive success story all around. Yeah.
When you can pick a, a domain where you say, is that really where we want to differentiate supporting another agent Yeah. Or another environment. And The answer is often no.
Yeah. Yeah. And in this case, it's almost, almost exclusively across the industry, but no, which is, which is good.
Good for everybody. So you're on the governance committee, governor's board, right? Yeah.
Yeah. So Open Telemetry is divided into, uh, many different groups at the lowest level. Lowest might not be the right word, but that's sort of the most sort of project oriented level.
We have sig special interest groups, and so those maintain the different parts of it. net or for Python. Each of those is different sig.
And then the collector is our sort of large agent for open telemetry that has its own sig. The sort of committees that sit a above that are sort of distinct from that is there's the technical committee, uh, formed a sort of, the most sort of senior engineers are sort of the, the people who can dedicate the most time to the project. They oversee the overall spec that says, this is what each of these language instrumentations should do.
This is what the agent should do. Or like, this is a metric, this is the, this is how you can operate on them. This is how a metric that describes, say, A-H-T-T-P request from a given service.
This is the, the sort of data, the payload, it should have. The technical committee does that. Then there's the overall governance committee that steers the direction of the project.
So I've been on that one, uh, uh, since the beginning of the project. Uh, and it's been, it's been very good. Plus you get to see the whole thing, the whole picture, right?
Yeah. Maybe helps shape the evolution. Yeah.
I think granted the ball, there's, There's maintainers and other people in the community who also have very broad visibility that's not exclusive to the governance Committee. Well, and to that point, you know, it's not like it's a top down effort. No, no.
Open source by definition. Right. And this is a, it's a funny thing 'cause people will come up a cube con and say like, what's the roadmap for open telemetry?
And we do have an official roadmap. You can go and GitHub and see it. It's there.
And, and we can talk a bit about the things on it, but the roadmap in open source is only so good as the, the willingness of the community to actually go and build those things. Like, it's one thing for the governance committee or, or me, 'cause I'm generally the one driving a lot of this to go and say like, you know, the big focus next year, one of them will be, uh, adding profiling as a four signal type, or this year logging as a third signal type. You know, it's one thing for us to have that vision and, and spell it out.
And in a company you hire people and have managers and chains of control that go implement that vision. But an open source, I mean, we don't, we're open's not a company and it's, it's not a single vendor project. There's dozens of vendors deeply involved, uh, which is a very positive thing, by the way.
And so, you know, it's one thing for me to say, this is our vision, but we need people to actually show up and do these. And, and so it's always been interesting building that roadmap. 'cause it's, it's this mix of like, this is what our end users want, our stakeholders want, but also this is what the community and the developers are actually willing and able to deliver and implement on it.
I almost equate it to a shared vision. Right? It is.
Because no one person has the full Correct. Or creates the full thing. Right.
And one Stakeholder or cohort of stakeholders even has that whole Other people wouldn't participate. Correct. I mean, if One Smoke or any other company was like, we're it, and you all kind of follow us.
Yeah. They're not gonna do that. Yeah.
And so it's, it's a, it's a interesting balancing act. And when the project started, I think there were concerns less for me, but I think externally of like, there's a lot of vendors involved in Open Telemetry. It's, and there were early on, uh, and so there was, people were concerned it'd be like this sort of vendor fest that hasn't really happened.
Uh, and I think that's because it's so focused on data extraction, again, like most of the companies are happy for, to get more data out of open telemetry. Mm-Hmm. Uh, and so it's been not, there's hasn't been any walking a type rope between vendors that's actually been a really positive relationship.
Uh, but still you have to balance like what people are actually willing to sit down and spend their time developing, uh, versus what maybe you think the project needs overall. I think generally we've done a good job of achieving that. Well, I mean, to have the support you do, I think you, that's a good testament to this working, right?
Well, we set a milestone. So we have 1,100 monthly active, uh, developers in the community. Wow.
So every month, how many? 1,100. 1,100.
Wow. So every month there's at least 1,100 people who are checking in, either checking in code, making pull requests, reviewing, poll requests, sort of activity like that. It goes beyond just basic comments.
And that's huge. Like, that's, I think almost exactly half the size of Kubernetes every month. And, and Kubernetes is again this massive, huge compute platform.
Yeah. And so to have achieved that for a project fundamentally is just focused on data extraction. Data collection data, pre-processing is very, very impressive.
And I think it does speak to the, to the community that we've built in the, the positivity and inclusiveness that, that we've, that We've delivered. So what are things that a governance committee deals with that you aren't gonna handle or, or couldn't only work on as an individual group or component? Honestly, not much.
And I say this in a loving way. Like, like the, the Governance. This is a good thing, right?
Yeah. Otherwise, you know, not everybody's gonna be happy if You had an open source project where the governors of it were constantly meddling in everything or dictating things. Like there are some examples for something that's so big where yes, at times you do need to like, bring everyone together and like drive a very specific direction.
But I think in, in open telemetry, as I've seen in other communities in c ncf FI think the fact that the governance committee generally doesn't have to step in that much is actually really positive. I think again, it speaks to the collaborative nature of the community, right? And, and as well as to the systems that we set up.
Like there are, you know, if you're maintaining the Java SDK right? You have a spec that you come and implement. Uh, but that spec is not defined by the governance committee.
It's defined by the sort of spec group, which is mostly technical committee members as well as all the people who are implementing it. And, and so the fact that it's not some super top down thing where the governance committee can focus a bit more on like project direction or, or, you know, end user feedback or showing up at conferences like this. I think it's very positive.
How about things kind of going back a little ways, but when, uh, o open tracing, you know Yeah. Bringing that role, bringing that, adopting that in. Yeah.
Does that kind of come bottom up as this something we should do and then approved by the government? Or has it come Well, that was almost the other way around. It's like, like open Approached.
Sometimes that happens. Not even, so open telemetry only exists because open census and open tracing merged. Okay, good point.
Yeah. So like, like, it was literally a, a group of meetings where, where I was working open census, we, we had like say like, like Ben Sigman and Ted from LightStep who were working on open tracing and, and various others. I don't mean to, to diminish the group, uh, where we're coming in and saying, there's two projects here.
They're maybe not technically in like direct competition, but they effectively are. Yeah. And because there's two different APIs, two different implementations, two different ways to extract this data from applications.
No one's building the integrations the instrumentation needed on either of them to make it really successful. Because if you're maintaining a database or a language runtime or something, and a end user comes and says, Hey, can you implement tracing or metrics or some way for me to extract my data and send it anywhere? And you see there's two competing APIs or implementations of this, you're often just gonna stare at this and just close the issue.
Like, it's not that you're gonna even pick one and that half the people will pick one and half pick the other. That that would be bad. But it's even worse where you just do nothing.
You stare at it and go, like, this is kind of a joke. Like, I, I don't, I don't need this right now. And you move on to doing something totally different.
And that's, that's where I think Open Telemetry, the fact that it was generated out of both of those and created the single standard for, uh, instrumentation, for data collection, data extraction, uh, made it very, very successful. Like, yes, there's the super positive community behind it. Yes, there's the 1,100 developers every month working on it.
Those are all massive and big. But if, if nothing else, open Telemetry made sure there was a single sane option in this space that people could rely on. And that alone accounts for a huge amount of the project success.
So you're wearing a couple different hats. Yeah. In some ways you could say sometime they might be in conflict, right?
What you might want is the director of product management for Splunk, eh, might not be exactly what the rest of the community once. And I think in theory that's very possible. And I mentioned like there was some concern about That.
I'm not, I meant not accusing you of anything. Oh, no. Yeah, no, I mean, I just, I could see a scenario where, yeah, I, I think the, the, one of the smartest things we did of Open Telemetry is kept the focus, uh, the project focused strictly on extracting data from applications and infrastructure.
Open Telemetry is not a backend, it's not a replacement for Prometheus or Splunk or Datadog or Dynatrace or, or Jager or anything else. It doesn't do that. All it does is take your data out, ensure it's nicely formatted, and you can send it to the locations you want.
And I think because of that, the commercial interests are generally not threatened because the commercial interest for these companies and for the open source solutions here, like Prometheus and Jaeger, is just like, we want to make this as easy as possible for people to get their data out. And we want it to be well structured, right? We want it to be sensible.
net, to have the same attributes and things on them. Open telemetry achieves that, that's all those vendors wanted. And the other thing is, I don't think any of them were satisfied with their own offerings historically.
Right? Like if you wind the clock back to like 2016 when I was getting into space, the incumbent APM vendors, like you'd go to them and if you, let's say a, a Java application or do net application, they'd say, oh, yep, we, we work off that. And you'd buy their product and they'd still have to send a bunch of field people out to like go make the integration work.
And meanwhile, if you came to them, I say like, PHP I'm making up example. They might've had good PHP support, but like PHP or for some of them may be Ruby. They would say, oh, weda, you can't use our products.
Like, it doesn't work with that. And so I don't think any of them were particularly happy with the own, with the solutions they had built. There were customers that they had to turn away, or customers for whom the experience was always kind of bad.
Uh, and secondly, they were spending a ton of money on agentry. Like these were big teams of very experienced engineers, like doing automatic instrumentation for a, a Java application or, I mean, pick your language of choice. It's not a trivial thing.
You need to hire people of deep intimate knowledge of the JVM or the T net CR for T net or the Python runtime for Python and so on. And so they're spending a lot of money on this. It wasn't working for them, and it was shrinking the overall market, or at least not letting the market grow open.
Telemetry is in many ways, sort of uncorked the bottle, right? Like those firms don't need to spend as much on instrumentation or in theory, they don't need to spend as much instrumentation as they did in the past. The, the data comes out in a beautiful, like, you know, like well structured way.
Uh, so it's, it's always, uh, consistent, uh, and better, better yet their customers can quickly adopt it and they can switch solutions if they want. And so it, we did find a way by scoping the project down to make it so that the, the interests of the, the, some of the vendors who are backing it aren't in conflict with it. I will also say with Open Telemetry in the last few years, like early on it was like Google, Microsoft, Splunk, white Step, like a bunch of observability or cloud vendors with observability offerings.
Some of the biggest contributors now are end users, like companies that are, you know, making money of open telemetry and that it makes their own services more observable internally. It makes them more reliable. But it's companies that are getting so much value out of Open Telemetry that they're putting engineers back on the project to make it even better.
And we've seen more and more of that every single year. It, it seems like, it's interesting how it's evolved to you think kind of an centric model, if you will, from APM to securities the domain. Yes.
Well, adopted Splunk's had great success there. It has, yeah. Um, domain driven design and applications and architecture adopting, you know, using open Telemetry within your application, through your stack.
Yep. Right. I, I'm curious, do you see AI as kind of another domain to start to, that can benefit from using Open Telemetry?
Or like, I, I'm not a AI expert by any means, but, but when you're analyzing reams of data, 'cause you're building a large language model or something, uh, or, or any other type of ai, right? Like so much of our current ML technology is based on just like massive, massive data being piped in for training. It helps a lot when that data's well structured.
And so, yeah, open telemetry probably isn't necessarily directly helping build like a large language model that's based off of like human text on Wikipedia or Reddit or something. 'cause that's human readable text. The whole point of that thing is to pars it.
But for any ML model or ai AI work that's modeling the type of data you would get from open telemetry metrics, traces, logs, machine data, really, then yeah, it's incredibly beneficial because that data's well structured. And, you know, I'll take logging as an example. Like we're announcing tomorrow at this conference that open telemetry logging is ga this is the third major signal type in open telemetry.
This is a big deal because we have metrics already. We had traces already, now we're writing logs. Those are the three data types that most firms are focused on capturing from their services, which means they can just exclusively rely on open telemetry if they want.
Uh, this is, This is particularly important for logs though, because logs historically are, generally they're unstructured. And when we started working on logging open telemetry, we wanted to only go into it if we felt we could actually improve things. And we talked to large companies that write a lot of logs or vendors.
They had two major problems today. One is the performance of the logging agents. Well that's somewhat fixable and not particularly relevant to the AI conversation.
But the other is logs often come in with wildly different formats, right? You have a service over here written by one team and one over here written by another team. Your timestamps come in differently.
This one might have the right host information on it, this one doesn't. Or it's structured differently, Level of specificity and it's All over the Place what the message is. Or even someone just structured the key for timestamp, like time dash stamp.
And the other is timestamp without the, without the hyphen. Right? And, and so you have to do a ton of pre-processing to make that work.
Or in many cases you don't and you can't query it properly. Well, from an ops perspective, this is incredibly annoying, but at least there's some fields where you can spend a bunch of money, do a ton of work, either in your queries and pre-processing to format it sort of consistently. But if you're trying to train an AI system, and again, not like a large language model that's trying to learn human language learning or something.
Yeah. But some like yeah. ML system that's gonna alert you to, to, uh, anomalies or something.
Well the fact that your data is totally wildly differently structured is gonna, it's gonna really gonna limit your, your effectiveness, right? With open telemetry. This data's coming in, even your logs is coming in properly structured.
It's part of, it actually has a data model in open telemetry. And so not only is it more performant to capture, but, and, and pre-process if you even need to. But it's coming in where all your timestamps, your machine IDs, your service information, all of the additional metadata you might have about certain interactions or events that generated those logs are all the same.
So now you can, you can have success much, much earlier. 'cause you can actually focus on training that I AI model training that system instead of spending all of your time formatting things and running outta time to actually do the real AI work. So no, I think it'll be, I think it'll be very effective caveat.
Again, I'm not, not, not some AI expert, but, but Well, And and I It's a massive benefit. Yeah. That's a whole nother domain.
But it seems like the common denominator for everything is everything that produces like me. Well there's, Well, and there, there's sort of examples I can draw to where there was a big push a few years ago amongst various AI companies, like nothing to do with cloud ops and stuff, but to do like AI analytics of health records to say like, oh, we'll look at these health records of a million people and we know the ones who had a certain type of cancer and the ones who didn't. And so we'll see if we can find a pattern in their health records that where we could bend in the future, look at other records and say, Hey, maybe you should get checked out for this certain type of cancer.
'cause you have some similar symptoms or things you went to the doctor for. And my impression is that most of those efforts failed because the health records are written by people. They're all over the place.
Have You seen a doctor's handwriting? Yeah. More or less what's in it?
Well, And even, even once you parse the handwriting, but like, like even once you parse it, it's, it's like short form notes, The details of what's put in in Yeah. And, and so my impression is none of those efforts are very few were successful, at least the ones I know about. And so it's kind of, you know, there's a corollary there.
It's a similar problem where if your data coming in isn't particularly useful either 'cause it's misshapen or it's just missing things to drive insights, even with a very effective AI model, or even frankly people looking at it is gonna be really limited. And so the same benefits that Open Telemetry brings to users of like Splunk or Prometheus or Grafana or whatever, like, you know, choose your observability solution. The same benefits it brings to them because of structured data making their queries actually work and everything else will also apply to any AI or ML models built on top of this.
Great. So I mean, I think it's pretty obvious, but if someone wasn't quite clear, like why would Splunk and all the other companies dedicate the amount of time and resource? Yeah.
Not, I mean yourself, but there's a lot more to it, right? Yes. It isn't like you're the sole contributor here to what's happening.
No, no. Neither I nor Splunk like, like Open Telemetry is proudly like a true open source, open license, open governance, uh, project. Uh, where there's a huge number of different contributors.
Like spl, Splunk is probably one of the, the more prominent ones. But you've got Microsoft and Google and Dynatrace and, and and LightStep, or I guess ServiceNow now and various other companies and vendors, Amazon, who are all putting in a, a ton, a ton of time and a ton of effort on the project. And that's good.
And it's, you know, the, the CNCF owns the license and trademark for it. So there's not gonna be any licensing shenanigans over it like we've seen at times in other communities. Yeah, no, it is intellectual Property issues, all that.
It's All above board. It's open governance, we're all elected. There's no, like, there's no permanent positions like it, it's, it's, it is a true success story.
And a few years ago that probably wasn't that uncommon, but I, I think given some of the changes in the industry recently, that's a very nice thing, very stable thing to point at. Yeah. Yeah.
Uh, and, and I think it goes back to what we were talking about earlier where the, the commercial interests are are I think per almost perfectly aligned with what the end users want here, which is, which is great. And it seems like carving out that lane, everybody's clear, this is what we're Doing. It's very Critical.
We're not gonna like venture into what we all do. Yeah, yeah. Makes a lot of sense.
'cause you see other projects, like you can take the extreme example of like open core projects where a company has an open source version that's kind of lightweight and then the enterprise version that, that has more features. If you go to the open source version and make like submissions, that'll make it too good. They, they kind of have an incentive to make those go away.
Yeah. Because it competes with a commercial offering if open telemetry tree, you know, not to belabor the point, but like, because it's not doing analytics or anything else, it doesn't really exist. Well, congrats on the success and longevity you've had with, with the project.
We've had a, you know, we're expecting even more growth I think over the coming year. Like I mentioned Logging's going ga or is GA now for a lot of companies. I think that was like, like first off, open country is super six on the industry, but I think there's a lot for whom they looked at it and said, there's tracing and metrics.
It doesn't solve my full story for logging. Maybe I'll adopt it for those two, but not the whole thing. But maybe I won't.
With logging it's the clear solution, right? It it, it's the way you can extract all of that data you have today and over the next year that's gonna get further extended profiling is getting added to the fourth signal type that allows deep analytics into, uh, compute into like where you're spending money on, on given function calls that might be costing you a lot of performance. We're also extending open telemetry to more and more platforms.
So projects historically have been focused on backend services and infrastructure, typically in cloud environments, but not exclusively. We're adding support for client applications for an Android and JavaScript for, for web browsers and others. Uh, and so it's really gonna extend Open Telemetry out of just the data center Mm-Hmm.
And into client applications, even theory into embedded systems. Yeah. Yeah.
And that's very, very exciting. There's work that's even started recently on mainframe support for open telemetry. And this to me is just really the full realization of the vision of you want observability, you want insight into your entire estate that you've built.
Well, that goes all the way from, from customer interactions to finding out why they're slow or why they might be broken all the way down to the smallest of, of backend surfaces or infrastructure. It's a great example of why constrain it to one person's vision, right? Correct.
You can get a shared vision And fairness. This is always part of the, part of the vision, But Yeah. Oh sure everybody does.
Right. But this is a great example when you have a shared vision. Yes.
Yeah. Shared outcomes, It's expertise. Like there's a lot of people from IBM and Broadcom who want to work on the mainframe part.
This is fantastic. They're the experts at that. Like I'm, I'm getting they're Yeah.
From my guys. So yeah. So there's a lot of great things happening.
Open telemetry. Hass already been very successful and I think over the next year and a bit like we're gonna see that grow even even faster. We look forward to more great things.
So yeah. Likewise. Thanks for joining us.
Yeah. Morgan McLean, who is, uh, director of product management as well as on the governing board. Yes.
Yeah. One of the co-founders of Open and one of the co-founders. Absolutely.
Of open, open telemetry. I want to say, Tracy, thank you and thank you. They appreciate you joining us.
We will be back with another great guest. I don't know as much longevity with an open source project, but another great guest nonetheless. We'll see you in a few minutes.





