OpenTelemetry with Splunk’s Morgan McLean at KubeCon Paris 2024
The OpenTelemetry project is now a huge part of every company’s observability strategy. At KubeCon Paris 2024, Morgan McLean and Mitch Ashley discuss its impact across the industry and how firms (including Splunk) are taking advantage of it.
Transcript
This is Textron tv. Hey, everybody. Welcome.
We are here day two at, uh, at, uh, C**n Cloud Native Con here in Paris, France 2024. Weather gets better every day. It's been beautiful.
And, uh, this week I think it's getting kinda even warmer. It's lovely. So great time to get out and do some good things.
And I tell you, the crowds here at Kup Con have been kind of blown me away. I mean, the attendance here is twice what I expected. I mean, you kind of think about the progression, you know, from the last couple of Kup coupons, and they were good.
They were building back. This one seems like we've jumped up there. Definitely.
Let me introduce you first. So, so, you know, you're just not, some random guy showed up and started talking here. Goodness.
Morgan McLean with Splunk, and also one of the co-founders. Yep. Right.
With CNCF and on the governance board. Or Open Telemetry. Open Telemetry.
Open Telemetry. Yeah. Yeah.
You just found the whole thing yourself. Right. Be quite Something.
Gimme a lot of credit here. Um, so anyway, well welcome. It's good to see you again.
Yes, likewise. Great to see you. Yeah.
Nice. You know, um, so we were talking, we're just a few days past the Cisco deal coming through Yes. On Monday.
Yeah. On Monday. So, um, I was, I was like, wow, that happened.
I didn't realize that was gonna happen that quickly. I, I, my impression is it happened sooner, sooner than expected. At least that that's what they'd said publicly and, and to Wall Street.
Um, so I don't have a huge amount to share on it as a result of also being here while, while most of the wheels have been turning. But it is exciting. I mean, I mean, Splunk certainly is, is an observability and, and, and a huge security vendor.
Uh, Cisco has its own, uh, uh, huge investments in things like AppDynamics, uh, Cisco Cloud Observability, uh, and various other, uh, uh, security solutions like thousand Eyes. So I think the, the combination of those is, is very, very exciting. Yeah.
When, When the, when the announcements was made about the acquisition, very clearly, Cisco positioned it as a security. You as a security company, we know you, you do security and A lot, I, I think if you ask And More the average person, I think the security part tends to be the, the bigger, more well known part. But of course, it's long histories and observability, and that's certainly what I work on as well.
Yeah. Do, do you anticipate being positioned primarily as a security in the security solution space or that being won won along with, you know, I'm the wrong person to ask. You're Okay.
All right. Like, Uh, I, I know we're working hard planning things for what's gonna happen next for observability. Uh, I'm sure Cisco or someone way above me on the pay grade, we'll have more to share on, On those decisions will roll out eventually here.
Yeah. Okay. Yes.
Probably in your email. You know, our emails stack up probably When I finally catch up with all my emails. Yeah.
Well, good. Well, good. Best of luck and congratulations.
Yeah. Well, I think it's exciting for us, exciting for our customer. It's exciting for Cisco, uh, certainly as well.
And, uh, we'll see what comes next. I mean, certainly in my neck of the woods, uh, Cisco is also a major contributor to open telemetry. Mm-Hmm.
One of course, you know, everybody knows, has been one of the biggest. Uh, and so that alone is, is a really interesting combination, right? Because we have two companies that have already made that major bet on this, on this open standard.
Yep. So, I'm sure just from Open Telemetry and our usage of it and the expansion of the standard itself, there's gonna be a lot of great things, uh, in the road head. You know, I was trying to think what's changed from the time when the deal announced to when it's now happened.
Yeah. And one of the things, not directly, but, but still related, is how much generative AI has taken off and the whole AI mantra. So tell us a little bit about, it doesn't have to be about Cisco specifically, or, or it could be about Splunk, but also what's happening with Open Telemetry.
Where is AI fitting into the whole picture? Yeah, like certainly at Splunk, we have our own, uh, uh, AI investments we've been making across the product portfolio. I think we've shown off a fair amount of what, uh, will be available from Splunk Enterprise or possibly Splunk Cloud rather, uh, for what people can use for that.
Not something I'm incredibly familiar with, uh, but, but, uh, the, the demos I've seen both internally and externally are quite exciting. Uh, Splunk observability is also making its own AI investments, but to me, like the, like, you know, just for reference for everybody, I mostly work on open telemetry instrumentation, getting data in. Uh, and so for me, the exciting part of that is, is how do we extract data from these, some of these AI systems, uh, and, and secondly like, what AI analytics can we perform if we have the correct data?
And I think last time we chatted, uh, possibly in Chicago Mm-Hmm. We were talking about like the value of open telemetry for AI based systems in that open telemetry. Yes.
People know it as a set of SDKs and agents and things that allow, allow you to extract metrics, traces, logs soon, profiles soon, other types of data from your systems, and go and put that, uh, go and analyze that somewhere. That's great. Uh, one of the other nice things about Open Telemetry is the data is really nice and structured.
And so if you have AI based analytics, much like if you have so traditional based analytics, having that data that's properly structured is gonna drive significantly better insights. Like an AI system cannot compensate for having data that's really, really low quality. Mm-Hmm.
With Open Telemetry, you're ensuring that your data's high quality. So we, we've talked about that at length before. I think it's relatively sort of self-evident, the benefits of that.
There's also work going on in the Open Telemetry community that's just started. So like literally in the last few weeks, people have been coming together on this. So there's nothing to show for it yet, but there's, there's lots of interest in this, and that's using open telemetry to monitor, uh, ML and AI systems, right?
So I, I believe one of the, the big, uh, goals of this amongst, amongst others is that if you are actually going and creating your own AI model, you are burning tremendous amounts of compute, right? Typically on GPUs to go do all the training for it. And Open Telemetry, certainly in the past could be used to extract metrics and traces and logs and other things from those systems.
But there were no sort of first class affordances for it. Like, it is not like people in hotel had sat down and spent hours, days, weeks, months thinking about like, Hey, if I was doing an AI training system, what data would I want? And how do I make this easier to use?
And so what's happened recently is a lot of people getting together, uh, to actually focus on that within the Open Telemetry community. So I think in the coming weeks and months, we're gonna see like what their actual plans are, and then they'll sit down and actually implement that. I would expect there'd be additions to our semantic conventions, uh, to at least sort of normalize the data coming out of training, uh, sort of feeding back on what I was talking about earlier, about having better structured data, but also various integrations or hooks, uh, into systems that might not exist today.
That mean that if you are building your own AI model, you can better monitor it, better understand where your costs are going, and, uh, perhaps optimize those. It's, it's, it's great to hear that news too, because AI can be such a black box to many folks. No kidding, right?
Yeah. Except the people that build it off the time. And sometimes even that is not, I even, it's not determinist exactly.
In a way that like most software thus far in human history has been, you know, very, uh, sort of imperative. Like, like very, like, you know, you're a developer, you write a bunch of code and you can read what it does, and you know what it does. And, and certainly for these AI systems, it's training something that like, yeah, we, we understand it's a giant, uh, But it's discreet, right?
It's discreet man, discrete logic, and you know what The Results should be under a certain set of conditions. Yes. And, and so you're right, it is more of a black box that, that people have sort of less visibility into, uh, whether that's the AI model itself, responding to various prompts or, or other things, generative ai or if that's the, um, the actual training process for the ai.
And so getting more visibility, that's gonna be exciting. It means you might be able to be wildly more efficient with your AI training process, uh, or more efficient with the compute and memory and everything you're expending, uh, on your models once they're actually, uh, in operation. You know, I, I wonder too, how much we'll be able to leverage, we should be able to leverage the learnings from applying observability to security, applying it to distributed and microservices cloud native applications.
You know, AI's got its own characteristics. We've been talking about it and more. Um, so rather than just throwing data into Splunk or an observability platform, right?
What are the ways to, to help make sure, you know, observable, observable driven design, I guess the term Yeah. Is if I have that right, so that, that way you just start to use those kind of tools together with AI systems. Correct.
You get some more meaningful, higher value context aware. And I, I think that's the excitement is like I, I've seen around the marketplace, like not just at Splunk, like a lot of firms have, have launched some level of AI features or features where, and particularly using generative AI to do things like make queries easier to write, right? So I, I know Splunk has investments in this, they've been showing off, and then certainly other vendors do as well, where you can write a, you know, you can, you can type into a prompt thing, sort of like a query, but more in plain English or, you know, whatever other human language you use, Natural language.
Yeah. Uh, yeah. And, and it actually turns that into a query underneath.
And that that's very nice. Right. And, and there's similarly I've seen, uh, uh, AI assistance from Splunk and from others where they're showing, um, uh, uh, sort of questions effectively for documentation or help, right?
People like, how do I set this up? It goes and gives 'em a guide. But I, I, you know, in this industry where we're analyzing reams and reams of data to give people valuable insights into it, whether it's for security, whether it's for improving the performance of their services or infrastructure or stopping outages or what have you, uh, the, I think the, the, the thing that people have a bigger, long-term interest in is applying machine learning and, and various AI advancements to actually performing analytics on that data.
Right. Analytics that they couldn't have achieved before. Uh, and you know, I, I don't work on AI at Splunk, but I, I, you know, I imagine that there's, there's sort of major investments coming that, and so there, there should be, uh, fair amount of excitement, I think, going forward about what's come.
Very cool. What's Come, tell us about kind the hotel side. Yeah.
The C-N-C-F-O tell what's happening. Yeah. Yeah.
Lots of excitement in open telemetry. So when we talked in Chicago last November, open telemetry just added logs, uh, that was very exciting for the project. So if you're not as familiar for Open Telemetry, uh, it allows you to capture data from your infrastructure applications in 2021 or late 2020, we launched, uh, support for distributed traces 2022.
We added metrics as the next signal type. Last year we added logs. Uh, there's numerous projects underway right now, probably the most visible is profiling.
So there's, uh, a large group of people, uh, working on, uh, adding profiling as a fourth signal typed open telemetry. Uh, so we have a number of contributors, like, like Brian and Dmitri from ANA Labs and, and a number of people from Splunk and, and various other companies who are going and working on this. As someone who much earlier in my career was working on a distributed profiling tool that was incredibly powerful.
But I never saw profiling tools generally gain, like mass adoption. This is very, very exciting and very, like, validating to things that I'd worked on a long time ago, finally hitting the mainstream. Uh, because you, the power of this, I think, is not quite understood by a lot of end users or, or, or just sort of the general public, like profiling the insights that can derive, particularly around like cost management, how to better optimize your compute, how to make things faster, is incredibly potent.
Uh, and a lot of the biggest firms in the industry, Google and, and, uh, meta and various others have made major investments in this long ago because they saw the benefit of it. Bringing this to the mainstream so that everyone can take advantage is, is huge. Like I say this as like, you know, Splunk, Splunk observability, we've had a profiling product for some time.
It's been very effective, been very successful. But when I look broadly, it's still not as well known as, say, metrics as a signal. It's not as well known as like metrics and logs or anything else.
I think Open Telemetry fully supporting, it's gonna give it that boost it needs and make it really easy, really ubiquitous for people and easy to set up. So that's probably the most visible thing that's happening in Open Telemetry. Uh, but there's tons of work going on at different things.
I mentioned the semantics for ai, but there's actually work overhauling our Semantic conventions right now for most different types of data. net and one from Java over here, they have the exact same semantics about latency and, and error rate and throughput, and the host are running on the service they're running on. And so that effort is perhaps less visible, less talked about, but it's really, really critical to the project success and, and to delivering value to, to anyone who uses Open Telemetry.
Uh, and there's, there's numerous other things we're working on in the community, uh, beyond that. Like obviously we have a huge number of developers. I think in November we said we had 1,100 monthly active developers pushing code into Open Telemetry.
I, I haven't checked the numbers since then, but one imagines it's gone up Pretty amazing numbers. But yeah, it's Incredible. It's the second most active project in the CNCF.
It's been that way since 2020. I'll be honest, when we started Open Telemetry, I thought maybe 50 people at most would work on it. That Would be a pretty good number that that's, you know, I would've been very happy with that.
We than two, two, uh, two beats a team, right? Yeah. And, and, and you could ship some stuff, but like the number of, like, the amount of things we've shipped, the number of languages we support now, the fact that we have full support for V Logs already, uh, and we have this huge community pushing features into the collector, like things like pre-processing that people now take for granted and use all over the place.
And that's great. Like a lot of this wasn't even originally in scope for the project, and it's amazing that we've delivered these things and now they're, you know, you walk wrong with floor here. They're super common.
And so like, like, just because I'm mentioning profiling and, and sematic inventions and other things, like there's over a thousand people pushing code into the thing every month. Like, like there is a ton of work going on in every single language, uh, support for every single language with the core features of the collector and the protocol and everything else with an open telemetry. There's tons and tons and tons happening right now.
How well is Open Telemetry, uh, tied into kind of the infrastructure's code community? You know, it can, everybody always thinks of Terraform is one of the largest, of course. Um, but with GI ops and automation and Yeah.
That's as much as the software deploying now as Yeah. As the app code and other Yeah. Infrastructure software along with it.
So it's, I mean, it's tied in, in the sense that it's just another artifact really with configuration that you would use those tools to deploy and manage, uh, which is probably sort of the highest praise you could give it, right? Like, it's, it's just another thing that you would use with Ansible or with Terraform or with ever, with whatever configuration, uh, or, uh, management tools you're using or deployment tools that you're using. Um, there is work, you know, it's, I neglected mentioned earlier, there is work going on Intel about, uh, improving, its, its ease of management, uh, both through the op-amp protocols, so you can manage it live, uh, but also through different ways of configuring things, uh, whether it's through files or environment variables.
But regardless, like today, since day one, really, like Open Telemetry has been a very, I think, sort of successful part of that story because it's not opaque. It's not an agent that only receives commands from a command and control server somewhere over here that you can't actually manage, right? Like, because it uses the collector, for example, uses YAML files for config or looks at environment variables for config.
And the same is true for the, the language instrumentation that makes it really, really easy to manage with any standard tool. Uh, and if you're using the SDK to instrument things, well, that's just, that literally is code. Uh, and so that makes it easy to manage through GitHub and through, or, or really whatever code repository you're using and whatever deployment system you're using.
So one last question is, I promise it's not a trick question. Sure. I saw, I saw on one of the placards at a booth, yeah, I won't say whose it was.
Um, but the, the tagline was, monitoring is dead. You know, we had, DevOps is dead, and, you know, everything else has been dead at one point. And now, so, you know, Elvis left the building while he is back, right?
Yes. When, when, when you hear people say things like that, is, is monitoring still one of the fundamental, fundamental elements of this? Or we kind of move past that into a next generation of Monitoring's?
A very expansive term? Mm. So if I, I assume the intent of, of whoever wrote this was saying like monitoring as in like pure infrastructure monitoring, as in you're running a large, you know, highly distributed web application, and the way that you're determining if something is wrong is staring at a bunch of dashboards of your CPU and memory consumption and inferring from that Issue, SMP traps that are Yeah.
And like, you know, flagging. Yeah. That's, that's dead in the sense that that's, I mean, you still have those graphs somewhere and you do have alerts on some of that maybe.
Uh, but really what you're using, what you're doing is you're looking at the performance of your application, right? You have, whether it's using SLOs or just generally looking at, uh, things like the latency or error rate or throughput or availability of your endpoints or of your client applications or of sort of critical processes that are happening. Like, yes, the pure infrastructure, that style of monitoring is dead, but it's still a thing you're doing.
It's more that's expanded, right? You've gone from monitoring infrastructure, from just staring at CPU and memory graphs and others to looking more in depth at the performance of your infrastructure and getting breakdowns of it. And then your actual alerting and everything else tends to be on the performance of the application.
Uh, and, and the reasons for this are relatively obvious, right? If you're staring at a graph of CPU consumption and it's in the nineties, that might mean something's wrong. It might also mean that your system's really well optimized and well utilized.
Yeah, it's good point. Yeah. Yeah.
And you're not spending more on compute than you need to be. Uh, and so really what you wanna look at is like, what is the end user performance, right? When they're opening an app or my website or somehow communicating with my, my APIs or using my, my services, are they responding the way I expect, or are they getting errors?
Are they getting the right, the right, uh, responses back? Is it actually available? Are they getting anything back?
What's the latency of, of those responses and various other factors like that. But you may still eventually jump to those, that CPU data or something. It could One of the, because it might be root cause of one of the data elements, one of the parameters what you're looking at, right?
Yeah. Maybe a health status or Something. I mean, it's the reason we use tracing, right?
Like it's, it's the reason that that, that a PM existed. But, so yeah, I, I would say like monitoring is dead in the sense that like, that act of staring at very specific dashboards of infrastructure data, it's not really relevant anymore. Maybe Standalone is the state of the art.
Yeah. That isn't all we use today, But you're using an observability solution that includes that and many other things, right? So it's monitoring isn't dead.
It just, it grew into observability along with it just up a PM and everything else up. Yeah. It Grew up into something else.
And it's, it's, yeah, I mean, it's, it's a controversial phrase because even if I use an observability tool half the time, I'll probably use the word monitoring, describe what I'm doing with it, and yeah, it's not technically accurate, but like, it still describes the human act of using my observability tool. Got it. Well, thanks for spending some time with us and yeah, likewise, sharing what's happening.
Cisco, Splunk world and yeah, open telemetry and AI kind of hitting a lot of topics of stuff Happening. It's, it's a very exciting time to be at Splunk, I I imagine at Cisco as well. Just, just given that all this, all these great things are coming together.
Uh, it's also just an exciting time generally in industry of all this, this, this, uh, sort of, I think, uh, well, excitement, I know I'm overusing that word, but about what's happening with AI and ML and, and other systems. It's, it's, uh, it's invigorating. It's, it's really, Really cool.
It's a great time. I mean, yeah. The, the attendance here and the buzz here is a good Yes.
Good data point of yeah. Of demonstrating that. So great to see you again.
Yes. Likewise. Look forward to our next co con or whatever event it is that we chat, likely K con.
Yeah. Likely coup con. We'll have to send each other a note on predicting what we might be talking about then, You know, might be more of the same, might be totally Different.
Know you never know. Great. I said, we have fantastic interviews, wonderful people, talent, you know, people that are, you know, really leading the charge in, uh, in, in observability, in cloud native in our industry.
So thanks for listening to our first interview. Stay tuned. We have a whole day.
We are packed. Our schedule's full, so don't go anywhere, you know, same, uh, bat station, same bat channel. We'll be right here.