Splunk’s Patrick Lin Explores the Need for Observability
Patrick Lin, SVP and GM for observability at Splunk, dives into why organizations, despite economic headwinds, are still finding a need to invest in observability as IT environments become increasingly complex.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Patrick Glynn, who's senior vice president and general manager for Observability at Splunk, and we're talking about what is the level of investment in observability, given the fact that there are a lot of competing priorities in this current, uh, shall we say, downturn in IT spending that we've been seeing for the last few months.
And who knows whether we will continue, but we'll dive in. Patrick, welcome to show. Hey, Mike, good to be here.
To what degree do you think observability is, shall we say, recession resistant? I know a lot of folks are kind of trying to figure out how to deal with the complexity of it, but we have monitoring tools and there's always resistance. So what's your take on the overall market at the moment?
Yeah, it's a great question. Well, I think the place that actually start is that I don't think observability is a set of tools so much as it is a set of practices for IT ops and, uh, engineering teams to make sure that the, you know, applications and critical infrastructure they run on, um, are performing well and highly available, right? Um, and so I don't think that that is something that they can get rid of as such, right?
It's the way they should be doing this Now when it comes to the tools that they buy to support that, right? Um, I think that, you know, the basic question is really about, okay, what is it that they're monitoring with this, right? What are they, what are they making sure stays, um, you know, at the right level of performance and availability?
And I think that, you know, over the past few years, um, it's a great level of investment that gone, that went into, um, you know, building, um, digital customer interactions, right? Um, and, and so many, many companies now have that as, uh, a core way that they actually reach their end users. And so if you assume that they need to continue to do that, uh, then it also, I think, uh, stands to reason that they wouldn't really be, uh, looking to cut back on that capability right now.
Will they look for ways to make it more efficient and optimize it? Sure. Right?
Uh, but I don't think it's something that they can do without, uh, assuming that they want to continue to reach their customers in the way that those customers are accustomed to being reached out to. We've been doing monitoring for a while, and that's basically tracking a set of predefined metrics at its core. Um, do I need to hit a certain level of complexity before I appreciate observability and the need to kind of dive deeper into root causes of things?
Where is that line between what I have today is good enough versus what I have to deal with tomorrow is not so much fun to work with? Yeah. Well, I think that a lot of the, uh, a lot of the kinds of things that I think we see people applying observability to are, uh, to your point, uh, a little bit more on the cloud-centric end of things, right?
So they may have adopted microservices, they may have adopted containers, uh, you know, they may have a, um, you know, mobile, uh, application, that kind of thing, right? And, uh, the reality is that, you know, for, for most enterprises, it's not just those investments, right? That make up the customer experience that they're looking to deliver.
Oftentimes it ties into, uh, some other, uh, system that they have, um, you know, uh, potentially running on, on premises and, you know, potentially not, uh, not, uh, you know, using one of these newer architectures, right? So it maybe that the data lives, uh, somewhere else other than, you know, uh, sitting on top of one of the, uh, public cloud. And so whether they are in that more newer kind of cloud native environment, or whether they've got that more, uh, complex hybrid, uh, environment, right?
I think either of those, uh, kind of necessitate, uh, what we think of, uh, typically as observability, right? Um, but I, I wanna emphasize that I think, uh, from Splunk's perspective, right? Um, I think we've always been in the observability business, right?
Like, I, I think if you go back to the kind of founding, uh, capabilities that were provided and the use cases that they were, um, built on, right? It was because, uh, the founders had built applications that they were having a hard time troubleshooting, right? Um, and so, you know, it's not that I, I guess I wanna go back to the point that, you know, this is not just because there are new environments, right?
Observability is, is a thing that we've been doing, and you have to be able to not just understand, um, you know, how things are going and when you have a problem, but also, uh, to be able to troubleshoot it, understand how to restore those services quickly as possible, right? Um, and, and so I think that while, you know, the, the newer spending and the newer category, uh, uh, as such of observability, um, you know, uh, gets associated with cloud, that's not kind of the, the total of it, right? How smart can the platforms get?
I think some of the folks that I talk to, they love the idea, but they struggle with figuring out what questions to ask. 'cause they don't know how to form the query. So at some point, will the platform just tell me what I need to know?
I think that is a great, uh, kind of north star for all of our, uh, all of our systems, right? I, I think that, uh, fundamentally what, what you're kind of pointing to here is the role of, um, ai, uh, when it comes to observability, right? And I think that, um, you know, uh, with, with Splunk, AI's been a thing that we've invested in for, you know, a number of years, right?
It's, it's not like AI is new just because of, uh, you know, the, the rise of chat GPT, right? Um, and then I think we've applied it to all sorts of use cases like, you know, anomaly detection, for example, right? Um, but I think, you know, kind of to your point about, uh, you know, potentially simplifying what users need to understand in order to get to, um, let's say the root cause of an issue or, you know, how to restore their service, there's definitely a role there for, uh, generative ai, right?
Um, I, I think our, our view is that gen, you know, it's gonna be something that can really turbocharge observability because it will help remove the barriers, uh, that people have to, um, being able to get the information that they want, uh, about their systems and their applications, right? Um, that, that they can sort of, you know, chat with their data as it were, instead of, uh, needing to, you know, parse that through a query language that they need to learn. Are we therefore on the cusp of democratization of observability?
Is that where we're heading? That is where we are headed. I, I think democratization, uh, is one way to put it.
I, I think, you know, the way I almost think about it is that we're making everyone an expert, right? Uh, rather than it being a, you know, lowering the bar, right? It's something where we say, E everyone can now, uh, get access to that information, and we don't have to rely on that one person who, you know, has mastered things, right?
Um, because it's one thing to, you know, understand that the systems themselves is another thing to understand the tools that help you interpret the information about those systems, right? Um, and I think that the information about those systems actually, uh, can be, uh, made more readily available as well, right? So you don't, you know, uh, have necessarily, uh, people who are like super, um, uh, specialized on one part of the system or another.
You can get a more generalized, uh, understanding, uh, through the tools as they improve as well. I know you guys are in the middle of an acquisition with Cisco, and there's probably not much you can say about that, but are we kind of also looking at the breaking down of all these different silos that we grew up with? It was, we had networking specialists, and then we had application specialists and cloud specialists, and yet our applications touch all these things and are dependent upon all of them.
So do we need a more holistic approach to observability in the first place? Well, absolutely, and I mean, I, I think that the whole premise for, uh, what Splunk is doing, um, as a, as a standalone business is that we are trying to provide a more holistic understanding of what's going on, right? And so whether you're talking about, you know, on-premises infrastructure or the cloud infrastructure and services, or third party applications that are part of the overall delivery of a business service or, you know, the cloud native microservices that you've just built, right?
Um, that kind of, uh, spread in the, uh, the environment and the complexity that comes with it, that's what Splunk is helping to, uh, diagnose. Now, you can that potentially get better, uh, as we, um, uh, become part of Cisco. Uh, sure.
Right? I, I think we're very excited about the potential, uh, that, uh, that might bring, There's a lot of vendors out there that throw around the phrase observability. What ultimately differentiates one platform for another.
What should people be looking for? Yeah. Well, I think that, um, I've spoken to a couple of them, right?
Um, I, I think one of them is certainly around the idea of having complete business visibility, right? I think, uh, sometimes when I hear people talking about observability, it's, oh, I do logs, metrics, and traces, right? And really all that's saying is, Hey, I can collect data in different formats.
But that's really kind of, uh, just the barest part of the foundation. What matters is, uh, actually the ability, uh, to, you know, knit that together and have people understand what's going on so that if there's a problem, uh, or if there is a change in their environment, they can understand, uh, the impact, um, of that, uh, of that, uh, uh, problem or of that change, right? And so that takes, uh, some amount of shared context, uh, shared data, shared workflows, right?
So that's one piece is this idea of complete business visibility. And then I think that's a lot of what we've been delivering, uh, you know, as I was just talking about, uh, with, with Splunk. Um, I think, uh, a second part of this, uh, is really about the ability to find, uh, and fix problems faster.
And this goes to your kind of ai, uh, related question or the, the prompt earlier, right? About, you know, how, how, how do we make it easier, uh, for people to root cause what's going on? You know, can we make it easier for them to more accurately and, uh, more quickly detect, uh, what problems they're having and when they need to pay attention, right?
And so that's an area, uh, that I think we stand out, uh, in. Uh, and then I think the third piece, uh, is really around the control of the data and the consumption, uh, of Right, having some level of flexibility and control and how to optimize, uh, how you're taking in that data and making the most efficient use of that, right? Um, so I think, uh, you may know that Splunk is one of the leading, uh, contributors to the open telemetry project, uh, which is, you know, a, a great way to bring in, uh, multiple classes of data using a single collector.
Um, and it's something that allows you to kind of own your data rather than have it be trapped, uh, behind a proprietary agent, right? Um, Splunk also has been investing quite a bit in making it possible for people to transform the different types of data they bring in, and also control, uh, what comes in versus say, uh, goes to, um, you know, uh, you know, the, the black hole because of data that you don't care about, even though it's being sent, right? Um, and so there are capabilities both in the, the Splunk enterprise and Splunk cloud platforms as well as, uh, Splunk observability, uh, for us to be able to provide that level of flexible consumption, right?
So those three things, right? Complete visibility, um, you know, finding and fixing problems, uh, faster, uh, and then having full control of the data. Those are where I think Splunk really shines.
And I think that that's what customers actually, frankly, are looking for. Do you think we will eventually reduce the level of stress and toil that IT professionals experience on a semi-regular basis every time there's a major incident and everybody gets dragged into a war room, um, and then I have to sit there and prove my innocence as it were? Um, is there just a better way to think about all this?
Yeah, absolutely. Right. I, I, and I think that, you know, toil is a, is a key word.
Innocence is a key word, war rooms, right? These are all things that, uh, we, we, uh, uh, kind of aspire to, right? Because I think that the, having the war room implies that you don't, uh, know actually what's going on, and you need people to bring their different versions of reality, uh, from different tools to, to, to, to the room so that you can kind of reconcile, right?
And so I think having a shared foundation on, on the data is pretty important for that, because then, uh, you get rid of that, you know, Hey, my tool says this, but your tool says that kind of thing, right? Um, and, and in terms of innocence, well, that's where having that shared, uh, data and context again, uh, matters, right? Because when you have the data, uh, more easily correlated across, uh, the different types and different layers of the stack, then, you know, uh, kind of what the leading suspects are, uh, for you to focus on, uh, when you need to troubleshoot it, right?
You don't need everybody to come and, and attest, uh, to, to what's going on. And if you can do that and you kind of, uh, reduce the level of effort in doing that correlation, then the toil and the stress, uh, goes down as well. So what is the big hurdle to overcome to get to observability?
Is it just the acquisition of the platform or their cultural issues? 'cause on the face of it, it would seem like a natural thing to do, so why aren't we seeing more people doing it faster, as it were? Yeah, so I think that there are, um, you know, obviously a couple different parts to this, right?
Because I think, um, observability, like I said, upfront, right? Is, is a practice, and it involves, you know, people process and, and tools, right? And so, um, having, um, the data platform and the analytical capabilities and kind of all the built in visibility, uh, that we provide, uh, is a key part of it, right?
But you do have to have, um, you know, uh, people who are, uh, incorporating it into, uh, their workflows, right? Uh, and also, uh, making sure that things are being, uh, brought in correctly from, uh, from the perspective of the data, right? Um, and, and you know, if you have that and you have kind of, you know, a, uh, well orchestrated process around making sure that that happens, uh, then, then it can really, uh, really work, right?
And I think this is a trend that we've started to see, right? Is the idea of a center of excellence, uh, in a number of, uh, uh, organizations where you have, um, you know, a shared team that that provides, you know, tools and platforms and so on, taking on, uh, observability as an area and helping to make sure that, uh, that the teams that, uh, uh, can and should make, uh, use of it actually do. There is of course, a lot of data floating around and data comes at a cost, and there's data engineering practices.
Are we getting better at managing the data? Can we be more efficient? 'cause some folks I talked to are a little overwhelmed by the cost of storing all the data and what goes into that.
Yeah, absolutely. Well, I think this kind of goes to the point I was, uh, mentioning earlier about, uh, the flexibility on the ingestion, right? And so, just to take one example, uh, in Splunk observability Cloud, we, uh, released a capability last year that we call metrics pipeline management, right?
Uh, and basically what it does is it allows you to look at the, the data that's coming in, uh, allows you to see what of that data is being used, let's say, in your alerts or your dashboards, right? Uh, and what parts aren't, right? And for the parts that aren't, uh, if you, uh, want to, you can, um, uh, simply say, let me filter those out because, uh, I don't need them.
Uh, or, uh, more likely, uh, we see a lot of people, uh, looking at ways to aggregate that data so that you can use them, uh, for alerting purposes. But then let's say if I wanna troubleshoot it using the traces or the logs instead, I can do that, right? Um, and so, uh, that capability, um, really kind of around the processing of the data, right?
It's not so much, again, the ingestion, but the processing of that data as it comes in, uh, where I think there's a lot of, um, interest in what we're doing. And I think that it kind of goes to your point, uh, okay. You know, can we help people manage it better?
For sure. Right? And, and we're definitely dedicated to helping people do that.
So what's your best advice to folks as they get down this observability path? There's clearly some maturity curve that goes with that. How do I kind of get in front of all this to have the best outcome on the back backend?
Sure. So I think that, um, oftentimes the first step, uh, that we see people taking is actually looking at, uh, how they get the data in, right? And so I think, uh, our advice there is to look at, uh, open telemetry.
Now, Splunk, uh, provides an open telemetry collector. It's also something that can be used as an add-on, uh, with the, uh, the log forwarding capabilities that we have. And so it's a great way to unify the data collection, right?
Um, and I think that's one big simplifying step. 'cause oftentimes, you know, getting, uh, the, the data off of the systems and the applications, uh, is a bit of work on, its on its own. Um, I think once you get, uh, past that point, um, you know, I'd, I'd say first take a look at what kind of out of the box, uh, capabilities are provided to you.
'cause oftentimes, uh, what we have is gonna be good enough so that if you've got, let's say Kubernetes, which tends to be newer for a lot of companies, um, take a look at, uh, all the work that we've put into making that kind of system, uh, um, pretty understandable and easier to troubleshoot, right? Uh, but then once you get past that, oftentimes there will be a need to, uh, say, well, actually I wanna be able to tune things. Uh, let's say I wanna change some of the dashboards that I use, or let's say I want to add in certain, uh, you know, custom metrics or some additional, um, you know, metadata, uh, by virtue of, uh, some, some custom tagging or whatnot to be able to provide, uh, better context, uh, on what I have so that it's very specific to envi my environment.
So I know that, let's say if this thing is, uh, having a problem that it affects, uh, this, you know, application or this service and this set of end user, right? Um, and so you sort of build up that capability over time, um, and you sort of continue to iterate on that as your, uh, applications, uh, and services you've delivered continue to deliver, right? But getting the data in, uh, first, um, making sure you understand what, what is there and then, you know, kind of, uh, customizing from there, right?
That's really kind of the approach that I think we see a lot of customers take it. All right, folks, you heard it here. Observability will make your life easier, but it all starts with the data.
So you gotta start someplace, and maybe it's some good old fashioned data management is where to focus your efforts in the short term to get the benefits in the long term. Hey Patrick, thanks for being on the show. Hey, Mike, thanks so much.
It's great talking to you. Alright, and back to you guys in the.