How to Build a Leading Observability Practice | The Six Five Summit
As engineering and ITOps teams have moved to the cloud and ramped up their pace of innovation, it often becomes harder to validate the impact of software and infrastructure changes – both to the business and to customers’ experience. Disconnected toolchains and an explosion of failure scenarios in cloud native environments have made this problem worse. As a result, it still takes too much toil, guesswork and expensive war room calls to find answers to problems or even know where to look. Join observability leaders at Splunk for a lively discussion on what it takes to build a leading observability practice. From boots on the ground experience working with our customers, we’ll share how high-performing engineering and ITOps teams are using observability to improve their digital resilience.
Transcript
Hello, and thank you for joining us at our six five Summit AI Unleash. Welcome to the session on how to build a leading edge observability practice. My name is Paul Nash.
I'm the practice lead for the application development and monetization practice at the Futurum Group. I'm thrilled to be here with Marla and Patrick from OB Splunk. Marla, would you like to introduce yourself?
Sure. Hi, Paul. Hi, everybody.
Uh, my name is Marla Pula, and I'm the GVP for OB Observability here at Splunk. Thank you, Marla. And Patrick.
Hey there, Patrick Lynn. Um, also from Splunk. I'm the SVP and, uh, general manager for observability.
Also, uh, with Splunk a, a Cisco company. I am really, really excited to have you on this today's session and to talk about the movement of observability and how it relates to AI. Today, we'll be talking to both of you on how engineering and I, a engineering and IT ops teams are moved from the cloud management to ramp up their pace and innovation also and how it becomes harder to validate the impacts of software and infrastructure changes.
So, let's get started. So Marla, let's start with you. What we saw in, in 2023 is a, a largely driven macro context where there's a lot of customers trying to rationalize their tool stack.
In fact, what we see in our own observability research, we see that 75% of respondents indicate that they're using six to 15 observability tools to gather organizational data. Do we expect us to continue in the near term, or would this be a long-term strategy look like? That's a great question, Paul, actually, in our observability research that we've done, the number of tools are much greater.
It's greater than 20 monitoring tools on an average that organizations have today. My view is this, right? I do believe this trend is going to continue, and as our customers mature, their observability practices, most often, they're also transforming their business models, and they're looking for more efficient ways to have a comprehensive view of the full stack across different telemetry types.
And also at the same time, optimize cost through rationalization. Some of the things that I hear from our CU customers consistently is they share similar challenges and opportunities as it relates to tool consolidation. And most often, I sum it up into three things.
One, what we hear often is the current macroeconomic climate is increasingly about improving operational efficiency within organizations. What does that mean? Drive more effectiveness with less solutions?
Second, what I hear is it's, it's around the ability for to be effective and have a comprehensive view across the full stack of infrastructure and application, as opposed to having different tools and different versions of reality. So when they do do that, what does that do? It it impacts the cost of downtime and service availability.
And the third most common feedback we hear is it's about driving cost optimization through rationalizing of tools and technologies and to eliminate that redundancy. So here's what I mean, right? On an average, like I said earlier, an organization can have greater than 20 monitoring tools.
So what does this mean? This makes it harder to diagnose when there is a, a system or service degradation. How do I find that root cause?
Which system, which application, which network or database is causing that degradation? And in some cases, customers are missing critical signals like failure or an alert or an outage, and it goes unnoticed. So I do see this trend on rationalization continuing.
Yeah, it seems like the, uh, the TCO for that approach is pretty high, right? When you have those multiple tools, and then, like you said, the integration has to work seamlessly in order to get that, that avail availability and, and, you know, just looking at how the systems are working. And so, Patrick, I wanna talk a little bit about, you know, when you look at the complexity and you look at the number of tools over the last couple of decades, we've seen a broad trend towards centralizing compute infrastructure in the form of cloud computing.
In fact, we see in our research, 94% of respondents are in, are using two or more distinct, uh, basically distinct cloud infrastructures and service providers. And, you know, and this means that there, they're more recently becoming clear that the certain applications in the use cases will require the infrastructures to be really closer to where the users are, right? And the data where they are.
So essentially when they're looking at computing at the edge, right? When we look at that edge and we look at clouds, um, what do you think, Patrick, from the, from the perspective of, of the trend and what does this mean for the observability at the Edge? Yeah, that's a, that's a great question.
Uh, Paul, so I, I think when we talk about, uh, things at the edge, right? There's a variety of, um, uh, kind of, uh, situations where we see that, right? Sometimes it's something as simple as, hey, it's retail, or, you know, a, a quick service restaurant or something like that where you fundamentally need some sort of, uh, in-store computing, right?
Uh, where it doesn't make sense to, uh, put all of that, uh, into the, the public cloud. I think there's also been a, a more recent trend of seeing people repatriate some of the workloads that they had previously moved into the cloud saying, well, actually, the predictability of those means that they can perhaps, uh, be in a different environment, right? And I think that in some cases, there were some things that never made their way, uh, into the cloud, right?
And so this sort of very distributed environment, uh, that you end up having, uh, uh, probably results in, in a lot of cases in, um, uh, critical business, uh, transactions or services being delivered across this very high ized environment, right? Um, and so it's not uncommon, I think, for us to see things like, you know, a new application that's been built where the front end is something that's in the public cloud, um, but, uh, it ultimately ties back to, um, a system or, or set of services that's still running on-prem, uh, because it never made sense to move it, or because they moved it back from public cloud whatnot, right? Uh, and so in order for that to be, um, observed and monitored, uh, properly, right?
Um, it's pretty important, um, to have a few different things. One of them is to have a pretty consistent way of getting, uh, the data in, uh, so that it is sort of, uh, consistent with, uh, itself, right? You don't want to have, um, silos of data based on where, um, a workload part of a service is being, uh, served from.
Right? Um, a second piece here is a set of, uh, capabilities that looks at that data and is able to show you, uh, kind of what's going on across it in a very consistent fashion, right? You don't want to have, uh, fragmented views again, uh, based on, you know, where, where the data's coming from, right?
Uh, I think the third piece, um, is that, um, because of that sort of distribution, you also need greater visibility into what's happening, um, across, uh, the network, right? Um, and, and, you know, in many cases across the internet, depending on how, uh, the, the application itself, uh, is structured, right? Um, and so it's more important than ever to have that be included as part of, uh, the, the visibility that you have.
Um, and then I think the last piece is that you need to have, uh, the ability to have, have, uh, the visibility both sort of in that on-prem, uh, or customer managed environment, as well as something that, uh, takes advantage of the public cloud, right? And I think, uh, overall, it actually, um, uh, is pretty consistent with what, uh, a lot of, uh, what ma was saying around, um, a, um, a sort of consolidated view across things and having full stack observability across that, right? Um, it is, by the way, one of the reasons why, um, I think the, um, acquisition of Splunk into Cisco makes so much sense because it's a way for, uh, us to be able to provide, uh, the connections across all those different source of information, uh, to use open telemetry as the, uh, common format for that, uh, information to come in, uh, right?
And for us to provide the, the sort of right tooling so that you can get to, uh, the root cause of issues, uh, as quickly as possible. Right? Uh, one last thing I'll add to this, by the way, is that I think, um, uh, sometimes, uh, there is a question about what data you wanna be able to bring in, uh, when it is not already centralized, right?
When the infrastructure's not already centralized. And I think the, the other piece that's useful there is the ability to, um, watch over the data as it's sort of making, its through the pipeline, uh, from let's say wherever that, um, infrastructure or application is located and where you ultimately are going to be, uh, doing your troubleshooting and monitoring and so on, right? Um, and so, um, having the ability to look at, uh, the data as this goes, goes through the pipeline, deciding whether you want to, uh, keep it, uh, drop it, you know, aggregate it, transform it, right?
That's another key thing, uh, that's important for people to understand, uh, how to use, uh, in, uh, in sort of the context of this trend toward, uh, some additional decentralization, right? The the commute infrastructure. Yeah.
There's a lot there to unpack for sure. I mean, when you look at everything you mentioned, um, I can back up most of what you just described with our research and our data, right? When we look at application, uh, portability, uh, 20% of respondents indicated that it's critical that their applications are portable.
But, well, one of the things you touched on was the repatriation point. And, you know, we do see that in our research. And when you, I talk about in the context of modernization, past, present, and future, you know, heritage applications moving to cloud native and such, and when refactoring occurs, um, that refactoring only 11% now is being done according to our research only done on-prem.
Um, when, and, and, and when we look at two years out, that refactoring on-prem is actually gone up over 30%. So there's a repatriation coming back from the cloud. So, but the, the point you were making, Patrick, about harmonization of the platforms and the tool stack, the tech stack to make sure you have that visibility across all areas is equally important.
Now, I do want to talk about that as we talk about modernization of applications. Patrick, I do wanna throw another question your way. When we look at, um, you know, there's a lot of interest around how security teams and developer and org operational teams can benefit from the signals from the each other's domain and their practice.
Um, what do you see in or hearing from customers around how they approach this? Yeah, that, that's another really, uh, highly, uh, uh, interesting topic. So I, I guess, um, maybe the place I'd start here is, is first by, uh, mentioning that, you know, when, when, uh, we think about, uh, security and, uh, development and operations teams, right?
Oftentimes we think of them as having completely, you know, uh, different, uh, objectives, right? And they're, they're sort of gold on different outcomes. Um, but you know, back, back in the day, maybe they weren't, uh, so separate, right?
Like, I, I think the kind of specialization that we see in larger organizations comes from, um, kind of the, the, the growth over time of these individual disciplines, uh, where you, you know, back, back in the beginning, they may have been using the same set of data, right? Um, and so I think that, you know, the, the fact that there's a lot of information that is captured in one context that tends to be quite, uh, useful, uh, uh, in the other, right? So one example of that might be you often find information about, um, assets, uh, that you have an identities that you have in let's say A-C-M-D-B, right?
That's managed by, uh, an IT department, right? Or, uh, that is being brought in, uh, to an observability tool, uh, because, uh, it's important for the development team to be able to sort of track all the, um, you know, whatever containers or functions or, you know, other things that are being used, right? Um, and on the flip side, I think on the security teams, they are often, um, kind of lacking, uh, very good visibility into assets and identities.
Um, and so being able to, uh, have that information, uh, be made available to the security teams, and then be augmented, uh, with more realtime information that typically comes in from observability, right? That's one, one example I think of how, uh, there's, there's some interest in bringing that, that data across, right? I think that the sort of, um, uh, conference of that is, is also true, right?
I think that oftentimes security is part of the context that the engineering teams need to operate in, right? And so if, uh, there's a, um, an issue, right? Uh, and there's a, uh, there's a spike in the traffic, uh, of some kind, right?
Then, you know, one of the natural questions is, am I, am I under attack? Right? And, uh, knowing whether something is being investigated, uh, from security side might be useful, uh, in those cases or in a kind of less urgent scenario, right?
Um, oftentimes, uh, dev teams, uh, spend a good chunk of their time making sure that their applications are secure and meet various, uh, compliance, uh, uh, needs and regulations, right? Um, and so, uh, to the extent the, uh, there's feeds of, of information that can be brought in, um, more from the security side, uh, to help contextualize that to say, well, you know, you need to prioritize your work in this way, right? It's more important to do that fix versus this one because, you know, your configuration here, um, is, is actually the one that is, uh, gonna be problematic.
Versus, yeah, there's a, um, something out there that indicates this area is problematic, but you haven't set it up in a way that actually is right. Having that very specific, uh, data that helps, uh, prioritize the work and ultimately let the dev teams get back to doing what they're supposed to be doing, right? Um, that's, that's great.
Uh, great information to have and to share, um, and ultimately kind of make the, uh, kind of collaboration, uh, across the teams better as well, right? Because they'll have a shared sense of reality, a shared sense of, uh, priorities. Yeah.
It sounds like, uh, a lot of what you were talking about is the whole shift left kind of nomenclature and kind of moving back to, you know, the security into, into the teams. And when we look at that shifting of, of responsibilities and, and how things are happening, you know, you see teams that, uh, or organizations that have DevOps, SREs, and platform engineering, but also DevSecOps kind of plays into it, and mala when we talk about maturity, because it depends on maturity and how these organizations are, are driving with regards to how they organize their teams. Um, observability in general is, is kind of, uh, in some, to some extent is really, um, um, an kind of an immature practice in a lot of organizations.
And then there's some organizations that really know what they're doing and they're really at the, the right end of the outliers. But like, when we look at the rapid maturity from alerting and monitoring to really actionable insights and how it's evolving rapidly, what guidance would you provide or have for the CIOs and CTOs as they think about their three to five year roadmaps plans for their organizations? Yeah, I mean, it is, it is an evolving space and domain observability, and it's a, it's an exciting space to be in as well.
I mean, just like Splunk, a Cisco company, uh, our customers are also evolving the observability practices, and it's very much a growing and thriving domain. Um, what we've seen is that rapid evolution in the past few years, both in the complexity in architectures and boundaries in businesses, and the rise for the need to ensure critical business applications, be it on-prem hybrid or cloud native, have reliability and are resilient and are also scalable. And I think one thing Covid showed us was the need to rapidly change business models to meet where your customers, where they are at.
So our customers are increasingly operating in this complex application landscape, like I said, be it cloud applications hybrid on-prem architectures, and not to mention the rise of AI applications recently. There is this need for CIOs and CTOs to drive standardizing absorbability tooling, uh, adoption across teams due to the speed at which some of these businesses operate. And fragmented ownership only hampers, I feel the holistic market viability of that solution.
So while working with our customers, we often hear, right, the adoption of observability when supported by product solutions across that, that varied application architecture and landscape. That's when value is realized. And most of the mo, most of the CIOs and CTOs we, we talk to, there are, there are on three value realization areas, typically, one, how do I improve developer productivity so that engineers can spend more time actually building and shipping code instead of managing and troubleshooting their tool chain.
The second one equally important is the ability to wrangle costs by maintaining governance and avoiding these runaway usage costs. And last but not least, right? Consistent practices to building digital systems, uh, that are built right from the get go to be observable.
So for CIOs and CTOs, my recommendation is it's just not about the products and technologies. They need to develop strategic relationships with technology players in this space that can effectively partner with them. As with the organizational roadmap, it is an evolving domain, like I said, and we are learning from our customers and, and vice versa.
So we, we, we expect a close vendor relationship with, with our customers. Yeah, that makes a lot of sense. And you know, I think the biggest challenge I heard when you were talking about is when you talk to these, these, uh, CTOs and CIOs, they're thinking about their roadmaps.
They also have skill gap issues, right? So working, like you gave the answer though, you working with service delivery partners to help get them where they need to go, will help augment their resources. And, and Patrick, you know, it wouldn't be a session if we didn't talk about ai, right?
We have to talk about AI, and you know, Marla talked a little bit about it, but like, when organizations are approaching AI and, you know, what does it mean for their observability practices? Well, Kyle, uh, I'm impressed we made it 18 minutes without saying ai, uh, because usually much faster for us to get to that topic. But, uh, all kidding aside, right?
I think, um, maybe to talk about ai, it's worth sort of stepping back and defining what we mean by it, right? Because I think these days, uh, the, the hotness is all about, um, you know, open AI chat, GPT, large language models, right? And the incorporation of that into, um, applications, the use of it for assistance and so on, right?
Um, but I think that AI kind of more broadly defined also includes a lot of the work that's been done over the couple of decades around machine learning and, uh, and so on, right? Um, and so I think that, um, where I would actually start, right, is making sure that, um, when, uh, we're talking about AI forward organization in the context of observability, right? Um, that first of all, uh, you're sort of making sure that you have, uh, the data necessary, uh, to be able to, uh, kind of identify when there are issues, right?
Uh, and to apply machine learning to that, uh, to make sure that, you know, when, um, there's something that is behaving in a way that you don't expect, right? And I think, um, you know, the machine learning part of that is just being able to do the, you know, what's wrong here in a more sophisticated way, right? Versus saying, oh, when it's above 90, that's bad.
Below 90 is good, right? That's way too simplistic for, um, the kind of real world scenarios that, that all of our customers have, right? Uh, and so, uh, so that's one piece, right?
And I think that, um, the, the other piece, uh, to this is about being able to take advantage of, um, uh, generative AI and what it's good at, uh, for helping to narrow down, uh, the source of issues, right? Um, it's actually pretty interesting, right? We've been doing some experimentation with this, uh, uh, internally, and it used to be that, um, you know, when we first launched our, our products, we would give a demo talk about how we would, you know, walk you through from, uh, one screen to another to find the root cause of an issue.
I just saw a demo the other day where essentially the demo I used to give that would take, let's say 10 minutes with me, kind of, you know, pointing and clicking stuff now, uh, could happen, uh, in less than a minute by simply asking the question, uh, of, uh, the assistant that, uh, that we've been working on, right? And, and so, um, you know, the fact that you could save, you know, those nine minutes there almost kind of pays for, uh, the, the service and the software already, right? Um, and so I think that's something that, uh, folks should, um, kind of think about as they go forward, right?
Is how do they make sure that, um, uh, they have, uh, the information into the system that actually allows the AI to be able to, uh, provide the, the better, an better and better answers to them, right? So they can run more efficiently and have, um, basically everyone be an expert rather than, you know, rely on like, Hey, you know, this person really knows this, this tool, this person really knows that one, if they're not here and there's an incident, what do we do? Right?
Um, one, one last thing I'll I'll just add is that, um, uh, you know, most organizations we talk to are in for some form of experimentation around actually using, um, large language models as part of, uh, their, their applications as well, right? And so I think, uh, it kind of goes without saying that, uh, like every other part of, um, the application, um, uh, stack, right? Uh, this is something that does need to have, uh, some level of observability and, and kind of security practices built around it as well.
Um, and so something, uh, for everyone to kind of keep, keep in mind as they, uh, you know, do the, the fun experimentation. Great. Absolutely.
Absolutely. There is a lot for the audience to consider in this conversation, and I know we just kind of touched on a, a brief topic here. There's a lot more to consider, a lot to think about.
But as we come to the end of our session, I wanna thank the both of you for your perspectives and insights on observability and providing guidance to, on the audience for their own efforts as well. I also wanna thank the audience for attending our session today. Thank you and have a great day.


