Elizabeth Lawler – The Obvious Evolution of Observability: Shifted ALL the Way Left
From managing custom-built bespoke servers in data centers to spinning up serverless applications in the cloud, monitoring and observability have undergone a generational shift over the past decade. Today, most companies implement expansive monitoring of their production applications to quickly respond to customer issues as they arise. The trouble is, by the time customers experience a problem, it’s already too late. In this session, we’ll talk about the difference between monitoring and observability, how the current generation of observability tooling is failing modern enterprises, and how to move observability directly into the developer’s coding workflow to surface performance, reliability and security issues before code is ever committed.
Transcript
This is texturing TV. Hi, thanks for having us today. We're here to discuss the obvious evolution of observability shifted all the way left.
Thank you. Who are we I'm Elizabeth baller. I've been working in the devops space since 2012 a little bit before it was cool.
com Pete. I am Pete cheslock. I am a technical operator.
So SRE devops admin. I got into cloud and devops back in 2009. So slightly before Elizabeth there were still some cool kids there not one of my own and then really do the luck in circumstance.
I got into this devops space. I've done the product thing. So I will air quote on the product do things in the past and worked in some local Boston companies, which is great and and working with Elizabeth at applian.
All right. So we hope we have a fun talk for you. Today.
We're going to talk about quickly talk about four Trends related to observability that we see in the market. These are trends that you as devops professionals should be aware of and they're an emergency focus of new areas for tools and efficiency. First topic we're going to talk about is The Greening of the developer population.
So we're going to bring a little bit more Dev into the devops discussions today by which I mean the market forces that are changing the composition of development teams. The second is the emergent software quality issues that represent new design constraints to Velocity in your organization namely those of understanding existing software design and the impacts of code changes in light of those designs. And this is the top topic.
We're labeling design is new constraint. The main focus of the topic of the talk will be observability how observability differs from monitoring and how it can relate to good software quality and delivery. And then lastly we'll talk about shift left, which is how does observability relate to the shift left movement Which is popular in security and what we've learned from that movement how we can apply it to make more progress in software development.
So the demand for this so I've made several shout outs in this talk to different speakers. I'm going to first shout out to my cookie from Pete just the other week. So the demand for software Engineers is on the rise for quite some time now and shows no sign of stopping software development employment It software developer employment is projected to grow at a rate of 21% by 2028.
That's more than all other occupations, which is 5% as a result many new software developers are joining the workforce. We have an explosion of green Engineers. These may be people who are just out of college who've trained in computer science or engineering they could be just out of a code academy or and join the field from non-traditional backgrounds.
But there's also High motility in the marketplace so jobs are highly competitive and people move jobs rapidly. So on average, why is everyone green? Well, they only stick around in any one Organization for about two years and the average tint if the average tenure is two years and it takes about three to six months to get up to speed on a complex perhaps.
I think one of the former speakers was talking about monolithic code bases. It's really everyone is brand new to the code. So no other how well organized or documented it is everyone is pretty green.
So the net result is that we have both experienced but new and inexperienced and new developers who are coming to work on code base is that they don't understand. This is a quote from charity Majors the CEO of honeycomb. She says we are shipping code.
We don't understand just systems. We have never understood and that is considered a common experience for developers working in complex work complex code basis. And moreover, you know, if you're coming whether you're coming to work on a piece of code that no one's touched in the last three years the person who wrote it no likely is no longer working there.
And this is one of the strongest I think Arguments for observability. These code bases could be are often highly complex large logo documentation. And so onboarding isn't really a point in time process the way you used to think about onboarding to a new job.
Actually. It's a continuous process that's experienced and we experience by developers while they're working. They can change teams.
They could move off of a project the project could come back up. It could be it's a life cycle of constantly onboarding rewarding to code. So these poor engineers in the photo have kind of attempted a umlish on the wall.
And you know what, they're probably trying to do while they're still looking at their laptops is create a shared understanding of either the code they're working on or potentially the impact that code changes may have on on the performance stability maintainability or security of whatever they're trying to ship. So lack of orientation and getting lost in the code base is a root cause a lot of good quality issues that make their ways into the later stages of this DLC and into production and I think that's really where we have an opportunity to use observability with developers to try and improve code quality and velocity. So you might ask because this sounds a lot like Mom and apple pie right?
Like everybody should know more about their code base. Everybody should be able to have Mastery of the domains that they're working on. But you know the question you might ask is people in the audience who might be buying things or looking at tools.
Is this a real business problem that we need to solve? Study by stripe called the developer coefficient that they published in 2018. They surveyed companies across the Spectrum about how they their developers spend their effort and what we're emergent barriers to Velocity.
And one of the most Salient data points that the undercovered was this developers spend almost half of the work week 42% in rework. What kind of work is this? They're either working on structural code quality improvement to existing software, you know working on Legacy monoliths trying to improve technical debt or simply reworking code.
They attempted to shift but didn't actually make it to through the production code quality Gates or resulted in incidents in production. It's a blocker to many modernization initiatives and trying to get stuff moved to the cloud and microservices. So when there's no mental blueprint to work from developers can spend up to six hours per event or ticket trying to Simply gather information.
They need to implement changes and fixes. Therefore writing a single line of code or attempting to iterate on tests. What's more most impressive about this whole thing is that it's $300 billion dollars of time and efforts spent in rework every year by Developers.
So it's big problem. It's happening at every organization your developers spending men's amounts of time. Just turning.
Um, so again, you know, so that's not a big enough number for you wasting half people's time isn't a big enough number do number for you. Maybe you have more specific quality concerns we can talk about whether or not this is like something urgent in the you need to address now. So as I some of you in the audience may know particularly those in the back row, I've come from a background in cyber security.
My previous startup was in the dev SEC up space that I build open source tools and these domains and so if you look, Recent Trends and security flaws. Let me try and encourage you to take a look at observability and what impact it might have on improving developer mental models and improving software quality. So this is a shout out number two.
This is to miter. I know we have another speaker today from miter. What's really interesting?
What's really interesting in is that minor every year? They published The Cup come weaknesses enumeration top 25 and those are the most impacting software security flaws that are affecting cybersecurity and sometimes performance in the world today. And what was the most most notable change?
So when I had my last company which was a Secret's fault for passwords plain text passwords things that were making themselves into Source control. That was the number one. That was one of the number one causes of breaches in that year fast forward to 2020 one and design flaws are actually making the list of the top security weaknesses.
Now, these are things that are not necessarily immutable to static code analysis and they're often familiar from owoss, but they're not something that you can really scan before in fact in their analysis of the top 25 report miter talked about how really the only way to remedy these issues and practice included deeper code reads by senior devs. I think report actually said you have to put your big boy pants on and read the code base to find these issues. Um, and I think we just talked about the greenness of of the devs entering the workforce, which is probably a root cause of some of these issues the use of design templates for consistency in implementation of these types of standard designs and practices and also better education, which is advocated by the OS Group, which is also known as the honor System.
So is a this is a prescient problem. It's important impacts many many dimensions of of code quality and velocity most developers are aware of these good design principles, but it's this growing complexity and lack of the right information at the right time that caused as degradation designs typically start out well and are Implement in the code, but they degrade over time. So why do they degrade?
Because understanding of how code will behave the things that impact those top 25 and and perform from simply just reading the code base is which is what most developers do to get oriented is particularly in particularly to unfamiliar code that you didn't write. It may not understand which we just established previously is really hard. It's easy to break things and all that complexity.
So the goal that we want to talk about today is how do we lower that barrier to understanding? How do we create interactive opportunities for sharing information runtime information behavioral information insights with as context for developers as the code, I think through observability. We can illustrate a lot of code both codes design and this and surface of potential runtime implications of code changes early in the process and it's the opportunity to do to extend the mental model of the developer while the codes being generated.
And validate earliest in the design process the understanding of the principles and best practices that are being are that are that they're being adhered to architected into the code quality change and attempt to reduce cycles of rework. So if you don't want to test in prod shifting left observability left can definitely help. So let's take a look.
Thanks. So this next train is something that's really near and dear to my heart as a former operator monitoring and observability has just been a big part of my life. And this concept of observability of everything is is really growing dramatically the cool kids call it Ali now you can see AKA zero.
It's oh one one why you'll see that a lot on Twitter and such but you know, what we're going to talk about is kind of where we've come from. What is really observability. Is it the same as monitoring?
Is it a buzz term and how I really think and generally how we see it. It's just a Natural Evolution. It's generally where a lot of folks are moving to So we'll do a quick audience participation.
Anyone here have any any thoughts about the difference between monitoring and observability anyone have opinions on this or is this just like carefully crafted buzzwords? Any differences between buzzwords? Okay.
Yeah. I mean I follow a lot of monitoring companies on Twitter definitely feels like a buzzword, but I really truly believe there are real differences between these two things and we're gonna talk about them. If you go to like the the very classic definition of the word observability the noun of it.
It's a measure of how internal states of system of a system can be inferred from knowledge of the outputs of that system. So if you kind of think of that from a software engineering concept, it's in contrast to monitoring which is something we actually do observability is again a property of that system. So if our our good old it systems and applications if they don't actually externalize their state and then even all of our best monitoring is really gonna fall short here, so I actually like this definition this kind of a charity Majors fan talk, but but she is the founder of a company called honeycomb who is probably one of the few monitoring.
Companies that is actually an observability company one of the few companies actually doing it and it's a project that actually is based on a paper out of Facebook slash meta called scuba, which is you're really basically capturing this High cardinality data think like all these different users billions of users you're trying to build for and getting answers from it. The unknown unknowns is kind of the term. And so this this updated definition.
I really love this this I saw it online which was monitoring is for running an understanding other people's code like you're infrastructure, but observe ability is for running and understanding your code that code that you write change and shift every day. It's the code that solves your core business problem. Maybe the code that people pay you for.
So yes, right monitoring observability. These are two kind of separate things. I personally believe that observability is again this Natural Evolution that you're probably starting off with your classic monitoring tools and then you eventually level up into but that's not always the case.
Sometimes people just Jump Right In So who's here have heard of this? Right? If you again follow the marketing speak the three pillars of observability which which I love this one you generally hear it as metrics logs and traces.
These are the pillars right and and you must have all three to be have observable systems. Now if I was a big monitoring company with with people who wanted my big monitoring company to grow and I had metrics logs and traces to sell you and I saw an observability buzzword flying away. I would want to make sure I would align myself and stay relevant in the space and I'm sure there's gonna be folks that disagree with me.
I think that these tools are actually the tools that you generally start with. I see a lot of companies starting with just logging right just logs and then centralized logging and then maybe metrics and then maybe they eventually move into things like Tracy and more advanced types of monitoring and observability, but some of these have low barriers to entry which is why they're so great to get started with but others have a high cost and a lot of different ways. So when talking about metrics logs and traces like you look at this dashboard, and it's so chaotic, but what this dashboard is really showing you is these are questions that I already know to ask.
I'm not actually learning about my systems. I'm asking very specific questions of my systems to understand and that's why a lot of these tools are great for understanding inside of a known boundary. I think this particular screenshot is very telling it's missing link to the code itself.
Right? What about the code that's happening? Yes.
Have enough time on a working lifetime. To know where to look for something important. Yeah.
So good question. Like how would you know where to look Overlook? Yeah, exactly.
That's that's the problem. I think of a lot of the monitoring tools. It's just it's too much and again, it's only useful like this.
Demo is great. I've demos it demos great but the usefulness really falls apart. So of course, you're like, I'm at an Enterprise.
I have money and I can go buy my way out of this problem. You sure can look there's a ton of companies and this is only a fraction of the companies that will happily deliver you a solution to your problems. You have monitoring vendors, they can automatically create graphs for you and they can set these thresholds Within These known boundaries when things go bad a Logging company will happily suck and up all your data and ingest it and they'll even let you index on all these Dimensions so you can again ask questions of your data and finally those APM tools that are out there.
They'll help you Auto instrument your code and they'll surface some of these well-defined unknowns, like maybe what endpoints are slow for example So I want to dive into these three pillars because I think they all have their own positives, but I don't really see them all together being observable because again, they're they're lacking this necessary code context, you know metrics these traditionally come pre-aggregated. So imagine you have hundreds of thousands of web servers. It can be costly and time-consuming to capture these metrics at a per service level or per user level.
And that added cost could be a monetary cost really so we aggregate them and we say I want to see how many people are accessing a website property and and create some metrics based on that. But then you lose all of these contexts you lose connection request data, you you don't understand your users, maybe fully because of it, but there's still useful right in that last giant dashboard. We had you can create these powerful dashboards and show things like Trend analysis and that's again a lot of important decisions are made based on that.
But again, these are helpful for when you're scaling inside and known boundary. Things fall into the unknown they fall apart pretty quick logs are my favorite because they're just these unstructured strings. We haphazardly write out to disk if you're maybe a more mature organization and you have a logging strategy maybe a shared Library.
Maybe you're writing logs in a structured way. You're usually a lot better off than than folks who are just getting started if you're lucky. Maybe your logs are a single line per request, but that's not always the case.
And so if you have just again these haphazard lines you're trying to correlate together and try to understand and finally as someone who's I've spent more money than I can even count on logging vendors and logging solutions to scale the all this machine generated data to get answers. There's there's a real monetary cost. If you've ever had to renew a Splunk or an elastic search license, right?
There's a real cost to this volume of data and you're essentially trying to find these needles and in these elastic Stacks, right? Now traces are we're kind of on the right path. Traces are very close to the code.
And that's what's important. Right? You can now visualize these code requests and the interactions between maybe your microservices or your serverless functions that are happening and understand where the failures are happening or the slowdowns then your users I mean impacted and so being close to the code is really important which is which is what we're really talking about today.
There's a high cost of implementing things like open Telemetry also known as otel and oftentimes net new projects get deep tracing embedded and then your monolith is kind of sitting here like a black box and no one can understand what's happening on the inside of it. But I think when it comes to traces and again just generally observability and tooling devops Centric teams should always be on the lookout for ways that they can actually demonstrate how to make their applications more observable and and provide that value back to the business and make it really evident. And so that's why you see like net new projects usually come with Tracy and involved.
If you're into the kubernetes world, there's again plenty of tooling to help provide insights into your kubernetes. Istio service. Mesh.
For example kyali lets you observe these Connections In Your Service mesh and it's really interesting, right? You can understand your request routing and circuit breakers request rates latency and and provide some insights into this, you know these apis and these systems that you may be don't fully understand from the outside, you know, looking at a yaml file and how it relates to the service delivery can be really challenging but mainly observability is what will help you understand how your code design needs to change whether it's for scalability or reliability maintainability security and monitoring basically tells you when you need to scale your systems. I need more of the thing to handle these requests.
So what is the problem ultimately we are spending big big bucks on this stuff. I have spent a lot of my career also in understanding Cloud costs people call it Cloud economics and I'll tell you that the number two spend for most businesses after their Amazon. Azure gcp bill is their monitoring and observability bills.
There's a reason companies like data dog and elastic are multi-billion dollar companies. There's a lot of money being spent here. So if we look at the sdlc process flow monitoring observability have shifted so far, right they're actually off the slide deck.
They're completely off. We have put all of our investment into production and that's after our customers are impacted and it's furthest away from development and the code itself. So, you know, there's a lot of reasons for this there's not good parity between development and production systems and so oftentimes the classic like we are testing and prod scenario happens, you add more insights and visibility to your code, but there's a lot of issues that we call, you know, anti patterns of code design.
That you can actually pick up within your deep code analysis that Elizabeth talked about before we can capture these earlier in the cycle. You know, but essentially the default position is Ops and SRE teams. They're pushing this burden of design issues into their observability tooling and those teams are at opposite ends of the delivery life cycle.
And so for Ops and devs devops SRE when all you have is a hammer, right? Everything looks like a nail they keep on putting more and more visibility and and metrics and data into their observability tools to understand why things are failing pushing more and more data into there and and raising their costs as part of it. So this leads to this this last Trend which is Shifting left.
So this term has been traditionally used for security. It was this concept of you know, people were scanning their production applications looking for security vulnerabilities and the message was shift. It left earlier in that sdlc flow that we that we you know, we showed you before but how we do this today in our cycle of a developer who's trying to debug a problem is you have that like gnarly dashboard at guinean you have SRE teams who are firefighting and they're seeing slowdowns and outages that are happening.
They there's oftentimes developers don't even have access to the Telemetry to debug their systems better. So a tickets created a slack channels created. So now you have devs getting involved.
They have to like become pulling all this context and become experts on this code base that they're new to and understand all these code paths and stack traces. It can take hours or days to load this mental context about how the code works so that they can actually make an effective change. So then they're in their IDE, they're pushing this change to production and and trying to understand it or if they don't have it fixed.
They add more Telemetry the add more monitoring and it's just a very very slow feedback loop. So what we're actually saying is that if the developers in the code base, their mental context is minutes old, right? They have the full map of how the code Works in their head how it connected and how it function calls what but when the the code hits pre-production their CI and CD system that context could be hours old at that point and it could be completely forgotten from that developer as they move on to new projects when it goes into production that context is now completely gone that that developer is a green engineer all over again.
They have to reorient themselves security tools proved that shifting left works. Right? We've seen that model work.
We've moved the information closer to the person that has the con the context and the ability to fix it while they're there years ago. You might have heard this called rugged devops and this is my shout out to Josh because he was a big speaker on this topic rugged devops. I like that turn better but shift left seem to the win over from a marketing standpoint.
So this is the new single pane of glass and this tonight Engineers may look completely chaotic, but for a Devon Ops folk, it's their curated workspace. It's where they live day to day. It has their code editing their terminal and an endless number of plugins that they can have all this adaptability to expand on so I've been in SAS software for almost my entire career the number of times of vendor including vendors I worked for would say no we have this tool it will replace your 12 other tools with this single pane of glass.
It'll be great. And then you just turns out you have 13 tools that you're gonna have instead. So we need to start meeting the developers where they're at.
They don't want to leave their code editor because the second they do that context we talked about starts to fade. This is already starting to happen. You have tools like sneak Stoner Cube open source projects like breakband all so at map that have ideations that surface the problems alongside the code right there.
So an engineer can see the problems and the insights right alongside the code in their editor and again, they don't leave their single pain of glass. I've even heard of Engineers don't want to go into slack which on blame them. They put slack messages and DMS to them in their IDE, so they don't have to leave and contacts which way So as technologists, you know, we want to choose tools that give our teams access to the data and information where they need it.
We want to shorten these feedback loops and we've done a miraculous job of this over the last decade plus of improving these feedback loops and shifting things left into the developer. We're saying you're stopping at the CI system. That's not far enough and actually you want to shift even further left.
So some research from Facebook and meta about an internal tool, they created called fausta, which I think I'm pronouncing that right. It provides a framework for dynamic code analysis runtime analysis. The end goal of this project was actually to help ends Engineers gain confidence in their code changes in an earlier stage and their research showed that when an issue was found in production only 20% of those issues were ever fixed and when it when they were fixed it took a long time up to 10 days to actually fix them and this is a company that has revolutionized how to deliver code extremely quickly and it still taking them this long When they shifted the code quality tooling left 74% of these flaws were identified and fixed in less than a day.
And again it dramatically lower cost per bug fixed. And what was amazing. They didn't incorporate runtime analysis as a security use case, although of course, you can use it for one.
They actually wanted to give the developers more confidence that when they shift shift to their code out. It wouldn't break. It wouldn't break users.
So that project brought this deep observability directly into the IDE with runtime analysis understanding the functions calling the databases and the callers of those functions. It opens this whole new world of possibilities. We haven't had before you can find things like an N plus one SQL query where some function will call a query 30 times and you don't even realize that when you call that function you don't understand the underlying And it can be really hard to find things like circular dependencies and things.
So these developer ecosystems are massive Visual Studio has over 13. 30,000 plugins and they're not only for day-to-day development. They're also the SAS vendors who are now starting to push their tools into the IDE.
So those teams can get insights and information. So observability is still an acient market for tools. There's still a lot to be figured out.
We're still in this infancy around tooling and techniques perspective, but the benefits of the investment in observability is very clear when we shift it left. So we're still focused on all the same kpis. You've got your slos.
You've got your okrs if you have Dora metrics and you're following those four key metrics ultimately that goal is to deliver high quality software. Our belief really is that shifting left is not just for security that was just the first use case that proved that it worked. It's shifting left should not stop at your CI and CD systems.
It's still too far away from the developer. We have to keep shifting left all the way down to those engineers and and those tools that they're using today. I think we actually have a couple minutes left if there are any questions we'd be super happy to chat about them.
But otherwise, yeah, really appreciate this. Question in the back. Let's do it.
Again, I can't do the math. In my head, I'm gonna say 36 years ago. I went through this kind of thing.
real production walk and it got me fired. Exactly what you just said. It's the death by a Thousand Cuts.
is by volume it's still going to take the operative third shift operator. Time to figure this out. Yeah, it was one building Empire.
Yeah, just having this context around, you know, understanding the shared context around how the code works right when you're coming in new to a code base that oftentimes people are actually managing up with like why is it taking so long to make this change to some higher level person? It's hard to describe what like load-bearing software looks like I mean this has been through so many different epochs over the years. I mean the, you know, the rational framework, I mean how many times have we tried to do this?
Right? But what's really interesting and what's really shifted is that the developer tooling end of the spectrum has gotten more flexible and more extensible. So now there's an opportunity to start to bring more tooling to them.
You know, there was first we started out with big platforms to build software on then we got to a myriad of constellation of SAS products that we all built software with and now it's coming back around to try to bring it down again to the end user to try and bring them contacts and I think that's really an interesting sea change. Awesome. Thanks everyone.



