Charity Majors – The Sociotechnical Path to High-Performing Teams
A chasm is opening up between elite, high-performing teams and the rest of us. According to the DORA report, elite teams get code to users 25,000x faster than the bottom half of teams, deploy 46x more frequently, and restore service from incidents in 2666x less time… and while the same resources are theoretically available to all teams, the performance gap is *widening* with each year. WTF Is going on? Are the majority of teams just doomed to mediocrity? Is this a skill gap (spoiler: no), a technical problem, or a business one? Most importantly, how can you help your team function at an elite level? Let’s talk through the social *and* technical strategies that great teams all of the world are using to be happier and more productive…and make their users happy too.
Transcript
This talk is called the Socio-Technical Path Type from Teams aka Observability and the Glorious Future. My name is Charity @mipsytipsy on Twitter and the co-founder, CTO of Honeycomb. And also in about six months, three months, those functions and I will have O'Reily Observability book coming out.
This is a talk about teams, teams or teams exist so that we can reason about each other in a scalable way. I like to think of teams. It's kind of great fight for people, right?
It's the team provides a framework for resiliency and dependency. And I feel like you underestimate just how much of an impact teams have on our lives and our careers. And I would argue that the teams that we join define our careers more than anything.
More than where you go to school, more than your industry, even more than your own personal experience, because I have a team with a couple of exceptions. There were a couple of startups. I did have two jobs that I really don't look back on, super fondly, and they were very different from each other, right?
One was kind of your traditional Silicon Valley startup where we got praised for pulling all nighters. Long commute didn't see my manager for an entire year and another was a rebound job from that one. I went straight to another that had like, you know, great people walking distance to my work and also, like, very out of date technology.
And I spent my days, like translating word documents into puppet recipes, very different jobs. But the way they made me feel very much the same. Right.
Even though they had almost nothing in common with each other. It let me let me. It left me feeling discouraged and burned out, honestly, I wasn't working more than I'd ever worked before, but I got more burned out than I've ever been.
Turns out I knew this experientially before. I really understood why these teams are good or bad. And I knew it because, on the high performing teams.
I had motivation. I had autonomy, mastery, and I had meaning, which is, you know, this is what elevates our work from labor into craft. This is what makes our work part of the fulfilling life.
This is what makes us think about it in the showers. We're getting ready in the morning. It's I feel like once you've experienced work, that is it just work, but it's work.
It's really hard to ever go back. It's also incredibly difficult to craft. These systems in a way that produces like party work instead of like, you know, labor work, which is why you can't really look at just the teams or just the systems.
You have to look at the whole thing, the socio technical system that you're embedded in because it's all interconnected. This you can go in a lot of white papers and stuff on the Internet about this these different terms. And I'm going to use all this means is that it's a feedback.
Your system is unique stuff like there's not you can't take anyone else's rulebook and apply it to your situation. You have to figure it out yourself from scratch, which is why it's fun, why this is interesting. There are, of course, some patterns that we can learn from each other and stories that we could tell.
And that's what I'm going to try and do in this talk. But the fun and exciting and terrifying thing of this is every system is different. Let's start with the socio part.
How well does your team perform? Now, we've been we have not had much science in this regard until Nicole, Jez and Gene went and did a bunch of science and came up with. There are actually four four questions.
I think every single manager, every lead, everyone in tech who cares and is tasked with these systems should be tracking constantly. How often do you deploy? How long does it take for your code to go live?
How many of your deploys fail? How long does it take to recover from outage? And I would add a fifth.
Everyone should be tracking how often you are alerted after hours, that's not all. If you read the stripe developer report, then you know a lot of time like. A lot of time.
42% is the optimistic self reported statistic that they came up with. 42% of our time is wasted. Doing stuff doesn't move the business forward or we're not learning anything.
We are creating anything new. We're just doing the shit that we have to do in order to get to the stuff we want to do that we need to do. It's orienting ourselves in time and space.
It's reproducing bugs it's working on the wrong thing and then having to go back and undo it and work on the right thing. And one of the biggest obstacles that we have in tech is we have just come to accept this is normal, this is just acceptable, that half of our time is wasted. It is not, it is not.
Back to the door report, though, like there's there's a really wide gap between the teams that have started to really get their shit in order. And the teams that haven't. Deployment frequency he bottom 50% of teams deployed to the most elite teams deployed many, many times a day.
And if you look year over year, what you see is like that elite, I don't like that term elite category is getting bigger and better while the bottom 50% is actually losing ground, because if you're standing still in tech, we're losing ground. What this tells us is, first of all, it really pays to be a high performing team like it really pays. And here's why.
Like, we often think that, like, how do you get to be a great engineer or how to get to be a great team? It's that you build a great team by hiring the best engineers. That is bullshit.
It's actually the other way around. You become a great engineer by being on a great team. And you just look at this like which one of these two kids is going to be a better engineer in two years, the one who gets three hundred opportunities to learn three thousand deploys per year opportunities to learn, or the one who gets five deploys per year and spends most of the time firefighting?
Firefighting is not creative labor, not fundamentally one different story. Our biases tell us that great, great people make great teams. But I've seen this over and over again where an engineer from one of the elite teams quote unquote joins a team in the mid performing, you know, levels.
And three to six months later, that's how fast they're shipping. Right. Because the rate at which you're shipping code is not governed by your personal skill level.
It's governed by the infrastructure, is governed by the best practices, is governed by the tools around you. It's governed by, you know, everything that has gone into constructing this environment. Your productivity will rise or fall to match that of the team that you join.
Setting aside the question whether you even can evaluate the best engineers people often like. Try to stuff their teams with like people from Google and Facebook, and that's just that's just backwards. By the way, also good managers don't hire people, they build teams, they craft teams, they grow teams, they don't expert to show up on their doorstep fully, fully formed.
They accept the need to develop their talent. Anyone invest a lot of effort into doing so. Every engineer actually has a dual mandate.
By the way, it it's not just the responsibility that you have to your customers, although absolutely there is that there's also the responsibility that you have to your team, the people who build it and who run it and who maintain it every day. I don't believe that these two are intention, which is sometimes how we tend to talk about them, like, oh, well, if we if we have a lower quality of service for our customers, that will make our people happy. Absolutely categoric.
We do not believe that that's true. I believe that, in fact, the only way that you can sustainably have either of those is, is they build on each other. Nobody likes to make their customers upset.
Right. It's actually the quality of life issue for us. As engineers.
We take pride in our business and our craft. We want to do want to do well for people. Right.
And. The way that we kickstart this feedback loop is by thinking about ownership and ownership begins with observability. So let's move from the social part to the technical part.
Tools, I don't know your systems, so I can't prescribe exactly what tools that you absolutely need to invest in, but I've seen a lot of systems and I feel like I can take a pretty good guess. Durability is where it starts for the for the same reason that I put on my glasses before I go and drive down the freeway at 80 miles per hour. If you can't see what you're doing, you're just going to waste a lot of time.
And I have yet to see a single company that has invested enough energy into their deployment software for deployment frameworks. Most companies haven't even really fulfill the promise of continuous delivery yet. It's like we got through CI and now we just say CI/CD but we don't actually do CD and that's a problem.
Observability is a term that I just threw out there. And you may or may not have heard me ranting about this already. You may or may not have already read about this, but what it means is this is borrowed straight, borrowed from control theory, mechanical engineering where observability means.
Can you understand what's happening inside your system just by observing it from the outside? And if you apply this to software, it means can you understand any system state? Without shipping new code, without shipping code to handle that system state, any asshole can tell what's going on.
Explain their system today if they could predict it in advance and. Right. Because they handle it.
Right. But it's about gathering your information at the right direction so you can slice and dice, ask new questions, understand new system states without having to ship in code. Now, I have written a lot about this and I'm not going to even try and dive into all of it here.
Suffice it to say that there is a lot of bullshit out there right now. People who are saying that they have observability tools that they don't, but not by my technical definition. And the stuff that I'm talking about requires some of that technical definition in order to be effective.
So some of the ways that you can you can pattern match if they're doing generic durability or tactical observability is by looking at the data types, if it's metrics, if it's unstructured, logs, not observability, anything that requires you to define indexes or schemas up front, not observability. If it if it's based off of these arbitrarily wide structured data blobs or unclogs, yeah, probably is. And there's more high cardinality, high dimensionality, there's more in that link, if you want to read about it, A lot of people stop at this point and they're like, yeah, that sounds great, but I don't really have time to invest observability right now.
Definitely it's on my list. At some point I'm going to get around to it. I would argue with that.
I would argue that not having observability with these complex systems, microservices where you've got, know, novel scenarios every day, it's like it's like taking off down the freeway without your glasses. It's just it's just I mean, it's your time to waste. But if you want to go for it.
Anyway, Liz and I wrote up this maturity model and we tried to make it like a choose your own adventure, right? So because we don't know your system state, but you can go up and look at you can you can look at each of the five categories and try to find yourself in them and see what your weakest in first is, strongest in resiliency or complexity. And in the ones where you're like, yeah, I recognize myself in the weal description.
Those are some areas to invest. Right. But I want to talk a little bit more about why why this and why now, because Observability hasn't always been this important, obviously, because we've gotten all this way without it.
Right. The reason it's happening now is because complexity is just off the charts. Right.
Like it used to be that we'd have the app. Right. And the database and all that complexity was bound up inside the application itself.
And if you really needed to understand what's going on, you go attach a debugger, step through it. Well, now we've got many services and it's hopping the networks. You can't write.
So a lot of these software issues have been sort of thrust into the domain of operational stuff very rapidly. And and the reason that the reason the observability is what unlocks the key to ownership is that. Technical observability results in a situation where you can slice and dice and take these rows here like these, these requests are different from those requests after I made this change.
Right. It allows you to compare that very granular level. And if you don't have that, you're just guessing.
Right? Maybe they're very good guesses. Maybe they're very educated guesses.
Maybe you've been doing this for a long time. You can pattern match very effectively, but I miss being able to go. I see your dashboard's.
It's retests. Like I miss that. Like I love playing God.
It was fantastic. But you can only do that if your systems are like rhyming in the same ways that they fail over and over again. And if you don't if you don't have those just feeling the same repeatable ways and if you don't have the right tooling.
You're going to really struggle to connect these feedback loops that are at the heart of a well constructed and well understood system. I want to give you an example, because I think that all this is kind of difficult to talk about at a high level. Let's let's take the example of.
Photos are loading slowly not for everyone, but for some people in the good old days of five years ago with a LAMP stack. Some of the scenarios would probably be like these we're out of capacity, but Dashboard's maybe we're currencies high or we've run out connections, whatever. You can build a LAMP stack look at it, size it up so I can predict in the way the system going to fail.
You can write monitoring checks for those things. That's great. And I'm not ragging and monitoring checks because you should absolutely have monitoring checks for all the things you predict will fail.
Great shortcut. OK, but now let's look at the same same exact scenario. But for microservices platform, like these are actually taken directly from outages we had a past and or Instagram like these are not things are going to happen over and over.
These are weirdo problems. They're going to happen once. Right.
And historically, we've put all this effort into we understand the outage if we're lucky and we were going to do a postmortem, we're going to craft this perfect dashboard so that next time this happens, we'll find it immediately. We put it in our run book. We educate everyone.
This is great. If your systems are failing repeatably in the same old ways, it is a complete and utter waste of time if it's never going to happen again. You have an observable system when your team can quickly, reliably diagnose any new system, state any new behavior without predicting it could happen, without understanding it is going to happen.
Having monitoring checks for it, it puts you in this continuous conversation with your code. And when you're when you're instrumenting and you're and you have this conversation with your code. Right.
It forms beautiful feedback loop that I think of as like observability driven development and no, knock on TDD. TDD is great most effective software movement in my lifetime, but it stops at the border of your laptop, right. Like it is.
It is effective because it abstracts away everything about reality. It's just like doesn't exist, which means that it's a very limited utility. And yes, we should still write tests, but that it used to be that your test you can catch like 80, 90 percent of all things would happen now, like maybe 20 or 30.
And so you have to start thinking about writing code and instrumenting it as you're writing code. You need to be instrumented with an eye towards how will future me understand this, right in an hour when I've shipped it and how will I know if something's not working? You should never accept a PR.
If you can't say how will this break and and explain it right. And and you should have. The reason why it's important to have CI/CD is so that that interval of time between you when you write the code and the code is live in production is as short as possible so that you still have all that original intent.
Stuffed in your head, right? So you're watching it so so that you merge to main you get coffee, five ten minutes later you're back and it's live and you go and you look at it and you ask yourself, is it doing what I expected it to do? Is working as intended.
Anything else look weird because you're in there every day. You know, you're not biased towards only looking at production. But its weird you understand what normal looks like.
If you break up that feedback loop, if you make it so that bugs don't get caught by the person who wrote them right after they ship them, if you make it so, if you make it so that the software engineer isn't even looking at their code or isn't even instrumenting their code, or if someone else is looking at it days later, you lose that beautiful opportunity where you can catch 80 to 90 percent of all problems before users ever even get a chance to notice that it exists. Right. That is most powerful moment in software development lifecycle is when you written it and you're looking at it right there.
Starting to yell at people, on call is where shit gets real, right? On call is for everyone who writes code. And I do not say that I know that.
So I come from Ops, absolutely, I will admit that we have a reputation for masochism that is well deserved. And the idea here is not to invite all of engineering done into our shabby hut. The idea is not make everyone suffer.
I swear to God, I'm over 30. I don't want to get woken up anymore either. Right.
The idea is that this is how we make it better. Right. This is the only way to lift us out of this react response react response firefighting mode into a mode where it is rare that you get woken up out of your hours, like vanishingly rare.
You cannot achieve that. If you have different people running the system than writing the system. You could only achieve that if you have a single tight feedback loop.
You can't gift your eyes to another team like, oh yeah, take all this context from my head. Go look at that shit that I just wrote. It doesn't work that way.
And it's on us in Ops to like to stop being all, stay out of production, right? We need to be welcoming. We need to encourage curiosity.
We need to build tools that build guardrails, that emphasize ownership and that we don't punish people. Progressive deployment is a new term of art for deploying all the things, and I'm not I wouldn't be clear here that I'm not saying that all staging environments have no value. I'm not saying that at all.
There is a value in stating. What I'm saying is that we've gotten the ordering wrong. We've taken engineering teams and gone, OK, go spend months building these elaborate staging and test and desktop and laptop environments for everybody.
And then we get to production and people like, oh, well, we're out of time. We don't have any more cycles to invest in. Right.
We'll let Ops handle it. All I'm saying is that is backwards. Please invest in production first.
Those lines that I showed you, the bubbles number from 2018 to 2019 where that 7% of elite teams are up to 20% in a bubble is getting bigger and higher. That is because those are the teams who are embracing this, this whole constellation of tooling that is around shifting the center of gravity to production, to reality, right. Feature flags, incredibly important, SLOs incredibly important.
Right. All of this tooling around, running your shit in prod safely. Is important.
I want to I want to quickly I want to show you how quickly this adds up and and gets us to that like 42% of our time is wasted. Right. This is the kind of thing that happens.
It happened most weeks, let's say, throughout my career. Engineer merges a diff and because it doesn't automatically key off of that deploy, you have to wait for somebody to come along and manually trigger a deploy, right? So as engineers merges the diff hours passed, more diffs get merged, right.
At some point, someone triggers a deploy. It might have hours or days worth of merges. At that point, the deploy fails or it takes at the site whatever pages on call.
Well, on call, probably a very different person on a very different team than the person who just triggered the deploy. They might not even know that each other are both working on this. Right.
So on call goes something's wrong, starts rolling back, starts investigating. You start get bisecting. Just try to figure out which one of these diffs was at fault.
Right. Maybe you cut another half dozen bridges or tried to play a bunch more times to to figure out which one is the error. And this eats up your day, it eats up time, from every person who's shipped a diff who might be the guilty one.
Right. Eats up on calls time. At the end of the day, like this might have taken four or five people can't actually get their shit done because they've gotten roped into something went wrong.
Supply production, everybody gets burned out. Everybody's bitching about how much on call sucks, how much they don't want to be responsible for deploys. You know, how much time is left there?
A lot. Ok, now multiply that by days, weeks, years teams. And you suddenly see that's where all our time is going.
Let's look at a virtuous feedback loop that is accomplished. This is achieved by nothing other than making it so that CI/CD is real, making it so that when you merge to main, it automatically triggers to deploy and automatically goes live. So engineer merges the diff, kicks off the automatic CI/CD and deployed a few minutes later, deploy fails, notifies the person who just did the merge to reverse the safety, knows exactly what you just did.
So she quickly fixes it, adds more test instrumentation with the fix it deploys again time. Timelapse, ten minutes. It really pays to be on a high performing team.
And you really have to internalize this truth, that speed is safety, speed is is is balance. This is like riding a bike. If you slow down, you fall off or slow down, you die.
Right. Speed is safety. And just briefly, on the build versus buy, you know, this is something there's no one right answer, but, you know.
It's pretty clear that more and more, more and more work is is up for engineers of all types. We're trying to spend more of our time on our core business differentiators and less time on infrastructure. Infrastructure is the stuff that we have to build in order to get to the stuff that we want to build.
And wherever possible, we should make that someone else's problem, because if that's their mission, they will do it better than you. You should do your mission best. You can tell him from Ops, because this is how I feel about software in general, build reluctantly, all code is legacy code.
Kill your darlings. My friend likes to say that the worst code of all is code you have to write making yourself the best code of all is code that someone else writes and maintains for you. Sorry I misquoted him.
The best code of all is no code whatsoever. Second best code is code someone else writes and maintains for you. The worst code is anything other than that.
Also, it's the job of senior engineers to amplify hidden costs. You need to point out when things go well so that they don't go unnoticed. Right.
When you've invested some time and engineering energy in something that is paying off. Keep bringing it up. Keep pointing it out so that you can reinforces the patterns.
And and when you can see we're heading down a bad path, point it up not just once. Right. I really myself do.
But we have to repeat ourselves, assume that decision makers are making the best decisions that they can given their limited information. And you often have access to instincts, intuition, information that would guide us in a better direction. But you haven't surfaced it to the people who need it.
Vendor engineering is going to be one of the biggest growth opportunities of the next decade. I guarantee you, it doesn't mean that you if you're if your business differentiators is like running a big website, it doesn't mean you don't need observability team. You do.
Right. But not to write you a time series database. Right.
You should get a vendor to do your observability like the core product, but you should have an observability team that is focused on being the glue between that product and all of your other engineering teams, making libraries and modules, making examples, making consistent interfaces, use cases, examples, docs, right? Constantly keeping up on what's out there and and bringing the team like a proposal. Any time that technology has leapfrogged, I feel like I really want to put a nail in the coffin of a resume driven development or people who want to be put on these very prestigious projects where they write absolutely useless software because it will get promoted to the next level.
And the only way that we can do that is by being very conscious that whatever we praise and promote people for is what we will get more of. If what you want to see is engineers who are ruthlessly efficient and who have the business in mind, promote on that, don't promote on the raw difficulty of algorithms that they write. That's bullshit.
That will get you writing. That will get you people who write so much bullshit code. Your system is a snowflake, high performing teams, both contribute to and are a consequence of a well-running sociotechnical system.
It is the job of every senior engineer, manager, everyone who understands how these systems start to work has has a stake in making sure that they are run well, most of us, most of us have systems that are just like hairballs that the cat coughed up that we have never understood. We've never understood them. And every day we ship more code we don't understand to these hairballs that we've never understood.
And then we wonder why it's a tire fire. It's messed up and it doesn't have to be that way. If we actually put our glasses on, if we invest in observability, if we invest in education, if we don't try and if we don't buy this line, our tools will tell us what we're supposed to look at and what it means.
No, they won't. Any machine can tell you when there's a spike. Only humans can attribute meaning to that spike.
Where are we going? Well, systems are getting more complex, exponentially more complex, and this means we need to rely on each other more than ever because none of us can fit it entirely in our own heads, in our own heads. And I feel like there's a real there's something kind of beautiful about this, because the next generation of systems, they aren't going to be built and run by people who are just checking and just putting in their eight hours a day.
It's gotten too hard. It's gotten too complicated. You have to care too much about what you're doing.
You have to invest too much of your creative self. You have to rely on so many other people on your team. Every time you're building one of these systems, you understand the part that you're on at the moment.
But when you're debugging, you have to understand everything that everyone on your team is working on. And we need to be building tools that help be our shared brain, take it out of our heads when we can't share reason about it. Put it into a tool we can.
You can't model these systems in your in your head and reason about them, and if you try, you're going to be outcompeted by teams with better focus, better tools, runbooks and canned playbook's just don't work anymore. Our systems are too unpredictable and your labor is a scarce and precious resource. If your heart isn't in your job, you won't do the best job you can do and you should probably lend it to those who are worthy of it.
You only get one career, if you like, in the same way that we need to raise our standards for our systems and not accept this hairballs anymore. Like you should possibly raise your standards for your career as well. You only get one of them seek out the high performing teams because that's where your career will take off by leaps and bounds.
You're surrounded by people who care as much as you do. Not mired in tech debt and shitty processes. I think this is a really neat moment in time where the underpinnings of technology are really moving towards a world that's more distributed, more egalitarian, more or more distributed.
And I feel like it's a real opportunity for us to to make things better for the humans who run it as well. We have a dual mandate, all of us, to our teams as well as our customers. And we can do better.
Let's do it.