SRE for High-Performing Software and Teams at SKILup Days 2024
Site reliability engineering (SRE) is a discipline founded at Google that is now widely practiced across the tech industry. SRE represents a set of principles and practices that applies aspects of software engineering to IT infrastructure and operations. In this talk, we will discuss the key principles and practices of SRE, and how they can be used to build high-performance software and teams. We’ll explore insights from the State of DevOps Report and how SRE can help foster the type of generative organizational culture that is a hallmark of high performing organizations.
Transcript
Hi everyone. I'm Jennifer Poff, and today I'm gonna be talking about site reliability engineering and how the key principles and practices of SRE can contribute to high performance software and teams. And specifically, I'm gonna take a look at the role that platform engineering plays in driving reliability.
I also plan to deep dive a bit on how learning and development programs can contribute to positive organizational culture and which is really a necessary condition for platform adoption. So let me start with a quick introduction. I like to do this at virtual events.
So, uh, hello, my name is Jennifer Pet Poff, AKA Dr. J uh, my nickname is Dr. J because I have a PhD in syn synthetic chemistry.
And, uh, today I do things very far removed from chemistry. I currently lead the global Google Cloud platform and technical infrastructure education team. Uh, we're a training team that focuses on scalable learning programs for an organization spanning thousands of engineers, S-E-E-D-U, and perpetuating Google's site reliability engineering culture is a key part of my, my, uh, and probably most well known outside of Google as one of the co-editors of the original s read book that we published in 2016.
And then, fun fact, if you, if there was a hallway track of this conference, I would tell you that my fun fact is that I'd love to travel and I'm a part-time travel blogger at Sidewalk Safari. Okay, so now that we've got the intros out of the way, let's start with some definitions. So, so platform engineering is a term, it's, it's bandied around a lot, but do we really have a common definition of what we need, uh, what we need, what we need?
Let me start there. So here's, uh, industry definition. So the industry definition of platform engineering, uh, is that it's a discipline that focuses on the design and development of internal developer platforms or IDPs.
These platforms offer offer self-service capabilities and streamline the software development lifecycle, aiming to improve developer experience and increase productivity within an organization. And, uh, I have Gemini to thank for coming up with this definition. On the other hand, we've got the definition of platform engineering for this event.
So there's, there's a lot written here, but, uh, a according, according to this event, uh, platform engineering is about fortifying systems against disruptions, uh, solutions that effortless effortlessly scale, and they, they maximize operational efficiencies. Those are some, some key tenets and things to look out for. So, so really I think these definitions invoke two types of platforms.
So you've got your internal developer platforms and you've got your production platforms, and I think IDPs and production platforms have a, a, a symbiotic relationship within modern development environments. So let's just take a quick, uh, quick comparison and quick look at the two. So IDPs, really the focus here is on a, a primary goal of improving the development experience.
So they aim to streamline software development and deployment within an organization. The target users are typically the organization's internal developers, and you can see here some key features. You know, self-service capabilities, abstractions in terms of hiding complexity through simplified interfaces for interacting with the underlying tools and infrastructure, you know, standardized golden paths, ways of doing things, uh, developer portals.
So, uh, providing easy access to the tools, documentation and, and those self-service options. And then integrated tool tool chains. So you've got a, you know, curated selection of tools that cover everything from CICD testing, monitoring, et cetera.
On the other hand, you've got your production platform and production platforms are optimized for running applic applications in a live environment, ensuring reliability, security, scalability to meet the demands of end users. And while developers typically interact with the production platform, it, it's usually, uh, primarily managed by operations teams like, like sre, for example. And the key features here of your production platform would be things like I, availability, scalability, performance, security, uh, monitoring and observability.
So, um, IDPs, again, in production platforms have a symbiotic relationship. And, uh, you know, they, they, uh, in, in terms of, you know, deployments, sorry, excuse me, development and deployment, the IBP provides a smooth path pathway for developers to build tests and deploy their code to production. Uh, the IBP often automates many, many of those tasks.
There's also a reliability feedback loop. So you can take the operational data and insights gathered from the production platform, feed it back into the IDP to help, uh, platform engineers improve, uh, the tools, processes, and address any issues that impact, uh, production reliability. And then finally, uh, thinking about it in terms of governance and standardization, IDPs help enforce security compliance and architectural standards and ensuring that applications are deployed to the production platform in a way that meets the necessary requirements.
Okay, so now that we have a common language, we're talking about, uh, production engineering, let's look at some drivers of organizational performance. So delivering software quickly, reliably, and safely is at the heart of technology transformation and organizational performance. And I wanna recap few statistics from the annual state of DevOps report that's, uh, delivered annually by, by jra.
So the annual state of DevOps report builds on survey submissions from thousands of practitioners around the world each year, and looks at the relationship between software delivery performance and organizational performance across dimensions like profitability, market share, and productivity. And the four dimensions that are shown here are, are found to be predictive of organizational performance. And, and these are, you know, deployment frequency, lead time for changes, change, failure rate, and meantime to recover from incidents.
And you'll notice here that two of these dimensions focus on speed while two focus on reliability. And I think, uh, IDPs and production platforms can drive improvements across both of these dimensions. So when looking at the state of DevOps report data, we see four clusters of performance across these four metrics.
And, uh, these, these clusters emerge when looking at software delivery. And, you know, performance runs from, from low, you know, medium high, all the way to elite performance. Top performing applications are able to get a change from a developer's workstation through to production in less than a day, and they're deploying on demands.
These deployments do not require immediate fixing about 95% of the time. And when they do need immediate attention, service is quickly restored, usually in less than an hour. So, thinking about this, so how does SRE and platform engineering fit in here?
So the state of DevOps report found that SE and DevOps are complimentary philosophies. So companies practicing site reliability, engineering report, higher operational performance, and teams that prioritize both delivery and operational excellence report the highest organizational performance. And, uh, SREs mission is to protect, provide for and progress software development and systems with an ever watchable eye on their availability, latency performance and cap, uh, and capacity.
You know, traditionally SRE use a set of foundational principles, practices, and culture to operate systems reliably in production. And then I like to think about platform engineering as a means to codify SRE principles and practices into the production platform so that the organization can standardize and scale with less overhead and, and boil. So the one key question to ask is, you know, if typical you build it, will they come?
So the, the executive summary of the 2023 state of DevOps reports, uh, states the following here, uh, it's a bit long to read, so I'm not gonna read the whole thing, but I, I've highlighted a couple of snippets here. So you wanna treat developers as users of the platform. You and, and users specifically need some kind of training to be successful.
Another element here is that, uh, there's a a a way to ident, you're looking for ways to identify and eliminate areas of friction and, uh, in terms of platform adoption. So this is another area where training can actually be, be helpful. So, so more on that, more on that in a moment.
But to realize, again, the benefits of the platform, engineering teams need to use it, but change is hard. We, we know that across any sort of, uh, organizational transformation and what platform engineering really entails is a culture shift. So now let's think a little bit about what drives organizational culture.
So what, what drives it? Is it a state of mind changing how people think or is it about behavior? So it turns out that, that John shook an industrial anthropologist, uh, is, is, is known for saying, uh, the way to change culture is not first to change how people think, but instead to start by, by changing how people behave, what they do.
And and this was, uh, published in, uh, MIT Sloan Review article that that basically discusses that the mechanisms to change, change culture and, and, and how to go about that stated another way. It's easier, your easier to act your way to a new way of thinking than to think your way to a new way of acting. And it turns out that the state of DevOps report research also shows that changing the way people work also changes culture.
So I, again, I like to think of it as confidence drives behavior, and that behavior repeated over time is what drives the culture. And to me, that begs the question, how do you generate confidence to start this whole virtuous cycle? So my job is to deliver learning programs to the engineers at Google, and I assert that learning and well-crafted learning programs can be a driver of a positive organizational culture and high performing teams.
So how, how is this learning? How is it that learning drives culture? So some people think about training and, and people learning things, uh, and getting them to learn things as being about cramming as much information into someone's head as possible, setting up this proverbial fire hose of information.
But the reality is that no matter how much you tell people, they're only gonna retain a small amount, especially in a lecture format. So the reality is training is not about the fire hose, it's actually about believing in yourself, you know, for adult learners, it's about building confidence, it's about fighting imposter syndrome. And this is especially true for people who are new to a team.
Going beyond confidence training is also about driving or perpetuating a desired organizational culture. So just like these hot air balloons training can give your culture a a lift. So going back to our diagram, I like to think about learning as the confidence engine, which in turn shifts the culture into high gear.
Let me share a quick study. So I'm gonna talk briefly about how we measure how effectively we move the needle on confidence with Google's SRE EDU orientation program. So in summary, we benchmark culture through confidence in a four step process.
We've defined a few reliability culture questions that are included in our SRE EDU orientation surveys, as well as in our general developer onboarding surveys. So these questions are, you know, please rate your ability to do the following tasks, explain the benefits of a blend list postmortem culture, measure my services, alignment with business objectives, so getting at the SLOs and error budgets. And three, using monitoring and alerting to make rational decisions about my service.
Now quick disclaimer, while thi this case study is about driving an SE culture, I believe the approach can be adapted to a platform engineering con context. So keep keep this in mind. You know, for example, you could consider asking a question like this, uh, you go, how much do you agree?
Or excuse me, please reach our ability to do the following task. Pick the best solutions and tools to minimize the time spent on production service management. So you could, you could get it, you could use this as a measure of how much the organization understands and is embracing the platform.
So, um, continuing through our e example of this case study, uh, the next step here, uh, in this benchmarking exercise is to use a grounded scale. So if you use a general five point rating scale without grounding, it really introduces variability. So without grounding one person's three could be another person's five, you know, depending on what mood someone's in on a given day.
So, so, so we use this rating scale to benchmark the questions that I shared previously and our rating scale runs from, I've got no idea where to start. I could do it with someone's help. I could read documents and figure it out myself.
I could do it mostly from memory. I could teach others mostly from memory. And you can see here that we're, we're, we're running from, you know, really low level of confidence to super high confidence.
I, I feel you're never more confident than when you can teach something to someone else. So to figure out how well a program is moving the needle on confidence, you need to look at survey data measured over time. So for our programs, we ask the confidence questions four times in six months.
So before orientation after orientation, one month after orientation, and six months after orientation. And in addition to looking at trends over time for SRE EED U orientation, we also look at comparisons between SRE EDU, where the learning objectives specifically revolve around reliability principles and best practices. And we compare this to Google's general developer onboarding, which has a more generalist curriculum that doesn't explicitly address reliability in the same way.
So here's the output of this exercise, uh, and and how we go about demonstrating that S3 EED U orientation moves the needle on confidence. So don't forget our assertion is learning drives confidence, confidence drives behavior and behavior repeated over time drives culture. So this slide assesses and aggregate the confidence of our learners and their ability to explain the benefits of a blameless postmortem culture at different points during their ramp up.
The blue line represents those who took general developer onboarding, and the red line represents those who took S-E-E-D-U orientation. So I, I've, uh, qualified an NLLS here. So it's a metric that we call NLS.
Uh, NLS is what we've been calling a net learning score, uh, or it's in some ways a, a a view of the broad level of confidence of our, our cohorts at different points in, in time. We define NLS like a net promoter score, where we take the top two buckets. You know, I could do it mostly from memory, I could teach others mostly from memory.
We subtract the, the bottom two buckets, I have no idea where to start or I could do it with someone's help and then divide by the total number of responses. So in theory, the scale hypothetically runs from minus one to plus one with plus one representing maximum confidence. And as you can see, the confidence of SRE EDU participants grows over time.
And SRE EDU participants are more confident about this topic at each survey point than their counterparts who took general developer onboarding. And this is important to show that with specific focus on reliability principles and best practices, we can drive confidence. And in the long run culture, we can see that this effect is even more pronounced when we assess confidence on measuring my services alignment with business goals.
So once again, we're seeing confidence growing over time with a bulk particularly pronounced between pre and post training surveys. Uh, so those who complete SRE EDU orientation are more confident on the topic of service level objectives and error budgets than those who completed general developer onboarding before. Our orientation confidence was essentially equivalent and low and then quickly diverged with SRE EDU attendees gaining more confidence more quickly.
So again, totally expected and working as intended since S3 EDU is specifically putting a focus on this subject. We also see a similar trend for using monitoring and alerting to make rational decisions about my service so that those who complete SRE EDU orientation are more confident on the topic of monitoring than those who completed general developer onboarding. Once again before orientation confidence was essentially equivalent and then quickly diverged with SV edu attendees gaining more confidence more quickly.
Now, the observant among you may have noticed dips in the confidence graph between orientation and one month post orientation. So we hypothesize that this is a correct, uh, correction for overconfidence. So I thought I understood this, but apparently I have more to learn.
There's really a lot of nuance these concepts and they can take full time to fully under understand, but we see that between one month and six months confidence increases once again and surpasses post orientation levels for S3 EDU participants, while general developer onboarding participants rate their confidence net negative even after six months. So now that I've hopefully piqued your interest by drawing this connection between learning confidence in culture, let's deep dive on some key considerations if you plan to spin up your own tra uh, training program. And while, uh, I'll talk about this in the context of SRE or DevOps, I think that general principles and considerations can be applied to any organizational transformation.
You know, for example, adopting, adopting a platform. So in light of all this, how should you go about training your engineers? There's really no one size fits all approach to, to training, but rather training methods that lie along a continuum from low to high effort.
So at the low effort end of the spectrum, you've got sink or swim, which means you just let people figure things out on their own with no guidance or support. Self-study means you point people at relevant resources, but they're ultimately in charge of consuming that material on their own. Pairing people up via the buddy system is another option.
Or providing new members of the team with a mentor to answer questions or to shadow and reverse shadow key operations. Ad hoc classes or whiteboard sessions require more efforts. And at the highest end of the spectrum, we have a systematic training program.
So this is a training program with well thought out learning objectives and content to support those objectives and tr uh, typically the curriculum will be offered consistently on a scheduled basis. Couple of tips right away. Avoid, avoid sink or swim.
If you value inclusivity, sink or swim can breed stress, frustration, even attrition. It can also contribute to imposter syndrome. Uh, sink or swim is also not likely to garner a lot of love for your production platform.
So for other, uh, for other options, consider the return on an investment of that effort invested. So some of the advantages of the higher touch options are organizational. So higher, higher touch or higher effort options, which culminate in this full fledged training program, can demonstrate things like leadership commitment to the platform and, and to the development of their employees.
Uh, they can also ensure that everyone speaks with one voice. So if you rely on the buddy system, the results you get may vary depending on who the buddy is and what their opinions are and their own, you know, preferences. Uh, the higher touch options can also help you imbibe the desired organizational culture and reinforce desired behaviors.
Okay, so the first question that probably comes to your mind is, what should I teach? What to include in your training really boils down to these dimensions. First of all, maur maturity of your organization in terms of platform adoption, for example.
Uh, the next dimension is familiarity. So that's more about the knowledge of individual engineers, uh, and and the knowledge that they have about your organization, your infrastructure and your, and your platform. And, and then there's experience.
So this is about the experience of individuals being trained, including their technical skills and familiarity with platforms and platform engineering. So let's start by considering the case of those just starting out driving platform adoption, for example. So how can you ensure that you minimize the pain on the, the long, the potentially long and bumpy road ahead?
So this is the low organizational maturity use case. You're just getting started and rolling out your platform. So, so you really wanna start by considering who you have on your team and how to tailor the message, the reaction of individual team members to proposed changes in ways of working.
It's gonna depend on how familiar they are with your organization and the current culture crossed with their level of platform engineering experience. For, for example, have they leveraged internal developer or production platforms elsewhere and seen the benefits of it. So people who've just joined your organization and who have a low or, or no low level or no experience with platform engineering, they're likely to just roll with it and go with the flow.
The people you really need to work watch out for are those who've been in your organization for a while, but they don't have personal experience with the benefits of, of, of, of, of, of platforms. So these folks may be resistant to the proposed change. And if so, you'll need to address, you know, what's in it for me?
Why is it different at this time? If you know the way you're asking them to work in this sort of organizational transformation is just seen as the new hotness that will, that will quickly fade for people who have existing, uh, platform engineering experience or seen the benefits of of of, of, uh, developing via platforms. These are your catalysts.
You really wanna empower these members of the team to tell their stories. So bottom line, tailor the format and content of your proposed training to address these unique needs. So now what about organizations that have already have a well established platform that is widely used now?
It's all about building confidence and reinforcing the practices and culture that you've built. So let's talk about how to go about building a training program for this use case. So this is the high organizational maturity use case.
You've already got this well established platform, engineering culture and practice. Now we need to dig in and evaluate the mix of people on the team. And once again, we're gonna do this long dimensions, organizational familiar familiarity as well as platform engineering experience.
So newbies, these are the folks who are just joining your company and are less familiar with with platforms. And, you know, working in a, in a platform engineering environment, internal transfers are expected to have high familiarity with your company, but may have less experience with platforms, especially if they're coming from a part of the organization that hasn't yet onboarded to it. For example, old timers.
So these are the folks who've worked at the company for a while, they know your company, they know the culture. And then we have industry veterans. So these are folks that have a high level of experience with platform engineering practices, but are perhaps new to your company.
So now let's look at the learning needs for each of these audiences. So once you've assessed, assessed this mix, figure out where to focus your training contents. So, uh, first off, newbies need the most attention.
They need to learn your infrastructure and your culture. Internal transfers may know your systems, but may be less familiar with your the platform itself. So focus, focus here, oldtimers.
These are people who are experienced on both dimensions. So go for technical depth to unlock their career growth. Also get the oldtimers to teach and share their experience, um, and and to to help others really understand the benefits that they're, that they're getting and understand these ways of working.
And then finally, industry veterans. These are people who need to ramp up on your infrastructure and and processes. They may need to unlearn some bad habits they've learned elsewhere.
So, so really focus on the specifics. Okay? So until now we focus on evaluating your organization and thinking about what you might wanna include in a training program.
However, there's a whole other slide we've yet to consider. I like to draw an analogy between software, develop the software development lifecycle and building and running a training program. In both cases, we need to consider the what and the how, the what and the how.
The what from a software development perspective is your shiny product features and the how is deploying to production in a reliable way to meet the needs of your users in the training program context, the what is your training content and the how is deploying a consistent and reliable training program that meets the needs of our students. So let's talk about the all important how of delivering a training program. So this is really about operations and there's two dimensions you should, you should consider.
So how big is your organization and how fast are you growing? So these operational dimensions inform how much effort it makes sense to invest in training. If your organization is small and growing slowly focus on one-on-one, knowledge transfer through mentoring or shadowing if your organization is, uh, yeah, large and sorry, yeah, large and growing slowly.
Yeah, small and growing slowly large and growing slowly, sorry. Invest in ongoing education as a retention group development tool. If your organization is small but growing rapidly, invest in auto boarding.
And then finally, if you're both large growing rapidly, it makes sense to invest in a a full lifecycle training program. So at Google, uh, the Google SRE team falls into this upper right quadrant. And as such, we've invested in a full-time team to develop this, this full lifecycle training program.
So to summarize, invest the most under conditions of rapid growth. Small but rapidly growing organizations benefit the most from onboarding training. But don't forget about your existing people.
Invest in ongoing education for their career development. So now that we've talked about the the what and the how, let's talk about the where as in where to begin. So where, where do you, where do you start when you're developing content?
So I have two words for you, ass plus batts, or rather one acronym as batts. Where as batt stands for a student should be able to, let's talk about possible as batts for a training program. So a student, student should be able to, your first thought might be something like a student should be able to understand the food service, but this is really written in a very passive and hard to observe way.
Instead, focus on the behaviors that you wanna drive. So a student should be able to use a certain tool to identify how much memory a job is using. Interpret a graph in your monitoring tool to identify the health of your service or, or move traffic away from a cluster using, using a drain tool.
In, in five minutes you observe, you really wanna observe and measure how the training is applied. And this is a great construct for, for doing just that. So there's a simple model you can use to aid in the development of your training content.
Instructional designers call this the adding model where ADDIE stands for analyze, design, develop, implement, and evaluate. And as you're doing this, you really wanna keep a very tight and iterative feedback loop to drive improvements in the process. And that's how I recommend using, using learning to drive confidence and ultimately the behaviors you want to drive to achieve hopefully broad platform adoption.
Okay, let me leave you with a few key takeaways. So training is an investment. It's an investment in your organization and in your people.
Make sure to evaluate the cost and benefits. You wanna make sure you're making the right level of investments and where to invest. This really, this really depends on the what and how of your organizational circumstances.
And in summary, you know, I hope you'll be able to use these techniques to drive a culture or, or platform engineering is embraced to achieve greater velocity and greater stability. Uh, you know, the win-win that we highlighted in the state of DevOps report. And again, quote for me here, uh, confidence is what revs your culture engine.
google/resources. And finally, I would be remiss if I did not give a shout out to the state of DevOps report. You can find and download the full report here and, uh, once more.
Yeah, just, uh, j dev is where you slash report is where you can find the full text. There's also Dora community if you wanna learn from other, other practitioners and other, other, uh, do enthusiasts. And finally, uh, this is my contact information on LinkedIn.
Uh, if you wanna learn more, if you have additional questions after the session, I'm always happy to connect on LinkedIn. And with that, thank you very much and uh, really appreciated the opportunity to be here today.