Shift-Left Ops Mindset: Why Devs Must and How to Include Observability From Scratch – SKILup Days 2024
It still seems to me that developers and pre-production engineers do not take into account non-functional requirements like observability from the beginning, such as during the development phase, focusing only on business requirements and trusting that it will be covered by someone else. Unfortunately it still happens even after 13 years of DevOps movement, which impacts the quality and reliability of the final product.
After my years training and mentoring teams, I have some well proven recommendations to address this problem on a continuous improvement path, and I am ready to share them with you.
Transcript
Hi, everyone. Welcome to this talk. Today, I bring some ideas on how you can shift left the operation mindset and why Debs must, and how they can come this from scratch.
Because from my experience, year, years, and years working with different companies, I have seen that developers, they are too much focused on the business side, and they sometimes they forget that it's important also the operations side to maintain this in production, to maintain this for customer. So I, I can explain some of the problem from the scratch for you. And then, uh, we gonna, uh, learn from my, my failures and my success, right?
Some ideas I have to share on how to shift that ops mindset. So, who am I? I am Ros Brazilian, but, uh, based in Chile now, uh, working at Citi, but, uh, across my career, I have been a DevOps human and a software engineer.
Uh, a month ago, I just joined the ambassador program of DevOps Institute. Uh, so I, I'm so happy to be here speaking as an ambassador. Uh, but also I participate in some communities like the DevOps days in Santiago, the Chile where I'm based now, uh, to bring this amazing event here.
And also some other communities, open source communities like open source, Santiago platform engineering, and the CNCF. There, you're gonna find me, right? Uh, here is my picture at the last Cucu in South Lake City was an amazing event.
Uh, you can also find me in my bank. So let's start, let's start, um, under understanding, uh, generating a shared understanding of observability. Probably you'll see a lot of this in the other, in the other talks today, but, uh, I want to use my own words right from scratch.
So to understand all the concepts, right? And have this can front communication between us, right? So let's start from the SDLC issue, right?
And in a simple value stream mapping of, of our development lifecycle, we're gonna have customer needs that will turn through this flow to be a product, a product that will be supplying this needs of the customer right to, to challenge the marketplace. And, uh, some companies will try to win through this process, growing this from understanding the, the need ideating that something, developing some software, some solution with technology to later verify and promote this production. If we go a bit deeper of the, over this, this pro, and we check of the, the outputs of this software development life cycle, we're gonna see that from the need.
We can have an idea, and that idea will be developed by someone, someone inside the company, someone inside the team having a white box where we'll be able to see everything inside, right? Because we have to code in the development phase. But later, when you move that idea that is being transformed to code to technology, we're cannot see that this white box will turn into black box where we will not be able anymore to see what is inside, right?
We will need to get back to the code to see what is inside. But, uh, definitely during the runtime will be hard to see what is happening inside. And this thing, black box from the development phase will go through prep production, later production to turn into a product and provide some service, some feature to our customers that can be, uh, internal external customer, right?
But the, at the end, this black box will be worked by different teams, right? So, uh, it's not just a matter of the development team that traditional we have in the development phase that will be able to know what is happening inside. Because also they are writing the code.
So it's easy for them to know if the software is falling is crashing, if something is not working as expected because they have the directly the code, it's the source code. And also they can back directly from the code, from the EDE, I can connect my software running and then go through each of the lines of code to understand what is happening there. But, uh, again, to, from a traditional organization, we are gonna see all these different roles, different teams working across the value stream to finally release this product as a black box to the approach where they wanna probably have some tool gates to ensure that this black box they are receiving is, uh, secure is safe, uh, is reliable, right?
So there is a lot of things, right? And also these barriers between the teams that traditionally we have well known from the DevOps movement as wall of infusion generating even more disruption across the flow. So going and deeper on this concept that is DevOps, that probably you are, you'll be saying, okay, with DevOps, I can solve all these problems, right?
But, uh, the depth of humans, usually they are too focused on breaking down these barriers, the walls of fusion between teams to make them communicate right, to speed up the flow from left to right, or even moving these tool gates and make four more to the left to make this, uh, fast and better feedback available in the areas stages. But, uh, the problem I have seen even as a DevOps human working as a DevOps engineering site economy or a DevOps lead, is that sometimes we focus too much on implement automation and implement a lean to remove the waste of the flow and make it fast from left to right using a lot of metrics. So there is a set of metrics called Dora DevOp assessment and re research and assessment.
That is why use it that focus a lot on the, on reliability of the, the software that is going through the software development lifecycle and also the productivity. But sometimes we are missing some aspects that is coming from the non-functional requirements, non-functional requirements that will help to, uh, even remove more of the wastes of the flow, generating a better solution at the end for our end customers. So I, I broke one of the practices I have used in the past.
com, that is called it non-Functional Requirements Map. Uh, usually I use this practice to help teams inside organizations to identify that there is a lock of concepts and also aspects of the software that sometimes we are not taking incur in the early phases, right? When we are ideating something to supply a need, where we are, when we are developing something from a idea like observability, if we turn back to the example I mentioned it before of the software development next cycle, and the outputs of each of these spaces, we are gonna have our black box, right?
Our black box that should be running introduction. But what happens if this black box start to fade, right? It start to, uh, not work as properly as expected, right?
Uh, and, but, and we have the development team, even in the organizations that already changed their ways of working and the infrastructure, uh, where the, the, the developers, they are more involved in the operations side. We still have all the business working and, and running, right? We need to keep the operation up as the time we are developing new features to keep competitive in the marketplace.
So not always the operators of the software, this, the engineers that are looking to the software will be able to understand what is happening there in this platforms because it's not easy to communicate with the software that was developed by someone else, right? Uh, I was developer in the past, so I know that this is a pain, right? Usually we have logs, we have metrics, a lot of monitoring tools, but, uh, if we don't take care of, of how this black box is communicating to the engineers, that was not part, were not part of the development pro process.
The the black box will be a challenge, right? And it, it's the norm. It, even when we, we said, okay, we are 13 years ago of DevOps movement, or even observability is not something new, but we still have a lot of developers with this technical debt, right?
Teams developing, just focusing on the business. And when we have the black box failing in production, it's hard to know what is happening there, right? It's just like a baby.
And, uh, if you, when you are parents or if you, you are part of only, you know that the babies, they, they don't, they don't, they cannot speak, right? And explain what they are feeling, uh, with the specific language we have, right? So sometimes as parents, maybe we will think, uh, uh, where is the handbook?
Where is the manual, the guide they can use to understand what is happening there? It's something similar, right? It's not the same definitively, but, uh, when we have this black box in production, we need to have a way to understand what is happening inside, right?
Or even in the early stages, right? Something you're gonna talk later, uh, it's super critical to keep this amazing idea that was supplied by the new product, developed it super fast up and running in production, because you can have, uh, the most valuable product. But if that product is not, uh, available and reliable in production, your customers will move to another company for sure.
So the observability comes with all these different aspects that should be teaching in cans, even in the early stages when we are ideating some solution. Because ideating is not just a matter of selecting the features that will be used by our customers, but also the architecture. The different standards I should be using to code is part of the ideation and also the development base.
So there, I should be looking to how my application is logging, it's generating proper information to my operators. I would be able to notify what is happening with my black box after it packaging everything and deploy in production. Same thing with metrics, RAC and, uh, why not other thing, right?
I need to have my other visible. And there are, uh, based on the behavior of my application, my software, but I know, I know there is a word of difference between what is knowing and what is actually needed to do, right? There is a chance from the, the needs and the products and even have the product up and running all time reliable and production with all the SLAs that's lower.
So we can set the define agree, right? So it's important to understand that gap, gap and work with that. It's part of the DevOps mindset.
I try to promote every time I sharing some idea or some learning, the continuous improvement, the continuous learning, the continuous, um, challenging of the flow, right? So show me how ship let, right, because we already, so that, uh, there is a, an important aspect of the software development lifecycle, which is how I can talk with my black box in production, my software, which is already packaged, and, uh, probably I will not be able to Trish, uh, the pack each of the nines of code when it's feeding the product. And obviously, I will not be there when it's failing properly.
So, uh, I need to be able to have the, the standards defined there and the part of the area stages, right? Because I don't want to fall in production, right? I want to be there reliable on the stable.
So Cayo, what do you, what do you learn from your failures? What do you learn from your, your successes? So the first thing is contracts.
This is something, um, it's not something proper only of the DevOps word. It's also something that is connected with a lot of other movements and aspects of the software development lifecycle, IE in, in general, which is, uh, how we are defining our conditions to work together, right? And, uh, using agile, uh, and one of the frameworks of agile, which is, uh, scram, we're gonna see the definition of done there, right?
So it is one of the, so if I can align with my developers or my stream, align a team, if we have the know team topologies in place, and we have teams focus on the value stream, we, and they have the mindset of, uh, if I code it, I build it, I run it, then I can bring the idea of this non-functional requirement, which is observability to the contract, which is the, the definition of that. And I can include some aspects of observability there. Like, uh, standards of logging.
Ensure that my developers, they are logging properly based it on the standards. They find that by, you know, an architect inside the organization, right? Same thing, using the right agents or the right technologies, right?
To, uh, integrate with the monitoring tools and ensure that my software, the, the, the, that code that will turn in my black box will be able to express what is happening inside if it fails or even before it fails. So we are gonna have the definition of that there, right? And if we are implementing fast iterations, we will be able to even improve the observability in early stages right before to go to UAT before to go even to, uh, uh, to, um, pull request from my feature branch to master, right?
Because I gonna have all these tool gates that usually I learn from the failures in operate in, in production, why we are operating, uh, learn from them from the beginning, right? The other thing is complement, uh, this, um, practice of, uh, ensuring that my software is speaking the same language and being able to express what is happening inside with the Charles's engineering, right? There are a lot of technologies in the marketplace, or even in open source communities we can use to challenge our software in a production like Environ.
Uh, and using this practice of, uh, injecting error there, I will be able to understand if my, my logs, my metrics are really showing the right information I'm expecting from them using the small cycles of P-D-P-D-C-A, uh, with channels engineering, cloud engineering, I will be able to, uh, learn from my observability aspect of my software and ensure that in production at least I will be prepared for the scenario. I define that during the my chosen engineering practices leader. We have a platform engineering.
It's another huge movement, which is, uh, there, right? We have a lot of, uh, material in the market, uh, talking about platform engineering, the state of DevOps 2023, were, was based on, on, on platform engineering. We already have, uh, a lot of community.
I'm part of platform engineering, do com, dot org, uh, community and, and by everybody to be there as well. Uh, so platform engineering is helping a lot organizations to reduce the cognitive load on developers understanding that developers, they sh they have an experience part as well. Or even, uh, if we go outside of this concept, concept of software development, like cycle the platform engineering needs, how cool in any aspect of the organization of the business, right?
But turning back to, to software development cycle, using platform engineering, I will be able to reduce the complexity of my developers to introduce observability aspects in the, in the beginning of the development phase, right? How with reusable components, tools, platform, services and knowledge base available for them to help them to adopt these technologies. Principles are even standards like, uh, using open source technologies that we have in, in the marketplace to supply these, uh, these needs, like OpenTelemetry that, uh, have a huge, uh, support for different languages to improve the visibility of the software inside this black box, right?
Uh, removing the cognitive load of developers to write the right logins, right? Uh, that, um, sometimes will make the code a bit more difficult to maintain, right? So using technologies like this, it's the, the developers or the different teams across the software development life cycle will be able to make the, the baby more understandable, right?
And, uh, using technologies like, uh, Kraken to do tower engineering, uh, they will be able even to challenge what they are doing with technology, like OpenTelemetry and the platform engineering will be providing the platform teams. They can provide these tools in a on demand solution, a doc solution for the, the internal teams reducing the, the, the, again, the, the cognitive load to start adopting this or even reducing the, the gap between know that I, I should be doing this and actually doing this, right? I had the experience in the past working with development teams or string alignment teams, actually, that they didn't talk before, that observability was a non-functional requirement.
Important thing to, to include there. And, uh, actually sometimes they didn't know how start doing anything with observability, which tools, which technology they can use. So with platform teams, we can provide the stable, uh, compliance solution inside the organization with well-documented and boilerplate solutions.
I mean, uh, templates, they can restart from, uh, with the right tools, the right configurations to integrate with other platforms inside organization to make you have visible information like Ana Pro, Grafana, Loki, jigger Ali, right? To make everything more visible, more understandable in production. And finally, some other things I want to share today is, uh, something outside of the, the use of technologies, right?
And, and practices, which is the communication inside organization, right? Uh, sometimes it's different teams. They are working for the value stream, but they didn't know where one start and where others start.
And also they are totally disconnected. Some of the open leadership, uh, transformation I have been doing when I was part of Red Hat was communicating these different things and generating shared goals. And shared goals can be related also with observability, ensuring that the, uh, the software we are generating together is really understandable in production, is prepared to generate, to provide the right service.
And observability is critical for that, right? I want to know that my software is properly logging the information. So I am, if it, it, it, if it crash, right?
I will be able to know what is happening there. It's the connection to the database. It's a third party.
A PII had just one issue today with, uh, some one team. They were asking me how I can increase the timeout of my Kubernetes deployment. So I asked him, why do you want to configure the timeout directly in your Kubernetes resources?
If, uh, probably is your software, which is taking too long, or it's, uh, integration that is taking too long, do you look to the logs? Do you look to the tracing? Do you, are, are you implementing tracing features to know what is happening inside or the, the, through the, the, the cycle of the process that from a client reaching your API and receiving a response, that is crucial.
So if we have shared goals across the organization, across the team, uh, everybody connected to the end, end, uh, outcome, which is have proper products supplying the needs of our customers, and, uh, have everything there reliable and stable, observability, observability would turn something important as well. And the part of this will be outcomes, metrics, right? OKRs, TBIs, uh, a lot of, uh, all I I will say most of companies better, uh, are trying to use metrics.
Metrics. It's super important because we cannot know where we are. We'll not be able to really improve if we don't know where we are, right?
So using metrics, I'm gonna be able to see I'm here, right? And Dora is one of the metrics I mentioned before, right? It's why use it and help us a lot to understand what is happening in production, because the four of them are based on what is happening there, right?
Uh, how much time I take to move something done to production, how much, uh, I'm deploying with production, how much take to restore something production and, uh, how much phase my deployment? So, uh, the, the specific metric, meantime to restore is, uh, from my, my experience is, uh, use it sometimes wrong to, uh, help teams or enforce teams to take care of observability. And, uh, I want to stop here because sometimes, uh, metrics, they are important.
Again, we cannot stop the use of metrics because if not, we'll not be able to really know where we are, right? And if we're improving or not. But metrics must be part of how we are progressing, where we are today, where we should be tomorrow, right?
Not part of our targets, right? I have seen a lot of teams pay because they set targets of, as, as an example, meantime to restore because the observability was not good in the beginning. So they defined it that meantime to restore must reduce 50%, right?
Uh, setting the specific numbers right was not this abstract outcome that I just said, which is not bad. But, uh, they just enforce from top down changes in observability, enforcing tools, enforcing practices to the teams that have been working a lot of time as they want, right? Uh, and uh, at the end of the year, they fail.
They have a lot of teams running in, sorry, uh, running in Pols because they didn't know how to pick something. They must show in one way, they, right? So improvement actions should be the target, right?
How are you logging? How are you implementing the standards of logging? Which kind of technologies, uh, you are implementing as part of the design or as part of the development process to all to ensure that the observability is good enough to reduce the meantime, to restore the time I need to solve something in production that is missing M series, right?
So this is more or less what I have prepared, right? So I gonna thank you for joining this conversation, this talk with me. Uh, if you have any questions, you can join me in the chat, use the code cure to, to, so to see me there, to talk with me, uh, maybe share with me your experience, maybe you have something related with, uh, what I exposed today and keep joining your, the skill, the skill app days and this event.
Thank you very much.