Harnessing Observability and Chaos Engineering for Next-Level DevOps Practices – SKILup Days 2024
This presentation aims to explore how integrating observability with chaos engineering practices can drive the best DevOps outcomes, enhancing system reliability and performance.
Key Points to Be Covered
–Defining Observability in DevOps: An exploration of observability’s role in modern development and operational processes, focusing on its importance for understanding system behavior in complex environments.
–Introduction to Chaos Engineering: Examining chaos engineering techniques, including their purpose and benefits for testing system resilience in controlled conditions.
–Real-World Applications: Presenting case studies that highlight successful implementations of observability and chaos engineering, offering insights into practical applications and outcomes.
–Practical Integration Tips: Providing attendees with actionable advice for integrating observability and chaos engineering into their workflow, enhancing system resilience and performance
Transcript
Hi, uh, welcome everybody. Well, my name is Alejandro Mercado, uh, rice and Burn in Mexico City. Um, this is a, my talk is about the intersection about cows engineering, um, a little bit of observability and eh, DevOps practices.
So, so very honored to be here. So this is a little, a little about me. Uh, I work as a senior DevOps engineer, and I have a lot of experience in, well, more than 25 years doing that, that activities and like, um, uh, girls engineer practitioner, and the Bob CICD and related things about that.
So, uh, I, I want to start with the end. This is, if you want to take a, a takeaway from this talk is, is this, uh, why do you want to do cows engineering and how does it benefits observability? Because of course you gain visibility, you identify failure modes and of course, connected with other tools, you can, uh, automate remediation and enhance collaboration.
So, so this is like some of the, uh, the best principles that, that we have been practicing this for many years. We can elevate our, the most practices and, and build more reliable, resilient, and responsive systems. That's the whole point of the, of, of, of this talk.
So it's, so we can start talking about a little bit about observability and telemetry and, and why is this important to our current systems? You know, we are gaming complexity, cloud native system on-premise system, hybrid systems, databases, network connectivities, uh, container serverless. We have, uh, uh, a lot of complexity in our current system.
So to observe things is not like a new concept, but it has gained ency in the subway industry recently. So this is just for a cultural observation. I, I, I think that when BLS made the first computer programming, uh, I, I am pretty sure that at this precise moment, she has started to think about how to improve the, the, the, the, the programming, the performance and, and detecting errors.
So, so I'm pretty sure, uh, I mean it's, this is something that has been with us and has bothered a lot, but I am pretty sure that, and the definition and ions of telemetry, because as I said, it's not a new term. It's, it's quite old. So we adopted this, this, this term from others in industry, like a lot of other concepts.
So telemetry is the process of collecting and transmitting data from remote or inex accessible locations. So we can see this, this slide, that, that can be traced back to the late 19th century in this STEAM era. So we have this motivation from lot of, many years ago, so to collected data to improve our systems, uh, uh, at that time it was of course another type of system.
But this is just like, uh, an example of the motivations of, and the relevance of this, uh, uh, about measure and observed systems. Uh, so this is another example of telemetry in this case is, is, is in the health, uh, territory. So, so that's why it is so important.
We have, we, we know we can see a lot of tools that have multidimensional analysis. So this are just a couple of example examples of telemetry. So we can request tracking, even login, resource monitoring, latency tracking.
So, so the idea is to have all, all, all of these access together. So you can think about a, a root cow. So that's very important.
So, so telemetry is, is a, is a, I can say this provides data need, needs to gain visibility, I mean to ingest data. So, so we can have these other tools that it's going to help us, uh, to improve many things in our systems. So if you want a brief tagline, take all about this slide.
So, I mean, we are talking about a lot of experience that we already have, uh, implementing observability, telemetry, monitoring system, observability, and now ca engineering practices. Uh, so, so it's not, I mean, it is, it, this is due to the, to the evolution of how, how we are doing things right now. Um, uh, prospect of, of the, at least a couple of years.
So we, you can see the, the last two, uh, milestones. Uh, we are not, not now doing this, this, this practices, these both practices with artificial intelligence and new concepts like AIOps and ops. Um, we are seeing now a platform engineer.
So we are gaining complexity. So we need, uh, other tools to, to help us to integrate, um, our practices, uh, in, in current systems like I, I mean this is a, a vendor agno agnostic tool, but you can see open circles and proprietary tools like, you know, this is a lot of, um, these are a couple example of these tools that we currently have available. Um, probably you already know about cows engineer, I mean the fundamentals.
But cows engineering is a discipline of experimenting on a system to build con confidence on the system's capability. So the idea is, is to, to know about the problems or issues before having it in, in, in production. So if, if we know about it before the user, I mean, we, this is observability.
So, so I mean, it is related the, the, the cows engineer practice practices with observability because sometimes we, we don't know, we have several doubts about the, the behavior of our system. So, so there must be a way of, of testing our systems, uh, in, in early stages or even in production. So, so of course there are a lot of cows engineers tools from website.
We have the, the traditional monitoring tools, the observability team tools and, uh, cows engineer tools. So we are seeing, like the evolution is overlapping in features, but we can see very specific, uh, tools for cows engineering, like steady bit open source or proprietary tools. It's depends of the business necessities.
Um, but there are just a few examples of the many cows engineer tools available. Uh, and the decision of course will depend on, there are specific infrastructure and testing requirements. So we have cows, monkey, lead, cows, uh, Grambling, uh, pba, steady be, et cetera, et cetera.
So the idea is to integrate observability and cow engineer, uh, I mean this is very important because, uh, it's going to be, um, a source of, of, of information, information that we be, be, will be valuable to our system to know if, if the system is re uh, we, we have this proactive res we with this integration. So there are a lot of companies that are, are doing that. So they are, uh, seeing a lot of benefits of, of observability and calcium engineer, like the, like the incident response time, the maintain to repair, you know, this, these numbers are, uh, decreasing or increasing depends of the, of the, uh, I mean, uh, we reduce the, the maintain to repair to say something.
The service reliability, the application performance, the infrastructure ization, we don't have, um, we, we obviously save money with all these practices. So I mean, if you want to, to think about it in the side of the business, well, this is, are some metrics that we are going to improve adopting these practices. And, and the idea of the whole idea is, is well just to make a ahan hypothesis or experiment about your systems.
So, so like, like what I mean, uh, when you deploy a a system to production, you, you may have some doubts because you, I mean, it is not in, it's different for, for the traditional testing, I mean from genetic testing or interracial testing. But you maybe want to know about what, what happen if, if there is, um, uh, a, a latency in the network is it's the whole system going down or, or just AC company is, is, is, is is the well architected system, or we need to do, uh, uh, something else. So we have a lot of, uh, range of experiments like simulate and server failures, uh, injecting neck or latency trigger and database error.
What happening, sometimes we, we don't, it takes too long to get a response from a query, uh, or injecting partial failures. And I mean, when the system is running for, for, for instance, think about this, evaluating a a, a bot that scaling well in a context of Kubernetes what's happening, maybe we, we have this doubt about it's going to scale our system. So I mean, uh, we are not going to wait until, until the black Friday, let's see, cyber Monday to, to have a lot of customers in our system.
So to, to know, or if, if the, if the system is, is sst. So maybe you, we can do at testing with cab engineering tools like stressing the CPU, scaling the deployment or, or make a specific deployment to see if everything is, is going well. So is there al HPA is al put out scaling.
So, so we can test with these tools, um, integrating with observability tools, uh, we can know we can handle the, the increase of load. So, well, this is why this is important. Of course, you, you can have a lot of questions, a lot of experiments, a lot of doubts.
Um, and we are having this other, um, I I can say another approach with, uh, generat with these new tools like generative ai. So we can also, uh, harnessing the power of, of, of cap engineer in the book practices. So, so we are seeing a lot of integration with, with, because when we do these experiments, we get a lot of information so we can correlate data, so we can get, uh, road cals analysis.
So maybe you have heard about at the dark depth, the, the dark depth. This, this concept about, well, maybe we have a hypothesis, maybe we can do some experiments, we have some doubts, but maybe we didn't know about a problem. It, this is like, um, the technical debt, but you are aware of the technical debt, the, the, the dark debt is that, that, that you don't even, uh, know that they exist.
Uh, some, some issues that may appear, but you, how come you fix something that you don't know that exists? So, so well, that's why it's important to the, uh, this, this practices. So cast engineer became a proactive approach to understand system complexities.
So unlocking the sec, the secrets of dark depth, so dark depth issues, dark deb issues, lurking in complex system, making their detection challenging and manifesting as, um, foreign anomalies in people, practices, processes, application platforms and infrastructure. So, well, this is why I think it's so important to, to, if you are not doing Cal engineer and using an observability, I think that, uh, you are losing competitive advantage. So what are some benefits of the proactive depth management?
Well, early detection of dark depth. I mean, not only, uh, our hypothesis for experiments, even the dark depth, we, we can detect on early stages so we can avoid problems and productions. So obviously this is going to, uh, this is going to be a, uh, a cost savings in the long run.
So I can say, if you are not doing cow engineer, you are losing money. Uh, so of course it's going to improve the customer experience. If we don't have any, any problem on production, uh, of course the customer experience will be better.
So in enhance financial visibility, we can save a lot of time, a a lot of resources, um, of course a lot of money doing these practices regularly, as at above practice. I mean, in the, in the software developer lifecycle cycle, we can integrate these practices the same way that we are doing the bs, and so we, we can reduce risk exposure. So yeah, maybe this is like the, the one of the conclusions or my conclusion is like that, uh, if we want to wrestling the cost of not doing cast engineer, uh, well resilient systems, so we can be pretty sure, uh, doing these practices that we are going to have resilient systems, uh, optimization opportunities, faster issue resolution, and of course the competitive advantage.
This is very important. So it's a crucial stepping ensure the long, the long term rest, license and profit feasibility of your business. So by proactively testing and improving your system, you can minimize down time, optimize cost, and stay ahead of the competition.
Yeah, that this is like a, a little redundant. So we, we are going to have this, this, um, takeaways from, from this brief talk. So observability provides visibility, gives you a comprehensive view of your system behavior, uh, allowing you to quickly identify or a solution.
So, so, uh, calcium engineers experiments or hypothesis practices improve alliances. So by intentionally injecting failures, we can have fault tolerance systems. So this is an integrated approach drives in innovation.
So combining observability and how engineer empowers your teams to continuously innovate and deliver high quality software with confidence and to have better systems, you are having a competitive advantage. You know, if you want, if you are in the business side, and if you want to talk with the, the CEO of the company, uh, you can talk about money. So, so you are going to have this, uh, uh, savings in, in resources, in time, in, I mean, at, at this time we're talking a lot of about finops practices, but we can improve our finops practices with, with this cow senior, I think they are very related.
We have a lot of new or not so new practices that are overlapping. So the idea is to have a full, uh, 360 degree view to have this competitive advantage. So, well, uh, that's it.
Uh, if you have any doubt or question, comment, you can reach me on LinkedIn or any feedback I, I will be around. So thank you very much.