Platform Engineering 4 FinOps with Axel Labruna – DevOps Experience 2024
FinOps is the practice of managing and optimizing cloud costs by aligning expenses with resources consumed, maximizing business value as north star. The methodology involves Finance, Operations and Technology teams working together to monitor and control cloud costs across the organization. As managers and leaders of technology teams, the role of Platform Engineer is becoming increasingly important, not only from a technical point of view, but also in the business.
This is an opportunity to consolidate cloud native strategies, involving software development, QA, Data, DevOps, Security, Architecture, Cost structures, budgets and corporate objectives.
To capitalize on this opportunity, it is important to stay informed and prepared to play the key role we can have in all of this.
Transcript
Hi everyone. This is Axel Launa from Argentina, and I'm glad to be here with all of you in the DevOps experience, 2024. Uh, I hope you like this talk.
We are going to, to chat about platform engineering and finops that, uh, they are both, uh, two trending topics nowadays. Uh, so what we can start right now, uh, we'll start with finops. Uh, it's a, a term we are using a lot nowadays.
Um, and what this is about, well, finops is the practice of optimizing a cost in cloud, uh, without sacrificing spill quality or security because, uh, it's pretty simple, uh, for anyone to, to add cost, right? Uh, without considering these three points that are quite important for our applications, uh, without quality or security, uh, there's no chance to get an application to production. So the challenge here is to maintain all of these three pillars, speed, quality, and security.
And at the same time, uh, taking care of the, of the cost we are, uh, we are having in the cloud, right? Because sometimes we just move our applications, we, we modernize them, or we just migrate to the cloud, but we are considering cost. And when the, the bill comes, it's like a shocking surprise, right?
Uh, well, this practice, uh, has that, that intention to, to cut those costs that are unnecessary or maybe, uh, that comes with, uh, for example, um, resources that are not used properly or that oversized, uh, well, those things and not more, uh, comes with, with finops. And just to deep, uh, a little, uh, more into the framework that finops give us. It's like Scrum, for example, scrum, uh, give us like a, a framework like some, uh, guides.
Uh, it's not a recipe for success, but it's just a guideline to, to follow, uh, and to have like, uh, better practices and a good roadmap, uh, to, to transition this cloud cost optimization process. So the, the first time, it's, uh, it's very important to maximize, right? The, the business value in the organization aligning those strategies that the business has with technology.
Most of the times they are treated separately. Uh, so, uh, that, uh, brings like a, a lot of issues on arguing because business wants something, technology says another. Uh, so the, the, the first recommendation, the first guideline here is to work together to align strategies, right?
Uh, and to have like a smooth process involving everyone. Uh, we'll see that, uh, Phillips framework give us like, uh, a couple of roles that are suggested from each area and each specialty. Uh, so the, the point here is to work, uh, as a team from the start and aligning those strategies.
The other thing is, once we have those roles or those people, uh, into this process, uh, and this exercise, uh, we can incorporate processes, skills, and tools, uh, for both engineering finals and business teams, right? Uh, the, the idea is to give all of them different skills, processes, and tools to make the the job done right. Um, we'll see, uh, in the couple of slides, um, some tool that we have to, to work with.
Uh, but if you enter to the finance, uh, do org, uh, page, you can see like, uh, a lot of suggestions, tools that are aligned with these practices. Uh, and select the, the best option for you, for your company, the, the one that is more suitable in the, in that specific time of, of your process, right? Uh, maybe one tool, it's useful in one particular time, but then, uh, a couple of months later, when you are like more mature in your process, you will need to replace it or to add some other, uh, into your pipeline.
So finally, the, the idea we were, uh, talking before is to, to work collaboratively, uh, and make data driven decisions. That's very important, uh, because sometimes, uh, this, this kind of process where business on technology argue because they don't know very well each other, and that the discussion, uh, is about we don't need security, we don't need quality, we don't need this tool. Um, it's, yeah, about feelings or sens sensations or, or, or maybe like it's, uh, like, uh, something that it's not valuable for us in, in our point of view, but it's for the other.
So, uh, to avoid all those things, uh, it's important to have data-driven decisions, and these tools that finops give us, uh, help in, in that way, right? We, we can also add a lot of more, uh, complex, uh, ideas, tools, and processes, but, uh, we can start small. Uh, that's an important thing, starting small, uh, but always having like a data driven decision process, right?
That that's the, the ultimate thing, uh, as very important to start with in, in this process. Then the, the other thing we are gonna talk about, and it's like very related to it's platform engineering. Uh, maybe you have, uh, came across with, with this term, maybe not.
Maybe you are just, uh, someone that's working in technology or business, and you start hearing about DevOps, DevSecOps, SRE, uh, platform engineering, CloudOps, master go, all kind of terms, uh, right. Uh, so what is platform engineering about really? What's the difference between the other roads?
Well, just to give like a, a quick idea. For example, let's say that, uh, we have three teams, uh, in, in the company, there are three squads, uh, software development, uh, teams or DevOps teams, uh, whatever you like. Uh, but the important thing here is that every team has its own process, its own tools, its own way to develop.
So, uh, when you have, like, in this example, three teams, it's like more manageable. Uh, it's easier, but if you think of that organization, of that structure in a larger company where you have like hundreds of teams, uh, it starts to, to become very hard to manage all those processes, and you have like the standardization issues. Uh, you have, you find it yourself like managing as like hundreds of tools, because one team uses, uh, AWS and GitLab, another uses Azure and GitHub, maybe some other GCP and Jenkins, and it's very hard to maintain standardized and ensure the quality and security in all teams, in all tools.
So, uh, that's where the platform engineering comes in, into the scene, right? Um, and what's the, the principle objective of, of this team? What is to standardize those tools, those processes?
It's like, uh, a center of excellence sometimes, uh, but the focus primary is to standardize tools and processes. Uh, this, this team is in charge of selecting the, the toolkit, the processes, the frameworks that are gonna be used, and then, um, start helping teams to develop their own, uh, for example, tools, scripts, templates, pipelines, uh, for their own teams. Uh, the, their work is about, uh, again, creating, uh, a platform where those resources are, and every team, it can be like a, a DevOps software engineer, a a quality engineer, uh, can go to that platform and, and use all those resources that they need for their squads, and also, uh, to personalize, to customize those resources and that process.
Uh, there's a lot of feedback from the teams that the platform engineering team can absorb and how, like, uh, this, uh, continuous process of evolving this platform and the resources they give. So at first, it, it can seem that it's like very hard. Uh, it's a lot of work to do.
There. Involve involves a lot of, uh, teams, a lot of processes and areas that have to, to agree. And, uh, in this moment, I would like to, to share, uh, our experience with, with my team, um, in fact, um, this is, this example is about starting small, right?
Uh, because it's the, the first steps are the, the harder to take and the ones that, um, are, is harder for us, uh, to step into, right? And so we started, as we can see, uh, in this small video we created for a Latin American client, this, uh, like platform may say, uh, we have a, a repo, uh, a repository of code where we started like this, uh, platform where all resources are, we have like templates for pipelines, for building blocks for Terraform. Uh, we have scripts, uh, we have guidelines for Docker, Kubernetes, uh, we have all that kind of things.
So every team can go to that repo, plan it and use it for their own good in the team that are working. Uh, again, the, the idea here is to start small, because there was nothing to start with. It was, we are a like, uh, in, uh, in, in zero.
So, uh, the idea of starting with a big platform with a, a big budget, it was, uh, like very hard, and it wasn't an option at that moment. So, uh, our idea was, well, let's start small. Let's start with something.
Um, and well, this is, was, uh, this was our, our first approach to, to this platform. We can see lot of code here that, that we have it. Uh, and this particularly is a tool that we start with sales in using, uh, it's called Infra Cost.
The idea of this tool that is recommended by the finops organization also, um, the idea is to know in every, deploy the cost of our infrastructure, right? Um, and what's the, the, the good point here. When we, um, integrate this tool into our pipeline of continuous integration and deployment, um, we are making every person in the team responsible for taking care of the infrastructure costs, right?
Uh, the idea is to not wait until the, the end of the process to gain, uh, this, uh, fins team and tell us, uh, hey, the, the costs are going up, what's happening? Uh, we don't want to wait until that point. So we took our, uh, our possibilities discussed with, with the whole team, uh, and our first approach was, okay, let's deploy this infra cost, uh, tool.
So in each pr, in each pull request that we made, uh, this pipeline runs previously just to check the differences in the, in, in the infrastructure that there will be with Docker, with Kubernetes, uh, it will see the, the manifest and the differences, and, uh, it will calculate the cost that we have and the differences between this new version that we're going to deploy and the previous one that it's been deployed. Uh, lastly, so when a person comes to this pool request can also see as a comment the differences in cost that infrastructure will have. So, uh, the team can see at firsthand the, the differences, the difference in, in cost, and decide in that very moment if it's going to accept the poor request or not, if the differences in cost is aligned with strategy with the objective in the sprint or not.
So in, in that very moment, in that point of time in the development process, we can know exactly how much will cost the infrastructure of our application. Uh, that's in a, a, like, uh, a great impact for us because, uh, there was, uh, like a small amount of people that, uh, that have that information that, uh, we, we worry in a, in a passive situation, uh, where we deployed our infrastructure or our application, and we started, uh, and we prayed, uh, not to have like, uh, uh, someone telling us that our, uh, RVs were to high. So we, we wanted to, to go from that passing situation to an active one and deploying, uh, this tool that it's, uh, open source.
Uh, it was, uh, a great, uh, a great, uh, start for us to, to change the dynamic we were having. Uh, and this was, this is a, a very good example of how we, we make the whole team responsible for the cost, and not just only like a few people in the organization. We have to work together as the rob say, you build it, you run it, you are responsible for your own, uh, for your own application.
So, uh, that's, uh, a very, a very good point, uh, at this, at this situation, right? Um, well, here is like the, a little demo as you have been watching, of how we, we deployed, uh, these changes with, with Docker, with Kubernetes, and well, uh, that's a, a, a small example of what we have done with this tool, um, as a bonus track also, um, maybe we, we always think as ops as just looking at the, at the bill and reduce costs and talking with people about what we are gonna, uh, reduce or, or to remove. Uh, and sometimes also with platform engineering, we just think about, uh, an internal development tool or platform.
Um, but there's space for a lot of ideas, a lot of practices that help us, uh, in this direction. One of, one of these tools is the chaos engineering practice or philosophy. Um, Kio Engineering is, is a discipline, uh, of running intentional experiments on, on our system or our application.
Uh, and, and the point is to inject, uh, precise and measure damages on our application infrastructure. The idea is to, uh, to measure the silence of our application and system and how it responds to different situations. For example, how can we connect this?
Will are talk about, uh, fine apps and platform engineering, for example. Um, one thing that we can experiment is the, the size of our resources in the cloud. Maybe some couple of resources are too big and they are too expensive.
Uh, and we have this doubt that what happens if we shrink it, we, we make it smaller. We change the, uh, the dimensions. Uh, what can happen if that, uh, thing is done Well, this practice, uh, can give us those, uh, experiment spaces and, um, those, uh, those metrics, for example, uh, we can test, uh, the, how the, the system responds, uh, how the, the request, the volume is managed with those different, uh, sizes.
Uh, so it connects directly between the chaos engineering experiments and this kind of, uh, intention of reducing cost, and always, as we said, uh, keeping in mind security quality and speed. Uh, so, uh, to sum up, it, it, it aims to, to help us improving our systems, uh, knowing with experiences with technical and objective experiences, how they respond, uh, with different types of failures. Uh, what happens if we inject, uh, latency in the network, how we reduce resources.
We, uh, we can, uh, remove some parts as of the application maybe that we, we don't need really, uh, and give us space to, to play, uh, a lot following like the method, right? And, uh, dipping into the, the cost thing. Uh, case engineering always, uh, keep us in mind the cost of down times, right?
Um, there's, uh, an equation, the formula to know the cost of our downtime, and it's basically composed of three parts. Uh, the ones that happens during the interruption of the service with the, the lost revenue and the productivity of the people because they are developing something new, something nice virtual for application, and they have to stop that work just to, uh, extinguish a fire, right? Uh, so that's, uh, another big loss we have for, for the business.
After the interruption, once we, once we are running again, we are productive. Uh, we have client charges, for example, for a regulations service level, agreements that are not fulfilled. We have another big cost there, and some others that are not ally measurable.
For example, the brand defamation. Uh, for example, if we are like a, a silver Monday, black Friday, uh, we enter 20 commerce to buy like that PlayStation five, I, I really want, uh, imagine that I will click on the buy button and it doesn't respond, and I have then a 4 0 4 error. I will be like, impressed, and I'm going to another eCommerce site, and I will buy it happily.
So next time, I'm not going to buy it to, to the, to the first site. Uh, I'm going to do the second one. So, uh, that's not easily vegetable, but that's, uh, a, a cost that implies the downtimes and not being prepared, or maybe focusing too much in the cost reduction and not thinking in quality security speed.
Uh, so, uh, just to, to have like a, a, a clear number in, in have, uh, for each hour of downtime, a typical company loses $3,000, $300,000 per hour. Amazon, for example, per hour losses like $13 million. So, uh, what's the point here again, uh, is always focusing in cost and in quality.
We have to balance those things to reduce RBLs, but always having in mind the application, right? So again, it has to be like a, every synchronized work between business and technology, some tools to practice chaos engineering, uh, some of them that are very famous like rambling. It's a, it's a great platform.
Uh, it has, uh, a lot of, uh, certifications, free certifications and trainings that I recommend. I have done them all. Uh, I like it very much.
Uh, I started with chaos engineering because of rumbling, so I recommend it. Uh, it's a paid, uh, software, but it really, um, gives value to that cost. Uh, on the other hand, we have, uh, some other tools that are open source lead most, for example, it's a great tool, uh, that I adore.
Um, some others are more famous like Case Monkey with Netflix. Um, and obviously, uh, cloud providers offers their own tools, uh, AWS Azure, uh, have their own. So those are also great options to start with chaos engineering.
But as I said before, uh, if you think that this is a, a good tool, uh, to start experimenting, uh, Grambling has a lot of documentation, very nicely, very well organized. So, uh, it's a good start to, to read, uh, how to start doing case engineering again, uh, from scratch with small steps. I know, uh, we are running out of time.
Uh, it's time to, to wrap up. So the, the key here is, uh, again, to align business with technology, right? Uh, it's very important, uh, to do that as we have seen with examples and the impact both on the applications and the business if we have downtimes, for example.
Um, and then lastly, uh, do we financially efficient? Uh, since the moment of the, the design and the development of the solution, uh, it's like a very good practice to ship left, uh, everything we can. So it's cheaper is easier for our valve to, to have that in line, um, just to, to have a, a nice closure.
Um, there's a, this, uh, this code I like very much that we're nothing certain, everything's possible. So, uh, when we have these, uh, new practices, new frameworks, and we don't, uh, we don't have clear or, or we're not certain of, of the next steps, uh, the good thing is everything's possible. Just, uh, you have to challenge yourself and challenge others, uh, to start practicing, to get into, uh, new tools and develop, uh, new ways, uh, to, to create, uh, value to, to the business, to the people.
Um, I hope this, this talk has been, uh, interesting for you. Uh, anyone who wants to keep talking about this can, uh, reach me in LinkedIn, in Instagram. There's the, the cure, uh, to, to have my, my contact.
So, uh, that's it. Uh, thank you very much. I hope you are enjoying the dev experience 2024 and I'll next time.