SREs and Platform Engineering at SKILup Days 2024
Generative AI is an enabler of hyper-productivity across requirement elicitation, code generation, testing case/script/data generation, generation of deployment and IaaS artifacts. It also supports additional use cases like self-healing of production issues. The predictive capabilities of generative AI combined with deterministic, platform-centric software engineering approaches are expected to revolutionise SRE’s automation and platform engineering capabilities.
Transcript
Hello. Good morning, good afternoon, good evening from Barba. You have tuned in.
I'm very happy to be here with you at the Skill Labs days for the DevOps Institute. Um, my name is Shiva GaN. I've been an ambassador of the DevOps Institute for a number of years, and here I am to be talking about the site liability engineering and the rise of a platform engineering.
I've been very fortunate to have authored to, to have been the lead author for the site reliability engineering practitioner course for the DevOps Institute and the observability course for the DevOps Institute. So today, uh, I will be sharing my thoughts on, uh, platform engineering. Why is it required and how does it really help the organization to scale out on SRE, um, practices?
Now, this is, um, I mean, I'm sure we've spoken a lot about cyber engineering. This is a small encapsulation of what SE is all about. So SRE is about resilience.
It's about system availability. It is to make sure that the business service risks are mitigated, and we do it in a very, very systematic manner, which is to determine the service level objectives that are pertinent to that particular business service. And then we, we bring it down, we express in, uh, what we call the SLI, so the service level indicators and all downstream systems services products, which, which contribute to this single business service has to meet the, um, you know, the, the, the service level objectives, right?
Um, so this could be through, uh, a number of, uh, indicators like, you know, the error rates, the availability, and the durability itself. Now, nothing is given a hundred percent of availability or, uh, reliability. So systems will fail when systems fail.
How do we take the learnings back to further improve on the processes, people, the technologies? How do we really make it, um, a very, you know, through the process of what we call the blameless postmortem, how do we bring in more and more learnings to make the entire ecosystem, uh, better? Right?
Now, this is the, this is about site level engineering. More what do you mean by platform engineering and why do you really need platform engineering? Now, the, the whole concept is about when we are, when we look at deployment of services, we go through what we call the operational readiness review, right?
This operational readiness review is primarily to uncover operational and resilience threats. And this is about, you know, looking at it from the lens of architecture, looking at it from the lens of release, uh, management, um, quality of the product itself. And then, you know, if there is an incident or an event, how do you really, you know, make it more efficient?
And how can we make sure that this, the, the entire learnings, um, you know, be it in any, uh, any level, be it an infrastructure or, you know, at a system architecture, how do we really take it back into the organization when you have, um, hundreds of teams, when we are talking about thousands of services, right? How do we really scale it out and how do we take the best practices across the board? How do we avoid patterns of failure?
And we leverage these lessons for operating at scale. Uh, this, this is, uh, uh, why, uh, we talk about, uh, platform engineering, and this is why, you know, you're seeing such a uptick on the rise of platform engineering, but it's, it's worthy to, to understand, uh, to remember that it doesn't mean that the platform engineering is about actually becoming a bottleneck, right? So the, the entire premise is that, uh, the, the individual, uh, application teams or service teams will, uh, will actually do, will be self-sufficient, you know, they're enabled to actually self assist the risks and mitigate the risks, but the platform engineering team will give the operational best practices and will make sure that there is reuse of these best practices across the board.
Okay? So how do you scale the operational readiness review best practices across the organization? Platform engineering is one of the ways we look at it.
How do we make sure that, you know, there is well architected, uh, uh, uh, reviews? How do we make sure that the, you know, the end-to-end review of all the components for the design patterns and for the architecture, and make sure that, you know, when you have third party providers, how do you make sure that these risks are managed, um, across the board through guidelines and through best practices? Yeah, so this is where, um, uh, platform engineering team comes in.
Now, the, uh, the biggest, uh, uh, advantage is when you start small. Of course, you know, you, you may, you may feel that, you know, the, the, the individual teams are, um, are well-versed to make sure that they, uh, they take care of, you know, the, the reliability and operational readiness aspects. But as you scale out, and as you have, uh, thousands of services being consumed by hundreds of applications, this is when, um, it becomes really very difficult to make sure that, you know, these teams understand the way they should operate with each other, the way they consume the services.
And there is a certain level of, um, organizational, um, you know, uniformity across the team. So this is where, at form engineering plays a very big role. Um, so finally, it is about the, you know, delivering value to the transformation itself to the customer, making sure that, you know, the, the operations are optimized and the people are truly empowered.
So when you have these individual product and service teams, how are they truly empowered? And they made self-sufficient through a digital platform to transform the product. Now, this is Gartner.
So Gartner says 90% of the enterprises, if they're trying to fail using, you know, traditional initiatives such as DevOps or DevSecOps, um, uh, uh, and if they do not adopt self-service, you know, they're going to fail. Now, the reason why, um, you know, you, you see such a big figure, 90% of the enterprises is even if you are a small organization, as you grow, uh, there, there is always a question of how much should we abstract away from our builders, or how much should we abstract into the platform so that we build inefficiencies and we don't keep building the same thing, or, you know, um, not reusing what is built by one application team or one service team. And you really have, you know, you're building this from scratch in other areas of the organization.
Now, why do you need platform engineering? So primarily to increase the developer productivity, to make sure that security and well architected infrastructure is there across the board, and also to lower cost and improve efficiencies to make sure system engineering principles and practices are followed. Uh, risks are mitigated and to preserve teams flexibility and agility, right?
So there are multiple approaches. One approach is you build it, you run it, course. This is the philosophy of DevOps.
Um, and then, you know, everything all, you can select the tools, you can select your SDKs, you can do whatever you want, your completely independent as the second, um, other extremists. Um, no, everything is going to be decided by platform, and all the applications will run on the platform, and we will run together. We decide what we, what I mean is the platform engineering team will decide about what kind of tools, technologies, platform should be used.
They abstract it depending on what choice you want, for example, serverless, or you want a Kubernetes cluster. You want infrastructure as code. You want deploy into development environment.
You want to deploy into production environment. You want, um, you know, look into, uh, uh, cany, um, deployment. No, there are guidelines.
There are, um, reusable packages. There are run books that you can use, which is provided by the platform team. And then you have the hybrid model, which is what we see, um, again, based on Gartner's way of how it does the, you know, the pace technology certainly do have some of the, uh, applications that may not change, uh, or may not, you know, deploy releases as frequently.
Uh, so depending on how your organization is. Um, so you, so you have systems of record, systems of innovation, systems of differentiation based on that you can actually adopt, uh, uh, an approach where you would want to implement platform engineering, or you leave it to the, uh, in individual, uh, plan, you know, application and, uh, product teams too to deploy, uh, and have full control. Now, platform engineering is to solve scalability issues.
It is to reduce fragmentation and to make sure that we reduce the waste. And as we increase the automation, and we, we, um, make sure that, you know, there is a lot of reuse from across the different teams, um, the platform team will, uh, uh, will, uh, will ensure that there is uniformity in the way. Um, the, the overall processes that can be abstracted out from the building process is managed at an organization there.
Now, I'll give you some examples here. So, so for example, right? If you really want to use some templates, and you wanna have it say, multi-cloud, you know, so you want to build something which is cloud agnostic, or if you want to build the landing zone for a particular cloud, um, uh, hyperscaler, you wanna use, uh, infrastructure code templates, DevOps, tool, chain code, and reusable patterns.
Um, we want to know how to govern all of this. Uh, we want to build security and shift, lift and build the automation into the DevSecOps pipeline, but we want to do this in a, in a very, um, scalable and uniform manner. So, so simply put, platform engineering is about providing simple, secure, and scalable products, which will enable rapid, repeatable, high quality, uh, deployment.
So this could be polyguard shared libraries. Um, so, so primarily the, the premise is to, to enable EPS and to provision and ecosystem that will enable self service for the application things and not to become important, uh, and to embed governance and controls and standards within the process itself. But the, but the, the, the platform engineering team is the team that will abstract, uh, all the undifferentiated heavy lifting and will provide them the value added its services, which can be consumed by the developer and improve their, uh, productivity, right?
So an example here is, you know, source code management tool or secrets management tool providing run books and playbooks, um, coming up with observability and tracing tools so that when these, when these services talk to each other, you have the distributed tracing that happens from across, uh, the board, right? So when you, when you have a business service, which is built from, let's say, five or 10 different services, and these different services are managed by, let's say different teams, and if these teams are going to be using different distributed tracing tools, then you know how complex it's going to be, right? They're not going to be handshaking and talking to each other in a seamless manner.
So this is, um, uh, this is why a platform engineering is extremely important. Now, if you see the core platform, the core platform, for example, if you want to provide, let's say, in a serverless, uh, run times, and you want the developer to just concentrate on the code, but really abstracting them from, you know, infrastructure, managing the infrastructure, how do you really provide that service? How do you provide a Kubernetes runtime that can be actually easily consumed by the, uh, developers?
How can you actually abstract them and give them CIC pipelines that can be just, uh, consumed? Um, we spoke about observability tools, source code repositories. How can you give a dev a portal that will allow them to have self-service, use the workflows, consume the APIs, and build, um, you know, UX enough, extremely fast, extremely flexible, and, and having, uh, the efficiencies that would not be possible if the platform is not there.
This is an example. So for example, you know, you want a, to build any service application you create by consuming some services. And these are templates which talk about, you know, uh, the tiered architecture.
How do you really consume a blueprint and build a web application, for example? Or if you wanna build a serverless application, a headless app, how do I really run code without, uh, you know, worrying about, uh, the under underlying, uh, architecture, uh, how can I deploy, uh, into dev development? How can I consume, uh, software components from a library?
So it's about creating new assets from ready-made blueprints, um, having a shared platform, uh, using code repositories, using easily configurable CICD pipelines, and do it in two or three minutes rather than probably, you know, two or three days or even up to a week, right? So this is typically how platform engineering helps. Platform engineering comes into play every, at every point, right?
So be it at the CI phase or be it at the, uh, continuous deployment phase. If you are talking about, you know, deploying to cannery, how do you really release it to UAT? What kind of metrics do you have to, um, monitor?
How do you actually configure the alerts so that, um, there is uniformity when we look across the board from across, you know, across multiple services that will constitute for, uh, for, uh, you know, the enablement of a single business service right now, uh, the DevSecOps pipeline, how do we shift left? What kind of security tools are we going to be using? Uh, how do we make sure that the artifacts are stored in a secure manner?
Uh, what kind of, uh, source code repositories, for example, you know, we do not have, um, multiple far flung, um, so source code repositories, everything, uh, is kept in a, in a manner that, um, there is reuse, uh, between the teams, right? So this is, um, where, um, observability is also a key component of the platform. Make sure that the monitoring, logging, and tracing and visualization is all built to scale in a uniform manner.
Now, this is just a simple team, uh, implementation. So here we can see that, you know, you have the platform team. The platform team can, can actually input s so the SREs are embedded into the different product teams, but the SREs will have, uh, um, either a direct connector, you know, dotted line connection to the platform team so that, um, the best practices from the SREs from the different products are cross pollinated through the platform three.
So for example, if this SRE in the product, which is product revenue and sales, is the, um, is the vertical, if there are some good practices, or if there are some good tools that are being brought in for this particular product line that comes into the product team and gets implemented into, let's say the group services abroad, ground operations, or different, um, verticals, right? Um, so trust allows, the platform team allows the automation to be taken to a higher level. It allows DevSecOps to be taken to a higher level.
Uh, it helps you to build frictionless, standardized, um, site reliability engineering practices. It also helps you to adopt, measure, and refine these practices, um, and to have a product mindset for the platform itself, which means that if there are changes that are happening to the components that you are releasing, the, the guidelines, the governance, um, the reusable components. Now, this also will have, uh, this also will follow up a product approach so that the, the teams that are using the platform are also well versed in actually consuming the product itself.
Um, this is, um, uh, so this is about platform engineering and about the rise of platform engineering. If you have any questions, um, please, uh, chime in. Please feel free to, um, actually drop the questions.
I will be there in the chat, uh, room. Uh, I'm very happy to take your questions. Thank you so much.