Real Use Cases in Platform Engineering at SKILup Days 2024
DevOps, SRE and infrastructure-as-code have become essential today as accelerators for projects and improving time to market by offering a lot of flexibility for developers. Throughout several projects where I have led the architecture or engineering management, I started to notice a gap between code, build and run phases at the infrastructure level. With the introduction of environment-as-a-service, I was able to reduce workload and configuration management problems by 30%.
Transcript
Uh, so, so thank you very much, uh, for, for inviting me like today, uh, for, uh, this interesting topics about EE and Rise platform engineer. Um, I'm Abdel, um, from France, Paris, and I'm DevOps Institute Ambassador. Uh, so today for our agenda, so we are going first to, to talk a little about, uh, the CE challenge.
And we, we see, um, like what's the difference between ECE and engineering challenge and platform engineering and what it's, um, going, trying to solve, and, uh, why there is a big rise of internal developer platforms today. Um, then real use case, uh, in, in platform engineering, like from my, my previous experience, uh, for famous platform tooling that we're using today. And, uh, what is the trend for platform engineering today by, in upcoming futures, the SAE ology engineering is practice, um, that focus on upline software engineering principle, uh, to maintain the reliability performance of complex system.
Uh, but however, SAE teams often encount challenge that, uh, hinder their effectiveness that we can see here our seven key principle, uh, of engineering team. So the first one is like embarrassing risk monitoring, distributed systems, eliminating TOIL and SAD release, engineering, ity and automation. And the three big sources of production stress today, uh, for system, uh, from the point of view, infrastructure and, uh, platform is the toy, the bed brain monitoring and immature incident handling procedure.
So I try to sum up, uh, uh, the most famous c and uh, today, uh, in the market. So for, for selecting the right metrics, I, I will adjust, uh, the, the, the most important one that I have, um, faced with, with the bold police. Um, so for selecting the right metrics, so the most important one is to eliminate that mono monitoring.
So it does mean yes, we should do collection, we should collect, uh, logs, we should collect metrics, but we should define, um, and the first, the ECA, so the, the indicator are before, uh, defining, uh, what is this, uh, what is this matrix, uh, trying to dispense to, and, uh, what is the, the most, uh, effectiveness story that we should use. So the, then the second one is to manage to, to manage incidents effectively. And here, uh, it's to promote blameless culture.
And the ensure continues reliability as the develop trinity principle. So here it is to define the SLO, the service level objectives. It doesn't mean we should already have the SLA, and from that you should, uh, implement the SLO in addition to and, and kios engineering.
Um, the, the, the, the, the, the fifth one is the resistance to change. So most, um, of the companies today try to have pilot program, uh, before doing any update, like, and deploying and take change, uh, for avoiding soil like ticket span, uh, mini change request that requires maybe, uh, support level four or level three and money applying, like small watch product change. So it's better to prioritize automation of a frequency, um, that's opinion to over complex to then establishing hit incident management.
So we try to make human find findable, and here, here we use iops like, um, and correlation between incidents and monitoring system and establishing escalation path. So if we go now to compare it with, uh, platform engineers, um, so we observe that DevOps here. Um, it's try, uh, like we have like the same product, um, like the same landscape of different products.
So we try like to, to, to, to, to bridge the gap between, uh, software development and operation and, uh, promote culture of collaborative collaboration, shared responsibility and interior, uh, software development life cycle. But the problem with this is that, um, for some pro product, we have specific, um, tools. We have specific technologies, uh, and this creates handoff and create a zero, create zero and mini, uh, manual task, mini, mini manual scripts.
And, uh, also, uh, a big challenge like for managing large scale system. So the solution engineer is to use, uh, work what monitoring system, uh, to improve reliability and effectiveness of service. And from the SE uh, the platform of, uh, platform engineer was right.
And this one, it's like, um, it's something that help us to build the f foundational infrastructure for values projects and, uh, and organization. Uh, uh, like for example, um, like for example, the Netflix, uh, for them, uh, the, the, the, the, the main difference between platform engineer DevOps and C3 O engineer is in, in their primary possibilities. So platform engineering here is responsibilities, creating infrastructure.
That's multiple business unit case that we can use. And SCE, the main responsibility still, like in platform tooling and from monitoring and scale application to ensure grid service for customer to achieve the high SLA and DevOps, still the focus and the live and application to production like quickly and efficiently by A and CD. For the, the, the today we have like, um, the more, uh, DevOps, uh, is, uh, stable and is high elu, um, um, evolution in society.
We, uh, remark that we try to analyze platforms and ize platforms, uh, and, uh, like, uh, at 50%, uh, of, uh, companies that implements DevOps use internal platform that we will see like platform sh like cycloid, et cetera. And why, because the needs come from, uh, this golden parts, uh, for many futures that we need between infrastructure and code, like adding governments in demand and changing configurations depending in the project, adding services and dependencies and rolling back debugging and rolling back also for deployments and spinning up new environments, refactor adding, changing error source configuration and enforcing airbag. Today, developer should be able to deploy.
Set up internal platform teams will be the internal, uh, developer platform. So developer can pick the right level of abstraction to run their application and service, depending in their preference. For example, uh, maybe they like mis messaging, uh, around helm charts, yamo, terraform with module and frauds end doesn't care.
Like if the application is running in the EQS in Kuber. And, and this what is fantastic. So we can just like, uh, have self service environment that can fully, uh, provision it with everything that they need to deploy and test their code without warring and without be considered where is the run and, uh, where their application are hosted and serviced and to, uh, like, uh, from, uh, from pounds of view of all the lifecycle, because we say you big it, you run it.
But this is true for divorce. And this why we have like the philosophy of the ICD today, but with, uh, engineering platforms, we don't need, like if we call, if we build it to run it, we just, we can code it and we have all, um, the landscape is, uh, un studied like here for building, testing, release deployment, monitoring and operate, and all the things are ized and normalized for all the projects. Um, so here we can see the principle of platform engineers so that it's clear mission and role.
So here, here we have like platform team and this platform teams. Um, so it's in, in this color pla it's the platform team, and we have stream aligning team and this stream align team. It does mean that we have like future team that is, uh, that's, that's considered as projects or product and enable enabling team like, uh, for common problems like networks, like for VPC, et cetera, and complicated subsystems like, like enabled for the, for, for, for like enabler project.
So we have here, so as principle of platform engineer, clear mission and role, and we should treat our platform as products and not products, and we should glue and glue here. It's, uh, like to have, uh, different tool chain, uh, to satisfy all our demands of code and hosting, uh, and the same topology, and also don't reinvent the wheel and try to use all tools. That is open source.
Big four real use cases that I face it, like in my experience is service catalog. Like, um, in enterprise architecture, like we have definition of technology status, and we have also the definition of mini projects. And in this project, we want to see them, like in GitLab's, we want to know in the portfolio management for gl what is them state today in the project, if they are in development, uh, what it is, uh, as INGO tasks, uh, what is the DORA metrics, for example?
Uh, what's the number of tickets, um, that we must open, for example, for, for deployed projects and production, et cetera. So service catalog, most of the company use it before SEC Lloyd and backstage for the, so to integrate with it, its service now like with HSM and enter enterprise architecture as artifacts. The second one comes as government as service.
It's like, uh, UCS for Amazon and plan service like in Azure. So the, the principle is, uh, uh, use it like, um, in popular like platforms, like platform is H and it's Behave environment as service like, like platform as service. But so, uh, I, we used this one in same SEGM, uh, to offer vari, uh, uh, um, values technical stacks that can serve, uh, standardized, uh, technology stacks for different projects like for e-commerce, uh, for c et cetera.
But the same government will will, when it's Java, the same ments in no Gs, and we don't reinvent the wheel. So we, we save a lot of time like by, by, by, by using, uh, government as service. So the third one, the third use case is to use standardized tool chain.
Uh, this one we find it like an Azure, an up service plan, like, or ECS or a KS, uh, and they provide runtime capability of scaling. But the problem is that here that we should, uh, like glue again, uh, to have like, uh, the same tool chain, uh, across different projects. And the final one is sandbox.
So we construct hybrid tool chain from CACD with observability, with management, with cross cutting concern, with governance, and with this sandbox, it comes like, um, in the organizational level, and we can do it like in platform in a pass or like in public pass. And we try to limit the services, we limit the access, we limit the policies, we limit firewalls and everythings and the number of service that we can use, and it becomes sandbox that can be used for every developer. And this, uh, can join like the summarize tool chain that try a platform engineer trying to offer.
So let, let's see, like the famous, um, like, like try to have, yes, uh, a view for the tooling platform tooling landscape. Uh, so we find here the development portal, we have like c we have backstage, and then we have version of controlling. And all this is gathered with, uh, with, with, with, with, um, the editor.
And we have application for ions, code and platforms source coding. And this one push the, the, the code like for say ICD for the image registry and for the orchestration platform orchestration like humanity, like rancher, it does the deployments. And we have like, uh, horizontally, we have all tools like for vaults, for security management, for security, for observability, for data matching data, masking data, encryption data and invest and process networking services and computing.
So what it's try to, to offer here to offer the same landscape for all the projects. Um, and for real example, sorry for the skill. We, we, we have four.
We have a platform we have and backstage. So, um, it, it's, it give us a view of all, um, projects that we have and every project is tied to environment. And this environment, we can see it as services here, and we can edit it in this template.
So we can have like infrastructure as code, and we can have template as code, security as code and everything. Uh, the second conent is check. Uh, SoTech is developed like around sandbox in Amazon, and it try to bring, uh, by them, uh, configuration code, uh, that proper SoTech.
And it does, uh, everything, uh, related to scoring the applications in terms of, uh, quality code of structures, uh, of security, et cetera. And by Terraforms, uh, it's, it's, it's, um, triggered the pipeline. Uh, that's being the image that's go directly to the registry.
And from NewTech, it's try to join, uh, the, the two phases between, uh, sorry, between, uh, the CI and cd, and then we use what we call it human tech deploy. It's a manifest for Yemen, uh, manifest. And this one, um, so, uh, like push in returns like in, in Amazon, EKS, um, and do, uh, provision for, for resources like data, for my skill, for, for everythings for Amazon, SKS, um, and, and all networking services.
And in addition to it, it use RCP Vault and, uh, Amazon CloudWatch and others for observability. Uh, so here we have another example is the same, like, uh, it's, it's like between, uh, if we compare it, cycl is something that's, um, very developed, uh, than backstage because backstage it's more for service capital and it is also, uh, implemented all the futures, uh, that humanity has today. So we have here all pipelines for projects that's normalize it.
We have infrastructure view that we, uh, uh, that we will see in the second screen, and we can see all the dependencies between, um, source code versions for all the projects and them states. So here like alpha view, so we can see it's like a blueprint, uh, for infrastructure as service for Amazon. And it show like different, uh, data flow and different services, um, that you use like from the CI to CD for Backstage.
And it use, um, like the same, it's like the same as it used also. So between, uh, services that we use them states, uh, them up futures, that's also organized and every data endpoint. Um, so this is more focused, uh, to, to have like service catalog for all services that we have in enterprise architects.
So if we want to talk about the upcoming futures today, um, so we don't have like, um, enough fine ops functions that is implemented platform engineer. We don't have green ops, but uh, maybe we'll have it, uh, like, uh, and the, and the next, uh, release of cycloid and iops also for all, uh, operation on and, uh, optimization automation, uh, metric as called like Dora, like accelerator. Uh, I mean more about indicator as code, like to implement SLA and SLO as, uh, declarative syntax, code security as code.
And maybe we can have like generative, ie. Uh, say, uh, for product engineering what is, for example, uh, the most used don't play it or less used cetera to have direct question. So this is for referral.
This, that's how we use, um, in this presentation. And I thank you lot.