Platform Engineering Meets Service Mesh: an SRE Love Story | Cloud Native Now 2023
Platform Engineering is the new black and we posit that service mesh is the perfect complement and an indispensable sharp tool in every SRE’ tool belt. Albeit been around for a few years, service meshes are now blossoming into a de-facto standard for providing much needed features to cloud-native platforms, such as zero-trust security enforcement, observability, traffic shaping and control over the application networking layer(s). This demo-heavy presentation will walk you thru the rationale behind adopting a service mesh and what it can do in practice to alleviate the pain of running at-scale platforms to deliver cloud-native self-service infrastructure to developers. We will focus on Istio (the most popular service mesh) but the learning extends to any other service mesh implementation; we will provide a workshop-style repository to follow offline.
Transcript
Welcome to my session. It's about two of my favorite topics, platform engineering and service. Meh.
Uh, the two things, uh, platform engineering is very, uh, very hype today. Everybody's talking about it. Service mesh has been around for a while, and I personally love it, and I think it's gonna be, uh, great.
If you stick with me, you're gonna learn what it is and why. Combined with Pro Platform engineering gives you, uh, the SRE superpower, and there is this perfect match between the two things and gives you your SRE team perfect, um, great tools to, to, to succeed. So, I'm Alessandro.
I'm a fast run advocate at Solo. We, we are a company that really, really, uh, invested in service mesh. Um, I do developer relations.
I'm based in Amsterdam. I'm a, I'm ba for the Cloud Navy Computer Foundation, which means that I'm really, really in love with all things cloud native. Uh, I run the local meetup, so I think I have, uh, some skin in the game to explain what is cloud native and why, why it's so relevant today.
So, I wanna start with some definitions. So we all know what we're talking about. You may differ.
I like to hear your opinion. Please reach out to me on Twitter or LinkedIn because I, these things are quite difficult to define. There's a lot of, uh, confusion.
So let's start with some definitions. So what is platform engineering, right? So, so platform engineer engineering, it's a complex, uh, topic.
There is a, actually a, a cloud computer foundation platform working group that just release a white paper that explained very well what this platform engineering. So what platform engineer is, in my opinion and in opinion of many others, is a self-service platform. So it's a self-service tool, and it needs to be, um, uh, coherent.
So it does need to have a self-service discovery onboarding for, uh, developers. Um, it needs to be integrated, so it's to be one single tool for, for everything. So not just a collection of, uh, these joint tools, but one coherent place where to find all the services, all the applications, and where developers can interact and can discover new and, and onboard new applications, um, has to be treated as a product.
So there is a platform engineering team that take care of it. Uh, there are stakeholders, there are product, uh, manager, there's a, um, the should be treated as a, as another product in your organization. And let's to be focused on developer experience, which is the magic world today, right?
So we all know that, uh, companies, organization lost, lose a lot of efforts and money and time in, uh, onboarding and not, not really exploiting the productivity for, uh, for developers. So a platform, an internal developer platform, d p as we call it, helps you earn nesting. This, this hidden then this, uh, untapped potential for developers.
So it's really important to, this is becoming quickly the, um, the defacto approach to improve the developer, uh, productivity and to just provide a better experience for your developers. So what are the values? Uh, reduce cony loads of developers.
So the person don't need to find the same information over different tools in the, the good of days of having a weekend system. And, uh, the source control controller, another system. These are, this introduced friction, right?
So you have to change and switch context between different tools. So it's not exactly, um, what you want. So reducing the connectivity load on the purpose is the, is is one of the values of, of internal development platform is a consistent review of value.
So all your application and services are now under the same, uh, the same tooling, the same platform. So it's easier to have a more consistency between them. So it's, uh, it gives you and your developers more consistency, improve sharing and reusing knowledge, which means no more duplication.
That's a, that's a huge waste of, uh, time and, uh, and money eventually repeating and, and, uh, creating the same content all over again. But if you have a single co platform, you should be able to, to reuse the knowledge and improve the, the collaboration between your teams. And of course, you centralize and automate.
So the core of every developer, prof developer platform is automation. So automation is the, it's been around, it is been introduced. It's one of the core talent of DevOps, of course.
That's why there's this very big overlap between platform engineering and DevOps. There's all, uh, debate. What is what and why They, uh, are they competing or are they helping each other?
In my opinion, it's just is the same, uh, definitely the same people, uh, people doing DevOps are probably also the people doing platform engineering. They just, it just overlap a lot. So automation is big and needs to be, um, in, in your internal develop performance from day one.
And you, you need to use it to centralize, automate security and observability from the beginning. And the old goal is developer productivity. So make sure that there is no blockage, there is no waste, uh, between the idea that comes in the brain of your developers and the execution.
So producing all the time to value, time to market. Um, so very important, platform engineer is a team sport. So you don't do it for yourself as a, as a platform engineer.
You don't work. Uh, you work on a product that serves, uh, a purpose or you really need to think in terms of, um, anthropologist. It's a great book.
Uh, I I encourage you to, to read it, of course, and there's a lot in there now. Site reliability, engineering, it's something that is not new. Uh, it's been around for a while.
There's this famous book from, uh, from Google. I remember when he came out, everybody was, uh, uh, rushing to download. It was free.
Uh, and it's still free. Um, and it kind of gave us a glimpse into how Google was, uh, providing, uh, reliable services to, to consumers and to businesses. So, so s r e is this practice.
It's, um, it's, I I compare it to, uh, a practical philosophy. Like, uh, so if you, I'm very big on seism, and so, so maybe you know what I'm talking about. It's, it's not just a culture is how you implement the culture in practice, right?
So it's a practical, um, uh, exercise of principles. S e principles, we can think about it. So, uh, these are some of it, of course, but, uh, again, like these things can be adopted various degrees.
So there's a book, of course, you can read everything in the book, but then you have to translate in your own culture, in your own organization. So it's not always black and white. Uh, and there is no institute that's gonna certify that your s e but it's something that you cannot adopt and, uh, use it for your advantage.
So some of the core values, again, automation, see the, the overlap with platform engineering. Uh, adopting S3 practices helps you adopting platform engineering as a, as a, as a team. Automation is everything.
So never, uh, do the same thing twice manually. Uh, even one time, it, maybe it's already time for, for automation, uptime is key. So the site data engineering started the Google, of course, and there uptime is absolutely the most important, uh, metric.
But it doesn't mean that for you, uptime means a lot. So if you're on a web shop, if you're running, um, uh, some, some online, uh, merchant based, uh, shop, then of course our time is, is king. Cause you're not up, you're not selling.
But for you, for your organization, that can be something else. It doesn't mean to be, it doesn't need to be one single thing like a time, but can be something else. But one, one principle is that you gotta find your own, uh, the, the value, the, the metric that brings value to your organization.
So that's, uh, and then that make, make that the core of your s e practices sharing, scaring, of course, nothing is done if it's not, nothing can be called finished or done until you share. Uh, this is something I, a lesson learned when I was at Microsoft. We really could not call a project finished until was shared.
There was a repo. Everybody was aware that we, we had some sharing session about the project. Uh, we with our colleagues in our site.
So this was one of the, this is one of the, uh, core practi core principles, sre and of course, be very rigorous in applying those principles. It's one of the, it it's something that I, it is part of the culture of sre, right? So, uh, unwavering un uncompromised, uh, application of the principle is, is what you sh should lead your, um, s e uh, practices.
One thing that is very important for any SRE team is metrics, of course, as we say. So you need to identify what metrics for you. One thing is everybody's talking about, of course, this DORA metrics or the, the DevOps research and assessment, uh, uh, institute led by nickel for, and of course, it's great book Excel.
You should read it already. If not, uh, this is not free, but it's very well spent money. Some of them are deployment frequency.
So it's part of the efficiency, uh, domain, uh, lead time for change. So these are all metrics that will help you defining your s e practices or something that you have to look at when you want to implement this. Uh, this idea.
So I just put there, we're not gonna talk about it in extensively, uh, but they're part of every SRE practice. Now, let's move focus to the cloud native application. So what happened in the last, uh, just to recap the last, uh, seven, eight years or even more, something really changed in how we we do things, right?
So how we deploy application, how we think about application, how we build them, deploy, uh, uh, monitor them, observe. So they, what happened was, is the cloud native idea, right? So cloud native is, uh, and I'm, I'm, I love the cloud native computer foundation.
There is a definition there, of course, uh, of what this cloud native, but it's basically a set of principles that, uh, that dictates how you build applications. So definitely we are talking about microservices, we're talking about containers, talking about a API based, um, communication, of course, all this in the DevOps, uh, salsa, so to say. Now, NCF F is of course, all about tools, right?
So it's the, the cloud native company Foundation is a big umbrella, big house where a lot of, um, projects are living, living and maintained by the, the foundation. So these are joke everybody talks about, of course, is this, uh, big, huge landscape of tooling, which are confusing, of course, but it's more of a meme of a joke between practitioners. We say, ah, just, uh, go to the C landscape.
And everybody's super, uh, super scared by all this, uh, plateau of tooling. In fact, you don't need all this, right? You need to find your own way.
That's why it's, uh, that's why I think, I think the, the, the big advantage of adopting CloudNet is that you can pick and choose your own story, right? So you, you can make your own, um, your own path. And that's, that's been important.
So, and one big part of, of cloud native computing foundation, of course, the first project was on boarded, is Kubernetes. Kubernetes is a platform to build platforms. This is the first time I heard this from, uh, from the late Danco.
It was actually quoting Kelsey Hightower. Um, what does it mean? So Kubernetes is just a substrate, right?
So it doesn't do much by itself, but it's a good building block for, for something else that you build on top of it, and that's where the value is. So this is resonates with platform engineering, right? So the platform you're building is built off tools, other platforms put it together, and Kubernetes is probably one of the best, uh, choice you can make to build, uh, coherent, uh, and useful platform for your, uh, for your developers.
Now, saying that, so part of the cognitive, cognitive applications are, um, uh, so we, we realized that in the last, you know, 7, 8, 10 years, um, lots of shift from monoliths to, um, to microservices, right? So we know we used to have this big, big Java blobs and jars, uh, running everything inside. And, uh, and then the 12 factor up and the, uh, and the move and the need for, for more agile, more, um, more flexible architecture, introduce microservices.
So microservice are the, the composition of, um, of mon monoliths, which is great because now you can write every microservice in, in tri zone, uh, with his own language. Uh, there is a clear interface between microservices through APIs, but the downside is explosion of complexity, right? So instead of a monolith talking to a database, now you have dozens or hundreds of microservices all having their own databases, all having their own interfaces.
It's not easy to manage. So, and that's why cloud native, cloud native application are a bit more complicated, but they give you all the advantages of microservices and so on. And so service mesh is one way to tame this complexity.
So what is a service mesh, if you service me, is really a mesh, is a topology, right? So if you're old enough to remember, we used to have these mesh networks between computers, literally connecting them through a token ring or some other, uh, networks. But really, um, it's a way to remediate for the, the famous eight fallacies of the serial computers on Wikipedia.
Uh, all these are not true, of course. Uh, so how to tame this complex, how to, um, manage the networking in microservices, you need a service mesh. Service mesh is really, um, a tool that another layer that you put on top of a platform, most likely it's Kubernetes, of course, but you can also use it to connect workloads beyond Kubernetes, beyond containers.
Uh, so you can connect virtual machines, you can connect even lamba, Lambdas, serverless, it doesn't matter. So, uh, service mesh is there to help you these four, uh, pillars, right? So to connect, to secure, to control, to observe your workloads.
So it's basically in, its very, uh, in its canonical form, like what we know is a service me for the last four or five years is issue, by the way. It's, um, five years old, uh, six years old, five years old. Uh, so what does, what does it mean to have a, to implement is issue or sales mesh?
It's really simple. You in if to help to deploy a proxy next to your application. So, and these proxies make a mesh, make a, um, a network of proxies or controlled and managed by a single central, uh, control plane.
And now you can do interesting things with it. So you can control what goes in and out of your mesh. You can control what kind of communication you have between the services.
You can do, uh, you can implement very easily principle of zero trust, security, like, uh, mutual tls and so on and so forth. There are many, many features of that you get when you, um, when you get, uh, a service me. So Iseo extremely popular now, six years old.
In fact, um, a CCF project now, uh, and even better, uh, soon to be, uh, to be graduated. And, and so one of the first thing you can observe, you, you can see when you adopt service meshes, you get this amazing visibility is like everything pops up and everything is light up, uh, from, from the darkness of your, uh, of your services, you can immediately visualize with talking to what, what kind of, uh, request per second you, you get, uh, is there, is the connection secure? See all the little locks?
So there's, that means that there's mutual pls between services. So these are one of the big advantages, and I, I believe this is a great addition to every cluster. And being Kubernetes is one of the, uh, best option for building platform engineering.
I think service mesh is really what you want to use. So just a little intermed, of course. So, uh, service meh used to be classically, uh, based on sidecar.
So a sidecar is, um, a little container next to your application. The con, the sidecar proxy will manage the traffic for you. So manage, of course, by the control plane.
Um, that's great. Uh, but there is a better way, uh, at least we believe, uh, it's a solo, but also in the history community, we believe there's a, there's an alternative way, which also can, uh, help you reducing, uh, complexity, simplify operations, save money, really, because you, you go from a number of, uh, proxies equal to the number of pods, uh, to the number of, uh, of, of your application pods or, or workloads to one proxy per node. So it's, uh, it is definitely a better way.
Uh, there is, uh, of course, this is still half, uh, beta, really, uh, but it's coming out very, very quickly. So, so this will be the next thing for service mesh. And I wanted to, to mention it because it's really something that we, we invest a lot and we really believe we will help, um, adopt more of service mesh.
So let's tie it all together. Tell the story of, uh, what I'm a platform engineer. Uh, I want to convince you to install meh a service machine in your classes.
I always say I'm on a mission to install a service machine in every cluster that I, that I see, because everything is all tars all the way down, right? So every layer is important. Every layer needs to be, uh, taken care of.
And service mesh gives you this superpowers, this, uh, x-ray vision inside your cluster. So why, how these things work together. So your platform engineering, you implement service mesh in your clusters because your application runs.
They are cloud native application, and they do run on Kubernetes, implement that you adopt SRE principles, and then you use, uh, the service, the s r e principles, uh, they, they combine very well with service mesh because you can use it for observability security. So I think this is, this is Avir loop, uh, everything helps each other to, to provide the almost perfect self-service portals and self-service, uh, um, platform for your developers. This is my take at the well architected internal develop platform.
Of course, you can, you can implement this in every cloud or on premise, but the, the basic building blocks, of course, uh, gi uh, GI could be GI doesn't matter, single source of truth for your application definition. So for your service definition. So, uh, this is the common backstage.
You store all the definition in Git, um, Kubernetes, of course, to run these applications makes sense. It's a very, um, it's a orchestration, platform oriented microservices. And today's, you probably are using microservice already, use Argo cd.
Uh, again, like it's just one opinion. You could use the other, uh, GI tops tools, but definitely we see the value of GI tops. We, I haven't touch about, I didn't, I didn't talk about it, but definitely is a way to deploy application more efficiently, more rapidly into your cluster backstage.
Again, this is one, just one example, but it's the core, the what ties in together everything else. And of course, service mesh for observability and, and, uh, control and manageability of the platform itself. Um, and then your user interface through the, um, through a portal, right?
So this is the, the core tenant of platform engineering. You have to go through one portal, could be a Google, could be a C L I, but it's something that your developer experience really benefits from when they have a single place where, um, uh, where they can interact and little plug. So we do have a product called Group Portal.
It's a compliment to backstage. In fact, we just released, um, um, a plugin for Backstage that that integrates group portal in backstage. So Group Portal is really an API management, uh, uh, tool that can help you, um, aggregate and, uh, and, uh, present those API to your, to your developers, to your users in a much better way than than just plain, uh, plain open AI spec, right?
So, so we just read this and it's pretty, pretty cool. So, and that's to the end of my talk. So this first service message really ties my platform together because I think it, it really helps you blending together the services, uh, having a coherent single way to, to, uh, to manage the services, having a consistent security level, TLS certificate, rotation, even external application you can do through, uh, service manage.
So it really makes, uh, makes this thing work together. Well, some tips to build a platform, because I, I want to leave you with some, uh, free tips for me. Of course.
Uh, and I, I got them from, from other people and from my own experience, stay humble. Start, start small and focus on the human experience, right? So we are, you are not building a platform for yourself or for some from some AI bot, but you are doing it for, for people, right?
So always keep in mind, what if I was a developer and I was using my, the platform I'm building? So, uh, this is a very common thing to do. Just take the very simple things, do them first, uh, get some, uh, buy-in, get some excitement about a platform, and let people, uh, come to you, get security observability right from the start.
That's no, no compromise on this, especially security. So there's a, we can talk about it another time, but there's a lot of, uh, uh, of course there's a lot of, uh, threat actors out there. They're coming to get you, you, if you're not smart enough to implement the very basic law, the very basic, uh, baselines for security.
And I'm sure somebody else who's gonna talk about like a, uh, secure supply chain, container scanning and so on, and observability is the core of every platform. That is, they can call itself a platform. Focus on your stakeholders.
These are the developers and plan for adoption. So this could be quite, you know, it could be quite surprised people start to get on board. So please.
So think about scaling your, your efforts and your, your platform quite rapidly and think about people again. So, and please, please use meh. You can use issue.
I would be so glad if you do. Uh, there are many more out there, of course, and you can find your own. Of course, this is the beauty of, uh, open source, and many of them are also part of C ncf because it's a big, big house where we can all live together and we can all, uh, have fun together.
But one of the, for me is the, one of the best, of course, is still so reference, uh, I'll, I'll, I probably can share, share this, this, um, presentation on Twitter or LinkedIn. So, and, uh, thank you very much for your time and, uh, I'll be live, I think, answering your questions. Thank you.





