#TeamCloudNative: Different Teams, Same Spirit | Cloud Native Now 2023
Lately, there has been a lot of confusion around the different teams supporting the cloud-native journey in the enterprise. Although containers and Kubernetes were first used within DevOps teams, other teams are now involved in achieving enterprise-level adoption.
In this talk, Mostafa Radwan will explore the different teams involved in the enterprise cloud-native journey to reduce complexity, optimize costs and improve the developer experience.
DevOps is not dead, it is evolving:
While DevOps teams have worked with cloud-native technologies for quite some time, it’s becoming challenging to perform all the tasks related to building, maintaining and operating platforms as cloud native is becoming mainstream and going beyond Kubernetes.
The Emergence of Platform Engineering:
As more and more development teams inside the enterprise adopted cloud-native technologies, platform engineering teams emerged. These dedicated teams of engineers are responsible for building clusters, keeping up with the latest trends and communicating internally with developers and other teams such as security, compliance, and FinOps.
SRE to the Rescue:
SRE teams define objectives at the service level. This can increase the ability to report and manage incidents. SRE teams collaborate with many other teams in the enterprise to keep the lights on and mitigate system failures.
Transcript
Good morning, good afternoon, or good evening to all of you. Welcome and thank you for joining Cloud Native. Now, this is Mustafa coming to you live from Chicago, Illinois.
Uh, uh, I help IT leaders navigate cloud native journey safely and securely. Today, I'll be sharing with you the different teams that make cloud native happen in, in the enterprise. Let's get us started.
We'll start off with the, uh, cloud native journey and why focusing on developers, especially now, uh, really relevant and we'll go through the timeline. How did we get here from dev teams to operations, infrastructure, DevOps, and so on. And then we'll focus on five teams that make Cloud native happen, not that we're only five teams, there are many teams that make cloud native happen.
Enterprise. Just for this session, we'll focus on those five teams and then we'll talk about the team cloud native, uh, hashtag team cloud native, and how to make your teams, uh, click. And then we'll wrap it up and, uh, happy to answer any questions you have.
Imagine you are going, uh, in a journey to climb, uh, a mountain of your choice, whatever your, your content. It doesn't have to be Mountain Everest or, or k2, but, uh, usually you start at base camp and get some, uh, training and get used to the, to the elevation. And then at some point you will, uh, you'll start taking some practice and you get the Mountain Sheba to, to guide you through the way.
And then I want you to envision all working to getting developers to the peak performance. We're not gonna climb all of us climb mountains in our lifetime, but if we can get to that peak performance and get developers to that peak performance, it's gonna really matter and help the organization move forward. All soft, all companies today are software companies, and hence the focus in developers.
If you remember this from the, uh, windows conference, I think in oh oh six by Steve Wilmer, he, uh, show a lot of, uh, showed a lot of appreciation and love to, uh, developers, uh, and he fully understood why it's important to make your developers heroes in the organization by focus in developers and having that ity, you can adopt any technologies you can imagine. Today we have a wave. Of course, we're living the cloud and cloud native.
Uh, we are, uh, going through the, the transformation when it comes to, uh, to a generative, uh, ai maybe, uh, a quantum computing next, who knows. But by focusing developers, you can always utilize and harness those technologies to deliver, uh, better results and delight your customers. Uh, eventually Let's go through some, uh, the timeline.
How did we get here? While s r e, uh, is not new concept, it took a while to take off and become mainstream. And, uh, Kubernetes changed how we build modern, uh, modern software.
Another major change, uh, uh, event was the, uh, the publishing, the release or publishing of the, uh, the Phoenix Project book, uh, by, uh, Jane Kim and others. And that kinda, uh, put the foundation for, uh, ba on on on real life, uh, use cases and, uh, scenarios how DevOps evolve it. And then we threw shift lift at, uh, dev devs seek ops and developers.
Uh, another thing to overwhelm them, while, while it's very important, of course, cause we found that the price of fixing those security vulnerabilities and issues early on can save 10 times, uh, can be 10 times cheaper than fixing them, uh, than production time. And then Kubernetes, of course, came along and we had the official s e books from, uh, from Google, and then came Topologist another, another great book helped us understand how to make your teams click. You might have different teams or gradations, but RD are working in harmony to get you to a fast flow, meaning peak performance to deliver what you're looking for.
And then we had the pandemic in, in, in 2020, and we have seen record adoption of Cloud native. And right now we are seeing, it's now going beyond Kubernetes. Kubernetes is mainstream.
Lots of us, the majority of us in this call use Kubernetes of some sort in your organization, even in production. Now we're looking at things like Observ, observability service mesh, um, eeb BF security related to cloud native applications, and how to kinda, uh, remove or reduce the complexity of coming with Cloud native. Let's talk for a minute about the different teams in, in any enterprise, not necessarily a software company, but as, as I mentioned, all companies are software companies.
Now building software of some sort or supporting or maintaining systems for a team of two, you will obviously have one communication pathway between them. Could be an email and it's a message or just someone will pick up the phone and call. If a team of five, those will go to 10, a team of 12, those will go to 66.
Imagine a small or growing enterprise of thousand people. The communication pathway between them will be around 500,000. And let's imagine for a second that I don't have to talk with every single person in, in, in the, the 999% of the organization.
I just have to talk to 10%. Just take 10% of the some 50,000 communication pathways. People in and teams in enterprise are overwhelmed with what's called communication overload, name, you name it, from NSM messages to, uh, to emails and so on.
An observation, uh, that became, uh, Conway's Law is that the, uh, organization design and build system that mirror their own communication structure. If you're talking about go microservice, for example, it's not necessarily about the technology. You wanna redesign your organization first before going with microservices.
If both the organization structure and the, uh, the team structure building the, the software, the software, if both are at odds, so we are going microservices, but we're doing top down, top down structure and we're not doing anything about it, the architecture of the organization wins, and you end up with whatever you have your organization. If it's top down, uh, designed with different business units and it's messy and bureaucratic, your software will will mirror that, uh, uh, right away. So always take a look at that and maybe invest some time reorganizing those teams for better outcomes and peak performance.
Looking at some of the organization, this graph might be slightly outdated because lots of, uh, transformation has, uh, ha uh, happened to many of those organizations. But, uh, if you're looking for example, Amazon wide, it's top down, uh, uh, structure, you will see they, they figured out a way to, uh, to have an organization of hundred thousand or more people and still innovate and, uh, lead in the marketplace by doing what they call the Amazon way. Having pizza size team, uh, idea number, I think it's around seven or so by minimize, by having a small team, they can communicate faster and deliver, uh, features and move, move quickly and limit the communication between the rest of the organization to only what is needed.
Facebook is, as you can see, mainly distributed teams, various small teams, but are totally distributed. There isn't much of a top-down structure. Google kinda took both together and mix it them in a, uh, you can see communication happening, uh, top down and horizontally vertically.
So there is still distributed and there's still top-down structure. Um, and you can see, for example, apple is still has central and as well as distributed in figure out what's worked for you. It could be decentralized, centralized, but what we're try aiming for here, especially for tech teams, is forming those teams for peak performance.
Not necessarily blaming on the technology that we didn't get. The results we're looking for. We'll talk about the different five teams that I believe, uh, make cloud native happen.
As I underscored earlier, there are other teams that are involved with, and their contribution is much appreciated. Just for the sake of this session, we'll talk about the five teams starting with the dev team. The main purpose of of a dev team is coding and shipping features, uh, on time.
Uh, the reality is that dev teams today are overwhelmed with the complexity, and you name it, cloud native, one of them shifting left, as I mentioned through a them earlier, and that is certainly important. And then learning new technologies, you name it, generative ai, quantum computing is gonna keep coming. Plus all the communication overload we talked about earlier, let alone dealing with the never ending technical debt release after release.
With our focus on developers, developers, developers, let's do anything we can to make their life easier, not harder by making their the developer experience more seamless. I came across, uh, tweet, uh, from Kin City Tower is a former Google engineer and the advocate for Kubernetes, uh, for the past five, seven years. As you can see, the C N C F cloud, uh, native landscape is huge.
We have, as of today, I think 200 or 30 or more projects. Um, some of them are open source and others, uh, com commercial solutions. The bottom line cloud native complexity is overwhelming.
And for a organization of 500 employees or a thousand employees, for example, um, the combination of tools you're gonna use and their team structure cannot, is not the same for a same size organization, maybe in the same business vertical. Um, they cannot pick the same solutions and expect the same result. It's not a one size fits all here.
Moving on to the, uh, infrastructure operations team. The main purpose of the, uh, IO team is installation configuration of infrastructure and day operations. IO teams today are responsible for a lot of things, including disaster recovery, uh, testing, uh, those, uh, those plans shifting lift to strong edma two, remember today, infrastructure is code.
And with that, there are, are vulnerabilities when it comes to infrastructure, uh, as code as well. And they are learning a lot of technologies, uh, because they can have to handle things. For example, like multi-cloud, the infrastructure is not that the same data center.
It used to be, uh, physically that we in a install, uh, uh, reactor, the different servers and, and, and storage. Now it's all virtualized, mainly, uh, cloud for lots of, for lots of organizations, and they have to figure out how to use and utilize those in different environments, in different clouds. Now, for those organizations, uh, that mature to have a deep ops team, or it's more of a mindset, not necessarily having a dedicated team team, they are in a better position since, uh, we're having more automation, collaboration and transparency to improve.
However, they're overwhelmed with many things. DevSecOps, of course, one of them. Imagine you want to secure your C I C D pipeline from, uh, supply chain vulnerabilities and a lot of technologies in a cloud native landing scheme.
Um, we all know of course, that containers, Docker started as a DevOps, uh, tool, but it it's now going beyond the just, uh, you know, one technology, one thing to do. There are lots of things, uh, risk DevOps team have to, uh, have to take care of. There is that school of thought that has been for, for a couple of years now, saying DevOps is, is dead.
And the platform team or platform ops is, is, is taken over. Now, as you can see, where we started with DevOps people, uh, process and technology, now we're dealing with, with cloud native DevSecOps, as you can see more, it's more evolving and it changing and there is gonna be a need for other type of teams, but not necessarily eliminating DevOps. We still wanna build ci cd bio plans.
We still wanna have monitoring, logging and tracing. We still wanna do, uh, automation of our infrastructure with the help of infrastructure and operations team. And those are mainly DevOps responsibility.
And the most important thing about DevOps here, that collaboration, it's not only DevOps, DevOps, platform ops, devs, sync ops, data ops, you, X Ops, you name it. There's gonna be always between collaboration between different teams to make things happen. You might have heard of SR team, uh, s e teams.
You, your organization might have one or think four on one. The responsibility here is the stability of your services. Uh, as you, as you, as you grow and build underneath different versions of your software.
You might have a, a service lever, service level agreement or sla that's a contractual agreement between your organization and your customers. And we all heard about, for example, the four lines. Imagine you are a business and you have an internet service provider, and, uh, they, they promise you of four lines availability.
That usually translates about one hour of total outage per year. From those SLAs, we can figure out the, the service level objectives, more of a three short for each service. For example, I can say my login page response time should be within five seconds if I do this and a whole bunch of other things, objectives that can allow me to, uh, to meet that service level agreement.
And for, for, uh, for operations and SRE team, that will usually translate to, uh, service level indicators. What mix of metrics, dashboards, uh, that can point out where the service is going. Is it healthy, is it unhealthy?
And if it's unhealthy, are we breaching those, uh, SLAs? The SRE teams today have loads on their plate as well because the, they have to keep the lights on, uh, with, with services, whatever the, your service level agreement is. And we also have to respond to instance, to instance, in, in the, you know, in terms of those, uh, service, uh, availability of those services handling things like postmortem and the like.
Now we can take a step back and maybe look at the, uh, services commonly used in the enterprise. For example, database. Almost all developers going to build a, a stateful application of some sort.
You might wanna run a NoSQL database or a SQL database. It can be in the data center or a multi-cloud. Now we can have a dedicated team to build and maintain the services for other developers to consume.
Developers are not necessarily database admins or database experts. They know enough to be able to store and retrieve data, uh, securely in the database. But having a dedicated team can reduce our cost and complex chill them from that complexity.
On the same note, due to the different reasons we mentioned earlier, including complexity of cloud native system, communication overload in the enterprise. Why don't we shield all the complexity from developers? By having a platform team, the responsibility is gonna be to set the golden path forward for developers to consume cloud native in the enterprise.
The platform team is gonna go behind the scenes and talk with the security team, DevOps team infrastructure operations. If you have a, a cloud center of excellence, probably you do by now. Um, the platform team will go figure it out in, in terms of building, for example, Kubernetes cost.
And as I mentioned earlier, cloud native now is beyond Kubernetes. So let's say you are already running clusters, but you wanna install a service mesh, you still wanna connect with the platform team. They might have those ser uh, service mesh solutions linked, for example, or SST already vetted, tested and figure it out which one to use for the enterprise.
Or they have their own version of it that can internal version of it, that they make sure it's security, uh, blessed by the security team and production ready. But why having a platform team is beneficial here in the enterprise? By doing so, developers can focus on the main purpose or main thing they do shaping code and features.
A dedicated team of experts will stay up to date and collaborate with the product. Mindset can always serve them better to improve resiliency of those platforms. If each and every develop development team in the enterprise is gonna build, for example, Kubernetes clusters their own way or use a service mesh their own way, we're gonna end up in a very, very messy situation, very similar to what we call encoding world spaghetti code.
But at the, at the enterprise level, we're gonna have different clusters, no standards, uh, whatsoever, let alone security as a concern. But by having a platform team who can reduce that cost by having investing in that team once, of course we have to maintain it ongoing. Those folks are expert on what they do.
They go and learn few things. They are up to date, they work with a product mindset, meaning they release products when they, for example, they are building Kubernetes cluster. They have their own versioning of it.
If they're releasing a service mesh, they are releasing different versions of it and give more options to developers to consume that technology. How about, uh, we, we flipped the coin. So what, having a platform team is beneficial for developers.
Why? As a developer, I care a platform I can do my own thing. One of the, the thing is of course, having self, self-service.
So same experience as you pick your phone and the pick an app from the mobile app store. This is gonna be, uh, with the help of something called portals, for example. So developers can log into a developer portal and then select from the different solution.
If you wanna change your container network interface, uh, you will have different options between caico, cilium or others. If you wanna pick a service mesh, you will have different options blessed by the platform team and in a golden path to move forward. Having a golden path is make it easy to consume all those technologies and hence reducing all that cognitive load and, uh, complexity and eventually will reduce the risk in terms of security and in stability by having, if all teams are using limited number of things that we are all, uh, okay with, and we all know what works for us, it's gonna reduce our risk of vulnerabilities and issues.
Of course, this is not for everyone. Some platform, some developer teams will opt out and say, I'm, I'm not interested in, in what we offer. And that's totally okay, as long as you get a a point and you get enough buy and that the majority of the borrowers go to the platform team for guidance and help in terms of those, uh, technologies you will be in Gucci every once in a while in the organization will see it in the field.
The office, for example, uh, of the CT o has some, uh, solution architects or developers. They work on the cutting edge of stuff and they are not willing to work with the platform team on the vast majority of the, of the OR work, for example, for cloud lead. They wanna be able to experiment and have more autonomy.
That's totally ok. And let's remember, if you have ever been to, uh, acute corn and Cloud Native con, you might have seen this hashtag on social media team cloudnative. It represents the collaborative effort of all maintainers and contributors to the C NCF landscape.
It's good to have structure of course, and different teams, but we have the same spirit, one spirit to serve developers and make their experience seamless. I'm gonna leave you with some, uh, great resources here. There are three books from, uh, from, uh, Google here in terms of s r e.
If you haven't started the s r e and your organization or trying to improve yours, those books can give you a good head of start, including our workbook here with plenty of examples. com. Uh, Matthew and Manuel did a great job of helping align technology teams with the business.
And we have seen earlier with the Conway Law how important that is to get to fast the flow and peak performance. To recap, uh, we wanna focus on developers. As I mentioned earlier.
Um, there are many more teams that I didn't have the time to go over here, including security, compliance, QA, leadership. I just focus on those five. But let's imagine, uh, let's don't forget, we are all part of one team, team cloud native, that can make cloud native happen in the enterprise.
And you work in your organization, work with your leadership and management to make your teams click. I wanna thank you for, uh, for joining Cloud Native, uh, now and joining this session. Please feel free to reach out to me on social media or email.
I will be around here until the end of the conference today. Thank you.





