Embracing Platform Engineering with Solo.io’s Keith Babo
Keith Babo, head of product for Solo.io, dives into the benefits and potential pitfalls of embracing platform engineering, as the pace of application development continues to accelerate.
Transcript
This is Textron tv. Hey guys, thanks for the throw. io, and we're talking about how platform engineering and security are gonna come together.
Keith, welcome to the show. Hey, Mike, great to be here. Looking forward to the discussion.
We've been talking about security as it relates to DevOps workflows and all kinds of things for a long time now with mixed success. And now we're seeing the rise of platform engineering, which in my mind is kind of a, a way to manage DevOps at scale. But can we bring DevSecOps into that conversation and, and how will that all come together?
Yeah, that'd be great point. And so maybe what we could do is just step back just a little bit, look at DevOps and platform engineering, and then talk about how security lays layers into that. Is that a cool Yeah, exactly.
So like DevOps as a principle is great. Like everybody's fundamentally on board. Like if you're gonna deliver an application to users, you gotta develop it, you gotta create it.
Then you gotta actually deploy it to production and operate it basically. Right? And like traditionally that throw it over the wall mentality of like, Hey, I created the application, now you guys support it, right?
Is is just, doesn't scale basically, uh, creates ticket based cultures, lot of, you know, uh, sort of, uh, unhappiness between developer development teams and operations teams. That's why DevOps was created. DevOps has in principle, is, is fantastic in application.
A lot of companies have struggled to deploy it, right? And, and really scale it because you're looking at like, Hey, I have development team to 10 people. You know, what it ends up being is about three of those people now become operators in most cases.
There's lots of other DevOps, anti-patterns, but that's one of the most common. Um, and so you lose velocity, right? And, and again, you can't take all of your teams and then just have them all have a database administrator, all have a network person, all have security person, um, platform engineering ca came along, sort of recognized the fact that what we really wanna do is marry these two groups together, developers and operations teams.
But it kind of views it as a two-sided market. The two-sided market buyers, sellers, right? Credit card issuers and, and stores, you know, and this kind of thing, right?
Or credit card consumers. Uh, what ties 'em together is a product in the middle. And so platform engineering is all about building a platform as a product that gives development teams self-service.
They're not filing tickets to bring applications to productions. They can all drive that on their own, but they have necessary guard rails in place, whether that's security, reliability, observability that platform teams need to operate and, and, and support those applications in production. So I think that's the key shift there.
When you mentioned DevOps and platform engineering is that platform engineering is a, is really a manifestation of DevOps that's a lot more practical for companies to sort of embrace and be successful with. You mentioned the word shift and we've been talking about shifting responsibility for security further left towards developers for a long time and they have complained about that for almost as long. Um, is this a way to kinda reduce the cognitive load for security from those folks as well?
Yes, in multiple dimensions and cognitive load, you hit the nail on the head there, which is basically, how do I help developers just focus on creating business value and focus on their business logic without having to worry about the technical stack underneath those applications. Everything that they, all the details that they have to internalize in doing that is cognitive load that they have to to onboard, right? So part of the idea is with like technologies, internal developer portals like Backstage for example, really strike a nice balance here where they give like a UI where the developer comes in, they say, Hey, create me a new spring boot application or take my existing application and deploy it to this environment over here.
Or maybe create me an ephemeral EKS environment so I can start testing my application in something that's representative of production. All of those can be self-service workflows where those people don't have to be infrastructure administrators or network admins or anything else. They can just stand it up, right?
So, but you mentioned, okay, so that's the cognitive load aspect. You mentioned shift left and shift, right? This is a very interesting pendulum shift, basically where we used to, everything was tested in production 'cause there was no testing on the development side, basically very little, especially from a security standpoint.
And so we ended up just learning and getting and, and, and learning about exploits in, in production. Now we're shifting left and saying, Hey, supply chain wise, let's contain, let's do some static analysis on our code. Let's, uh, scan the container images that we're we're laying on top of and this type of thing.
Those are activities that can shift left, but we still want the ability to actually do a little bit shifted right? As well and enforce security things, whether that's MTLS right, or authorization policies for accessing services. We wanna build those into our platform and have them be enforced and developed on the right hand side of that, shifting that responsibility, right?
So developers don't have to bother with that. I think platform engineering sets up a really good balance of allowing, uh, developers to use the self-service portal and the platform teams can build in the guardrails for security underneath that, that allows those, uh, constraints to be satisfied in in the production environment. Um, as we kind of work our way through those issues.
Um, what would be the role of the cybersecurity folks in this conversation? Are they part of the platform engineering team? Because I think one of our issues we had was we couldn't really pull them across to the DevOps team because well, there just wasn't enough of them to go around.
Yeah, Totally. And so one of the things we see there, whether it's security or you know, from an API management perspective, how you're, you're creating your APIs in, in across pretty much every domain. Um, we've seen at companies like you make the the same mistakes over and over again.
How do you prevent that? You establish a center of excellence again for whatever, like, you know, software design, you know, software patterns, architecture details, networking, security, whatever. That center of excellence then has like the people that have been there and done that and they know the, the best practices and they do their best to layer over a lot of individual teams that are implementing those.
Then you implement a lot of manual workflows on how those people get involved and they kind of review what you wanna do and then they approve it and this type of thing. One of the opportunities with platform engineering is to take that knowledge that those uh, folks in the center of excellence have and encode them inside of rules and policies within the platform itself, right? So now you have very uniform and consistent enforcement of those policies and you're getting those best practices are effectively encoded in the platform, which means that developers don't have to worry about it.
Those people that have that expertise now scale much better across teams and having to engage with them directly, they can just build that directly into the platform. Where, um, do we get to the point where the developers kind of buy into this whole thing? And I'm raising that question because a lot of folks embrace DevOps to get out from under centralized IT in the first place.
And a lot of times platform engineering, when you're explaining to them, it kind of smells a little bit like, you know, the revenge of centralized it. Sure. You get a little wiki It totally.
No, no, no, I totally agree with you. So here's a very common pattern, a surprisingly common pattern that I've seen in the DevOps movement, which is that, as you mentioned, most development teams got very frustrated classic on-premise workflows of they can't bring an application to production or, you know, launch a new environment without going through a whole bunch of, you know, support tickets and workflows to get that done. That was generally measured in like weeks, very, very painful kill development and velocity.
And so when the organizations where those teams were operating started to move and adopt cloud, um, they got the keys to the kingdom, they're like, Hey, here's our AWS credentials, right? Or here's our Google Cloud credentials or Azure and like go nuts basically. And then you go in and you're like, oh, there's databases and serverless and all this kind of stuff I can just provision on my own and we can go wild.
But then what almost always happens is those DevOps teams, they start approaching production and they realize that they actually have to support this thing. So like the best practices and architecture that were built into that whole ticketing workflow they miss that went straight to development. And when they realize that they're gonna get paged in the middle of the night and they have to support this as a production app, it gets incredibly taxing, right?
Again, they have to stay, start peeling off developers to just operate and support that environment and that also kills their velocity. Ultimately how that ends up is they declare operational bankruptcy, but the difference is in the old model, they were just throwing their application over the wall, right? In this model they built their own stack, you know, top to bottom and they're throwing their stack over the wall to the operations team, right?
Which is like way, way worse. So I think what we're really trying to do here is just strike a nat natural balance with those development teams and say, listen, you want self-service, we get it. You wanna be able to start an environment, you wanna be able to deploy your application.
We're not standing in your way. In fact, we're facilitating that. We're gonna make that very easy.
So you are driving the car still, but we're gonna do it where there's the guard rails in place to make sure it's supportable in production and development teams. Don't worry, you're not gonna have to necessarily carry the pager. We're from a platform team standpoint, good to be able to support this application because we know it's been built with the right security, the right resiliency, and the right observability controls.
We need to support it Outta curiosity. How often do you think that there's a scenario where a developer goes, I got an idea, and then, um, they think about what it takes to stand up the infrastructure to go do that idea and they just shake their head and go, eh, I'll go do something else. Y yes, like very much I think there's been, it stifles sort of an intrapreneurial and, and sort of creative aspect of I have an idea I want to get to production quickly with it, but as soon as I start realizing all like the VPC config and all the other infrastructure details, I mean in cloud is very complex.
The com the infrastructure details, you're not getting a free pass there. You're just not running it in your own data center anymore, right? And so development teams, that's all still operational and infrastructure stacked to them and that does hurt their velocity in terms of like trialing new things, POCing new things and development platforms support and, and sort of encourage that more rapid ideation and standing up things quickly.
So what is your best advice to folks about how to get a platform engineering team together in a way that will stick, Start small and prove value fast, right? So if you approach this from the standpoint of I'm gonna build a wonderful castle and it's gonna have all of these things and we're gonna big bang migrate everything over to what we do, it's just not gonna work, right? It is, you need to prove value very quickly.
'cause as you mentioned, there could be skepticism in the organization. Development teams might be like, are you, is this like a rub pull? We're trying to go back to the old method where you controlled everything.
And so you have to prove to them very early on that no, that's not the case. We're giving you what you want, right? To other stakeholders and executive stakeholders, there's a team, there's a platform team that's gonna run this, that's budget.
You're applying people to this. How are you showing them that your meantime to resolution, you know, like failure reduction, your overall security postures all better as a result of this initiative. There are small chunks of things you can do immediately to prove that and then just incrementally build off of it In your mind, is platform engineer a title a thing that I go hire?
Or is platform engineering something that, uh, the existing software engineering teams just kind of do as a methodology? Funny question because I've seen customers along this adoption curve of, of deploying this and many times it starts as like a skunkwork project, right? An internal team will pick up something like backstage and they'll say, Hey, we feel this pain.
Like let's start actually standing this up maybe in conver uh, in, in, uh, collaboration with the ops team of like, let's start this small and begin it here and just start using it for this one application or this one environment and then it grows from there, right? Um, it, it could be the development team could be like someone on the operations side that's looking to, to sort of, um, uh, to sort of evolve. But ultimately it always becomes, in my opinion and my experience looking at well, uh, working with customers that you end up with a platform team that owns this infrastructure because you truly have to treat it as a product.
Like many times there will actually be a product manager that owns this environment overall because it is basically your DevOps stack for your company, right? And you're constantly looking at new things you can facilitate in this platform as a product to help development teams and ops teams get value out of it. Yeah, There is of course this little thing on the horizon called ai mm-Hmm.
Do you think that as we kind of think about the building of AI models and we integrate them into our applications, and then we think about the security aspect of that, that that all might wind up being a forcing function for platform engineering? I think it will because, and you already see it that whether it's, you know, like corpus exfiltration, like I have some sensitive data, I've trained models other people get access to that I have escapes where I'm actually accidentally giving, uh, data to, um, uh, to outside LLMs, um, like hallucinations. If I'm integrating in a, like a chat bot or an agent, uh, based infrastructure with an LLM, like all of these are preventable scenarios, right?
That if you have the right guardrails in place, but teams are just like, many teams are just trying to understand and come to terms with what their first steps are with ai and they will absolutely make mistakes there. And we can get in front of that now. Like, like as we were just discussing around like the COE and encoding best practices in the platform.
These patterns already exist for LLM consumption and Gen ai and it's just a matter of building the guardrails into the platform. So absolutely platform engineering is going to be a key aspect of how companies use ai, gen AI specifically safely and securely, which is pretty much the top concern that most c-suite folks have with Gen AI initiatives. Everybody of course is talking about how do we secure our software supply chains.
And as you look at everything we just talked, how big a journey is this, do you think? How long will it take for us to get to the point where we can continue to build and deploy software at the rate we do, but it's just more secure. It it, so it's funny because it's not going to slow down, right?
Like it is still, we are still, so like I'd say that we've crossed the chasm clearly in terms of cloud adoption, right? Like we are, like now most organizations have some footprint and maybe a majority of their footprint in public cloud as they're doing that, the granularity of applications and services getting smaller and smaller, right? So we're getting more and more of these things and then as we just talked about with, with LLMs and Gen ai, like entirely new classes of APIs and services are entering the scene, which is more we need to incorporate into our sort of IT landscape.
Um, so I think it's, it is what we're effectively chasing here is how do we do what the same thing we're doing now or do better at what we're doing now, but scale it at, at sort of an exponential growth curve and not a linear growth curve, right? And automation is gonna be the key aspect there. You mentioned like when secure software supply chain and we talked about earlier container scanning, static analysis of code, uh, you know, policy enforcement, NTLS and the network at runtime.
All of these like egress controls of the services you're consuming outside your network. All of these are, are really important things to automate and have fully in place. So they just become part of the infrastructure and allow you to basically keep deploying and integrating with new services without the cognitive load of considering those things.
All right, folks, you're hearing it here. It's a lot like car racing. We, we got get that car and still wants to go fast.
We just wanna make sure nobody gets killed while it happens. Hey Keith, thanks for being on the show. Hey, Mike, great being here.
All right, and back to you guys in the studio.