DevOps and IT Service Management – Ming Gong, Blameless
Ming Gong, newly appointed vice president of product management for Blameless, dives into how the role of the SRE is evolving as IT organizations seek to bridge the divide between DevOps and IT service management (ITSM).
Transcript
This is Textron TV. Hey guys. Thanks for the throw.
We're here with Ming gong. Who's the newly appointed vice president product management for blameless. And we're talking about what is going on with all things sres Ming.
Welcome the show. Nice to be here with you today. My oh playlist has been at this whole effort to create a platform specifically for sres to help them manage their workflows and everything that goes on from one end of the process to the other and the question I have for you is you know, Are you seeing people actually become sres and then look for a new platform or is it they kind of just accidentally fall into this role and they wake up one morning and somebody starts calling him an SRE and they're like, oh, well, that's great.
Thanks for letting me know who I am. But I guess I'm trying to figure out is you know as SRE becoming standard job title or is it still more of a description of a function? That's a really good question.
What we're noticing in the market today is that SRE is very much a job title still. but as blameless evolves and as the market evolves, I think the Practice of SRE is goes beyond a title. It goes beyond the role.
So more and more you see teams. Operate and develop software all within the same team. So SRE is very much a practice that's being adopted in the software realm also.
What do you think is the relationship between an SRE who manages the it infrastructure and a lot of the processes for deploying the application and the rest of the it organization, which is typically made up of Administrators who use more graphical tools. Does there need to be some sort of meeting in the minds of these sorts or is the SRE kind of operating and isolation from all of those Mmm Yeah. That's a really good question.
I think about this question in terms of teams how teams operate and collaborate together to ultimately satisfy the needs of their customer. And typically in a software development like life cycle. The landscape is made out of a lot of different roles a lot of also a lot of different tools and best practices.
So if you thinking about Satisfying your customers you have to operate as the as a team and oftentimes these tools and these practices. We'll have to be aggregated and integrated into a single platform that way the teams can collaborate more effectively and understand what's going on in a wall holistic way so that SRE is not isolated. But it's the whole entire team that's good that gets involved in particular if you think about a good incident response scenario.
It's important to aggregate all that's happened in the development operating pathway. So that you can get the right people involved the right time and better manage expectations with your customers and stakeholders. You can better resolve we can better learn as a team.
So in essence, I think that SRE you shouldn't be isolated and it all comes from the tools and all comes from integration and aggregation to a single platform. So regardless of role in the organization you kind of want to have a view of the same thing that everybody else is looking at and yes re is the leadership of that. Effort you still need a view for the it administrator and understand what it is that they're supposed to go fix.
That's right. That's right. Ultimately, you really want to centralize all that's happened along your tool chain.
So you can have this vantage point to assess these hidden points of friction across not only your tool chain, but your Aussies and your teens and across your life cycle so that you can make sure the same incident doesn't happen twice. We have been having this debate about devops versus traditional it service management tools and there's different Frameworks and different levels of agility and speed and reliability concerns. And I wonder as a Time come to kind of put aside that conversation and kind of just move forward and say, you know agility doesn't have to come out there expensive reliability and vice versa.
I think that's right when you talk about agility. And reliability really you're talking about going as fast and as safely as possible, I blameless we think about that in terms of resilience. And resilience ultimately comes from the ability to learn and ability to prevent and also that goes back to my earlier point in how we have to have that Vantage Point both learn and prevent so you can create an environment of resilience and why that's really important.
Is because that so that teams. Can see you know based on past data past incidents. They can set better goals going forwards, you know to better manage resources to improve collaboration improve safety improve velocity, ultimately improve trust with their customers.
No matter how fast they go. And yeah, that's what we think about in terms of resilience. Yeah, the other animal we see in the devops zoo these days the so called Full stack developer.
And my question to you is does that function exist? Because there wasn't somebody who was an SRE to manage the infrastructure for folks or developers kind of took that on themselves, but it seems like it's a rare development that actually wants to manage infrastructure. So we have a moment in time where maybe you know full stack developers were necessary for a moment, but maybe everybody just wants to go back to writing code because there's an SRE to take care of it.
That's a good question. I'm when I think about full stack developer I think about the need for context switching What I mean there is this might be a little spicy but really I think businesses really want to move as fast as possible. And when you have somebody who is a full stack developer.
They don't necessarily need to cross train. They don't have to switch context. They know how to do the whole entire stack and I think We we think about operating it's not about not wanting to operate is the need for the business to go as fast as possible and having the least amount of context switching.
What is your sense therefore of the level of maturity that we're seeing in organizations? I mean are they just beginning to get sres or they further down the path? I mean and as you think about this how automated can automate it get because some folks are like well we're seeing AI so maybe we don't need SRE so, you know, we're real journey.
Yeah, that's a that's a lot to unpack there. Okay. I think that AI Ops is still quite nascent because tools and processes differ from Team to team.
Secondly the there's another Force at play here is that as developers a full stack developers transition into the operational role. You're going to see a maturity curve that will make it so that some organizations will be more advanced than others in some customers. We're here today.
They believe that incidents are still very fraught. but in other cases customers have transformed to a state where these fraught moments and Incident Management scenario become an opportunity to learn. So I think it's it's still emerging.
We also hear a lot of debate about the word certification these days. So do I need a certification to be an SRE or is it more just a question of my level of experience and you know people should know that I have a capability just by looking at my resume, but I don't need somebody to give me a piece of paper when I hang on it that says I'm an officialist or You know, I'm not necessary myself. I've only observed from the from the benches from the bleachers so to speak.
In my opinion and this is not you know industry standard. I think that it's just about how effective a team might be. So, you know, if you you can operate better with certificates or official roles and titles.
Maybe that's that's a best fit for that work, but for some other orgs, I see advantages for the SRE practice being applied across all disciplines. Well, it does seem like if you do call yourself an SRE, you're more likely to get a raise, right? I think so, but also I think that that's a That's an area that I'm not qualified to say.
All right. So what is that one thing that you see your customers doing when they first get started with the platform? That means you shake your head and go can't believe that that's their launchpoint.
What do you wish that they knew before they got started? That's a really good question. I'm still early in.
developing an opinion here But I believe that customers are getting started in all different ways and I'm supportive of all different ways. So I see some customers starting from Incident Management perspective where they are frantically reacting and trying to be better at triaging their incidents getting their faster improving their mttr or scores. But also I see other customers starting.
And from a objectives standpoint meaning they'll want to set some slos to begin with and then work backwards from those objectives. I really like that approach but also sometimes that approach means that they're forcing the SLO across their teams. regardless of how much are the practices might be So do you think we're on the cusp of being more proactive about the management of it?
Because despite the last five Decades of effort. It seems like we're still in this reactive mindset. But yeah, are we on the cusp of changing something fundamentally about the way the it culture operates?
Well, I don't know we're becoming more proactive but I can say that being proactive is really important. You know. Again, when we think about proactive making sure the same incident doesn't happen twice.
You really thinking about resilience you really thinking about going faster and going safe for at the same time really earning that customer trust. So I have to believe that if you're in the soft development and in the space where you're trying to satisfy customers and moving as fast as possible you're going to be Interested in becoming more resilient and more proactive at the same time. All right.
I think you just hit on a great metric there folks. If you're encountering the same problems over and over again chances. Are you need an SRE?
Hey Ming. Thanks for being on the show. Thank you.
All right and back to you guys in the studio.