Prashanth Nanjundappa on Evolving DevOps with AI and Metrics
Prashanth Nanjundappa, VP of product management at Progress, discusses the evolving landscape of DevOps, highlighting the importance of maturity models, platform engineering, and security integration in modern software practices. He also shares insights on the role of AI in DevOps, emphasizing its potential in testing and observability, while stressing the need for organizations to measure key metrics to drive continuous improvement.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Prashant, who's vice president of product for Aris, and we're talking about, well, DevOps and how it's all evolving from here.
Prashant, welcome the show. Hey, thanks Mike for having us. Hi, me.
Every now and again, we have this moment where everybody has this conversation about, well, DevOps best practices, and what are they and what do we need to do to achieve that goal? And I kind of sometimes take a step back 'cause I sometimes wonder if each organization kind of has a slightly different definition of DevOps and they kind of meld it to fit whatever they're trying to accomplish. And of course, you know, the latest buzzword in the DevOps landscape is platform engineering, which, you know, may not fit everybody, but what's your sense of where are we right now as we're kinda looking at the next year?
Uh, that's a very interesting question because, uh, something that we deal with on a regular basis, as, you know, um, uh, you know, uh, I run Chef, which has been one of the early, uh, early pair and, uh, DevOps. And, uh, the way we work with our customers is actually looking at how mature they are in their DevOps journey. And we have, uh, built out, uh, maturity metrics as well with which we can access and tell them where they are and, uh, give them some pointers and being a trusted advisor and see, uh, you know, so that they can improve in their maturity model one.
But to your question on where they're, we're honestly seeing it's across spectrum and, uh, uh, as compared to like four or five years ago when I started, um, taking over chef to now we see a big change in a positive side. So, uh, I think six or seven years ago, agile was a norm, uh, that everyone was adopting and it had taken adoption and that was to give predictability and repeatability in software development aspect. And at that time, DevOps was still picking up.
And now we see DevOps is a practically a norm in any organization who wants faster type of market, who wants, uh, to have a high quality software, reliable software and production. So they apply, they choose these practices, but like you said, uh, the maturity varies, uh, from team. So the, the something that is constant that we are seeing is the spirit of, uh, you know, embracing or the, the fact that they're embracing the spirit of it, which is to reduce the friction between the teams to increase, um, productivity or to reduce time to market.
And that is the common goal that we see everyone working towards. But, uh, the way they approach and how they have structured team is different in different organizations. Some organizations are still kind of for me, whereas some organizations are much, much ahead.
They have identified key metrics that they need to track. They have a very clear path of measuring a, publishing those metrics and also have a plan to improve. And some of them have gone further ahead on integrating security into the DevOps, and that's also, uh, accepted as DevSecOps.
And which gives not just, uh, reduced time, which does not, I mean, just, uh, not just reduce time to market, but also gives that safety net of security ahead of releasing the product, right. So we see, uh, across these spectrums, One of the things that strikes me a little bit is a lot of the conversation always seems to assume that we have an infinite appetite for building and deploying more applications. And I just wonder, you know, are some organizations kind of try to figure out what is the top end of that number that they can actively manage and support?
Yeah. Um, that's, like you said, it's a very, uh, so there are two ways I I look at it. One, does the business need, uh, to deploy those many applications, you know, in the frequency that they're expecting?
Another, is it just because, uh, you know, the technology can, uh, meaning, you know, if I have taken, if I have adopted cloud native tools, if I am using, uh, uh, modern technologies, containers and Kubernetes, for example, which gives ability to deploy applications faster just because I have that ability, do I want to deploy faster? We see both of these. And the second one is more of an enthusiasm, which I'm also a engineer at heart, and I also do want to do things, uh, as fast as I can.
Um, but, uh, the risk there is breaking things or not really having alignment with other parts of the business. I think that's where it, we need to focus on, uh, the first aspect. Does the business need a faster deployment of application or does business need multiple applications to be built and managed?
So if you look at, uh, some of these super apps on the mobile, well, uh, where it is actually a con, uh, you know, it's a collection of let's say a taxi booking, food ordering, grocery shopping, uh, you know, anything that you can think of, those applications are designed in a way that it meets multiple needs and the market is as such. So there is a business need to actually build those applications and manage those applications, um, in the lifecycle or the fast lifecycle they want. Whereas if you look at banks or any financial sector, they, their customers use two or three core applications and they value stability over minor upgrades or new feature every month or every day or every day.
So I think the, we have worked with organizations who have that past need as well. It comes at a cost, right? Uh, it can be implemented, but it comes at a cost.
And the cost is, you know, as the saying says, you know, we only see the tip of the iceberg. The real cost is underneath. So what are some of the costs that they need to keep in consideration cost and foremost, the operational cost.
It is not just software development, but operation. What does it mean? Uh, how do you, how do you, what, how do you test that software that the application that you build, how do you integrate it into your build pipelines?
How do you release, uh, those build pipelines? How do you monitor whatever is released to make sure that they're up and running? And how do you ensure that sec And then the next, uh, aspect is security.
Security cannot be afterthought. So if you're releasing those applications, if you're operating in, let's say if you're using accepting credit card payment, you need to get PCI DSS compliance. If you are operating in us, you need, uh, so Europe, you need GDPR compliance.
So are those considered? Who is going to manage that? Who is going to maintain that?
And there are, there are responsibility to report audit and report periodically. So all these things needs to be considered. And along with that, if you have to operate at that pace, keeping all the security and compliance constraints, you have to automate, you have to embrace some of these methodologies that DevSecOps don't, uh, you know, uh, prescribes.
Are we trying to find some middle ground these days? 'cause I feel like if I think about the history of DevOps, a lot of it got started as a reaction to centralized it, and there was too much restrictions and people didn't feel they had the level of freedom, and there was this whole shift left mentality and the developers would be able to manage everything. I think in hindsight, that is not feasible and the developers are kind of choking on a lot of the things that they're now being asked to manage.
And we're seeing the rise of platform engineering, but can we get to some level of centralization that all the stakeholders involved can get behind? Uh, in fact, you kind of cured up the answer for me in your question itself, and that's how, uh, we are also seeing things evolve. So like you really said, like you said, um, uh, at some, you know, the one extreme where a team is asked to do all the task, right?
On the other X extreme, uh, which was what we had many time ago, long time ago, where there was totally compartmentalized team, it was almost like a waterfall model. Software developer creates a bundle and enhance it over to ops, ops, pick out how to deploy it, and things get stuck. Uh, you know, and we don't know where it was stuck.
Now both of them have, I mean, clearly the second one is disadvantageous, but, uh, uh, everyone doing everything works in small organizations, but, uh, does not, does not work so well if as you start scaling, because it is hard to implement governance, it is hard to create policies and monitor them and ensure that people are actually doing the right thing, uh, and also the right way. And that's where the platform engineering, uh, discipline is evolving. I think Gartner coined it, but each of, uh, ma many organizations are using different terms.
But, uh, now we are seeing designations also crop up, uh, in LinkedIn and in our customer base where they're identifying their teams as platform engineering. So they have three or four responsibilities, one, identifying the tools required and, um, standardizing the tools and the models of upper end of those tools. Second, defining a policy, um, uh, across, uh, application security compliance, codifying that and putting it as part of their software development lifecycle.
I'll, I'll take an example, uh, uh, and see if, if, uh, if I can kind of, uh, you know, explain that better. So, you know, as you know, developers have access to A-W-S-G-C-P Azure Cloud accounts, and they can go click a few buttons and spin up PC two S3 or whatever services they want outta these cloud accounts. But, uh, the organizations want to regulate how they use these services.
So they, there are, uh, these platform teams create policies that if you use any of these cloud providers and use database or storage, they should be encrypted at trust and encrypted at transition, right? Um, so then the policy is created. So whenever a new developer, uh, or an engineer, uh, even he or she goes clicks on AWS the based on the policy that is implemented by this, uh, platform team, platform engineering team, the, the services, when they get provisioned, they get provisioned as per policy.
So the, this is a kind of a, uh, kind of balance that, uh, teams are can getting to where they're giving adequate level of, uh, uh, autonomy for individual developers, but from an organization level, they have control over, uh, how, how diverse things can be. So they kind like get the benefits of self-service and their guardrails. But I'm not, um, you know, creating a ticket hoping that somebody is gonna come around and fulfill my ticket in timelines for measured in, uh, days and weeks rather than hours.
Yeah. That, that's the, that's the approach there all been to, uh, and there is an interesting model that is coming up, um, which was, um, used in software development and, uh, it's making, its in way here, which is, uh, internal open sourcing. So all these platform engineering teams, they're not really taking the burden of implementing all.
So r permutation and combinations, they create, uh, templates, they create, uh, the basic policies and they, they also create a framework with which let's say a Java developer, um, or some eso uh, programming, let's say, um, uh, you know, developer, there is a small team who is using lan and there are no policies that are, uh, defined by this core platform engineering team. They create the framework where this team of airline developers can actually submit a pool request for, uh, Orion, and the, hence the auto, the centralized platform engineering team knows and accepts, um, whatever change is coming in at the same time, they don't have to be SMEs, uh, for all these esoteric things, right? So this is another model that is emerging internal open sourcing, even in platform engineering or in DevOps DevSecOps practices.
And this is also seeing a lot of this is giving us a lot of traction in security space because security is hard for ops people, and ops is hard for security. So this kind of, or internal open sourcing is giving, uh, uh, flexibility for some of the tech folks within security to contribute into automation. So they also enjoy that, or they, they, they want to contribute, but they don't, they don't want to get, uh, you know, uh, involved full ffl.
So we are seeing, uh, this also getting traction. Of course, this is all happening in the background with ai. And so what's your sense that, well, just how big an impact is AI gonna have on DevOps and what will be the job of a software engineer going forward?
So let me tell you what it won't be, at least for the near future. And we, uh, because that is what, uh, we have seen, like, we have also tried, and we, we operate with large customer, customer base, especially operating in banking, software, uh, uh, software, financial segment, and federal space. And, uh, so in DevOps, um, especially operate, uh, tools like share, uh, the scripts that they write or the code that they write operates at a very elevated level partnership, meaning it's almost a system user, so you can't have room for error.
So consequently, um, the code that is generated using generative AI cannot be trusted. And we have had instances where they have used it for, um, uh, you know, curiosity and they thought it did good enough and put it on production. They had downtime of hours or critical data was wiped out.
So I don't think, uh, gen AI is going to replace code generation for, uh, critical applications or, uh, you know, code that, that are required to manage critical infrastructure. However, it is, it can be like a copilot, uh, how we are, how we are seeing in document generation, um, and many other cases. It can co it can, uh, coexist with a developer and help, uh, improve their productivity.
And, and I think that is what we are seeing. Um, and there is a very, uh, very, uh, you know, nuance. What what I learned is these elements are really good in natural languages, but when it comes to code, and especially when it comes to, uh, code base or code, uh, coding, which is not that big, uh, in, uh, let's say leite, chef, puppet, Ansible, any of these, or Terraform, there isn't such a vast database that it can learn and it can be trained.
And, and there is a lot of effort that needs to be put in or train, uh, uh, as compared to language. If you think of, uh, English or any other language, there are petabytes of data, uh, uh, you know, that is available in world by web, and the data that we have for real coding is less. So consequently, this code generation with, uh, high level certainty is going to be, uh, slow process In my mind.
It might have a bigger impact on things like testing than it would on actual the code that we're gonna use ultimately in a production environment. Exactly right. And not just testing, testing, you pointed out testing.
Uh, and that is where we are seeing a lot of usage. In addition to that, uh, observability, that is another place where, uh, pattern recognition is something that, uh, AI is able to do very well. Uh, we had predictive analytics for a while, but that, uh, the, the with, uh, with, with the technology enhancement, those predictive analytic models have become much, much more retro.
So we have seen now, and it can, it can consolidate data from hundreds of sources and, uh, we can train those models faster. And we are seeing a lot of our customers use it, uh, not just for, uh, creating alerts, but actually, uh, going through all the alerts and prioritizing the alert that they should work on. And in some cases, it even has helped identifying a pattern that, hey, some developer has been making these changes, and whenever this developer commits changes, there is a downtime, so maybe you might want to put some additional, uh, monitoring on this developer.
So we are seeing those, uh, level details also come out, uh, through ai. To your earlier point, we have made a, you know, a, a significant amount of progress when it comes to security, but we got a long way to go. And I can't help but wonder as we kinda look at code and AI and governance that somehow or other maybe that will improve the overall state of security.
'cause we'll be looking at the code closer than we have in the past. I mean, if you look, uh, going back to our maturity model that I kind of touched upon one, one, the last top stage of stage, uh, poor as we call it in that we see security policies codified. And it was surprising for us to see how, uh, how good the traction that is.
Um, so there are a lot of organizations who have, who have invested significant amount of time and effort in codifying, first of all, writing down a, uh, organization-wide security policy, and then codifying it and using tools to automate. And this is, uh, this is possible not just because of tools, but also some cultural, uh, uh, you know, tweaks that we have done. One common thing that we have seen is actually identifying champions, uh, security champions across the teams.
So, uh, you know, we, we in progress have also adopted our CISO is really a one person, uh, which, which sounds very, uh, rare because whenever you hear from a ciso, you're like, oh, I'm in trouble. But, uh, you know, getting that, uh, getting that feeling out of people's mind is the first step. You, you want to, you want to work in an organization where you see your security team as a partner, and that can be done, you know, uh, in various steps.
One is educating us, educating people who are of, uh, who are in different disciplines. Um, so that is fast. And second, creating champions.
So we have seen organizations where a security team has a formal training program or, uh, program around cre identifying champions and training them, and third, and reviewing the architecture and, uh, early access, let's say alpha, beta releases from security perspective and bringing it as part of the, uh, release process. And, and that way things are incremental and, uh, they see a lot of, uh, advantage going through. For example, one of, a couple of our customers include some of the, they do pen tests in early stage of the product release and whatever tests that were done.
They automate those pen tests and include that as part of a, uh, as part of their incremental release between let's say alpha to beta, the general availability. So at the type of general availability, they have proof that they can provide to security teams saying that, Hey, you did a hundred tests, all of them are failing, so that means our system is actually doing great. Uh, so if you really want to test, go find a better tester so that, uh, he or she can actually look at a different perspective, uh, then what they have already looked at, because we have covered our basis on that.
So ultimately, what's your best advice to organizations right now as they kinda look and evaluate all these aspects of DevOps? I think arguably there's more stuff up in the air than in recent memory. So what to focus on.
Yeah. Uh, so the general guidance that we gave is look for a few business metrics. And, uh, if you don't have metrics, uh, then the effort is very fix.
You can't really justify the effort that you put on. So there are a whole lot of metrics that one can look at. Um, so some of the sta uh, simple metrics that we recommend to start with is uptime sla, do you have an uptime?
SLA? If not measure, start measuring that. And, you know, the industry, industry benchmark says you have to be at least three nines if you are operating a software as a service.
But it can, uh, you know, many of them offer offer up to finance. So see where you are. And another metric is, um, meantime resolution, meantime, P-M-T-T-R, meantime resolution or a response.
So if a customer responds or requests or, uh, reports issue, how much time do you take actually to resolve that? And that is that it looks like a very simple metric, but that gives through, that throws a lot of light onto the amount of disconnect, uh, we have across teams. We had, uh, a very, uh, interesting experience that, uh, our customer reported that they, that change of la the change that was required to resolve the issue was, let's say, uh, a punctuation mark, but it took three months for them to put the release out because there were approval process, and those approval process were not automated.
And those things had different priorities. So just looking at that simple metric like MTTR can be very, very eliminating. And a couple of other little advanced metric is, uh, uh, change failure rate, uh, as in, if you are deploying a software, how frequently is it changing?
And, uh, uh, you know, what is the lead, uh, time for change? For example, if someone asks for a change, how much time does it take for us to bring those change? So there are a whole lot of metrics, uh, I don't wanna bore with, uh, you know, list of metrics, but identifying four or five metrics and measuring them and being drilling down, uh, into the details of why is it bad, how can I improve it, and how can these practices that hundreds and thousands of companies following, can I, how can I inculcate that into my organization to improve this one metric or three metrics?
And I think that is the way we recommend and we help our customers to be successful. All right, folks, you heard in here this old thing that says, well, you know, things measured or things done, but if you're measuring the wrong thing, it might not matter at all. Hey, prate, thanks for being on the show.
Thanks, Mike for having me on the show. All right, and back to you guys in the studio.