Building Complex Software at Scale with RapDev’s Tameem Hourani
RapDev founder Tameem Hourani dives into what it really takes for software engineering teams to develop and deploy complex software at scale.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Tamim Ani, who is founder of Rap Dev, and we're talking about how to maintain this blameless culture that's kind of at the core of our philosophy of DevOps.
Even though the volume of software continues to increase, things are more complex than ever, and there's more dependencies than ever. So we seem to be working at maybe opposite goals here in some ways, but we'll see how we go. Tamin, welcome to the show.
Thanks for having me, Mike. Uh, good to be on. So what do you tell people about how to kinda maintain their sanity when a thousand things can go wrong at any given moment?
Do, do go wrong, right? Um, I think it's important to understand that things will break and, um, nobody wakes up, uh, one morning and says, today's a good day to break something. And, uh, I think figuring out how to enable, if you enable a true blameless culture or as close to it as possible, I think you get the most outta your engineers.
A lot of it comes down to understanding how to enable that, right? And, um, when we say blameless culture, um, if something breaks, you don't wanna find the person that broke it. Uh, it's generally, uh, a proxy of a system that isn't resilient enough.
And that's the theme around a blameless culture from an engineering perspective, is making sure you build a system resilient enough so that if mistakes happen, if bad deploys go out, um, you can quickly identify what the problem is and how to roll it back. And there's kind of a few different segments to that. It's, um, historically we've always done a really slower, uh, command and control prevent changes.
That's how the industry tended to work 10, 15 years ago, even five years ago, right? Change is bad and you're always looking at these metrics of 90% of outages are due to change. And well, yeah, great, but if you don't change, you're not improving in our building product.
You're not innovating. Um, so how do we balance the two? And I think we're starting to see things swing over to promote change.
Um, more change is better as long as you have the right boundaries and the right, um, culture is one of them. Uh, I hate process, but I checklists in place to make sure when changes break your systems, you can, uh, make them better the second time around. In theory then, if we're trying to make sure that the systems are resilient, maybe we should celebrate the fact that somebody broke something to highlight The fact identify resilient.
You got it after you fixed the problem. Yes. Uh, don't celebrate until, until you rolled it back.
But that's, that's kind of the, that's kinda the point, right? So like you go, there's like the first version of hit this is, how do I make sure that no matter what it is that's been deployed, can be rolled back, can be rolled back quickly. You get into different things around ab deploys, blue green deploys, uh, feature toggles.
There's n number of ways to put the right tech in place to prevent outages from lasting too long. They will happen at a five minute outage is better than an hour outage, and a one minute outage is better than a fi five minute outage. Um, how quickly you can put those systems, systems in place and use them becomes super important.
And then the second thing is, yes, once you've used that system to prevent or to roll back an outage, how do we make that system better? Why did it break in the first place? Is it bad code?
Is it bad testing? Is it, uh, a some part or some, uh, outlying component of our platform that doesn't behave the way we expect it to? Uh, is it a capacity?
Is it a a scale auto scale issue? There's a number of things that could go wrong, but once you've found a problem, you make sure that it's, uh, identified, resolved before your next deploy goes out. I feel like though rolling things back is harder than people like to admit.
And maybe that's also part of the resiliency issue. So how do we make things easier to roll back so that we can feel confident in experimenting with things? Well, ro rolling back can mean mul multiple different things, right?
Um, rolling back doesn't mean you have to literally roll back the package you deployed. Uh, that's where the different deploy methodologies come into place. But feature toggs is a great example.
Um, I'm not, I'm not sure, uh, how, how granular we wanna get. Uh, but essentially you select what percentage of traffic goes to the new code and you start with 5%. Um, and you have both, both versions of your codes deployed in production and you can slowly start to send traffic over.
Uh, and that's a great test. Uh, that's a great, um, business test as well. You're, you're measuring the business impact of the CodeDeploy.
A super simple example I like to use is you change the color of the checkout button from green to red. Do you lose people? Do people stop seeing red?
Uh, do people see red more? 'cause they're colorblind and green's harder to see. All these things come into play, but you do them with a very small percentage of traffic.
So if that percentage of traffic is negatively impacted, all you gotta do is toggle it back to zero. You're not really going in and rolling back your code. You're just saying don't send any more users that way.
That's one method of, uh, toggles, uh, of, sorry, of directing traffic. Same concept applies at the whole package level. So you've got two different clusters running your application or running your service.
You just start to send users from one cluster to another that that's AV deploys. Um, and then you can, you can roll through so many different variations, but to your point, actually moving that, uh, that package or that block of code outer production is so much harder than just moving the direction of your traffic from one subset of, uh, service to ne to the next, or pods or name spaces or whatever. It's, It sounds like I need to be able to orchestrate that.
So what is the, for lack of a better phrase, a, a control plane that enables me to kind of manage that traffic flow and make sure that the components are, aren't overloaded? 'cause somewhere along the line, I need some way to manage this thing. Totally.
There's this, we, I mean, obviously wrapped up, we're gonna have a bias, right? We work with ServiceNow Datadog, um, and both are super important. ServiceNow from a, um, a kind of record perspective.
So what is going on? What is happening? Let me keep track of all this stuff so that once this, uh, blast radius is sorted, I can go back out and look at everything that's happened.
Uh, so that's kinda one piece of it. Um, on the record keeping, on the actual control of, uh, traffic and where it's going. We use Datadog extensively.
Um, uh, Datadog is a really good indicator of the health of your traffic, the health of your application. Um, every time a deploy goes out, am I seeing an impact in response codes, access logs, uh, HGTP codes, right? Um, and at the same time, that's coupled with some sort of feature toggle tool.
Could we launch darkly? Uh, could be homegrown. We've, we've done a lot of homegrown custom feature toggle tools with some really solid caching that accomplish the same thing, right?
It's not really that hard, but measuring the impact is what's important. And that's where observability comes into play. That's where data comes into play.
Um, anytime something blows up, everyone's gonna say, it's not my fault, right? It's that team. It's that team.
It's not me. And that's the, that's the opposite of blameless culture, right? You're trying to say, Hey, let's figure out technically what broke so we can technically improve our systems.
And that's where really solid observability comes into play. Do you think that some of these challenges might get worse in the age of ai? 'cause we are now generating more code than ever, and a lot of that code, um, may be suspect 'cause it was created using models that were trained using code that was probably flawed, Worse and better.
Um, I think even just because you're using, uh, some sort of models to generate code, it doesn't mean you should circumvent all your automated testing and automated, uh, uh, scanning security scans. That should all still happen. Um, you're not saying I'm generating code through some sort of copilot, therefore this is safe code.
Uh, another kind of layer on top of that is even though you're using a third party generate code, perr, et cetera, um, a human is still behind the screen watching what's going, what's being merged, what's getting prd. So on that front, um, it's helping you get faster. It's helping you get more eff more efficient.
It's not replacing the need for anything that's currently in place on the production side or on the deployment side. You can actually generate code to fix your bugs. So if you do detect something in production, you do know what merge went out, you do see a slew of alerts coming in through observability tools.
Um, instead of having a human review it and determine what's going on, you could use models to say, Hey, here's my error log, here's the commit that broke it. Um, what do I need to modify in that commit? So you could actually generate, and we're, we're doing this, uh, for customers already.
We generate a whole new PR with the fix as a commit in the PR to essentially resolve the outage that's happening. And that can save a ton of time, right? That goes back to, um, do I want to turn my toggle off and go back and look at it as a human and take a week to come back?
Or do I just wanna look at what my agent, right, or whatever buzzword you wanna use from for LLMs is suggesting the solution is, and that could very well be your fix. I mean, as a human, you look at it, you say, oh, duh, I should've thought of that. Um, but when it's, when it's generated on the fly, it does save a lot of time.
So you could use, you could use it on both the, the dev side and the production side, uh, pretty effectively. So in effect, I am using AI to help heal the ai. Yes.
Um, and I always use, yes. Uh, ironically, I always say we use AI to generate, uh, a 70 page deck to mail to someone to use AI to summarize the 70 page deck into three bullets. There's a lot of that going on.
Um, it still makes you go faster, right? At the end of the day, if it helps, if it helps you go faster, if it helps with momentum, if it helps with velocity on the engineering side, I hate PowerPoint. I don't think there should ever be a world where we're using AI to generate 400 pages of PowerPoints.
And I think that's one of the areas in business that will get impacted very quickly, very heavily. But with, from an engineering perspective, you're building features, you're putting features on your platform. If you can do that five times faster and maybe take a hit every now and then, but you can resolve that, hit faster, then why not?
Um, nobody's gonna be able to, no human is gonna consume 400 pages of slides. It doesn't matter how good you are, you're gonna summarize them. That brings you back to scroll.
How is this gonna evolve as we go forward? We hear a lot about AI agents, and I can imagine a world where what you just described, some of those tasks are being handed off to an AI agent that's been a member of the DevOps team, as it were. Um, I wouldn't think of them that closely as like, I wouldn't, I wouldn't translate an agent to a human.
I think an agent is a subset of functions or methods that will execute. Um, and re like agents are essentially, you make a prompt, you get a response, you run that prompt through another prompt, and you try to, you try to, uh, improve upon the response you're getting from a model. It's just three or four steps instead of one.
Um, it's very unlikely that you're gonna get a very accurate response from a model on the first prompt you send it. Um, so I think of agents as a refined, uh, kind of conversation, if you will, with a given model. Uh, but yes, there definitely is a world where, especially on the kind of the lower end skills, things like password resets, things like add me to an ad group, things like really low, low level help desk is I think is gonna be very impacted by this.
Uh, call centers will be our, we're already seeing, um, tens of thousands of people being put out of call centers because that's a very easy thing, uh, to manipulate and to move over to generative AI models, right? Or, uh, generative audio, just not wireless language models or audio models. They video models.
Um, that can be very easily done in real time, uh, especially with voice tokens instead of text tokens. So you can start to run things in parallel. Um, but all that is to say absolutely all the low level stuff, um, that's historically been, uh, offshore near shore model.
'cause it doesn't cost as much when it gets offshore, will be replaced with, um, a lot of different AI models in the next two to three years. And anybody that's not saying that, um, I think it's crazy. We're definitely gonna start to see that.
So let me bring this full circle a little bit. If we think about blameless as a culture, it was always kind of the, the high end of the DevOps h engine mark. Yep.
It was, it was in the sense that, you know, I had to be pretty mature in my DevOps workflows to get to that kind of blameless mindset and kind of feel that if I have AI and I'm starting to automate more stuff, will more organizations get to that level of, let's call it DevOps nirvana, because they're gonna understand that the system itself is designed in a way that enables them to maybe stop pointing fingers at Each other, be resilient. Yeah. I don't know if it, it, you don't have to be super mature to how I blame this culture.
I think it's probably the opposite. You can be very immature, but the way you approach your half the maturity could be very wildly different between a blameless organization and a non blameless organization. And what I mean by that is you can work somewhere with a ton of bureaucracy and red tape and be immature.
Um, but the way you try to become more mature in your platform, your code base, your microservices, your application, whatever it is, is very much, Hey, every time something breaks, we're gonna sit in a room and we're gonna meet and we're gonna find out who wrote that code and why they didn't take their training and because they didn't take their training, we're gonna blame them for writing bad code versus a a, a platform in the same maturity stage. But every time something breaks, you're gonna say, Hey, what can we, what guardrails or what tech can we put in place that prevents this from happening a second time? Both of those are immature and both of those will find their way to maturity one through a different culture, faster culture than the other.
Um, historically the tech industry has very much been a world of slow down, don't go fast, don't break stuff, don't work. Let's do everything on weekends. Let's do everything on Friday night.
If you break something on the weekend, it's gonna take you eight hours to get the right team on board. If something breaks at one o'clock on a Monday, everyone's already online, you can fix it a lot faster. And that's just the mind shift in, uh, in that culture.
So AI being a benchmark sure is gonna help, but you can still very much, uh, accomplish that benchmark without having, um, without having to leverage AI to get to that blameless culture regardless of the maturity of your platform, maturity of your team, organization, et cetera. So thinking this through a little bit, um, you know, you hear the phrase over the years software factor. Yeah, I understand the concept, but I also feel like, you know, it winds up taking people out to that woodshed every time there's a problem.
And that necessarily doesn't necess create their culture you're looking for. So what is the balance between art and science and the world of software engineering? It's a, it's a very open and, and question and also just software factory is everything factory, right?
Is let's build a factory, let's build a t-shirt size factory, let's build if, if things are that simple and that, um, reproducible, you wouldn't need that many people working on whatever project you're trying to build a factory for, um, chances are they're very low likelihood of things being that, um, uh, repeatable, right? In an environment where you need to migrate thousands of VMs or get out of a data center or refactor from a monolith to microservices, there's no factory model. Um, I do think the higher up, the higher the high, the more complex engineering problems require a little more art.
I think that's where you start to differentiate between what a copilot can do and what a very experienced, um, software engineer can do regardless of, I'm not gonna say someone of a bachelor's or masters, any of that. 'cause that's also kind of irrelevant, but it's how much have you seen, right? And the difference between, uh, a software architect or somebody with a ton of experience that knows how to design patterns or how to design libraries or frameworks is gonna be a lot more relevant in two, three years than somebody who just knows how to write a function that will very quickly be replaced by copilot.
So I think we're gonna see a push to really force engineers to become a lot more, just, just think a lot more, uh, creatively in the way they write their code versus just write a prompt that gets the job done. Um, and then performance comes into play and scale comes into play. Those are the things that it's gonna take a little longer for some of these models to catch up to.
Uh, versus a human that has seen this for a long time, that'll become the differentiator in my opinion. So what is that one thing you see DevOps teams doing over and over again that just makes you shake your head a little bit and say, folks, we can be better than that? Um, that's a good question.
I think a lot, there's a, there's a really, really big misconception that, um, infrastructure patterns for as code are gonna solve all your scale problems. Uh, that's not true, right? Uh, just because you got, just 'cause you moved to Terraform or you're writing Ansible playbooks, um, you still gotta think about the way you're scaling up and scaling down your services.
Um, autoscale groups and helm charts, uh, are not the final end all be all. Um, you can, you can get pretty far with a lot of the kind of standard auto-scaling stuff, but that starts to incur a ton of cost and being able to balance those two, um, you're not done when you've moved everything to, uh, to infrastructure's code. You're still a lot of work to tune that down to make sure cost is under control and you don't have long lived, um, workloads that don't need to exist.
Uh, scaling back tends to be where things get a little more hairy. Um, and that's, I think I've seen that over and over again. That's, that's not something, security is another interesting one.
We're starting to see a lot of, um, secure. I mean, I think we will see more of this, but we're already starting to see more security shift into DevOps teams. Um, and you're really self-servicing your security needs through, uh, infrastructures code versus having to go to security team to do what, what, what we historically used to do, right?
Gimme access, gimme firewall rules, gimme traffic patterns that'll continue to move, uh, in the, in the way of, uh, kind of shifting to the developer, just the same way infrastructure shifted to the developer of the past five years. I think security will follow the security team. Then we'll just be focused on policies, um, procedures, making sure, um, guardrails are in place, but the way they get implemented and changed will definitely move, uh, more into the gi GI ops model.
So things are clearly pretty fluid and we hear a lot of phrases like platform engineering being one of them in the final analysis. How do you see DevOps kind of evolving from here? I think we're gonna see more.
Uh, I mean, DevOps kinda means DevOps is used for a lot of different things. I think generally thematically we'll see more infrastructure being managed by developers as things continue to scale out. And as a lot of the infrastructure management tools and platforms, uh, become more, um, stateful and code focused, I think there, there will be a layer underneath for shared services, which is platform engineering, which is, uh, your caching, your DNS, your, uh, traffic patterns, your network layer.
That stuff doesn't really need to be managed by developers. But I think a lot of the infrastructure stuff, the autoscaling, we will see more and more of that shift over. Um, I don't know if that's exactly what DevOps is gonna be in five years.
The, the, the term DevOps has frankly just morphed. Uh, and we'll continue to morph. DevSecOps is now a thing, right?
And GI ops is a thing and all those things continue to change. Um, but I think we'll see more, more control, uh, in the hands of developers, uh, than we have in the past. I don't think that pattern's gonna change.
Alright folks. Aaron in here, one way or another, we wanna deploy more software safely, faster than ever. Each organization may get there slightly differently, but that's where we're all going.
Tamim, thanks for being on the show. Thanks for having me, Mike. All Right.
And back to you guys in studio.