Credit Karma’s Jeremy Unruh on Boosting Developer Productivity with Argo
Jeremy Unruh, head of developer efficiency for Credit Karma, explains how his company is employing the open source Argo continuous delivery (CD) platform, developed by parent company Intuit, to boost developer productivity.
Transcript
This is Textron tv. Hey guys, thanks for the thrill. We're here with Jeremy Ra, who's head of developer efficiency for Credit Karma, and we're talking about how they're using the Argo continuous delivery platform to drive a lot of their DevOps workflows and what went into that decision.
Jeremy, welcome to show. Thank you, Mike. Thank you for helping me.
If you wanna just kind of get us started, but what is the relationship between Credit Karma and the open source project? I think you guys are contributing to that. What got you here and kind, you know, why this project over any other?
Yeah, so we originally started out, um, managing our deployments with what's called bluegreen, uh, on the service mesh. And so that's where you have two act, you know, two colors you deploy to say a green color and then you're manually shifting traffic, 1%, 5%, 10% to start to, you know, make that color the new, the, you know, the, the, the new control. Um, over time though, we started really getting more aggressive with our CICD and we wanted to have that fully automated.
You know, developers should be able to just merge their PR walk away unless they get notified, everything's successful. And, um, in order to do that, we had to get move away from Blue Green, um, and we looked at, uh, various solutions and Argo was one of the top of the list. Um, we do kind of a canary deploy rollout with Argo.
Um, what that allows us to do is Argo will, you know, start with a canary, run, a bunch of analysis tests that you've defined, and if everything looks good, it adds more and just keeps ramping on its own. Uh, and so we started evaluating that and it worked. Uh, it started working really well.
We, it, it also helped upload our reliability because when you're working with service owners, you really are starting to poke holes in how, how they roll out their deployments, what errors they're checking for, things like that. And when you automate that, you really wanna make sure it's thorough because you wanna trust the system, right? You wanna trust what it's, it's monitoring for you as it's ramping traffic.
And so, um, it was a large effort, multi quarter to get adoption, um, because we had to kind of go through and really inspect to make sure few people felt comfortable with their, their definitions on how their metrics are defined. And, um, as we started progressing, we started realizing we really wanted to have a lot of visualization, um, around Argo. And so the developers could know what's happening, they could really dive, you know, dive in and dissect, you know, you know, why something failed, things like that.
And, um, to back up a little bit, we, we have our own product we built internally, which manages all of our services, front end packages, everything. It, it, it basically allows teams to scale pods, you know, traffic manage, uh, CR tabs and Kubernetes, all that stuff. And so we, we embedded our goal into that and we built a UX on top of that.
And, um, we, as part of that journey, we started, uh, working with the Argo team at Intuit. 'cause you know, we are an Intuit company and showing them kind of what we've done. And they were really, um, intrigued by our UX that we built and asked us, would you mind contributing this back to the open source project because this would be great to have for everybody, uh, to build to see kind of how the ramps are happening.
And so that's how that journey got started as far as open sourcing r UX over to the Argo community and, um, Intuit's now using it themselves internally. So, um, Yeah. What were you doing before you discovered Argo?
Um, 'cause a lot of people have been talking about sometimes, you know, we use CI and CD in the same breath, but they're kind of different functions and processes to a certain degree. Um, what's your sense of where are we and how did you discover CD as kind of a distinct discipline? Well, it's something we've always wanted.
Um, you know, when I came into Credit Karma, we were, what I, I like to call it a, we were a decentralized model. Um, uh, you know, we were a smaller company at that time. It was six years ago, uh, and we were in hypergrowth.
And so a lot of teams were kind of doing their own DevOps. They were managing their own Jenkins jobs. And, um, so I came in with the approach of, you know, with our rapid hypergrowth, let's centralize, let's have, developers shouldn't need to know how their services are deployed.
They shouldn't need to write code on how it's deployed, like it should just be kind of turnkey. Uh, you know, and, and then that way we have a little more control, uh, of what's happening. And so as part of that journey, you know, we originally did ci, cd, you know, from merge to say our test environments.
But in order to get the, you know, leadership to buy into trusting it in production, we needed to start to make a lot of bets. And one of 'em was Argo because, you know, Argo allows you to define your rollout plan and all your metrics and how it's gonna be analyzed, how long to let it bake, all these controls that you want in place. And so we needed that piece of the puzzle to really solve CICD from PR, merge to production.
And Argo seemed to be the best fit. And, you know, contrarily at the time, we didn't realize that it was in open source buy into it. And so it actually made that partnership even stronger 'cause we had that internal connection to go to ask questions and, and you know, where we can um, you know, integrate that into our product.
Uh, the, um, so it's been a journey. Um, you know, like anything else, we had to have a lot of other gates and policies and things like that in place to, you know, outside of Argo to really say, oh, is there any security vulnerabilities? How's your quality?
Like, have test automation ran? So we had to build a lot of that up over the years to get us to where we are today. And um, today we are moving forward and all hands on deck moving everybody to what I call zero touch the ICD, meaning hands off, I've murdered, I urge my PR and I can walk away.
And, um, that's what we're rolling on now. Is this part of a push towards platform engineering for you guys then? 'cause it seems like you're centralizing some of the DevOps workflows with a sense of automation.
Um, what's driving that? I mean, do you use that term? Is that a, something people recognize or is it just kind of like a logical progression till you wake up one morning and go look ma no hands platform engineering.
It's funny you say that. So we are part of platform engineering. So platform engineering is quite a large organization.
They have efficiency. You can think of it as we're building the products and the tooling on top of the, what the infrastructure teams are putting out there for managed services. Um, and so, um, as part of product engineering, our whole, our whole charter is to, uh, you know, improve the dev experience for developers and make them faster, you know, improve efficiency.
And so, um, we didn't want developers to have to do fine grain training on how to work with Kubernetes and what they need to do. We wanted to kind of abstract that away with some knowledge, you know, that, that they should have on how to diagnose, but make it more turnkey. And what we've noticed is by doing things like that on top of platform, and it goes beyond just deployments, um, with our other tools, it's actually made folks onboard faster.
We've been able to ramp teams, you know, new developers can come in and actually quickly start going through a workshop and actually know how to manage and build code and deploy it. Um, but it's also giving us a lot of knobs and controls on things like, oh, we are having a site incident, we wanna pause everything from making change in production, you know, just so we can kind of, uh, reduce the impact. And so, um, that model has worked really well and it was, I'm glad we did it, you know, earlier on as we were growing 'cause it would be very hard to do now with the size of our company.
What's your best advice to folks about how to approach that? 'cause on the one hand, developers will say, uh, I don't wanna know all this stuff. The cognitive load is too high and you should automate that for me.
And then the next breath they'll say, but I want to use any tool I want to do. So how do you kind of strike a balance between letting the developers innovate and experiment with whatever they want to do versus bringing some adult supervision to the conversation? So we don't try to, we don't try to, you know, we have what's called like a paved road and you know, you do it these ways and this is the most effective way and you can run fastest.
However, though the paved road allows for you to customize along the way. There are obviously some stops and things that you can't control. Um, but I would say that initially, um, we did get that early on in the company, folks were like, well, why do I need to do this?
But what happened was I did something very different. I did this in my previous company too. I brought in a designer in the very beginning and you know, so a designer and platform who's ever heard of that?
And people thought, you know, why do you need to hire a UX designer? Well, if we really wanna understand the needs of developers and what their experience should be, then you need a designer that can kind of jump in and understand like, how does, what does this mean to you? How would you use this?
You know, these types of questions and build the optimal design. And once we started pushing this out, the, the developers actually really embraced it. And we do a lot of, you know, internal surveys and stuff.
And our products that we're offering are very highly rated because it's actually day-to-day. They actually seem to like that, that they don't have to think about them. They can focus on delivering business, uh, feature sets.
And, uh, so I'd say pushback in the very beginning a little bit. But then over time, when they saw what was coming out and how much it made their lives improve, um, it flipped the other direction. And, um, we're very active, uh, my organization within the engineering community.
And we are always, you know, like I mentioned, setting out surveys, but also asking for feedback and where can we improve? What, where, what additional flexibility do you need? And so we really take that feedback seriously and, um, help, you know, kind of provide that happy medium.
How does one measure developer productivity these days? Because it's not lines of code anymore, so it's gotta be something else. And I'm asking the question 'cause a lot of times you can give developers all the time in the world they want, but unless inspiration strikes, they're just not gonna write code.
So how does this kinda get measured in something that the business folks will understand or they just kind of nod their heads and, you know, along for the ride? Great question. So, um, one of the other products that we built was called Flare.
And so Flare is a data warehouse for everything that happens within developer lifecycle. So anytime there's a PR merge or a comment on a PR or a deployment, or I've done this action in Kubernetes, all that data goes into flare. So we have a vast amount of data.
That data is used for what we call in adding policy. So our security can say, Hey, we want a policy that says if you have a P zero vulnerability that you cannot get to the next environment, great. You throw that in as a definition and the system automatically blocks that, you know, new version from going out because it's, it's not met the policy.
On the other hand, it's provided us, um, very rich UX on visualization for leadership as well as developers. Um, so back to the efficiency question. Um, we started out actually, uh, implementing Dora metrics, um, you know, which, uh, something Google put out and that's, that's great.
It kind of tells you how fast your throughput is and your stability. Um, but that wasn't enough. You know, it's a good early signal, but it's still not enough when you're really asking questions.
And so we started evolving and we started saying, okay, you have Dora. Now we're gonna have a new score for your team or your service called quality, which will basically be like, what's your coverage? Like how, you know, how fast are you closing bugs?
You know, things like that. So you can kind of start to say, okay, here's your quality pillar. We did the same for security.
And right now we're actually working on the efficiency, which is additional to dora. And how we're measuring that is we're looking at things like, how long did it take for you to, um, from first commit to creating the pr? So what was that development time?
How long did it take from time to first comment, meaning your team engaged and started reviewing your PR to when it was merged, and then when was it finally released into production and ramped and not rolled back, right? And so we're kind of looking at that whole timeline. We're also measuring things like, are you using the recommended tools and practices and patterns we've put out there that we know make you more efficient?
And if you do, then you gain more points on your score, for example. Um, so all these little things make up the efficiency score, which gives us an early signal. One interesting example was we were looking at, um, you know, throughout, you know, right after the holidays and we took a few teams, um, and we looked at like, you know, right into the January and we were like, oh wait, it's interesting because this team's been very consistent with how fast they reviewed prs.
But in early January it was a huge timeline and we noticed this across various teams on how long that PR review took. Well, a lot of folks were still out from, you know, the holidays, they extended time in January. And so it was really reflective.
'cause obviously the team has a makeup of maybe certain seniors need to review the PR and maybe they were out. And so we're starting to see early signals of this pattern, um, evolve. And it's also allowing us to dig in and going, well wait, why is this team shipping things faster than this team?
And we're not really looking at it that way, but it's more about for leaders to kind of say, wait, why is this taking longer than it should? And you can start to dig in and going, well, it looks like the team's really taking while to actually respond to a PR when it's created, or it's this thing over here. They're not using this tool that should be used.
So then it's up to the team leader to kind of dig in and go, Hey, let's try to improve this and see if it makes a difference based on these metrics These days. You cannot now walk down the street without somebody leaping out to tell you about their great new AI thing. So you're collecting all this data.
Can we throw algorithms at DevOps and what might that look like? It's, it's funny. So my, my org also is, uh, is driving the, um, the what we call the dev ai lifecycle.
So we've, for example, like most companies, we were the ones that pushed our GitHub co-pilot to all the engineers. We have also leveraged, um, Intuit's, uh, Genos studio, which is, you know, uh, basically like a pro private open ai. And we've built, you know, tooling on top of that, um, almost like a mini AI platform.
And the reason why is because as a developer, I'm authenticating internally with, you know, my corporate credentials, right? But our users, our customers, they have a different authentication. So, you know, we had to kind of bring that bridge in on making it seamless.
We know who this developer is, we know what they're accessing and how they're using ai. Um, so from there we started, once we built this kind of mini platform, um, on top of, you know, just like a layer on top of Intuit's, uh, Genos, we ended up, um, creating, uh, one of the biggest areas first we we delved into was knowledge discovery. Because I mean, I'm sure you've heard in lots of companies, there's the whole Slack game, Hey, which team do I contact for this?
Or where, how do I do this? Right? So we ended up, um, taking a vector database and we took a bunch of documentation from one of the, uh, for progressive delivery actually, and Falcon, which is the product.
And we, we basically embedded it and threw it into a Slack bot. And it was amazing what that did it any, almost any complex question you could throw at it was able to answer it and tell you exactly what you had to do and where the documentation was if you wanted to get in further details. And so we're expanding that because that really starts to reduce time that the team spending answering questions, but also reduces the, the cognitive load of an engineer not knowing who to ask, right?
Uh, because you have this kind of one knowledge bot out there. We're also doing it for pull request optimizations, being able to, um, you know, refine descriptions and things so when reviewers come, they actually really understand what's, you know, what they're looking at, uh, from, from, you know, this optimization. And so it reduces PR churn and we're embedding it into a lot of the data that we have in flare and looking for opportunities for the eye to help give leadership recommendations on what they can do.
Hey, you do this and your quality score will go up a whole grade. You know, this is a very easy thing for you to do. So, um, it adds a lot of power in basically driving change, um, by bringing that in.
Do you think we're gonna struggle with managing the amount of code that we're gonna see generated by all these AI tools on the front end? And are our backend CD processes able to cope with that? Because the software bills seems like it's gonna increase in size and there'll be more of them.
So what should we be thinking about long term? That's a great, that's a great question. We haven't ran into it yet, but at the same time, we, we have built every one of our systems, even CICD to scale.
All of our systems are in Kubernetes. We can easily scale out if we have more workloads. So it's all automated that way.
Like, oh, our queues are building up scale more pods, right? And so we can accomp, uh, accommodate the load. But even rolling out GitHub copilot, even though the surveys we did, you know, a lot of engineers are very satisfied using it, and it has really helped with their, um, you know, boilerplate code.
I like to put it like tests, running tests, things like that. We still haven't seen like a massive jump in, uh, frequency of deployments, for example. It hasn't really, the, the needle hasn't moved that much, even though we have I'd say 50% or more of engineering using AI for code generation.
Um, and I would expect that I would see that. But what we're seeing is, um, is teams are able to add better, um, better testing and things that maybe they didn't have time to do before because they're getting pushed from the business to do this. And so our quality and our stability is actually going up, believe it or not.
And so we're starting to see that needle move, which I didn't think would be indirectly related, but I'm, I feel that because they're not having to spend time writing this boilerplate, they can let AI do it, and that helps. Um, that being said, there are still gaps because copilot doesn't know our internal ip and a lot of this, you know, um, technologies we use in our, you know, in our platform and our stack and, you know, um, it'll be interesting 'cause as we start to incorporate more of into its Gen OS and we start maybe replacing some functionality copilot does with our own, because we can train it on our, you know, our code pattern, security centers, things like that, then maybe that will change because it'll be more accurate with some of that code generations on internal, you know, uh, technologies that we use. All right.
So clearly you're down the path with Argo and Kubernetes and Cloud native. What do you know now that you wish you knew a year ago, two years ago when you were first starting this whole journey? Um, well I wish we would've looked into Argo earlier.
Um, something like that, even if it wasn't Argo, something that can do, you know, a non bluegreen and we, and you know, more automated ramping, uh, I feel that if I knew AI was gonna hit so fast, there would've been things I would've done with maybe having our, you know, the way our data was sorted and things like that, which would've been a little easier to embed and, you know, throw into a vector database for training and things like that. Um, and so now it ended up being a game going, oh wow, this hit fast. These are all the things we gotta do right now so we can at least get the data in a format that AI's gonna appreciate, right?
So it would've been, it would've been less time on my team to kind of scramble to make that happen if I was ahead of that and ready to go, you know, knowing AI was gonna come. So I would say those are probably the two biggest things. All right, folks, you heard it here, Argo's here.
It's happening. And we're all taking the next step in the DevOps journey because arguably we kind of did a lot of stuff with small teams, and now we need to figure out how to do all this stuff at scale. Hey Jeremy, thanks for being on the show.
Thank you, Mike, for having me. All right. And back to you guys in the studio.