AI’s Impact on Software Development with Tracy Bannon
Tracy Bannon, a senior principal in MITRE Corp’s Advanced Software Innovation Center, dives into why DevOps teams should participate in a study of the impact generative artificial intelligence (AI) is having on software development.
Transcript
This is Textron tv. Hey guys, thanks for the throw. We're here with Tracy Bannon, who's with the Mitre Group, and they're starting a study on this whole gen AI impact that will have on DevOps, and they're inviting folks to participate.
But we're gonna dive into, well, why Tracy? Welcome to Share. Hey, thank you so much for the invite, Mike.
I've been excited to talk with you about this stuff. We've got a whole bunch of, uh, effort going on around researching this. Everything from looking at the literature, what does the literature say, the scholarly literature say, and then an inter, and we have, uh, interviews that are going on.
We've got case studies and we've got an anonymous survey. Think of it like your Doras or your stack overflows, but we are not for profit, so we have no motives other than identifying what's really going on. Um, the goal is to figure out where should we be applying ai.
Generative AI is the key focus right now, but where should we up be applying it across the entire software development lifecycle? It's more than coding. It is really more than coding.
So, but I, you know, Mike, I've even sent you a, a copy of the survey to fill out because you, uh, play in this space as well. We're not looking for just engineers. We're looking for anybody related to the software development lifecycle to give us their thoughts on this.
Takes about 10 minutes. Takes about 10 minutes. com that it kind of explains the purpose of the research?
So we're gonna refer everybody over to that. But it seems to me we're not quite making the distinction between, well, something that helps me maybe write code faster, doesn't necessarily result in software building faster. And so my question to you is, you know, we see all these surveys that say developers are more productive, but there's also some studies starting to emerge that suggest that the DevOps workflows themselves are not getting much better.
And what's your sense of where are we on this journey? I can tell you what, uh, six months of research has showed me. It has shown me that there's no change right now today to the SDLC organizations are churning and trying to figure out, because there's this fear of missing out, this fomo, this hype, we need to use this, but how am I gonna use it?
How am I gonna use it? Well, our developers should use it. Oh, we need the, we need the engineers to use this.
Well, that's okay for code generation, but there's so many other places that it potentially could help. The problem when you're looking at developer productivity is that it's a concept called, um, perceived productivity. And what that means is that the data, the observed information about what you're doing is different than you self-report.
Why though? Why would people report that they're being more productive? Because they perceive it?
Because there's energy and excitement in getting to use the button and it generates this, or I'm chatting with the back and forth, but voting is only one small part of it. And quite frankly, right now, developers spend, from the data I'm seeing, they spend about 10% of their time writing code and they do all of these other amazing things that they have to do. So we've gotta gotta get out of this habit of focusing on how fast an individual developer does their bit.
And start looking at how is the software development lifecycle being improved so that you're getting value delivery. Do you need higher quality? Do you need it faster?
Do you really need more deployments? What are you getting after? What are you trying to accomplish?
Imagine if you map to that, uh, instead of what we're doing right now, which is fomo, let, let's grab it. Let's go, let's go. It's Gen More code.
Um, It's not clear to me that even writing more code is progress, because if you look at some of the analysis that goes on the lms were trained on code pulled from everywhere. And that's a varying quality. And it looks like in a lot of cases, folks aren't really looking at that code too closely.
They trusting the machine too much. So there, um, there are a couple of different things to consider there. You're right, that the ability of a language model to generate something, right?
That's what they're for. They are to generate, they're really two use cases and I'll come back to that. But if they're to generate, um, it doesn't necessarily mean they're gonna reduce the amount of of code.
If you have it looking at your code base, is it going to suddenly say, Hey, you need fewer lines here? Generally not. And you're right.
There are studies coming out that should worry us. One is from a group called Get Clear, uh, and they track a number called Code Churn. And Code Churn is ha that that number has been relatively static, um, by domain, by business area FinTech versus military versus medical, et cetera.
And what it represents is, as a developer or as an engineer, code gets checked in, then it gets checked out because you tinker with it and you check it back in and you make a change to it. That cycle is called code churn. Well, there's a reason that it was pretty even from 2021 to 2023.
And now they are saying they are on scale right now to double code churn. Hmm. Double code churn.
Why is that? Are we magically doing a better job that we're refactoring? No, there could be a correlation that more generated code that less people are paying as much attention to means more tinkering.
There's also and, and report that came out yesterday, it is an academic report that came out from MIT and they really simple, they got three teams and they gave them each the same problem. And I don't have the details in front of me, but I'll sit tell you that each team was asked to solve the same problem. Team one was to use chat, GPT.
Team two was to use Code Lama and Team three was to use just Google. Team three had to decompose what the problem was and go after researching different pieces and parts. Who was the fastest?
The GPT folks, second in line code lava third were the people who were using Google. Then they gave 'em another task, which was essentially to leverage all of the knowledge that they had gained to solve another task. The only ones who could solve the task were the ones who'd used Google.
So there's merit. We have to, we have to really be thinking about what happens when we're asking for a leg up, but we're no longer, uh, authors. We're becoming editors.
Think about that, Mike. You write a lot, right? You work with a lot of folks that are that, that are writing things.
There's a difference between writing things creatively and being the editor on top of that, right? You pay attention to different pieces. You don't necessarily have to be an expert in that area.
What could happen? That's, that's part of what we're trying to get after with this, with this research effort. What are the barriers to adoption right now?
Is there a way to help you measure when you should and when you shouldn't? And there are, uh, and that case study I was telling you about, we have a decision tool that can help some folks do that. What happens with human machine teaming?
Should I trust it? Should I not trust it? When should I trust it?
When should I throttle back? But what's gonna happen in the future is where I'm super interested. S DLC is not changing right now 'cause we're having too tough a time figuring out how to use these tools and inject them into our software development lifecycle right now.
But as we look at a Gen X, as we look at the ability to provide more trustable ai, what will that look like? I believe we're gonna see a dramatic, a dramatic change, but we're gonna have to figure out how we, how we create a workforce that can thrive in this new, in this new environment. Yeah.
There's been studies that have shown over the years that if you simple, if you write something down or you take notes on something, you internalize it more than you do if you just listen to it. And yeah, I think that may be playing out here with coding. I think you're spot on.
You're spot on. I was thinking about, uh, mentioned it to a friend yesterday. I was thinking about, uh, calculus in college and having to on paper, right?
Doing those serums and proving those things out and, and at times failing and needing to go in and talk and walk through the exact steps and then being shown where I went awry, right? Think of this like debugging. Where did I make that mistake?
You remember it, you remember it. The next time it's experience that you carry forward. We are at an early stage where we might be able to head this off, but it's gonna change how we educate people, especially software architects, software engineers, testing individuals, anybody related to software we're, it's gonna change the way that we educate them.
I am. I got some concerns about it. Mm-Hmm.
Are you worried that we might be injecting more vulnerabilities into our applications? Because a lot of these things are trained on code from everywhere and a lot of that code is faulty and the next thing you know, we're having more vulnerabilities than ever. It's a, it's a yes.
And if you have an organization that has matured, CICD, that they have a safety net in place, I'm less less concerned there because generally in those teams are taking full responsibility for the code base. If you don't have that, then you're probably in a potential world of hurt. Exactly.
To your point, there is research that came outta Stanford that showed something in the neighborhood of 56% of the code gen being generated using GitHub's copilot at that point. And this is a couple months old, had had security flaws in it. I know that I did my own research, sat down and went through an afternoon workshop, um, worked with, um, um, sync and use their tool base, and I was able to generate code.
Uh, and when I generated that code, I could then use this tool to double check if I had created, uh, security flaws. I was able in three sentences to create two different owas pens. And I wasn't trying to, I was just trying to get data from a screen to a backend API.
So yeah, yeah. We gotta teach people about what it means to have secure coding practices, quite frankly. Is in your sense though, that we may be on some journey here because we see the latest revs of the large language models that people are putting out there or have better reasoning capabilities in there.
And it's one thing to write code, but maybe the reasoning capabilities will help us with the other aspects of DevOps and maybe help automate some of that and maybe we'll get some balance in the system. Well, I don't need generative AI to help me with DevOps. I was in a workshop on Friday and we went around the horn talking about all the different ways that we were using generative AI in specific in the DevOps pipeline.
And I can tell you that we're really not using it in the DevOps pipeline because it is non-deterministic. We can use machine learning algorithms, we can use, there's so many other types of AI algorithms that have been around and that are deterministic that we can count on. Those are the things that we need to be inserting.
There are really two buckets that I tend to club things that are coming out of generative AI models and out of large language models, it can generate some kind of content code, whatever you wanna call it. So it will generate that for you. It can also help with what I'll call is automated reasoning or automated analysis.
It doesn't do it for you, but by feeding in enough contextual information and asking for themes and looking to rationalize across a corpus of text information, it can really help you. We used an example of putting somebody who had to create a test plan, but they were doing manual testing. So they took their requirements, they took what they understood that they took the test plan, the samples that they had, and they were able to come up with a bunch of really awesome ideas, wasn't being generated on their behalf.
It was augmenting, right? It was going back and forth to help them with that rationalizing. So that's really where the two buckets are.
But think about that with your DevSecOps pipeline, which is all about repeatability, auditability being able to absolutely count on it to repeat itself again and again and again. Doesn't sound like generative AI to me. So thinking this through, then to its logical conclusion, um, should we just be more careful about the use cases for generative ai?
Because it seems to me that there's a rule of difference between writing a script that I may just need to kind of connect to things in business logic. Mm-Hmm. Um, it's another yes.
And the, the corpus of reviews that I've been doing, the interviews that I've been doing with folks show that those that are more senior in their career are able to take true responsibility for anything that's generated and move it into their SDLC in whatever way they want to. If they need a script, that's great. They will use that, they will tinker with it and they'll toy with it.
So yes, it can be helpful. I'm not a hater on any of this. I actually believe that it can be wonderfully helpful and people need to be thoughtful about the potential ramifications it does in, it does have these hallucinations, or we could also call it a vivid imagination, right?
Where it just comes up with something. I personally like to think about lang large language models that I don't like the term hallucinations. I prefer to say this vivid imagination because that's a benefit.
That's what they're supposed to do. Statistical mathematical likelihood that these words, that these phon ees, that these symbols are somehow connected to one another, that's what they do. So hallucinations or vivid imagination is actually a feature and not a bug.
So we have to think about that when we're applying it. Where, what's the risk? If I have something that's not immediately recruitable, what's the risk?
Okay. Am I ever going to allow that current day technology? Am I ever going to allow that to flow forward without human eyes?
Mm-Hmm. No, no, not yet. The, the part I always find concerning about all these hallucinations is they're presented with such enthusiasm and authority.
And I wonder if, uh, what we really want is an assistant that kinda says, here's the suggestion and here's the level of confidence that I have in it. But you should, you know, evaluate it yourself. And as opposed to, I think we have a tendency in as humans to rely on a machine when it's, you know, presenting us something very confidently.
Well, and it's, we also tend to personify, uh, the generative AI interfaces right now. Especially when you add on that I can now use chat GPT and I can talk with it. And the voice that comes back is not perfect.
It has us and stutters and restarts. They've in, they've inflected, right? They've put human characteristics against it, which makes it even more believable.
And you're right, the just because something is really well formatted and appears correct and might even work if we're trying to compile something or work if we're trying to run a script, doesn't necessarily make it right or appropriate. Um, but that's, it gets back to that risk. If it's wrong, what's the risk?
Well, what if we say that those people associated to the software development lifecycle are responsible for the software that's in the software development lifecycle? I mean, that's how it, it's always been. We need to not cut people a break and say that they're not responsible for it.
That's, I think one of the big, the big warning signs is when we are at that inflection point of saying, oh, well Trace made a mistake because she was using this, this thing. It was obvious that things fall. It's the tools fall now.
Right now it's humans, humans in the loop. So folks want to participate in this study. Where do they sign it?
Um, they can hit, hit up two different ways. org. org.
org. So, uh, they're able to get at either of those emails, and like I said, we have both a case study, uh, with, with the decision tool. Takes about 90 minutes to participate in this before and after research.
Um, or simply shell out the survey. Alright folks, you heard it here. Even in the age of ai, it takes a village to get it right.
So help us out. Hey Trace, thanks to being on the show. Hey, thank you so much.
All right. And back to you guys. And Steve.