Copilot to No Pilot: The Rise of Autonomous DevTools with Ray Myers at AIE 2024
Earlier this year, Cognition released the viral demo Devin, marking a new generation of AI coding tools with unprecedented autonomy. Open Source competitors quickly followed. While previous “assistants” operated simply as an auto-complete in an editor, these new tools promise to carry entire tasks through to completion! Will this save us from an avalanche of tech debt, or just create a new one?
Transcript
Hi, I am Ray Myers. com in resilience engineering area. And today I'd like to talk to you about autonomous dev tools from copilot to no and kind of untangle what many of us believe is a very, uh, interesting direction for DevOps automation and artificial intelligence as well.
So what we're gonna talk about here, we're gonna talk about the viral demo, Devon, and many subsequent developments we're gonna touch on that are in a similar direction. And we're gonna try and untangle hype versus reality. So we are gonna do a, a splash of cold water here, but I think it is all in the service of this actually being a really interesting and important direction that deserves to be taken seriously.
And, uh, what actually, once we've done that, what actually is this? How do we even think about it and, and reason about it? And there's much of this is gonna be, uh, speculative forward looking, but as you're gonna see, there's actually some tools in this direction that you could be using right now if you chose to.
And then we're gonna talk about what's next and how open source fits into this picture as, as kind of a, a vitally important part in the long term. I believe so lots to do. You may have heard that in on March 12th, cognition Labs came outta stealth mode, this startup with a series A round announcing this kind of viral demo, Devon.
Now, I think the most important part of the announcement that they shared and go over some of the other parts, uh, in a moment. But it was simply, Devon is an autonomous agent that solves engineering tasks through the use of its own shell code editor and web browser. There were a number of different pictures of what, uh, coding assistance or even a more autonomous version of that might look like before this.
But when cognition dropped, this things started to come into focus and many people kind of rallied around this. Now, if you haven't seen this briefly, this is the sort of thing that was in the series of, of demos. They, they had, uh, you would be talking to Devin as you would talk to, um, chat GBT, but it's got its own development environment.
It's able to see the results of things that have happened, make decisions in this case that, oh, it should add some debugging output in order to, you know, take the next action so forth. So this is able to see non-trivial, perhaps tasks through on its own. And that is a very interesting leap from the coding assistance we've already been been using.
Millions of people have already been, been using it in other forms, right? So what happened next was even more incredible to me because in the next, just about six weeks after that, you, you see, um, just this array of other announcements. Now you see, I've got a little notation here with a, uh, a ruler and a bar graph if it's being measured on the, uh, the current academic benchmark for this kind of tool, which is called SW Bench out of Princeton.
Uh, you, you can look that up. I'm not gonna cover that in this talk. And then, uh, the ones that are open source offerings have a little globe.
So, uh, and this isn't even everything that's happened, but this is a, a decent sampling of it that we've had in the, including Devon, since three different proprietary offerings announced in this direction. And, uh, over four different open source projects. The, some of them that I've selected and put here, and, um, some of them in the benchmarks are competitive with, with Devon, and some of them have gooeys that are similar to Devon.
And of course, just a few days ago you may have heard that, uh, GitHub copilot has entered the race as well. Now, this is changing so fast. I actually have, uh, started a website to help follow this, which is no pilot dev.
We can get to that later. So if you wanna see the latest, that's where you'd go. Clearly the amount of interest that this has generated, the amount of activity this is going somewhere, but where, where is it going?
What even direction is this taking? That's what I wanna wrestle with today. Where is it going?
Is it going somewhere good? Is it going somewhere not so good? Uh, is it like when Mickey Mouse discovered DevOps?
And if you've seen searcher's apprentice, you know that, uh, automation can be treacherous. At first, it was very nice that Mickey discovered he could use some of the wizards, uh, magic and get the broom to do some of his job for him. But what soon happened is that there were too many and they were out of control and the whole place flooded from all the, the water that was being, uh, pushed in the wrong way.
And if you've, you know, we're on the DevOps track here, right? We have experienced this with previous generations of automation, uh, the, the same DevOps versus Dev, oops, rules still apply. So certainly we, we have reason to be optimistic about what a new form of automation can do.
We have a reason to be very cautious and skeptical as well. So let's get skeptical right now. This is the remainder of the announcement and I wanna just go over Cognition's announcement as I have in more detail in a, in a YouTube video that I'll, I'll tell you about a second.
But, um, really this is to give you the cognitive tools, the mindset that you should be approaching everyone's announcement with tools like these. So, sorry to pick on the people who entered the conversation first, but we need to be very, uh, skeptical about this. So what, what have they, what have they said?
They said that the thing that I think is really important earlier that it's an autonomous agent and it has its own integrated, you know, tools to interact with its environment. That's great. They talk about the SW bench benchmark, which I think is very important.
But, um, let's just look point by point. And as I say, there's a video where I go into more depth on this, but some of the things that are being said here, it, its whole tagline is that it is the first AI software engineer. I think it is absolutely absurd to call any of the tools that have been released.
An AI software engineer that personifies them, anthropomorphize them to an absolutely hazardous degree. It, um, creates a, a, a narrative that this is actually a replacement potentially for a human engineer. What, whether or not that's what they intend to send in terms of the messaging, that is how it is bound to be received if you go around talking like this.
So I think we need something else to call 'em, which is what we're gonna get to in a moment. What can we call these things? Then, um, does talk about the SWE bench coding benchmark.
I think they could have been a little more careful here 'cause they didn't run the complete benchmark. They ran 25% of it. They were not able to run the full one.
But, uh, nonetheless, they were the first people to put a legitimate respectable score on the board for that. So I think, I think we should at least give them partial credit. It was a, it was a important advancement.
Then they say this, they say it is successfully passed practical engineering interviews. Like that's a completely meaningless claim if we're talking about a computer program program over a year ago in, in the Microsoft's research, uh, sparks of AGI I paper, we saw that G PT four prompted well could excel at a variety of professional interview exams, but it was no closer to being able to do those jobs. If GPT-4 can pass a bar exam, it doesn't mean that GPT-4 can practice law.
So that's absolutely spurious and I think a little bit deceptive to be talking about it that way. 'cause it sort of implies that it was actually being interviewed for a job that, that this is a meaningful statement and it's, it's not even if it's true. Um, and then there's another claim here.
Even completed Real jobs on Upwork. This has been under more scrutiny lately. And, and, uh, cognition as to their credit admitted they didn't actually do the task that was in the Upwork job that that demo was about.
They did something sort of related to it. They interacted with the repo that that task involved, but they didn't actually do the, the job on Upwork. So all of the things that are a little more sketchy are pointing in the direction of this being a replacement for a person, which is not, it's a dev tool.
Dev tools are great. We like dev tools, we buy dev tools, we evaluate them as we would, uh, you know, the quality of any other product we might buy. Um, we shouldn't go around talking about these things like their people.
So if you wanna see this in, in more detail, this is the video I I refer to. It's on my channel Craft versus Cru. So if it is not an AI software engineer, what do I say that it is?
So there's a few things we can meaningfully call these, the, the previous generation of AI coding helpers, right? We've been calling assistance. So that was kind of the, the baseline.
And maybe this is something more than that. Uh, the distinction I would suggest is that, uh, an assistant, and this is something like GitHub copilot, like cursor, like continue Dev is kind of a fancy auto complete. It is taking, um, about one action per user interaction with it, often by prompting, uh, an LLMA large language model, something like chat, GPT, and, uh, so it's like chat GPT is sitting in your editor and is able to kind of interact with your, uh, um, IDE or, or, you know, make updates to code, things like that.
So then, then what do we have? Well, we've been calling them coding agents, which isn't really a product name, but that's been kind of the jargon name. One of the things these are called are, are coding agents.
I would say the difference between an agent and an assistant in this context could be just called, it's instead of committing just one, uh, isolated action, it's doing a series of actions that are all pursuant to its plan to complete a, a meaningful task of some kind. And it is going about doing that in response to environment feedback that happens as it takes this or that action. That's really useful in a coding agent because it's able to try something and get maybe an error message because what it did didn't work and then already be, um, trying to recover from that error and suggest other ways rather than a human having to see every single failure that this is going through, right?
So there is a lot of potential here, given that we've seen the assistance work well, adding environment feedback. Um, I think very good move. Now I'm introducing autonomous dev tool.
It's kind of the productized version of a coding agent or another dev related agent. And so that I, I would say would be an agent that is a pilot's product and it is ready and responsible to integrate into your software development life cycle. So basically when these grow up, they're gonna be autonomous dev tools.
That's how we should start thinking about them. That's how we should be evaluating, um, you know, how well they're doing at any given time. And I think they will do quite well at the right time.
Um, just if you're interested in how these things tend to be implemented, um, you could probably go write a coding assistant if you wanted to, uh, just take an LLM that you can call such as GT four, such as anthropics models and, and so forth. Uh, or even locally host something like LAMA three, which is, um, quite good at meta and take a prompt that gives you, you know, good LLM behavior for the task in question. And, you know, you, you do your prompt design or your prompt engineering if you like to, to figure out what a good prompt is.
But basically this is kind of the same workflow of using chat GBT to help you with, with code, except that you've, uh, you've integrated it with a user interface. So getting it to be polished and, and, and nice to use, certainly very subtle, but we've kind of digested how these work agents are a little more complicated, but not that much more at the, the base level because they're an an LLM that you're calling in a loop and it can use tools in your environment. So making them, you know, in a way that works very well is still very, very subtle.
We're still learning new ways of doing it, but just to implement an agent, it's almost surprisingly easy to get one that at least does something and we'll see that in a second. Um, so you may be skeptical, you know, because of all the, um, uh, real problems that you have trying to get LLMs to reliably do work, that if you have them steering the ship for too long, they're not going to do well. They're going to, you know, hallucinate or they're gonna, you know, fail to reason properly or they're not gonna have the appropriate context.
These are all valid concerns, but the reason that agents become plausible is they can interact with their environment and, and sort of have a source to ground them away from the various wrong assumptions that the LLM might proceed to make if it were just operating on its own. Um, so it takes careful system design, but it is very possible to do this. And in particular, the, the fact that lms, uh, can use tools that is more powerful than is sometimes realized because you can give them safe tools.
If you were to go to the extent of giving an LLMA tool that no matter how bad your input is to it, it will not actually hurt anything. It it will be at worst case neutral than you've pretty much solved the problem, right? Like that that now mitigates the fact that LLM output is sometimes useful but unreliable.
Uh, what might that look like? So here's a, a very constrained example, and this is one that I've actually used on production code successfully at Indeed, um, adding type hints to legacy Python. We have a bunch of old Python code, for instance, and we would like greater type safety.
We'd like to start running my PI and, and, uh, make sure type errors, uh, don't exist in the, in the code and aren't introduced going forward. Well, it's very toilsome to introduce all, all these hints, right? That's one of the bigger reasons people don't do this more often is it's just, it's just a pain.
Um, but one way not the only way that you can, uh, automatically get type hints to be generated would be to call an LLM that does work fairly well. Not always, crucially, you must handle the cases where it's not. But we can, in the case of, of uh, type errors, because if it's introduced a bad type, my pi will catch it.
The exact, uh, mechanism we wanna use to check our types. So by simply writing a loop and a and an if condition cleverly, we can say, oh, we've got, um, an update here. We'd like to add a type hint we will use, we won't rewrite the code directly out of the LLM output because LMS aren't so reliable at that always.
But we will use something else to incorporate in this case, tree sitter. Um, you can use something to directly manipulate the syntax tree of the language in a safer way, and then having added those type hints at the syntactically appropriate place, run my pi. And if it fails, just don't, don't keep that, only keep the ones that validate and this is quite reliable.
And even if it were to, um, end up with a bad hint, at the end of the day, they don't even impact the runtime behavior of the program. So on, on a lot of levels, this is an almost completely safe way to, um, to use a coding agent on legacy, uh, code to improve the safety of it. And this whole thing without even code golfing was only 400 lines.
The Python script that does this. So there is at least, you know, there's an existence proof in a very small amount of code, you can implement a, uh, an autonomous dove tool workflow that that is safe. Uh, if you wanna do things that are more aggressive than, um, adding type ins, obviously you'd have to come up with more ways to, uh, to check between automated feedback and human feedback.
But those are essentially the types of games that you'd wanna play. Um, so promise to get to where open source fits into this. Um, many of us predict that autonomous dev tools will become a very important part of the software lifecycle, uh, on par with other kinds of automation such as, uh, inform management or, you know, uh, maybe some say to the extent of that a compiler is a core part, maybe not that far, but it will be some, uh, very important part of the tool chain.
And you may have noticed that there are strong open source offerings for every important part of the software tool chain. Every com, uh, language, every, you know, um, top language has a compiler that is open source, uh, there for every, uh, of the popular types of databases. There's a strong offering that is open source.
There's room for proprietary offerings as well, of course, but there, um, we've kind of arrived at the equilibrium where there needs to, at any given time be an open source option. This is in the collective interest of the industry. So we believe that autonomous dev tools will become, uh, a member of that club sooner or later.
And, um, perhaps the sooner the better because it will help steer the, um, development of these tools in a more safe, uh, productive direction. dev, where I'm following the, the progress of this category of tool and, uh, you know, sort of collecting the, the ideas about, about how to build them well. So if you want to know more that's a little more up to the minute than whenever you're watching that, you might check that out.
Uh, remember this is a new wave of automation, not magic, but we've seen automation be quite important and it doesn't replace the discipline of software engineering. It will be effective to the extent that it aids in the successful execution of software engineering discipline. And, and I think that's how we should be approaching using this.
Uh, you should look at your use cases and see what is safe and what is valuable in your context. Um, and above all, don't pretend programs are people because they're not. Thank you very much.
I look forward to your questions.