Agentic development

Adopting agentic development on a legacy codebase

Not enough test automation, an architecture nobody would defend, and a queue of advice telling you to build feedback loops, CI/CD, test automation, guidance files, skills and model routing. All of it is right eventually. Almost none of it pays this month.

Four phases, in the order that pays. Each one says what to do, why that comes now rather than later, and what you are allowed to ignore until it does. Each ends at a gate that either happened or it did not, so you can date it and you cannot argue with yourself about whether you are through.

Where are you, for the code you are about to touch

  • I use it for the boring parts, but I still write anything that matters. Phase 1 →
  • I spend more time reviewing its code than I used to spend writing it. Phase 2 →
  • It finishes faster than I can decide what it should do next. Phase 3 →
  • I have four running and I have lost track of what is in which one. Phase 4 →

The map

Down a column for what to do now. Across a row for how one strand changes. Click any cell.

1 Hand over the typing You read every line it writes 2 Hand over the checking You build the tests that say no 3 Hand over the feature It builds the change. You drive. 4 Hand over the loop One command, several at once THE STORY everything below follows from this You hand over the coding but keep the review, and update the guidance to make the coding better. You move on when the code written for you is mostly right. Reviewing has become the bottleneck. To go faster you have to stop reviewing yourself, which needs a harness you can trust. Building that is this phase. The reviewing bottleneck is gone, so an engineer has more time to build and does more at once. Product now has to scope the right thing, faster. Scoping, coding and reviewing can all be trusted now. What is left is to automate the whole process: hand over the loop, and run even more at once. WHAT TO DO in this order 1Hand over coding, easier areas first 2Review everything 3Correct defects, not your preferences 1Write your test levels and the rule 2Write tests for the code you touch 3Monitor and optimise your harness 1Write tickets in product language 2Size the work, let a model split it 3Build a command to help you scope 4Run 3+ tracks in parallel locally 1One command, ticket to merged 2Scale the same setup to 5 or 10 3Make the product strategy hold ARCHITECTURE as far as the harness can catch you Define your technical design. Make the boundaries you care about fail the build. No big refactors yet. Small local ones, only to make that area's tests cheaper. Bigger refactors are possible now. The harness can catch them. Nothing new in this phase. FEEDBACK LOOP what you write down, and whether it worked Set up a session review skill, to turn each correction into a rule, and keep the markdown files and the design current. Set up a periodic review skill: the agent learns from the fixes made after the merge. Harness review skill: manage the cost. Extend the periodic review skill to count how often a feature changed after it shipped. Bugs, design, scope. Watch what parallel work does to your process and your architecture. Both will crack. Pruning guidance: still open. PROCESS AND ROLES the bottleneck moves one step each phase Tickets, ceremonies and estimation stay as they are. One thing starts: product builds POCs to scope with. QA moves to acceptance testing, what must be proven, and harness efficiency. Less structural testing. Product builds POCs on the product. The PM stops writing an outcome and scopes the work fully. Engineers may do it instead. Same scoping skill, either way. The product strategy has to be solid. Much more, much earlier. Development got cheaper. Being wrong did not. THE GATE it happened, or it did not THROUGH WHEN The code it writes is mostly right, 80 to 90% THROUGH WHEN For one area, you trust the tests over your own read THROUGH WHEN A ticket goes to merged with no questions asked No gate. As far as this map goes, for now. Phase 4 is not better than phase 3. These are levels of autonomy, not quality, and a team may stop on purpose where its risk model needs a human gate.

Scroll the map sideways to see all four phases.

One product, four speedsexample

You do not finish a phase for the whole codebase. You finish it one area at a time.

PHASE 1 PHASE 2 PHASE 3 PHASE 4 Payments API NOW Checkout NOW Search NOW The old admin NOW A change that touches Payments and the old admin is a phase 1 change. It sits in the lowest phase of any area it reaches.

You go through each gate once per area, not once for the product. What moves per area is coverage, and therefore how much you trust the harness where a change lands. The plumbing does not: your pipeline, your test runner, your guidance file and your build command are one thing for everybody, which is why the first three steps of phase 2 happen once and the fourth happens forever.

A change sits in the lowest phase of any area it touches. Reach into phase 1 code and you handle the whole change like phase 1, however good the rest of it is. That rule is what makes running at different speeds workable instead of wishful, and it is computable from the diff rather than a judgement call.

This holds without service boundaries. A monolith has areas too; they are just not drawn in the folder structure. They are the clusters of code you keep touching together, and you already know what they are. A service boundary makes them visible, its absence makes them invisible, not absent. What a monolith really costs you is the rule above: changes cross areas more often, so more of them fall back to the lowest phase, so the whole thing moves slower on the same mechanism.

And the way to cheat at this. Advance the area you change most, not the one that is easiest to advance. A phase 3 area you touch twice a year is the wrong thing to have optimised, and it will make your progress look better than it is.

1

Hand over the typing

You read every line it writes.

You hand over the coding but keep the review, and update the guidance to make the coding better. You move on when the code written for you is mostly right.

You are here if

You use the agent for pieces of work and paste or adapt what comes back. You run one or two sessions at a time. You decide everything, read everything, and run everything yourself.

What is keeping you here

You still write a lot of the code, because you do not trust the agent. The parts you keep are the hard ones. Those are exactly the parts where it would have got things wrong, so you never see it get them wrong, and you learn nothing about what it is bad at.

Your runs are too short to leave you idle. Work you hand over in one prompt comes back in under a minute. There is no gap to fill, so a second session would not help. If two feels like your ceiling, that is usually why.

The idea behind this phase

You hand over the typing and you keep the review. You are not trying to go faster yet. You are finding out how your agent fails on your code, and you cannot learn that while you still write the hard parts yourself.

Watching is not enough. If you keep making the same correction week after week, you will never trust the agent enough to stop checking its work, and you will not leave this phase. So you do two things at once: hand the code over, and start writing down what you correct so it stops coming back.

You are done when the code it writes is mostly right. Not perfect.

What to do

  1. Hand over the coding, easier areas first. Start with the tests. They are the safest thing to hand over, because they are easy to judge and they show you what the agent misunderstood. Then move to code, beginning in the areas you know best. You are here to judge whether it got the change right, and you cannot do that in code you do not understand yourself.
  2. Review everything. Every line, every change, for as long as this phase lasts. It is the only phase where that is the right thing to do.
  3. Fix defects and design. Leave the style alone. The agent writes code differently from you. That is not a problem, and correcting it is expensive twice over: it costs you the time now, and anything you write down becomes a rule it carries on every future request. So push back on two things only. Something is broken or will break. Or the design is wrong: the change sits in the wrong place, or the shape of it is wrong. How the code is phrased is not your business any more. What earns a place in writing is Guidance.
  4. Define your technical design. Which services exist, which boundaries matter, what lives where, which pattern is current and which one you are leaving. The agent can work most of it out by reading the code, but it pays for that in tokens every session and still guesses wrong at the edges. See DO control the technical design.
  5. Make the boundaries you care about fail the build. Take two or three you would be angry to see crossed and write each as a lint rule on imports. It is the only check you can have in a phase with no tests yet, and it catches two of the three things an agent gets wrong on a legacy codebase: changing more than it was asked to, and putting the change in the wrong place. Your architecture is probably not the one you want, so record the violations that already exist and fail only on new ones. You are not claiming the structure is right. You are claiming it will not get worse.
  6. Set up a session review. At the end of a working session, read the conversation back for every correction you gave, work out which rule would have prevented it, and write that rule down. It keeps your markdown files and your technical design current, so the next session starts from what the last one learned. Cockpit runs it as a skill: session-review. See DO catch a mistake when the session ends.
  7. Let the product side start building proofs of concept. Your tickets, ceremonies and estimation all stay as they are in this phase. This is the one exception. A question about what to build, answered by something you can click, is worth a week of discussion.

Through when

The code it writes is mostly right. Say 80 to 90%.

Waiting for perfect keeps you here for nothing. The last ten or twenty per cent is what phase 2 is for: a harness catches it more reliably than more correcting ever will.

OpenThat last claim is a theory rather than something measured, and it is the assumption the whole sequence rests on.

Back to the map
2

Hand over the checking

You build the tests that say no.

Reviewing has become the bottleneck. To go faster you have to stop reviewing yourself, which needs a harness you can trust. Building that is this phase.

You are here if

The agent writes most of your code and it is mostly right. You read every change before it merges, and that is now where your day goes. You would like to run more work at once and you cannot, because you are the only thing standing between a change and production.

What is keeping you here

You wrote the rules down instead of building tests. This is the common one. A markdown file does not stop anything: nothing fails, nothing goes red, and the agent can ignore it. So you are still the only thing catching mistakes, and it feels like progress because the file keeps growing.

Your tests are stuck at the expensive level. Without a clean architecture you cannot test a small piece on its own, so everything has to go through the browser or the whole system. Those tests are slow. A slow suite is one people skip, and a skipped suite catches nothing.

You want to clean up the code first. You cannot yet. A refactor with no tests around it is a change nobody can check, and an agent doing it does not make that safer.

The idea behind this phase

You are building the one thing that lets work leave your desk: something other than your own eyes that can say a change is wrong.

You are not building a test suite. You are building something you will trust. That is a higher bar. It has to catch what you would have caught, it has to be fast enough that nobody skips it, and you have to be able to see what it costs, or in six months you will have a forty-minute pipeline and no idea which part of it is worth having.

Which sets the order. Decide the levels before the agent writes a single test, because it will write thousands at whatever level is easiest and you pay for each one on every run. Then build, one area at a time, for months. Then watch what it costs you, because it only ever gets slower on its own.

What to do

  1. Write your test levels, the rule for choosing one, and what the suite may cost. One file, which the agent reads before it writes a test. Name the levels. Say what each one may reach, as a table rather than a judgement call. Give one rule for picking between them. Put the run time you are willing to pay in the same file, so the budget is in front of the agent when it decides. Cockpit's testing strategy is the worked example. See DO decide the test levels yourself and DO write down what the suite is allowed to cost.
  2. Write tests for the code you touch. You could ask an agent to write tests for the whole codebase this month, and you should not. Tests written against code nobody is changing only record what it does today, bugs included, and nobody reads ten thousand of them closely enough to notice. The tests you end up trusting are the ones written while you were changing that code and could still tell whether they were right. So they arrive one area at a time, as part of the change to that area. At first the only ones you can write go round the outside, through the browser or the API, because the code will not let you in lower. Write them anyway. Never keep one you have not seen fail: see DON'T trust a test you haven't seen fail.
  3. Monitor and optimise your harness. It only ever gets slower and more thorough on its own, and a harness you cannot afford is one people skip. You need four answers on demand, with no reporting stack behind them: how long a change takes from first commit to merged, how long CI takes and which job is slowest, what a change costs in CI minutes and tokens, and how often a check fails for a reason that is not the change. Then run the cheapest thing that can safely judge a change, and do the same with reviews instead of one depth over everything. I left this late. My pipeline went from under a minute to eleven over three weeks, on 246 runs in a single week, and I still cannot tell you how much of that was worth paying for. See DO build a lead time report, DO publish the pass rate of every check and DO let the change decide how deep the review goes.
  4. No big refactors. Small local ones, only to make an area's tests cheaper. Once a slow test is around an area, you will have to change the code so a faster test can reach the logic. That is a refactor, and its purpose is the speed of your tests rather than the quality of your design. The limit: the test you just wrote has to be able to catch it going wrong. See DO let agents fix old architecture, behind tests that can catch a broken refactor.
  5. Set up a periodic review, so the agent learns from the fixes made after the merge. Corrections you make during a session already reach your guidance. What never reaches it is everything you fixed a week later. Every few weeks, read what merged, the review threads and the closed issues, and file one issue for each thing that happened at least twice. Cockpit runs it as a skill: periodic-review. See DO spend a session every few weeks reading your recent work.
  6. Set up a harness review, so you can manage what your harness costs. Nothing removes a check unless somebody goes looking. Take each required check, ask what it has actually blocked and what it costs to run, and file one candidate for each one that is not earning it. I ran a remote security review on every pull request for weeks: it found nothing on 30 of 31 runs, after an earlier sample of 39 that also found nothing. Cockpit runs it as a skill: harness-cost-review.
  7. Move QA, and do not cut it. Away from structural testing, towards acceptance testing, deciding what has to be proven, and keeping the harness efficient. This is the phase where that judgement is worth the most, and the phase where most companies cut it because the agent writes the tests now. See DO agree the product's rules before the work starts.
  8. Let the product side build proofs of concept on the real product. In phase 1 those were throwaway. Now there is a harness, so a rough version can go into the actual codebase and the tests will say whether it broke anything. That is the first thing the harness buys anyone outside engineering, and it arrives long before the loop does.

Both review skills arrived on day 33 of my own project, after the loop was already running, and my own cycles still mostly watch the harness rather than the guidance. This is something I learned rather than something I did in the right order.

Through when

For one area, you trust the tests more than your own reading of the diff.

Not when the suite is green, and not when coverage hits a number. The test is whether you have stopped double-checking a kind of mistake yourself, because the harness keeps catching it first.

For one area, not for the product. You pass this gate once per part of your codebase, and the first time will be in whatever you touch most.

Back to the map
3

Hand over the feature

It builds the change. You drive the process.

The reviewing bottleneck is gone, so an engineer has more time to build and does more at once. Product now has to scope the right thing, faster.

You are here if

Your tests catch things before you do. You hand over whole changes now, but you scope each one yourself, start it yourself, and remember the review yourself. Two or three at a time at most.

What is keeping you here

You answer the agent's questions while it works. It asks what you meant, you tell it, and the change comes out fine. So you never find out that the ticket did not contain the answer, and the next ticket does not contain it either.

Scoping takes you about as long as building takes the agent. If a change takes it forty minutes and a ticket takes you thirty, you cannot keep two running, whatever your tooling does.

The idea behind this phase

Reviewing is no longer the bottleneck, so an engineer has time they did not have and starts running more than one thing at once. Both of those land in the same place: the work has to be defined faster, and better.

A ticket written for a person is not a ticket an agent can build from. A person reads between the lines, notices what you forgot, and asks. An agent decides for itself, and you find out at the end. So defining the work stops being paperwork in front of the real work and becomes the thing that decides what comes out of it.

Get this right before you automate it. A loop on top of good tickets gives you leverage. A loop on top of vague ones gives you merged mistakes, faster. That is why the build command waits for phase 4.

What to do

  1. Build a command that helps you scope. You are about to write far more tickets than before, and the shape of a good one is not obvious. A skill that asks the questions in order, decides whether the work has to be seen before it is built, sizes the slice and writes out the rules the change has to satisfy gives every ticket the same shape, whoever wrote it. Cockpit's is scoping, and it runs before any code exists on every piece of work.
  2. Write tickets in product language, not implementation. That is what keeps someone who has never opened the codebase out of the how, and still lets them start the work. See DO write the ticket in product language.
  3. Size the work, and let a model split it. On a legacy codebase keep it smaller than feels necessary. The smaller the change, the less there is for the agent to get wrong about where things belong. See DO set the limit on how big a piece of work can be.
  4. Run three or more tracks in parallel locally. A worktree per track, each with its own branch, ports and environment, is the usual answer, but the worktree is only the mechanism. The requirement is that nothing is shared: two agents in one checkout is the first thing everyone tries and the first thing that breaks. See DO give every track its own worktree, branch and environment.
  5. Take on the bigger refactors you have been putting off. The harness could not catch them in phase 2 and it can now, because you have been adding tests to every area you touched ever since. The limit moves from one seam to whatever your tests cover. See DO let agents fix old architecture.
  6. Extend the periodic review to count rework. How often a feature had to change after it shipped, through bug fixes, design changes or scope changes. That is your scoping quality, and it is the number nobody tracks. A ticket that needed no questions and then three rounds of rework was not a good ticket. See DO spend a session every few weeks reading your recent work.
  7. The product manager scopes the work, not just the outcome. Not a ticket saying what should be true, but the whole piece of work shaped through the same scoping skill an engineer would use. Where it makes sense an engineer does it instead. What matters is that both use the same skill, so the result is the same shape whoever produced it.

Through when

A ticket goes from written to merged without the agent having to ask you anything.

Not once by luck. Routinely, on ordinary work. It is the precondition for phase 4, because an agent running unattended cannot come and ask.

Back to the map
4

Hand over the loop

One command, several at once.

Scoping, coding and reviewing can all be trusted now. What is left is to automate the whole process: hand over the loop, and run even more at once.

You are here if

You write a ticket and the agent builds it without coming back to ask. You are still the one starting each change, checking on it, and remembering to run the review.

What is keeping you here

You run every change through the steps yourself. Scope it, start the build, come back to see where it got to, run the review, merge it. None of those steps is hard. There are just five of them on every change, and they all need you. That works at two changes at once and falls apart at four, and what fails is that you forget a step on the one that mattered.

Your setup was built for three tracks. Ports, databases, preview environments and CI capacity are all sized for what you do today, and each one becomes the constraint in turn as the number goes up.

The idea behind this phase

Scoping, coding and reviewing can all be trusted now, so the last thing you still do by hand is the sequencing. Hand that over too, then run more at once than you could ever have watched.

What breaks is never the agents. It is the infrastructure underneath them, and it fails in ways that look like the agent misbehaving: a port collision looks like a flaky test, a shared database looks like a bug that will not reproduce. Most of this phase is that work.

And development getting cheaper does not make being wrong cheaper. At this rate a vague product direction stops being survivable, because you now build the wrong thing faster than you used to build the right one, and every check will pass while you do it.

What to do

  1. Build one command that takes a ticket all the way to merged. It runs the steps you have been running by hand: scope, build, test, review, ship. Do not restate your rules inside it. See DO build one skill that takes a ticket all the way to merged and DON'T restate your rules inside the command.
  2. Scale the same setup from three tracks to five or ten. What you built by hand for three does not survive ten. This is infrastructure work and it is most of the phase. You will also need a way to ask where each session stands without reading back through it: see DO ask each session where it stands.
  3. Watch what parallel work does to your process and your architecture. Both will show cracks, and the cracks are the roadmap. No automated review will tell you the team is stuck. That one you have to look for. See DON'T expect any automated review to tell you the team is stuck. OpenWhat comes out of the guidance file, which has been growing for months and is read on every request, is still unanswered.
  4. Make the product strategy hold. A weak one starts to hurt here, and it is worth treating as a target rather than a consequence. Far more has to be decided and prepared ahead of the engineers, because what they get through has gone up sharply. Development got cheaper. Being wrong did not. See Roles.

No gate

This is as far as the map goes. Not because there is nothing after it, but because I have not run far enough past this point to say what the next constraint turns out to be.

Back to the map

These phases describe autonomy, not engineering quality. Phase 4 is not better than phase 3, and nobody has to reach it. A team building medical devices, banking infrastructure or anything where being wrong is expensive may deliberately keep a human gate forever, and be right to. Stopping on purpose is a different thing from stalling, and the difference is whether you can say which risk the gate is there for.

The same holds per change rather than per team. A CSS fix can go all the way while a schema migration stays behind a human gate in the same week. Move each class of work as far as its blast radius allows, and no further.

The sequence is one I ran on a new codebase. The parts specific to legacy codebases are what I saw on one of them over nine months, plus reasoning. Anything marked Open is a question I do not have an answer to yet.