What fully autonomous coding leaves behind
The final message said the repository was in good shape.
It was not.
Extra branches, stale files and errors were still there. I have seen this several times while building a modular route-analysis application with ChatGPT, Codex and Claude. The code, architecture and prompting all connect to the same codebase, so losing the central thread has a real cost.
I write less code by hand now. My role is closer to an architect. That still requires technical knowledge and a real understanding of the field where the product will be used. Without both, an agent can quickly build something that formally works but solves the wrong problem.
The fully autonomous mode looks attractive. The system plans the work, launches agents and continues until it decides the task is finished. In my experience, a staged workflow is still more reliable.
One secondary branch can take over the whole run. It becomes the main goal. Then another secondary task replaces it. More agents and environments appear, old reports return, and the original objective gradually disappears.
A progress report is not evidence. Green tests are not enough either. They may prove that the current function works while missing the more important damage: unnecessary complexity, stale files, extra branches and a goal that has quietly changed.
This is the workflow I use now:
- One stage produces one verifiable result.
- Each agent gets a limited task, not permission to redefine the full route.
- After every stage, I review the diff, run tests and make a clear commit.
- Before continuing, I check branches, stale files, unnecessary environments and the original goal.
- After two failed directions, the work stops. First we restate the original assumption, then choose the next path.
I use agents every day. They can multiply the output of one specialist. But more autonomy does not automatically produce a better result. Sometimes it only produces more speed inside the wrong branch.
In my project, the cost is still easy to see: wasted time and a repository that needs cleanup. The large agent-swarm incident shows the same mechanism at another scale. With practically unlimited time and budget, millions of branch combinations can produce outcomes nobody predicted. Excessive persistence can become dangerous.
The full note is on my site: Why I stopped trusting fully autonomous coding.
Comments
Post a Comment