The setup
Our coding agent has a mode in which it edits its own source code. It picks a weakness, changes the code and runs the tests that belong to the change. If they pass, the harness keeps the change and the next iteration builds on it. If a test fails, the harness discards the change. If the agent changes nothing, the log marks the iteration as empty and does not count it as a success.
In September 2026 this mode ran with Qwen3.8-27B in the four-bit NVFP4 version, served by vLLM on an RTX 5090. The context window held 32,768 tokens at the time, and the model wrote about 24 tokens per second. No request left our own card; the cost journal shows zero for all runs, not counting power and the card.
The runs
| Run | Outcome | Duration |
|---|---|---|
| Iteration 1 | kept: the first change written by the agent, two files | 119 s |
| Iteration 2 | kept: guards against an empty search string in the edit tools, with tests | 198 s |
| 60 minutes, 16 iterations | 13 kept, 2 without a change, 1 discarded | 13 to 542 s per iteration |
| Run after a fix to the harness | refused: the working snapshot was stale | — |
| 4 further iterations | 2 kept, 2 without a change | 10 to 117 s |
The experiment is the hour with 16 iterations; the rows before and after it show the lead-up and the follow-up. The 13 kept iterations added about 248 new lines, roughly 190 of them tests. The discarded iteration broke an existing test, and the harness rolled it back without anyone stepping in.
What the agent improved in itself
The agent mostly reworked error messages from which it works out its own next step. A provider error now separates “too many requests” from “server fault”, because the first calls for waiting and the second for a retry. Search says what is wrong with a malformed regular expression. If the file to be read does not exist, the message adds a hint at what to try. If listing a directory hits a file, the message points to the read tool. The agent secured each of these changes with a test that pins the old failure.
A developer rewrites messages like these after the same unclear output has stopped them three times. Here the model found the spots itself, changed them and wrote the tests, without a request leaving the building.
Three faults in the harness
The run exposed three faults, and none of them lay with the model. All three sat in the harness, the program around the model that executes tools, starts tests and manages changes.
The first was an overflow. Twice the server rejected a request because it held 32,769 tokens where 32,768 were allowed. The harness compacted the history only after sending and exempted the most recent output, so a single test output of about 14,000 tokens burst the window. Since then context management knows the hard limit and, in an emergency, also trims large outputs inside the protected range, though never the newest one. In the last run the error no longer occurred; Context that does not overflow describes the details.
The second hit the step that keeps the work. A passed iteration never reached the main branch, because the harness tried to merge a branch into itself and tripped over uncommitted leftovers. The finished change stayed sitting next to the main branch.
The third sat in project memory. The file with the project notes had grown to about 30 kilobytes, and the harness loaded it in full at every start, into a window of 32,768 tokens. The harness now caps it at 8,000 characters.
The “refused” row in the table also traces back to the harness. After the overflow fix it did not start the next run, because its working copy did not yet contain the new code. The log recorded the abort, and the fault was fixed before the last run.
Limits
The full test suite takes a good five minutes and would have eaten the hour if run in every iteration. Each iteration therefore ran only the targeted tests, and the full suite ran on the final state, with 1,785 tests passing. If a change breaks a test outside its own neighbourhood, that shows up only then. We measured on the agent’s own project, which it knows well and whose dense test suite checks every change. An unfamiliar project without tests lacks that check, and the number of kept iterations then says nothing. The changes were small; whether the same loop holds up for larger refactorings we have not measured. Since this run the same card works with a larger window and faster, as the article on speculative decoding shows.