1,508 lines, not a single run

A build by our coding agent wrote 1,508 lines across fourteen files in 25 minutes, called the shell exactly once, before any code existed, and reported that it was done. Two modules imported relatively beyond the package, so four of the eight main modules failed to load and the program’s entry point was dead. A single import command would have shown this at any time.

The loop noticed none of it. The agent had a shell, the brief asked it to report once the tests pass, and the turn ended without an error. A language model writes the sentence “done, everything works” as fluently as any other, so something other than the model has to decide whether it is true.

The rule: code that was written must have run

The first version of the rule deliberately asks for little. Writing code sets a flag, and running anything afterwards clears it. If a turn tries to end while the flag is set, the harness sends it back once to run what it wrote. The rule asks for neither good tests nor passing ones, only that the code has run at least once. That would have sufficed for the dead build, because the very first import would have failed.

A cap on consecutive writes followed. Another build wrote ten files in a row, the whole product including its tests, before a single line ran, and 26 of 39 writes sat in streaks longer than three. Since then the harness refuses further writes after three files written without a run, until the agent has executed something.

Two holes, both measured

This rule had two holes, and we found both only in the logs of the runs.

Every shell call opened the first, because even ls -la cleared the flag. Ten files that had never been executed counted as checked because the agent had looked at them. Since then only a command that executes code counts.

The second went deeper. The harness did not read the result of the run, so a red run cleared the flag just as a green one did. A quirk of the shell hid the fault: in one of the builds all eleven test invocations had the form pytest -q 2>&1 | tail -N, and without pipefail such a line reports the success of tail, not that of pytest. At the only place the loop looked, a red test suite could not be told from a green one. Since then the harness reads the result of every run and executes command lines with pipefail, so the status of pytest survives the pipe.

As a third tightening, the last run now counts and not the best one. If an agent fixes a bug, sees “15 passed”, then changes something else and gets “1 failed”, it has nothing verified any more. A red run after a green one sets the flag again, and the note at the end of the turn quotes the red output.

In long builds the plan becomes state

That is not enough for a build that runs for many hours. A longer build with a local model reached the third of thirteen steps in 236 calls and about six hours. Along the way the history was compacted fifteen times, and after every compaction the agent re-derived from a summary where it stood, read the plan again and repeated its exploration. The loop recorded nowhere that the agent was at step 3 or which command checks that step.

Since then the harness keeps the plan as state, where the agent previously only read it as text. Every step forms a section in a file, and the commands that prove it sit in a code block beneath it. Such a step looks like this:

PLAN.md · example
## Step 3 — Write-ahead log (WAL)

Every write lands in the log first; after a crash, start-up restores
the last state from it.

```bash
python -m unittest tests.test_wal -v
python -m kv.cli selftest --crash-recovery
```

The harness reads the steps and their gate commands from this file and records in a file of its own which ones are done, when and at which commit. The agent gets a tool with three actions for this:

milestone
milestone(action="status")
    current step, its scope, its gate commands

milestone(action="done", step=N)
    accepted only if every gate command of step N ran green
    in this turn after the last file write

milestone(action="reset", step=N)
    reopen a step, e.g. because a later one broke it

The middle action closes a step. The harness accepts it only if it finds every gate command green in its own log, and only after the last write. The loop thereby checks completion as a claim and no longer has to believe it as a sentence. After every compaction of the history the harness injects the current step and its gate commands again, so the agent does not have to re-derive where it stands.

Our harder benchmark includes a task of its own for this, a key-value store in five plan steps. The agent passed it with the local model on the first run, after 38 model calls and 509 seconds; that is a single run.

Limits

The rule checks that something ran and what it returned. It does not check whether the tests are any good, because an agent that writes a test that always passes satisfies it. A person or the agent itself writes the gate commands of a step into the plan, and a command that is too weak makes the step cheap. A gate command counts as run if it appears in the log as part of a command line; embedding it in a longer line that swallows a failure, for instance with a trailing || true, defeats the comparison. What the rule is worth on an unfamiliar project without tests depends on the plan alone.