What accumulates in a long run

At every step a coding agent sends the entire history so far to the model: the task, every file read, every command and every output. In a benchmark of harder tasks our agent read 2.57 million tokens and wrote 158,000. The history therefore forms the main item in memory and in the number of tokens processed, and it grows with every tool call. A model on your own hardware has a smaller window and a hard limit. Once a request exceeds the window by a single token, the server rejects it and the turn is lost.

Our context management therefore works in two tiers. The first checks at every step and trims what is old and bulky. The second kicks in at a threshold and replaces the old history with a summary. We rebuilt both several times from run logs. The program around the model, which executes tools and keeps the history, is called the harness below.

Tier one: trim without inviting repetition

The first tier replaces old tool outputs with a placeholder. The original wording was, in essence: “output removed, call the tool again to see it.” The model complied, and in a single turn the agent read the same file 55 times, because every placeholder asked it to. Since then the placeholder names the call, keeps the first 240 and the last 120 characters of the output and asks for nothing.

The second fault lay in the direction of trimming. In one log, 152 kilobytes of call arguments sat untrimmed in the history while the results of those calls had been cut to 38 kilobytes. The history thus kept what the agent had typed and discarded what it had learned. In another session call arguments filled 55.6% of the window, and 51.4% came from write and edit calls alone. Yet the content of a written file sits on disk and can be read at any time, whereas the agent merely sends the copy in the history along again at every step. Completed write calls now shrink to a note that they were applied.

The third lay in the protected range. Tier one does not trim the most recent messages, and this range was fixed at 12,000 tokens. With a window of 130,000 tokens that amounted to roughly the last three files read. A task that needs more in view at once could never hold it, and the agent read in circles. The protected range now spans a third of the window.

Tier two: summarise what happened, assign nothing

From 80,000 tokens a separate model call without tools summarises the old history. Its instruction demands a record in the past tense and explicitly forbids next steps. A summary containing a list of open items would keep demanding those items long after they are done, and the agent would do them a second time. Open work therefore lives in a separately kept task list that is re-inserted verbatim after every summary, together with the agent’s notes and the current plan step from milestone mode.

The instruction also requires findings to survive with their value: a measured number, a location with file and line, a corrected formula. A sentence like “the margin is too small” gives the agent nothing to work with after the summary, and it measures again.

Here, too, a setting was wrong. The summary kept the last 24 messages verbatim, and in a long build those came to 20,000 to 35,000 tokens. After a summary the history therefore still stood at 50,000 to 65,000 tokens and reached the threshold again about fifteen calls later. In 236 calls the summary ran fifteen times, and each time what it had just removed grew back. For local models context management now caps the verbatim remainder at an eighth of the window, so it summarises less often and more deeply.

The emergency: one token too many

That leaves the case in which neither tier acts in time. In the first longer run with a local model the server twice rejected a request of 32,769 tokens where 32,768 were allowed. Trimming ran only after sending, and a single test output of about 14,000 tokens was protected as the most recent output. Since then context management checks against the server’s hard limit before every request and, in an emergency, also trims inside the protected range, though never the newest output.

We corrected the threshold for this twice. Context management estimates the token count with a tokenizer that is not the model’s own, and the estimate measured just under ten percent too low, at 83,000 estimated against 91,000 counted by the server. With a ten percent margin below the limit the emergency fired only at about 128,000 real tokens of 130,000. The margin is now 20%.

The second correction hit the room reserved for the answer. If a request reserves room for the answer, the server requires request and reserve together to fit into the window. A turn of 270 tool calls grew to 105,425 tokens, plus 24,576 reserved, together 130,001 where 130,000 were allowed. Since then context management subtracts the reserve exactly.

Trimming costs the cache

Every cut in the middle of the history costs something that appears in no token count. The server keeps the processing of a request’s beginning for as long as that beginning stays the same; in the benchmark mentioned, 74% of the tokens read came from this cache. Rewriting old messages invalidates everything after them. For local models tier one therefore trims only if it reclaims at least an eighth of the window. If a trimming pass falls short of this threshold, tier one undoes it in full instead of leaving half-applied changes behind.

Limits

Every summary loses something, and what was lost only shows when the agent derives it again. We set the thresholds on a few long runs, mostly with one model and a window of 130,000 tokens; for other models they are starting values. The token estimate remains an estimate, and the 20% margin is paid-for space that normally stays unused. If the summary call fails, a mechanical short version takes over, marked as such. Nobody checks whether the summary itself claims something false.