An agent that reads, changes and runs code
An agent with shell access needs protection that does not depend on the model. Here permissions, sandbox and undo sit as code between model and operating system.
The agent works in a loop of reading, editing and running. Three mechanisms bound every step: a permission engine that denies, asks or allows in a fixed order, a sandbox for every shell command and a snapshot before every change.
It runs on cloud models as well as on an open model on our own GPU server. In local operation no request goes to a model provider, so it can also work on highly sensitive source code and confidential data. In a one-hour run on its own source code every request went to our own graphics card, and the cost journal showed zero. The run in detail →
- In useinternally, among other things to build TippelPi
- Tests2,839 passed, 22 skipped
- Suite runtime3 min 32 s, no model call
- SandboxSeatbelt · bubblewrap · Docker
- Local operationQwen3.8-27B · vLLM · one RTX 5090
- As of9 October 2026
The suite runs without a model call. It checks the code between model and operating system, meaning permissions, sandbox and undo, and says nothing about the quality of individual model answers. The sandbox backends bubblewrap (Linux) and Docker have not yet been proven in continuous use, and tracing across nested agent runs is missing.
$ python -m pytest -q 2839 passed, 22 skipped, 1 warning in 211.62s (0:03:31)
From an empty project: a CSV explorer
One prompt, no starter code. The agent builds a tool that runs entirely in the browser: drop a file, sort and filter the table, a profile and chart for every column, export. Parsing and statistics live in tested modules. At the end the recording walks through the tool: sorting, stacking two filters, exporting the filtered rows, loading a file of your own. Every step is a real click and is checked while filming. That is how the first walk found a bug the tests had missed: the filter field lost focus after one character. The agent fixed it in a second session and added a test for it.
Nothing is staged: the agent did the work, and the terminal scene shows the real output of the commands in the repository. The terminal window itself is drawn for the recording.
Notes on the coding agent
- An agent improves itself: one hour on one graphics card6 min read
- Our coding agent designs a humanoid robot in 24 days5 min read
- Context that does not overflow7 min read
- Speculative decoding: from 24 to 100 tokens per second6 min read
- When a coding agent may say done7 min read
- Open models: months behind the frontier, absent from the regulation debate9 min read
- Running Coding Agents Safely7 min read
- Detection Engineering with LLMs8 min read
More projects