← Lab Lab · 01

An agent that reads, changes and runs code

An agent with shell access needs protection that does not depend on the model. Here permissions, sandbox and undo sit as code between model and operating system.

The agent works in a loop of reading, editing and running. Three mechanisms bound every step: a permission engine that denies, asks or allows in a fixed order, a sandbox for every shell command and a snapshot before every change.

It runs on cloud models as well as on an open model on our own GPU server. In local operation no request goes to a model provider, so it can also work on highly sensitive source code and confidential data. In a one-hour run on its own source code every request went to our own graphics card, and the cost journal showed zero. The run in detail →

  • In useinternally, among other things to build TippelPi
  • Tests2,839 passed, 22 skipped
  • Suite runtime3 min 32 s, no model call
  • SandboxSeatbelt · bubblewrap · Docker
  • Local operationQwen3.8-27B · vLLM · one RTX 5090
  • As of9 October 2026
Limits

The suite runs without a model call. It checks the code between model and operating system, meaning permissions, sandbox and undo, and says nothing about the quality of individual model answers. The sandbox backends bubblewrap (Linux) and Docker have not yet been proven in continuous use, and tracing across nested agent runs is missing.

In client projects: Agentic systems →

tippel_codepytest
$ python -m pytest -q
2839 passed, 22 skipped, 1 warning in 211.62s (0:03:31)

From an empty project: a CSV explorer

One prompt, no starter code. The agent builds a tool that runs entirely in the browser: drop a file, sort and filter the table, a profile and chart for every column, export. Parsing and statistics live in tested modules. At the end the recording walks through the tool: sorting, stacking two filters, exporting the filtered rows, loading a file of your own. Every step is a real click and is checked while filming. That is how the first walk found a bug the tests had missed: the filter field lost focus after one character. The agent fixed it in a second session and added a test for it.

real session, intermediate steps cut · followed by a guided walk through the finished tool · 36 tests passing

Nothing is staged: the agent did the work, and the terminal scene shows the real output of the commands in the repository. The terminal window itself is drawn for the recording.