A throwaway Docker sandbox for agents.
Run code or shell commands in a resource-limited, network-isolated, single-use container and get back a structured result β stdout, stderr, exit_code, timed_out, duration_ms. So an agent can verify its work, reproduce a bug, or check output without ever touching the host.
Part of tools-for-agents. Zero npm dependencies β drives the docker CLI directly.
Every run is launched with:
--network noneβ no network unless you opt in (network: "on")--memory 512m --cpus 1 --pids-limit 512β bounded resources--cap-drop ALL --security-opt no-new-privilegesβ minimal privileges--rmephemeral container, work dir is a fresh temp mount, removed after- a hard timeout (default 30s, max 300s) that
docker kills the container - optional
secure: trueβ read-only rootfs + tmpfs/tmp
node src/cli.js check # docker version + presets
node src/cli.js run python 'print(2**10)' # β 1024
node src/cli.js run node 'console.log(6*7)' # β 42
echo 'print("hi")' | node src/cli.js run python -
node src/cli.js sh 'apk info 2>/dev/null | head'anvil is stateless by default. Set ANVIL_DB to record every run, then browse them in a dashboard:
export ANVIL_DB=./.anvil/runs.db # opt in to the run log
node src/cli.js run python 'print(2**10)' # β¦runs are now recorded
node src/cli.js serve # β http://localhost:7930 (--port to change)A zero-dependency forge log of what the sandbox executed:
- Run list β every execution with its language, status (ok / failed / timed out), exit code and duration, colour-coded at a glance.
- Run detail β the exact code or command, full stdout and stderr, and the resource limits it ran under (network, memory, cpus, timeout). Each block has a one-click β§ copy so you can lift the code or its output straight to the clipboard.
- Compare two runs β hit β compare, pick any two runs, and see a side-by-side line diff of their code, stdout and stderr, plus every limit that changed (
took 258ms β 559ms,mem 512m β 256m). - New run β hit οΌ new run to type a snippet (bash / node / python), execute it in a fresh sandbox right from the dashboard, and watch the result open in the detail. Runs go through the same guarded, network-off container as the CLI (β/Ctrl+Enter to run). Pick a resource profile in one click β π strict (network off Β· 256m Β· 0.5 cpu Β· 15s), β default (512m Β· 1 cpu Β· 30s), or π networked (network on Β· 60s) β and the sandbox's memory, CPU, timeout and network are set for you (recorded with the run).
- Run code against something β
run()has always taken{ path: content }files and written them into/workbeside your snippet (agents use it), and the form never offered it. Now οΌ file adds fixture files β a module and its test, a parser and a sample payload β and οΌ stdin pipes input to the process. A snippet on its own can only prove syntax; a snippet with a fixture can prove behaviour. Paths that try to climb out of the sandbox are refused before a container starts. - Run a real repo, from the browser β the new-run form used to do snippets only, so the dashboard couldn't do the thing anvil is for. Switch it to command in an image, point mount at a directory, and it runs any command in any image with your repo mounted read-only at
/repoβcd /repo && npm testβ without the repo ever leaving your disk or being writable by the container (touch /repo/xβ Read-only file system). There's a read-only rootfs switch too. A mount that doesn't exist is refused in the form, before docker gets a chance to fail at you. - Watch it run β the sandbox writes as it goes, so anvil now says so as it goes: hit βΆ run and the container's stdout/stderr stream into a live console (with an elapsed timer) instead of leaving you staring at a spinner until it returns. A 30-second run no longer looks identical to a hung one. Streaming is opt-in on the API (
POST /api/exec?stream=1β an event stream ending with the logged run), andrun({ onData })gives the same live chunks to any caller β the CLI, an agent, anything. - Re-run β hit β» re-run on any run to re-execute its exact code and limits in a fresh sandbox; the new run is logged and opened, so you can compare it against the original.
- Keep a run in cortex β a run is evidence: this code, in this sandbox, produced this output. Hit π§ β cortex and it becomes a note in your second brain β status, exit code, duration and the resource limits it ran under, the code as a fenced block, and its stdout/stderr β carrying a
#run=<id>link straight back to the run in the forge log. anvil never writes: your browser POSTs to cortex's own/api/capture(point it elsewhere withANVIL_CORTEX_URL). - Link to a run β opening a run puts it in the URL (
β¦/#run=<id>), so any run in the log can be linked, bookmarked or handed to someone else. - Prune the log β the forge log grows with every run, so β delete drops a single noisy run and π clear empties the whole log. Both arm on the first click and only fire on the second (clicking away, or waiting, cancels), and both live behind
DELETEon the API β a strayGETcan never prune anything. - Search β a live search box filters the log by code, command, language or captured output as you type (matches highlighted, with a running
3 / 7count);Escclears it. Composes with the status filters, so you can find "thatcurlrun that failed" in a long history. - Filters + stats β narrow to failures, see ok/failed counts, average duration, and a by-language breakdown of everything the sandbox has run (
python 3 Β· node 1 Β· bash 1). - Keyboard navigation β press
j/k(or β/β) to move a cursor through the run list and open each run as you go β browse the whole forge log from the keyboard without touching the mouse (it stands down while you're typing in the search box or in compare mode). - Keyboard-accessible β every control has a visible focus ring, and the run rows open with Tab + Enter (not just the mouse), with aria-labels throughout.
Logging is opt-in, fire-and-forget and fully guarded β it never slows or breaks a run. Try the demo without Docker: ANVIL_DB=./.anvil/runs.db node scripts/seed.js then anvil serve.
| Tool | Use it to⦠|
|---|---|
anvil_run_code |
Run a bash / node / python snippet, get structured output. |
anvil_run_command |
Run a shell command with supplied files and/or a host dir mounted read-only at /repo (e.g. run a test suite). |
anvil_check |
Docker availability + language presets. |
{ "name": "anvil_run_command",
"arguments": { "image": "node:22-alpine", "mount": "/path/to/repo/src",
"cmd": "node --check /repo/core.js && echo OK" } }{ "ok": true, "exit_code": 0, "timed_out": false, "duration_ms": 178,
"image": "python:3.12-alpine", "stdout": "...", "stderr": "" }anvil is the run safely leg of tools-for-agents β an operating system for agents.
Nine zero-dependency, MCP-native tools that form one loop, with a self at its centre:
| π°οΈ | agent-hq | coordinate β The company's work, made visible. |
| π | lens | read code β Read code without reading files. |
| β | anvil | run safely β Run it before you claim it works. |
| π | keep | hold secrets β Use a secret without holding it. |
| π§ | cortex | remember β A second brain that outlives the context window. |
| π§ | scout | read the web β The web, ~90% lighter. |
| π» | prism | read data β Read data without reading the blob. |
| β | recall | recall it all β One query. Every store you have. |
| π | iris | see β Look at what you built. |
| π» | ghost | the self at the centre β A self that persists across sessions. Not a tool: it is what the agent is while it calls these. |
Reading this as an agent? /llms.txt is the map, and
/tools.json hands you all 79 MCP tools β every name, every
description, every install command β in one fetch, without cloning anything.
MIT licensed.

{ "mcpServers": { "anvil": { "command": "node", "args": ["/abs/path/to/anvil/mcp/mcp-server.js"] } } }