Skip to content
98 changes: 98 additions & 0 deletions docs/agents/claude.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,98 @@
# Claude Code

Claude Code is sandcat's default agent (`--agent claude`, or simply omit the
flag). The sections below cover authentication, the host paths sandcat mounts
for it, and its RTK hook; everything generic — network policy, secret
mechanics, stacks — works the same for every agent.

## Authentication

Claude Code supports two authentication methods inside the container:

- **API key** — add an `ANTHROPIC_API_KEY` secret to `settings.json`. The
entrypoint detects the key and seeds `~/.claude.json` with
`{"hasCompletedOnboarding": true}` so Claude Code uses it without interactive
setup.
- **Subscription (browser login)** — omit `ANTHROPIC_API_KEY` from
`settings.json`. On first run Claude Code will display a URL and a code. Open
the URL in a browser on your host machine, enter the code, and authenticate
there — the container itself cannot open a browser.

**Autonomous mode.** The bundled `devcontainer.json` enables
`claudeCode.allowDangerouslySkipPermissions` and sets
`claudeCode.initialPermissionMode` to `bypassPermissions`. This lets Claude Code
run without interactive permission prompts inside the container. The trade-off:
sandcat already provides the security boundary (network isolation, secret
substitution, iptables kill-switch), so the in-container prompts add friction
without meaningful security benefit. Remove these settings if you prefer
interactive approval. See [Secure & Dangerous Claude Code + VS Code
Setup](https://warski.org/blog/secure-dangerous-claude-code-vs-code-setup/) for
background on this approach.

**Host customizations.** The example `compose-all.yml` bind-mounts
`~/.claude/CLAUDE.md`, `~/.claude/agents`, and `~/.claude/commands` from the
host (read-only) so your personal instructions, custom agents, and slash
commands are available inside the container. Remove any mount whose source does
not exist on your host — Docker will otherwise create an empty directory in its
place.

**Multi-line prompts.** Composing a multi-line prompt with `⌘+Enter` does not
work on macOS — the terminal reserves the `⌘` modifier and never transmits it
over the PTY, so `sandcat attach` (and Claude Code) only ever receive a plain
`Enter`. This is not sandcat-specific and cannot be fixed inside the container.
Use one of these instead:

- **`\` then `Enter`** — inserts a newline in any terminal with no setup. The
simplest option.
- **`Option+Enter`** — Claude Code's macOS default. In Apple Terminal, first
enable *Settings → Profiles → Keyboard → Use Option as Meta key*; iTerm2 sends
it out of the box.
- **`Shift+Enter`** — the most familiar combination, but Claude Code only
receives whatever bytes the terminal chooses to send for it, so it needs a
one-time mapping in the **host** terminal. Claude Code's `/terminal-setup` is
meant to install this, but it has two traps in this setup: it configures the
host terminal, so running it from Claude Code *inside* the sandbox does
nothing; and it caches an "installed" flag, so a second run reports *"already
enabled"* even when the terminal was never actually changed. The reliable route
is to map the key by hand:
- **iTerm2** — Settings → Keys → Key Bindings → `+`, record `Shift+Enter`,
choose *Send Hex Codes* and enter `0x1b 0x0d` (this is `Option+Enter`, which
Claude Code treats as a newline). GUI bindings take effect immediately. To
confirm it worked, run `cat -v` in the sandbox shell and press `Shift+Enter`:
it should print `^[` instead of a blank line.
- **VS Code integrated terminal** — add to `keybindings.json`:

```json
{ "key": "shift+enter",
"command": "workbench.action.terminal.sendSequence",
"args": { "text": "\u001b\r" },
"when": "terminalFocus" }
```

## Host paths and mounts

**Claude paths** (host `~/.claude/`, read-only when mounted):

- `CLAUDE.md`, `agents/`, `commands/`

Project-local configuration (`.claude/` in the repo) and the isolation
semantics of these mounts are described in
[Customizing optional volume mounts](../configuration/volume-mounts.md).

## RTK hook

Works out of the box, zero configuration. `sandcat init` generates an
`app-user-init.sh` block that runs `rtk init -g --hook-only --auto-patch`
on the first container start; the hook lands in the sandbox's
`~/.claude/settings.json` (inside the `agent-home` volume, not
bind-mounted). Subsequent starts are idempotent no-ops.

See [RTK — LLM token compression](rtk.md) for what RTK does and how to opt
out.

## Convenience alias

Every claude sandbox ships a `claude-yolo` alias (= `claude
--dangerously-skip-permissions`): the sandcat network isolation is the
security boundary, so bypassing in-container permission prompts is the
intended workflow.
47 changes: 47 additions & 0 deletions docs/agents/codex.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,47 @@
# Codex CLI

Sandcat installs [OpenAI's Codex CLI](https://github.com/openai/codex)
into every codex-agent sandbox and wires `OPENAI_API_KEY` through the
mitmproxy secret substitution layer. Codex reads its config from
`~/.codex/config.toml` (per-sandbox, agent-home volume) and picks up
the API key directly from the environment — no `codex login` required.

**Setup:**

```bash
sandcat init --agent codex --ide vscode
# Edit ~/.config/sandcat/settings.json — set secrets.OPENAI_API_KEY.value
sandcat run
codex "explain this codebase"
```

**Bash alias:** `codex-yolo` (= `codex --yolo`) is available in every
codex sandbox for parity with `claude-yolo`.

**Host config sharing** (optional, default on): `~/.codex/AGENTS.md`,
`~/.codex/skills/`, and `~/.codex/commands/` are bind-mounted read-only
from the host into the container, matching how `~/.claude/` is handled.
The rest of `~/.codex/` (config.toml, credentials, history) lives in
the container's agent-home volume — per-sandbox persistent, per-sandbox
isolated. Opt out with `SANDCAT_MOUNT_CODEX_CONFIG=false`.

**RTK integration:** works out of the box. On first container start,
sandcat seeds `~/.codex/AGENTS.md` (from the host bind-mount if
present) and runs `rtk init -g --codex` to write `~/.codex/RTK.md`
and add an `@RTK.md` reference to `AGENTS.md`. Idempotent: skipped
once the reference is already there. Disable with `--features
no-rtk` or `SANDCAT_RTK=false`.

Note: because rtk needs to patch a writable `AGENTS.md`, sandcat
mounts the host's `~/.codex/AGENTS.md` at `~/.codex-host/AGENTS.md`
(a helper path) — the user-init step copies it into the writable
`~/.codex/AGENTS.md`. Host edits to `AGENTS.md` take effect after a
`docker compose down -v` (or manual rm inside). Skills and commands
directories are bind-mounted normally at `~/.codex/skills` and
`~/.codex/commands`, so those live-reload as usual.

**Auth model:** first iteration supports `OPENAI_API_KEY` only.
ChatGPT sign-in (`chatgpt.com` / `auth.openai.com`) is not in the
default allowlist — users who want that flow can add the hosts to
`.sandcat/settings.local.json` and run `codex login` manually inside
the container.
61 changes: 61 additions & 0 deletions docs/agents/copilot.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,61 @@
# GitHub Copilot CLI

GitHub's [Copilot CLI](https://docs.github.com/copilot/how-tos/copilot-cli)
(`@github/copilot`) is available as a first-class sandcat agent. Sandcat installs
Node.js 22 and the Copilot package into every copilot-agent sandbox and wires
`COPILOT_GITHUB_TOKEN` through the mitmproxy secret substitution layer.

**Setup:**

```bash
sandcat init --agent copilot --ide vscode
# Edit ~/.config/sandcat/settings.json — set secrets.COPILOT_GITHUB_TOKEN.value
sandcat run
copilot "explain this codebase"
```

**Authentication:** Copilot CLI requires a GitHub token. Choose one of:

1. **Fine-grained Personal Access Token (recommended):** Create a PAT at
[`https://github.com/settings/personal-access-tokens`](https://github.com/settings/personal-access-tokens)
with the **"Copilot Requests"** permission (Read and write). Then add it to
`~/.config/sandcat/settings.json`:
```json
{
"secrets": {
"COPILOT_GITHUB_TOKEN": {
"value": "github_pat_...",
"hosts": ["api.github.com", "*.github.com", "*.githubcopilot.com", "*.githubusercontent.com"]
}
}
}
```

2. **GitHub CLI OAuth token (quick setup):** If you already have `gh` CLI logged in,
run this once to write the token directly into `settings.json`:
```bash
export TKN=$(gh auth token)
yq -i -o json '.secrets.COPILOT_GITHUB_TOKEN.value = strenv(TKN)' \
~/.config/sandcat/settings.json
```

**Note:** Adding Node.js 22 and Copilot to the base image increases its size by
approximately 120 MB. The image is built once and cached locally; rebuilds are
fast.

**VS Code integration:** When the IDE is `vscode`, the bundled `devcontainer.json`
includes the `GitHub.copilot` extension. Note that the VS Code extension
authenticates through VS Code's own GitHub sign-in (not the `COPILOT_GITHUB_TOKEN`
env var used by the CLI), so you may need to sign in the first time you open the
extension.

**Placeholder:** Sandcat automatically sets the placeholder to
`gho_SANDCAT_PLACEHOLDER_COPILOT_GITHUB_TOKEN`. The container sees only the
placeholder; the real token is injected by mitmproxy only for allowed Copilot
hosts. No manual configuration is needed.

**Bash alias:** `copilot-yolo` (= `copilot --yolo`) is available in every
copilot sandbox for parity with `claude-yolo` and `codex-yolo`. `--yolo` is
equivalent to `--allow-all-tools --allow-all-paths --allow-all-urls` — the
sandcat network isolation is the security boundary, so bypassing in-container
permission prompts is the intended workflow.
139 changes: 139 additions & 0 deletions docs/agents/cursor.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,139 @@
# Cursor CLI

Cursor CLI support is available via `sandcat init --agent cursor`.

## Authentication and CLI configuration

Cursor CLI support is available via `sandcat init --agent cursor`.

- The current template uses temporary compatibility defaults for auth/network:
- **Auth passthrough via placeholder substitution.** The container sees only
`SANDCAT_PLACEHOLDER_CURSOR_API_KEY`; the real `CURSOR_API_KEY` is injected
by the mitmproxy addon only for allowed Cursor hosts.
- **HTTP/1 compatibility bootstrap.** On startup, Sandcat forces
`.network.useHttp1ForAgent = true` in Cursor CLI config to avoid known
proxy/TLS instability with HTTP/2 streaming through mitmproxy.
- **Proxy command defaults tuned for Cursor.** The generated proxy config uses
the Cursor addon and keeps mitmproxy HTTP/2 enabled (`http2=true`) (plus
streaming-safe mitmproxy
flags such as `stream_large_bodies=1m`, `connection_strategy=lazy`,
`anticomp=true`, and `timeout_read=300`).

Those streaming-safe flags are **Cursor-only** — they are intentionally
omitted on the Claude path (`sct_agent_mitm_streaming_flags`). With
`stream_large_bodies` unset, mitmproxy buffers request bodies up to ~1 MB
before forwarding, which lets the addon's `_substitute_secrets` run a
body-content scan for placeholder leaks. Setting them on Claude would
weaken that defence-in-depth check; on Cursor they are required to keep
Connect/HTTP-2 streaming responses stable, and the body-leak check is
instead enforced via header/URL scans plus the textual-only body-mutation
gate (binary protobuf bodies are left untouched).
- **Streaming detection is path-only.** The Cursor addon decides whether a
request is streaming purely from the request path
(`/agent.v1.AgentService/Run*`, `/aiserver.v1.RepositoryService/...`).
A client-supplied `content-type: application/connect+proto` header alone
is **not** sufficient — accepting it would let any request with the right
header bypass body substitution and the placeholder leak check.
These defaults are conservative and may be relaxed when Cursor proxy behavior
is consistently stable across environments.
- **Authentication:** put the Cursor API key in `secrets.CURSOR_API_KEY` in
Sandcat settings (not in `cursor.cli`). The agent container receives only
`SANDCAT_PLACEHOLDER_CURSOR_API_KEY` via `sandcat.env`; mitmproxy substitutes
the real key on allowed Cursor hosts (see placeholder substitution above).
Do not use `agent login` in the sandbox unless you accept that Cursor may
store session state under agent-home outside Sandcat's placeholder model.
- **Cursor CLI settings via Sandcat:** add a `cursor.cli` block to
`~/.config/sandcat/settings.json` (or project `.sandcat/settings.json`) using
the same JSON shape as Cursor's global `cli-config.json` (permissions, model,
network flags — not API keys). Sandcat merges settings layers at mitmproxy
startup, writes `/mitmproxy-public/cursor-cli-config.json`, and the agent
deep-merges that fragment into `cli-config.json` in agent-home on each start.
Sandcat-owned keys win; other Cursor-written keys in that file (model choice,
permissions allow/deny lists, etc.) are preserved. The Cursor user template
defaults include `cursor.cli.network.useHttp1ForAgent: true` for mitmproxy
stability.
- `SANDCAT_MOUNT_CURSOR_CONFIG=true` mounts host Cursor config into the agent
container. Customization paths are read-only: `AGENTS.md`, `rules/`, `skills/`,
`commands/`, `hooks.json`, `hooks/`, `agents/`, and `mcp.json`. Runtime state
for this sandbox is read-write on the host under
`projects/<workspace-id>/` only (`workspaces-<project-name>` — agent
transcripts, terminals, MCP session state). `chats/`, `plugins/`, and
`subagents/` are not host-mounted (they live in `agent-home`). On
`sandcat init`, missing bind sources are pre-created on the host (directories
via `mkdir`, JSON files with minimal valid defaults, markdown files empty) so
Docker mounts a file instead of materialising a root-owned directory.
- **Config precedence:** `~/.config/sandcat/settings.json` governs network
allowlists, secret substitution (mitmproxy), and Sandcat-managed Cursor CLI
settings (`cursor.cli` — not credentials). Host Cursor customization mounts
are read-only user config. The workspace-scoped `projects/<workspace-id>/`
mount is read-write on the host. MCP servers in `mcp.json` still need
matching mitmproxy allowlist entries before they can reach the network from
the sandbox.
- **Cursor CLI TLS through mitmproxy.** The Cursor CLI bundles its own Node.js
binary with compiled-in Mozilla CA roots. Sandcat sets
`NODE_OPTIONS=--use-openssl-ca` so the bundled Node.js uses the system CA
store (which includes the mitmproxy CA) instead of its built-in roots.
When Cursor honors that environment setting, mitmproxy can intercept Cursor
API traffic and perform `SANDCAT_PLACEHOLDER_CURSOR_API_KEY` substitution
transparently.
- Provider-specific onboarding/bootstrap logic is intentionally minimal in this
first iteration and can be extended in project-level Dockerfile/scripts.

## Host paths and mounts

**Cursor paths** (host `~/.cursor/`):

| Path | Mode | Typical use |
|----------------------------------------------------------------------------------------------|------------|---------------------------------------|
| `AGENTS.md`, `rules/`, `skills/`, `commands/`, `hooks.json`, `hooks/`, `agents/`, `mcp.json` | read-only | Shared customization |
| `projects/<workspace-id>/` | read-write | This sandbox's transcripts/terminals |

Sandcat mounts only `projects/<workspace-id>/` for the current sandbox
(`workspaces-<project-name>`), not the whole host `projects/` tree. `chats/`,
`plugins/`, and `subagents/` stay in `agent-home` so other workspaces' runtime
state is not exposed.

Cursor CLI keys Sandcat manages (`cursor.cli` in settings) are **not**
host-mounted — see the Cursor section below.

Project-local configuration (`.cursor/` in the repo) and the isolation
semantics of these mounts are described in
[Customizing optional volume mounts](../configuration/volume-mounts.md).

## RTK hook

Cursor's rtk hook lives in `~/.cursor/hooks.json`. Sandcat bind-mounts
that file **read-only** from your host (`SANDCAT_MOUNT_CURSOR_CONFIG=true`
default) so cursor customizations are shared across all your sandboxes.
Because the mount is read-only, sandcat cannot install the rtk hook
into the container's copy of `hooks.json`.

**Setup — run once on your host:**

```bash
# Install rtk locally (needed once on the host)
brew install rtk # or: curl -fsSL https://raw.githubusercontent.com/rtk-ai/rtk/b34be37caf3796b69a50952a28e60e32b5daad43/install.sh | RTK_VERSION=v0.45.0 sh

# Register the cursor hook in your host ~/.cursor/hooks.json
rtk init -g --hook-only --auto-patch --agent cursor
```

That writes the rtk hook to your host `~/.cursor/hooks.json`. Every
sandcat cursor sandbox from now on bind-mounts that file into the
container, so Cursor CLI sees the hook and calls `rtk hook cursor` on
each `Bash` tool invocation. The container's own `rtk` binary
(installed by sandcat) executes the hook — you never need rtk on the
host for anything except this one-time init step, and you can
uninstall it afterwards if you like.

Bonus: the same host hook is picked up by every cursor sandbox you
start on that machine (and by host Cursor CLI, if you use it directly).

**If you skip the host init:** the container prints a one-time warning
on start (`sandcat: rtk hook not found for cursor. Install rtk on your
host …`) and Cursor CLI runs without the hook. The rtk binary is still
on `PATH` inside the container, so you can invoke `rtk grep`, `rtk ls`,
etc. by hand.

See [RTK — LLM token compression](rtk.md) for what RTK does and how to opt
out.
30 changes: 30 additions & 0 deletions docs/agents/rtk.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
# RTK — LLM token compression

[rtk-ai/rtk](https://github.com/rtk-ai/rtk) ("Rust Token Killer") wraps
shell commands invoked by AI agents and compresses their output before
the agent reads it, cutting token consumption 60-90% on typical dev
commands (test runs, grep output, build logs). `sandcat init` installs
the `rtk` binary into every sandbox by default and wires the agent
hook so the agent picks it up automatically. Setup differs slightly
per agent — see below.

**Opt out** if you'd rather run without it (e.g. debugging a shell
command's raw output):

```bash
sandcat init --features no-rtk ...
SANDCAT_RTK=false sandcat init ...
```

Both are equivalent — the env var is the scripted counterpart of the
interactive/CSV feature flag. When disabled, the rtk binary is not
installed into the image and no init hook is emitted for any agent.

## Per-agent setup

| Agent | Setup |
|-------|-------|
| Claude Code | zero configuration — see [Claude Code → RTK hook](claude.md#rtk-hook) |
| Cursor CLI | one-time host-side hook — see [Cursor CLI → RTK hook](cursor.md#rtk-hook) |
| Codex CLI | wired automatically at init — see [Codex CLI](codex.md) |
| GitHub Copilot CLI | wired automatically at init — see [GitHub Copilot CLI](copilot.md) |
Loading