Skip to content
Closed
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion DESIGN.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,7 +65,7 @@ The rule that ties the layers together: a package or tool belongs in the *lowest

**`/tmp` is node-local scratch space, deliberately not the shared home volume.** Every workspace's `/tmp` used to be whatever the container's writable overlay layer gave it for free - fast, but unbounded, and on the node's root filesystem. On the operator's own long-lived workspace that grew to several GiB (dominated by Claude Code's own scratch directory, `$TMPDIR/claude-<uid>/...`, which agent sessions use for downloads and experiments) and pushed the node toward the kubelet's disk-pressure eviction threshold - a risk to every other pod on that node, not just the workspace that caused it. The home PVC has ample free space, but is NFS-backed, which is a bad fit for what actually lives in `/tmp`: build caches and compiler intermediates are exactly the write-heavy, latency-sensitive workload NFS handles worst. A plain `empty_dir` would keep `/tmp` fast but doesn't fix anything, because `empty_dir` lives on the same constrained node root filesystem the container overlay already did. The fix is a Kubernetes "generic ephemeral volume" on a Longhorn storage class that is both node-local (so still fast) and backed by a separate, much larger partition on the same node than the root filesystem is - see the comment on the `tmp` volume in `deployment.tf` for the specific class and why it's the non-replicated one (scratch data costs nothing to lose) rather than the default replicated class other PVCs in this cluster use. Its size is a fixed ceiling rather than left unbounded, so a runaway consumer now fails predictably inside its own volume instead of eventually pressuring the node. Because this volume's lifecycle is tied to the Pod rather than the container, something has to wipe it on every container start, so a container restart within a live Pod doesn't just inherit whatever the previous container left behind.

**That wipe runs in the container entrypoint because `/tmp` is not only scratch space - the Coder agent lives there too.** The first attempt put the wipe in `script-agent-startup.sh`, which is the wrong side of the ordering and shipped a broken template: Coder's generated bootstrap downloads the agent CLI into a per-boot `mktemp -d -t coder.XXXXXX` directory under `/tmp`, makes it the agent's working directory, appends it to the PATH of every session and script the agent runs, and *then* runs the startup script - which deleted it, leaving every `coder stat` metadata panel reporting `coder: command not found` while the agent process itself carried on from an unlinked binary. The tempting repair is to exclude the agent's own paths from the wipe, and it is the wrong one: that set is an implementation detail of whatever agent version the control plane happens to serve, and it is not even uniformly named - v2.35.3 owns a random-suffixed `coder.XXXXXX/`, `coder-agent.sock` (a hardcoded absolute path, not a `TMPDIR`-relative one), rotated `coder-agent*.log` files, `coder-script-data/`, `coder-screen/`, and `boundary-audit.sock`, which carries no `coder` prefix at all. An allowlist would stop matching on some future upgrade and fail exactly as invisibly as the original bug. The workspace container's `command` is therefore a small entrypoint script that wipes `/tmp` and `exec`s Coder's bootstrap: it is the only hook that runs on a container-only restart within a live Pod (init containers run once per Pod), it runs before the agent exists so there is nothing to exclude, and `exec` keeps the agent as PID 1 for orphan reaping and the liveness probe. The wipe there is best-effort rather than fatal, because that entrypoint is the only path to a running agent and a hard failure would be a `CrashLoopBackOff` nobody can shell into; it records its outcome instead, and `script-agent-startup.sh` asserts both that outcome and that the agent CLI still resolves and runs - the check the original change lacked, which turns a recurrence into a failed startup script in the workspace UI instead of eight quietly broken metadata panels.
**That wipe runs in the container entrypoint because `/tmp` is not only scratch space - the Coder agent lives there too.** The first attempt put the wipe in `script-agent-startup.sh`, which is the wrong side of the ordering and shipped a broken template: Coder's generated bootstrap downloads the agent CLI into a per-boot `mktemp -d -t coder.XXXXXX` directory under `/tmp`, makes it the agent's working directory, appends it to the PATH of every session and script the agent runs, and *then* runs the startup script - which deleted it, leaving every `coder stat` metadata panel reporting `coder: command not found` while the agent process itself carried on from an unlinked binary. The tempting repair is to exclude the agent's own paths from the wipe, and it is the wrong one: that set is an implementation detail of whatever agent version the control plane happens to serve, and it is not even uniformly named - v2.35.3 owns a random-suffixed `coder.XXXXXX/`, `coder-agent.sock` (a hardcoded absolute path, not a `TMPDIR`-relative one), rotated `coder-agent*.log` files, `coder-script-data/`, `coder-screen/`, and `boundary-audit.sock`, which carries no `coder` prefix at all. An allowlist would stop matching on some future upgrade and fail exactly as invisibly as the original bug. The workspace container's `command` is therefore a small entrypoint script that wipes `/tmp` and `exec`s Coder's bootstrap: it is the only hook that runs on a container-only restart within a live Pod (init containers run once per Pod), it runs before the agent exists so there is nothing to exclude, and `exec` keeps the agent as PID 1 for orphan reaping and the liveness probe. A newly formatted volume may itself contribute a protected, root-owned `lost+found`; the unprivileged entrypoint preserves only an entry with that structural type and ownership, so a user-created namesake cannot become unwiped storage. The wipe there is best-effort rather than fatal, because that entrypoint is the only path to a running agent and a hard failure would be a `CrashLoopBackOff` nobody can shell into; it records its outcome instead, and `script-agent-startup.sh` asserts both that outcome and that the agent CLI still resolves and runs - the check the original change lacked, which turns a recurrence into a failed startup script in the workspace UI instead of eight quietly broken metadata panels.

**Unprivileged by default.** The workspace itself runs as an unprivileged, non-root, fixed-identity container. Anything that genuinely needs elevated privilege (installing packages, preparing shared volume state) is scoped to a narrow, short-lived setup step that runs before the workspace shell exists, not to something the workspace user can reach into.

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -39,8 +39,14 @@ TMP_WIPE_STATUS_FILE="/tmp/.tmp-wipe-status"
# can fail loudly in the workspace UI without taking the workspace down.
wipe_tmp() {
echo "Wiping /tmp..."
local errors
errors="$(find /tmp -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + 2>&1)" || true
local errors preserve_lost_found=()
# A fresh filesystem may expose a protected, root-owned lost+found. Preserve
# only that structural entry; a user-created namesake remains scratch data.
if [[ -d /tmp/lost+found && ! -L /tmp/lost+found ]] &&
[[ "$(stat -c '%u' /tmp/lost+found 2>/dev/null)" == "0" ]]; then
preserve_lost_found=(-not -path /tmp/lost+found)
fi
errors="$(find /tmp -mindepth 1 -maxdepth 1 "${preserve_lost_found[@]}" -exec rm -rf -- {} + 2>&1)" || true
if [[ -n "${errors}" ]]; then
echo "ERROR: failed to fully wipe /tmp:" >&2
echo "${errors}" >&2
Expand Down