Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 1 addition & 3 deletions DESIGN.md
Original file line number Diff line number Diff line change
Expand Up @@ -63,9 +63,7 @@ The rule that ties the layers together: a package or tool belongs in the *lowest

**The template is authored in OpenTofu; the live deployment's provisioner is a separate repo's decision, not this one's.** Coder has no concept of a provisioning tool beyond "whatever binary is named `terraform` on `PATH`, in a supported numeric version range" — there is no vendor check, so a `tofu` binary standing in for that name applies the template identically. This repo can only control the provisioner in its own disposable test control plane (`.github/compose/compose.yaml`), which bind-mounts an OpenTofu binary over the official Coder image's own `terraform` path — both to dogfood the toolchain this repo now authors in, and because that swap is what surfaces a real Terraform/OpenTofu behavioral divergence, if this template's HCL ever grows one, before it reaches a real deployment. The live deployment's provisioner is owned by `homelab-ops-kubernetes-apps` — `homelab-ops-kubernetes-apps#3900` makes the same swap there, via a Kubernetes-native init-container-plus-`emptyDir` mechanism instead of a Docker bind mount — that cluster's containerd doesn't support single-file subPath mounts on image volumes, only whole-directory.

**`/tmp` is node-local scratch space, deliberately not the shared home volume.** Every workspace's `/tmp` used to be whatever the container's writable overlay layer gave it for free - fast, but unbounded, and on the node's root filesystem. On the operator's own long-lived workspace that grew to several GiB (dominated by Claude Code's own scratch directory, `$TMPDIR/claude-<uid>/...`, which agent sessions use for downloads and experiments) and pushed the node toward the kubelet's disk-pressure eviction threshold - a risk to every other pod on that node, not just the workspace that caused it. The home PVC has ample free space, but is NFS-backed, which is a bad fit for what actually lives in `/tmp`: build caches and compiler intermediates are exactly the write-heavy, latency-sensitive workload NFS handles worst. A plain `empty_dir` would keep `/tmp` fast but doesn't fix anything, because `empty_dir` lives on the same constrained node root filesystem the container overlay already did. The fix is a Kubernetes "generic ephemeral volume" on a Longhorn storage class that is both node-local (so still fast) and backed by a separate, much larger partition on the same node than the root filesystem is - see the comment on the `tmp` volume in `deployment.tf` for the specific class and why it's the non-replicated one (scratch data costs nothing to lose) rather than the default replicated class other PVCs in this cluster use. Its size is a fixed ceiling rather than left unbounded, so a runaway consumer now fails predictably inside its own volume instead of eventually pressuring the node. Because this volume's lifecycle is tied to the Pod rather than the container, something has to wipe it on every container start, so a container restart within a live Pod doesn't just inherit whatever the previous container left behind.

**That wipe runs in the container entrypoint because `/tmp` is not only scratch space - the Coder agent lives there too.** The first attempt put the wipe in `script-agent-startup.sh`, which is the wrong side of the ordering and shipped a broken template: Coder's generated bootstrap downloads the agent CLI into a per-boot `mktemp -d -t coder.XXXXXX` directory under `/tmp`, makes it the agent's working directory, appends it to the PATH of every session and script the agent runs, and *then* runs the startup script - which deleted it, leaving every `coder stat` metadata panel reporting `coder: command not found` while the agent process itself carried on from an unlinked binary. The tempting repair is to exclude the agent's own paths from the wipe, and it is the wrong one: that set is an implementation detail of whatever agent version the control plane happens to serve, and it is not even uniformly named - v2.35.3 owns a random-suffixed `coder.XXXXXX/`, `coder-agent.sock` (a hardcoded absolute path, not a `TMPDIR`-relative one), rotated `coder-agent*.log` files, `coder-script-data/`, `coder-screen/`, and `boundary-audit.sock`, which carries no `coder` prefix at all. An allowlist would stop matching on some future upgrade and fail exactly as invisibly as the original bug. The workspace container's `command` is therefore a small entrypoint script that wipes `/tmp` and `exec`s Coder's bootstrap: it is the only hook that runs on a container-only restart within a live Pod (init containers run once per Pod), it runs before the agent exists so there is nothing to exclude, and `exec` keeps the agent as PID 1 for orphan reaping and the liveness probe. The wipe there is best-effort rather than fatal, because that entrypoint is the only path to a running agent and a hard failure would be a `CrashLoopBackOff` nobody can shell into; it records its outcome instead, and `script-agent-startup.sh` asserts both that outcome and that the agent CLI still resolves and runs - the check the original change lacked, which turns a recurrence into a failed startup script in the workspace UI instead of eight quietly broken metadata panels.
**`/tmp` is node-local scratch space, deliberately not the shared home volume.** Every workspace's `/tmp` used to be whatever the container's writable overlay layer gave it for free - fast, but unbounded, and on the node's root filesystem. On the operator's own long-lived workspace that grew to several GiB (dominated by Claude Code's own scratch directory, `$TMPDIR/claude-<uid>/...`, which agent sessions use for downloads and experiments) and pushed the node toward the kubelet's disk-pressure eviction threshold - a risk to every other pod on that node, not just the workspace that caused it. The fix is a Kubernetes "generic ephemeral volume" on a Longhorn storage class that is both node-local (so still fast) and backed by a separate, much larger partition on the same node than the root filesystem is.

**Unprivileged by default.** The workspace itself runs as an unprivileged, non-root, fixed-identity container. Anything that genuinely needs elevated privilege (installing packages, preparing shared volume state) is scoped to a narrow, short-lived setup step that runs before the workspace shell exists, not to something the workspace user can reach into.

Expand Down
9 changes: 9 additions & 0 deletions templates/kubernetes/homelab-workspace/scripts.tf
Original file line number Diff line number Diff line change
Expand Up @@ -57,3 +57,12 @@ resource "coder_script" "vscode_server_gc" {
cron = "0 0 4 * * 0"
script = "/bin/bash /scripts/script-vscode-server-gc.sh"
}

resource "coder_script" "supervised_services" {
agent_id = coder_agent.main.id
display_name = "Supervised Services"
icon = "/icon/terminal.svg"
run_on_start = true
start_blocks_login = false
script = "/bin/bash /scripts/script-start-services.sh"
}
Original file line number Diff line number Diff line change
@@ -1,16 +1,6 @@
#!/bin/bash
set -eo pipefail


# Written by script-container-entrypoint.sh, which wipes /tmp before the agent
# starts.
TMP_WIPE_STATUS_FILE="/tmp/.tmp-wipe-status"


# The check that would have caught the regression where wiping /tmp deleted the
# agent's own CLI: every "coder stat" metadata panel reported "coder: command
# not found" for a released template version, and nothing else noticed.
#
# Coder's bootstrap unpacks the agent CLI into a per-boot directory under /tmp
# and the agent appends that directory to the PATH of everything it runs, so
# resolving "coder" here goes through exactly the same lookup a metadata script
Expand All @@ -33,28 +23,7 @@ assert_agent_cli() {
echo "Coder agent CLI: ${cli}"
}


# The /tmp wipe cannot abort the container entrypoint without risking a
# CrashLoopBackOff, so it reports here instead. This keeps a partial wipe a
# loud failure rather than a silent leak on a fixed-size volume.
assert_tmp_wiped() {
if [[ ! -f "${TMP_WIPE_STATUS_FILE}" ]]; then
echo "ERROR: ${TMP_WIPE_STATUS_FILE} is missing - the container entrypoint" >&2
echo " did not run. Check the workspace container's command in" >&2
echo " deployment.tf; /tmp is no longer being cleared on restart." >&2
return 1
fi
if [[ "$(head -n 1 "${TMP_WIPE_STATUS_FILE}")" != "ok" ]]; then
echo "ERROR: /tmp was not fully wiped on container start:" >&2
cat "${TMP_WIPE_STATUS_FILE}" >&2
return 1
fi
echo "/tmp wiped on container start"
}


main() {
assert_tmp_wiped
assert_agent_cli
if [[ ! -s ~/.bashrc ]]; then
echo "Setting up starter bash rc scripts from /etc/skel..."
Expand Down
Original file line number Diff line number Diff line change
@@ -1,58 +1,7 @@
#!/bin/bash
set -eo pipefail


# Where wipe_tmp records its outcome for script-agent-startup.sh to assert on.
# Kept in /tmp on purpose: it is written after the wipe, so its presence is also
# evidence that this script ran at all this boot.
TMP_WIPE_STATUS_FILE="/tmp/.tmp-wipe-status"


# Restore the "a fresh container gets a fresh /tmp" property that /tmp lost when
# it moved off the container's writable overlay layer onto a per-Pod volume (see
# the "tmp" volume in deployment.tf). A Pod recreate still gets an empty volume
# for free; a container-only restart within a live Pod - an OOM kill, a liveness
# probe failure - does not, and would otherwise inherit whatever the previous
# container left behind.
#
# This has to run in the container entrypoint, before the Coder agent exists,
# and not in the agent startup script. Coder's own bootstrap (the generated
# /scripts/workspace-init.sh) unpacks the agent CLI into a per-boot mktemp directory
# under /tmp, chdirs into it, appends it to the PATH of every session and script
# the agent runs, and only then runs the startup script - so a wipe from the
# startup script deletes the CLI the agent installed moments earlier, which is
# exactly how every "coder stat" metadata panel came to report "coder: command
# not found".
#
# Excluding the agent's paths from the wipe instead is not a fix. The set is
# version-dependent and not even consistently named: on the deployed agent it is
# a random-suffixed coder.XXXXXX directory, coder-agent.sock, coder-agent*.log,
# coder-script-data/, coder-screen/ - and also boundary-audit.sock, which does
# not carry the "coder" prefix at all. An allowlist that silently stops matching
# after a Coder upgrade reintroduces this same failure just as invisibly. There
# is nothing to exclude here, because nothing of the agent's exists yet.
#
# The wipe is deliberately best-effort rather than fatal: this script is the
# only path to a running agent, so aborting here turns a stale-scratch-space
# problem into a CrashLoopBackOff with no way to shell in and look. Visibility
# is preserved instead by recording the outcome for the startup script, which
# can fail loudly in the workspace UI without taking the workspace down.
wipe_tmp() {
echo "Wiping /tmp..."
local errors
errors="$(find /tmp -mindepth 1 -maxdepth 1 -exec rm -rf -- {} + 2>&1)" || true
if [[ -n "${errors}" ]]; then
echo "ERROR: failed to fully wipe /tmp:" >&2
echo "${errors}" >&2
printf 'failed\n%s\n' "${errors}" > "${TMP_WIPE_STATUS_FILE}"
return 0
fi
printf 'ok\n' > "${TMP_WIPE_STATUS_FILE}"
}


main() {
wipe_tmp
# Hand off to Coder's generated agent bootstrap, replacing this process rather
# than spawning it: the agent has to stay PID 1, both because it reaps orphans
# in this container and because the Deployment's liveness probe pgreps for it.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -55,4 +55,3 @@ else
fi

touch "${state_dir}/applied"
/bin/bash /scripts/script-start-services.sh
Original file line number Diff line number Diff line change
Expand Up @@ -6,6 +6,15 @@ if (( SUPERVISOR_SERVICE_COUNT == 0 )); then
exit 0
fi

dotfiles_state_file="${HOME}/.local/state/dotfiles/applied"
timeout 180s bash -c "until [ -e '$dotfiles_state_file' ]; do sleep 5; done"
if [ $? -eq 124 ]; then
echo "Timed out waiting for dotfiles to be applied. Please check the logs for errors."
exit 1
else
echo "Dotfiles applied successfully. Proceeding to start supervised services."
fi

state_dir="${XDG_STATE_HOME:-${HOME}/.local/state}/supervisor"
mkdir -p "${state_dir}"

Expand Down