Run a fleet of unattended opencode agents on one box — and know within minutes when one dies, instead of days.
Two problems every autonomous-agent operator hits:
- Loops die silently overnight. Cron-fired
opencode runsessions crash, hang on permission prompts, or get killed by an OOM — and nothing tells you. - Ad-hoc loop scripts rot. One
while trueper job means no locking, no timeouts, no log rotation, no shared state between your agents.
This kit ships both fixes as small, readable files you can audit in ten minutes: a battle-tested continuous-loop template set (per-loop locking, hard timeouts, memory guard, one kill switch) and a dead-man-switch plugin so silence itself becomes the alert.
Provenance: sanitized from a production fleet that runs five such loops
continuously on a single VPS, chaining opencode run back-to-back around
the clock.
Prerequisites: Linux with systemd user sessions, opencode installed and authenticated.
# 1. Scaffold a fleet home
mkdir -p ~/fleet/logs ~/fleet/prompts && cd ~/fleet
curl -fsSL https://github.com/aniripsaretro-max/opencode-fleet-kit/archive/refs/heads/main.tar.gz | tar xz --strip-components=1
cp templates/supervisor.sh templates/run-worker.sh .
chmod +x supervisor.sh run-worker.sh
# 2. Give one loop its instructions (example included)
cp templates/prompts/builder.md prompts/
# 3. Install the user unit (no root needed)
mkdir -p ~/.config/systemd/user
cp templates/fleet-worker@.service ~/.config/systemd/user/
systemctl --user daemon-reload
systemctl --user enable --now fleet-worker@builder
# 4. Keep it alive after logout/reboot
loginctl enable-linger "$USER"
# 5. Watch it work
journalctl --user -u fleet-worker@builder -f
tail -f ~/fleet/logs/*.jsonlStop everything at once: touch ~/fleet/FLEET_STOP. Remove the file to resume.
Add more loops by adding prompt files: cp templates/prompts/scout.md prompts/ && systemctl --user enable --now fleet-worker@scout.
Two things to know before you walk away:
- Pre-seed permissions. An unattended session that stops to ask "allow this command?" is a hung loop. Configure opencode's permissions so your loops' routine tool calls run without prompting — this is the #1 silent-hang cause.
- Run a fire drill. Kill a loop deliberately (
systemctl --user stop fleet-worker@builder) and confirm your heartbeat alert fires within the grace window. Death-detection you haven't tested is a hypothesis.
Drop plugins/heartbeat.js into .opencode/plugins/
of any project your loops run in (or ~/.config/opencode/plugins/ for all of
them), arm one variable, create a monitor, done:
export HEARTBEAT_PING_URL="https://your-heartbeat.example/ping/<key>"Every completed session pings once (hooked on the session.status busy→idle
transition — provider retries don't ping, a retrying session is still
alive). Crash, hang, provider outage, dead box →
pings stop → you get alerted (a wedged session never reaches idle, so its
silence is the alert; the tick timeout kills it so the next tick proceeds).
Works with any healthchecks-style endpoint;
Golemreach Heartbeat has free self-serve
monitors with public status pages and webhook/email alerts (no account).
Full guide — including grace periods, the setting that prevents false
alarms, and a wrapper-script shape for single loops:
docs/unattended-fleet.md.
Running claude -p loops instead? The same dead-man switch arms with a
native Stop hook — a JSON snippet, no script: copy
adapters/claude-code/settings.json,
paste your monitor key, done. Every completed turn pings; silence pages you.
Guide with the upstream-bug workarounds (why not SessionEnd/SubagentStop),
cron examples and the matching kill-switch one-liner:
docs/claude-code-loops.md.
| Path | Purpose |
|---|---|
templates/supervisor.sh |
Continuous runner: fires the loop back-to-back, no idle gaps; FLEET_STOP kill switch; memory guard |
templates/run-worker.sh |
One fresh-context session per tick: flock locking, hard timeout, per-tick prompt snapshot + JSONL log, 14-day retention |
templates/fleet-worker@.service |
Systemd user-unit template — one unit per loop (fleet-worker@<name>) |
templates/FLEET_STATE.md |
Shared-brain state file: ranked queue, blockers, append-only logs |
templates/prompts/ |
Example builder/scout loop prompts |
plugins/heartbeat.js |
session.status busy→idle → heartbeat ping (dead-man switch for any opencode session) |
adapters/claude-code/settings.json |
Native Stop-hook snippet — same dead-man switch for Claude Code loops, zero scripts |
recipes/unattended-loop.sh |
Single-loop wrapper: timeout + explicit success/fail pings |
docs/unattended-fleet.md |
The unattended-fleet setup guide |
docs/claude-code-loops.md |
Dead-man your Claude Code loops (Stop hook + gate pairing) |
Design notes:
- Fresh context per tick. Each run starts a new opencode session reading the shared state file — no drifting conversation memory, full audit trail.
- Locks, not hopes.
flockguarantees overlapping ticks can't double-run. - Timeouts everywhere. A wedged tick is killed at
FLEET_TICK_TIMEOUT(default 55 min); the next tick proceeds cleanly. - The state file is the interface. Loops coordinate only through
FLEET_STATE.md; prompts are instructions, never coordination channels.
The kit is MIT-licensed and free, and that won't change. If it saves your fleet's night, SUPPORT.md describes Fleet Kit Pro ($19 one-time: our production prompt library, coordinator standing orders, postmortem templates).
MIT