Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,8 @@ npm install

# Start the server (keep this running)
npm run dev
# ...or on macOS, once, to keep it running across reboots:
./scripts/install-launchd.sh
```

### Verify (in a second terminal)
Expand Down
9 changes: 9 additions & 0 deletions docs/TROUBLESHOOTING.md
Original file line number Diff line number Diff line change
Expand Up @@ -21,6 +21,15 @@ cd ~/Documents/ensemble && nohup ./node_modules/.bin/tsx server.ts > /tmp/ensemb
Then re-run `/collab` from the same shell where `claude --print "ok"` and
`codex login status` both succeed.

On macOS you can hand the service to launchd instead, once:
```bash
./scripts/install-launchd.sh
```
It then starts at login, comes back after a crash, and restarts on demand with
`launchctl kickstart -k gui/$(id -u)/dev.ensemble.server`. The agent carries the
PATH of the shell that installed it, so install from a shell where the agent CLIs
work. When preflight finds the service older than 24h it does the kickstart itself.

### Why the preflight catches it
`scripts/collab-preflight.sh` runs before every team spawn, and checks only the CLIs the run
actually needs (`collab-preflight.sh codex,claude,grok`, or `COLLAB_AGENTS`):
Expand Down
3 changes: 2 additions & 1 deletion docs/architecture.md
Original file line number Diff line number Diff line change
Expand Up @@ -65,7 +65,8 @@ ensemble/
│ ├── collab-livefeed.sh # Continuous live feed
│ ├── collab-status.sh # Multi-team dashboard
│ ├── collab-replay.sh # Session replay
│ ├── collab-cleanup.sh # Temp file cleanup
│ ├── collab-cleanup.sh # Finished + abandoned runtime dir cleanup
│ ├── install-launchd.sh # macOS launchd agent for the server
│ ├── team-say.sh # Agent message send
│ ├── team-read.sh # Agent message read
│ ├── ensemble-bridge.sh # File→HTTP message bridge
Expand Down
8 changes: 7 additions & 1 deletion docs/collab-scripts.md
Original file line number Diff line number Diff line change
Expand Up @@ -196,13 +196,19 @@ Shows: team name, status (active/finished/stale), message count, last message, d

## collab-cleanup.sh

**Remove finished team runtime directories** from `/tmp/ensemble/`. Dry-run by default.
**Remove finished and abandoned team runtime directories** from `/tmp/ensemble/`. Dry-run by default.

```bash
./scripts/collab-cleanup.sh # list what would be deleted
./scripts/collab-cleanup.sh --force # actually delete
```

Finished directories (with a `.finished` marker) are removed after 24h, the latest
three are always kept. Abandoned directories, without a marker and without a single
message (a launch that died before the agents spoke, a stray lock directory), are
removed after 24h as well. A directory that holds messages is never touched: the
team may still be running, and `collab-history.py` reads those messages later.

This does not disband a running team. To end one, press `d` in the monitor or
`POST /api/ensemble/teams/<id>/disband`.

Expand Down
5 changes: 5 additions & 0 deletions docs/getting-started.md
Original file line number Diff line number Diff line change
Expand Up @@ -102,6 +102,11 @@ npm run dev

You should see: `[Ensemble] Server running on http://127.0.0.1:23000`

On macOS, `./scripts/install-launchd.sh` registers the server as a launchd agent
instead: it starts at login, restarts after a crash, and logs to
`/tmp/ensemble-server.log`. Run it from a shell where your agent CLIs work; the
agent inherits that shell's PATH. Remove it again with `--uninstall`.

### 3. Verify (in a second terminal)

```bash
Expand Down
78 changes: 71 additions & 7 deletions scripts/collab-cleanup.sh
Original file line number Diff line number Diff line change
@@ -1,13 +1,18 @@
#!/usr/bin/env bash
# collab-cleanup.sh — Clean up old finished collab runtime directories.
# collab-cleanup.sh — Clean up old finished and abandoned collab runtime directories.
# Usage: collab-cleanup.sh [--force]
#
# Finished: has a .finished marker (written on disband). The latest 3 and anything
# under 24h are kept. Abandoned: no marker and never a message (a launch that
# died before the agents spoke, or a stray lock directory); removed after 24h.
# COLLAB_RUNTIME_ROOT overrides the root, for tests.
set -euo pipefail

SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
# shellcheck source=./collab-paths.sh
source "$SCRIPT_DIR/collab-paths.sh"

ENSEMBLE_ROOT="/tmp/ensemble"
ENSEMBLE_ROOT="${COLLAB_RUNTIME_ROOT:-/tmp/ensemble}"
KEEP_RECENT=3
MIN_AGE_SECONDS=$((24 * 60 * 60))
MODE="dry-run"
Expand Down Expand Up @@ -55,6 +60,29 @@ finished_entries() {
done | sort -rn
}

# Newest mtime of a directory and everything in it.
newest_mtime() {
local dir="${1:?dir required}" newest=0 ts
while IFS= read -r -d '' entry; do
ts="$(mtime_epoch "$entry")"
[ "$ts" -gt "$newest" ] && newest="$ts"
done < <(find "$dir" -print0)
printf '%s\n' "$newest"
}

# Directories without a .finished marker that never held a message. A team that
# spoke keeps its directory whatever its state: it may still be running, and its
# messages are what collab-history.py reads later.
abandoned_entries() {
[ -d "$ENSEMBLE_ROOT" ] || return 0
find "$ENSEMBLE_ROOT" -mindepth 1 -maxdepth 1 -type d -print0 |
while IFS= read -r -d '' runtime_dir; do
[ -f "$runtime_dir/.finished" ] && continue
[ -s "$runtime_dir/messages.jsonl" ] && continue
printf '%s\t%s\n' "$(newest_mtime "$runtime_dir")" "$runtime_dir"
done | sort -rn
}

for arg in "$@"; do
case "$arg" in
--force)
Expand All @@ -78,7 +106,15 @@ while IFS= read -r line; do
ENTRIES+=("$line")
done < <(finished_entries)

ABANDONED=()
while IFS= read -r line; do
ABANDONED+=("$line")
done < <(abandoned_entries)

TOTAL_FINISHED="${#ENTRIES[@]}"
TOTAL_ABANDONED="${#ABANDONED[@]}"
ABANDONED_ELIGIBLE=0
ABANDONED_REMOVED=0
PRESERVED_RECENT=0
PRESERVED_FRESH=0
ELIGIBLE=0
Expand All @@ -98,12 +134,13 @@ if [ ! -d "$ENSEMBLE_ROOT" ]; then
exit 0
fi

if [ "$TOTAL_FINISHED" -eq 0 ]; then
echo -e " ${Y}No finished collabs found.${R}"
if [ "$TOTAL_FINISHED" -eq 0 ] && [ "$TOTAL_ABANDONED" -eq 0 ]; then
echo -e " ${Y}No finished or abandoned collabs found.${R}"
exit 0
fi

for idx in "${!ENTRIES[@]}"; do
# ${ARR[@]+"${ARR[@]}"} keeps set -u happy on an empty array under macOS bash 3.2.
for idx in ${ENTRIES[@]+"${!ENTRIES[@]}"}; do
entry="${ENTRIES[$idx]}"
finished_ts="${entry%%$'\t'*}"
runtime_dir="${entry#*$'\t'}"
Expand Down Expand Up @@ -142,18 +179,45 @@ for idx in "${!ENTRIES[@]}"; do
fi
done

for entry in ${ABANDONED[@]+"${ABANDONED[@]}"}; do
newest_ts="${entry%%$'\t'*}"
runtime_dir="${entry#*$'\t'}"
runtime_name="$(basename "$runtime_dir")"
age_seconds=$((NOW - newest_ts))
age_hours=$((age_seconds / 3600))

if [ "$age_seconds" -lt "$MIN_AGE_SECONDS" ]; then
echo -e " ${C}skip${R} ${runtime_name} ${D}(no messages yet, ${age_hours}h old, may still be starting)${R}"
continue
fi

ABANDONED_ELIGIBLE=$((ABANDONED_ELIGIBLE + 1))
if [ "$MODE" = "force" ]; then
if rm -rf "$runtime_dir"; then
ABANDONED_REMOVED=$((ABANDONED_REMOVED + 1))
echo -e " ${G}remove${R} ${runtime_name} ${D}(abandoned, never a message, ${age_hours}h old)${R}"
else
FAILED=$((FAILED + 1))
echo -e " ${Y}failed${R} ${runtime_name} ${D}(abandoned, ${age_hours}h old)${R}"
fi
else
echo -e " ${Y}would rm${R} ${runtime_name} ${D}(abandoned, never a message, ${age_hours}h old)${R}"
fi
done

echo ""
echo -e " ${BD}Stats${R}"
echo -e " finished dirs: ${TOTAL_FINISHED}"
echo -e " kept (latest 3): ${PRESERVED_RECENT}"
echo -e " kept (<24h): ${PRESERVED_FRESH}"
echo -e " eligible old dirs: ${ELIGIBLE}"
echo -e " abandoned dirs: ${TOTAL_ABANDONED} (${ABANDONED_ELIGIBLE} older than 24h)"
if [ "$MODE" = "force" ]; then
echo -e " removed dirs: ${REMOVED}"
echo -e " removed dirs: ${REMOVED} finished, ${ABANDONED_REMOVED} abandoned"
echo -e " failed removals: ${FAILED}"
echo -e " reclaimed: $(human_kb "$REMOVED_KB")"
else
echo -e " would remove: ${ELIGIBLE}"
echo -e " would remove: ${ELIGIBLE} finished, ${ABANDONED_ELIGIBLE} abandoned"
echo -e " reclaimable: $(human_kb "$REMOVED_KB")"
fi
echo -e " finished footprint: $(human_kb "$TOTAL_KB")"
Expand Down
27 changes: 24 additions & 3 deletions scripts/collab-preflight.sh
Original file line number Diff line number Diff line change
Expand Up @@ -92,11 +92,32 @@ if [ -n "$SERVER_PID" ]; then
if [ -n "$AGE_SECS" ]; then
AGE_HRS=$((AGE_SECS / 3600))
if [ "$AGE_HRS" -gt "$SERVICE_MAX_AGE_HOURS" ]; then
fail 2 "Ensemble service is ${AGE_HRS}h old (>${SERVICE_MAX_AGE_HOURS}h threshold)
# Under launchd (scripts/install-launchd.sh) a stale service is a restart
# away, so do that here instead of sending the user to a shell one-liner.
LAUNCHD_LABEL="${ENSEMBLE_LAUNCHD_LABEL:-dev.ensemble.server}"
LAUNCHD_TARGET="gui/$(id -u)/$LAUNCHD_LABEL"
if command -v launchctl > /dev/null 2>&1 && launchctl print "$LAUNCHD_TARGET" > /dev/null 2>&1; then
warn "Ensemble service is ${AGE_HRS}h old — restarting it through launchd ($LAUNCHD_LABEL)"
launchctl kickstart -k "$LAUNCHD_TARGET"
RESTARTED=0
for _ in $(seq 1 10); do
sleep 1
if curl -sf "$API/api/v1/health" > /dev/null 2>&1; then RESTARTED=1; break; fi
done
if [ "$RESTARTED" = 1 ]; then
ok "Ensemble service restarted (fresh process under launchd)"
else
fail 2 "Ensemble service did not come back within 10s after launchctl kickstart; check /tmp/ensemble-server.log"
fi
else
fail 2 "Ensemble service is ${AGE_HRS}h old (>${SERVICE_MAX_AGE_HOURS}h threshold)
This is the 2026-05-08 stale-state issue: agents will spawn with broken auth.
Fix: pkill -f 'tsx server.ts' && cd ~/Documents/ensemble && nohup ./node_modules/.bin/tsx server.ts > /tmp/ensemble-server.log 2>&1 &"
Fix: pkill -f 'tsx server.ts' && cd ~/Documents/ensemble && nohup ./node_modules/.bin/tsx server.ts > /tmp/ensemble-server.log 2>&1 &
Or install the launchd agent once (scripts/install-launchd.sh) and preflight restarts it for you."
fi
else
ok "Ensemble service age: ${AGE_HRS}h (within ${SERVICE_MAX_AGE_HOURS}h limit)"
fi
ok "Ensemble service age: ${AGE_HRS}h (within ${SERVICE_MAX_AGE_HOURS}h limit)"
else
warn "Could not determine service age (no usable ps) — stale-service check skipped"
fi
Expand Down
107 changes: 107 additions & 0 deletions scripts/install-launchd.sh
Original file line number Diff line number Diff line change
@@ -0,0 +1,107 @@
#!/usr/bin/env bash
# install-launchd.sh — Keep the ensemble service running as a macOS launchd agent.
# Usage: install-launchd.sh [--uninstall] [--no-load]
#
# Writes ~/Library/LaunchAgents/<label>.plist that runs `tsx server.ts` from this
# repo, starts it at login and restarts it whenever it exits. After a code change:
# launchctl kickstart -k gui/$(id -u)/<label>
# collab-preflight.sh does that by itself when the service is older than 24h.
#
# The agent inherits the PATH of the shell that installs it, so the agent CLIs
# (claude, codex, grok) and tmux resolve the same way they do in your terminal.
# Label override: ENSEMBLE_LAUNCHD_LABEL (default dev.ensemble.server).
set -euo pipefail

SCRIPT_DIR="$(cd "$(dirname "${BASH_SOURCE[0]}")" && pwd)"
REPO_DIR="$(cd "$SCRIPT_DIR/.." && pwd)"
LABEL="${ENSEMBLE_LAUNCHD_LABEL:-dev.ensemble.server}"
PLIST="$HOME/Library/LaunchAgents/$LABEL.plist"
LOG_FILE="/tmp/ensemble-server.log"
DOMAIN="gui/$(id -u)"

UNINSTALL=0
LOAD=1
for arg in "$@"; do
case "$arg" in
--uninstall) UNINSTALL=1 ;;
--no-load) LOAD=0 ;;
-h|--help)
sed -n '2,12p' "$0" | sed 's/^# \{0,1\}//'
exit 0
;;
*) echo "Unknown argument: $arg" >&2; exit 1 ;;
esac
done

if [ "$UNINSTALL" = 1 ]; then
if [ "$LOAD" = 1 ]; then
launchctl bootout "$DOMAIN/$LABEL" 2>/dev/null || true
fi
rm -f "$PLIST"
echo "Removed $PLIST"
exit 0
fi

if [ "$(uname)" != "Darwin" ]; then
echo "launchd only exists on macOS; use a systemd unit or your init system instead." >&2
exit 1
fi

NODE_BIN="$(command -v node || true)"
if [ -z "$NODE_BIN" ]; then
echo "node not found on PATH" >&2
exit 1
fi
TSX_BIN="$REPO_DIR/node_modules/.bin/tsx"
if [ ! -x "$TSX_BIN" ]; then
echo "$TSX_BIN missing; run npm install in $REPO_DIR first" >&2
exit 1
fi

xml_escape() {
printf '%s' "$1" | sed -e 's/&/\&amp;/g' -e 's/</\&lt;/g' -e 's/>/\&gt;/g'
}

mkdir -p "$(dirname "$PLIST")"
cat > "$PLIST" <<EOF
<?xml version="1.0" encoding="UTF-8"?>
<!DOCTYPE plist PUBLIC "-//Apple//DTD PLIST 1.0//EN" "http://www.apple.com/DTDs/PropertyList-1.0.dtd">
<plist version="1.0"><dict>
<key>Label</key><string>$LABEL</string>
<key>ProgramArguments</key>
<array>
<string>$(xml_escape "$NODE_BIN")</string>
<string>$(xml_escape "$TSX_BIN")</string>
<string>server.ts</string>
</array>
<key>WorkingDirectory</key><string>$(xml_escape "$REPO_DIR")</string>
<key>EnvironmentVariables</key>
<dict>
<key>PATH</key><string>$(xml_escape "$PATH")</string>
<key>HOME</key><string>$(xml_escape "$HOME")</string>
</dict>
<key>RunAtLoad</key><true/>
<key>KeepAlive</key><true/>
<key>ThrottleInterval</key><integer>5</integer>
<key>StandardOutPath</key><string>$LOG_FILE</string>
<key>StandardErrorPath</key><string>$LOG_FILE</string>
</dict></plist>
EOF
echo "Wrote $PLIST"

if [ "$LOAD" = 1 ]; then
# A loose `tsx server.ts` from a terminal would hold the port; hand over to launchd.
launchctl bootout "$DOMAIN/$LABEL" 2>/dev/null || true
pkill -f 'tsx server.ts' 2>/dev/null || true
sleep 1
launchctl bootstrap "$DOMAIN" "$PLIST"
for _ in $(seq 1 10); do
sleep 1
if curl -sf "http://localhost:${ENSEMBLE_PORT:-23000}/api/v1/health" > /dev/null 2>&1; then
echo "Service is up under launchd ($LABEL); log: $LOG_FILE"
exit 0
fi
done
echo "Service did not answer within 10s; check $LOG_FILE" >&2
exit 1
fi
Loading
Loading