Skip to content

fix: reap orphaned Daytona sandboxes from the sync media-fetch path - #49

Merged
SomeRandmGuyy merged 1 commit into
mainfrom
devin/1789013883-media-fetch-sandbox-reaper
Sep 10, 2026
Merged

SomeRandmGuyy merged 1 commit into
mainfrom
devin/1789013883-media-fetch-sandbox-reaper

Conversation

@SomeRandmGuyy

Copy link
Copy Markdown
Contributor

What

Adds a periodic reaper for orphaned Daytona sandboxes leaked by the synchronous (non-job) crayo.run_autoclip path, and wires it into the existing /api/cron/ops backstop.

Why

While investigating a stuck live AutoClip run, I found a separate, recurring automated caller (Hermes, via the API channel, actor label crayo-autoclip) hitting crayo.run_autoclip every 2–50 minutes and failing every time ("Agent activity" log, client m2deZvboUP1FpkCjjPaWIhoLHDxKEpXT).

Root cause: crayo.run_autoclip calls that arrive without an AgentRun runId (i.e. from autonomy-actions.server.ts, used by external API/Hermes callers — as opposed to the Agent tab UI, which always has a runId) fall into the older, synchronous fetchPageVideoToCrayoAsset() path in media-fetch.server.ts, not the tracked background-job path (media-fetch-job.server.ts, fixed for hangs in #43).

That synchronous path creates its own Daytona sandbox and does the whole yt-dlp install/probe/download/upload inline within one serverless function call, budgeted just under Vercel's function-execution ceiling. Its cleanup lives in a finally { sandbox.delete() } — but if the function is killed on timeout instead of throwing, that finally never runs, orphaning the sandbox until Daytona's own 10-minute autoStopInterval catches up. With this external caller retrying every few minutes, multiple orphaned sandboxes can be alive concurrently, plausibly exhausting the Daytona account's concurrency limit — which is consistent with why a fresh, unrelated, manually-triggered AutoClip run (via the safe async path) was hanging indefinitely at "Sandbox started. Installing yt-dlp..." waiting for capacity.

There was no sandbox reaper/cleanup job anywhere in the codebase before this change.

What this PR does

  • Adds reapStaleMediaFetchSandboxes() in media-fetch-job.server.ts: lists Daytona sandboxes labeled purpose: media-fetch (the untracked synchronous path only — not media-fetch-job, which is already tracked by tickMediaFetchJob/abortMediaFetchJob and has its own 60-minute auto-stop), force-stops/deletes any older than 8 minutes and not already stopped.
  • Wires it into /api/cron/ops, which already runs every 15 minutes as a backstop, alongside the existing Linear queue sweep and media-fetch job ticker.

This does not change the fast-path/hang-fix logic from #43 — it's a safety net that stops orphaned sandboxes from silently eating account-level capacity, regardless of exactly why an individual synchronous call failed.

Verification

  • npm run typecheck — clean.
  • npm run lint — 0 errors; only pre-existing warnings (confirmed via git stash diff — none newly introduced by this change).
  • Manually traced the code paths (autonomy-actions.server.ts → handleCrayoAction with no runId → runAutoclip → fetchPageVideoToCrayoAsset) and confirmed the finally cleanup gap and the Daytona SDK's daytona.list({ labels, limit }) shape used elsewhere in daytona.server.ts (collectLabeled).
  • Did not have direct Daytona API credentials in this session to enumerate live sandboxes and count actual orphans before/after — Ove, worth eyeballing the Daytona dashboard once this is live to confirm the media-fetch-labeled sandbox count stays low.

Follow-up (not in this PR)

Pausing the failing Hermes crayo.run_autoclip playbook step itself should still happen separately (I couldn't cleanly identify which of the several "Hermes Agent" API keys in Settings → Automation & Hermes is the one calling in as crayo-autoclip, since the UI's "Last used" timestamps didn't match the very recent activity-log entries) — this PR only stops the resulting sandbox leak, it doesn't stop the retries themselves.

🤖 Generated with Claude Code

https://claude.ai/code/session_01LYohqvCYqttwGukV8MDRYV

External API callers (e.g. a Hermes playbook step) that hit
crayo.run_autoclip without an AgentRun runId fall into the older,
synchronous fetchPageVideoToCrayoAsset() path instead of the tracked
background job. That path's `finally { sandbox.delete() }` never runs
if the caller's serverless function is killed on timeout rather than
throwing, so a failing/retrying external caller can leak one Daytona
sandbox per attempt well before its own 10-minute autoStopInterval
catches up — and a burst of these can exhaust the account's sandbox
concurrency limit, stalling unrelated, properly-tracked AutoClip runs
(including ones started from the Agent tab UI).

Add reapStaleMediaFetchSandboxes(), wired into the existing
/api/cron/ops backstop (runs every 15 min), to force-stop/delete
sandboxes labeled purpose=media-fetch older than 8 minutes. Scoped
only to that label so it never touches media-fetch-job sandboxes,
which are already tracked by tickMediaFetchJob/abortMediaFetchJob and
have their own 60-minute autoStopInterval.

Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LYohqvCYqttwGukV8MDRYV
@vercel

vercel Bot commented Sep 10, 2026

Copy link
Copy Markdown

Deployment failed for project clippyos with the following error:

Hobby accounts are limited to daily cron jobs. This cron expression (*/15 * * * *) would run more than once per day. Upgrade to the Pro plan to unlock all Cron Jobs features on Vercel.

Learn More: https://vercel.link/3Fpeeb1

@SomeRandmGuyy
SomeRandmGuyy merged commit 7069c25 into main Sep 10, 2026
1 of 6 checks passed
@SomeRandmGuyy
SomeRandmGuyy deleted the devin/1789013883-media-fetch-sandbox-reaper branch September 10, 2026 04:21

This branch was successfully deployed

1 active deployment
Preview – clippyos — 4b0403ff Deployed Sep 10, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant