fix: reap orphaned Daytona sandboxes from the sync media-fetch path - #49
Merged
Merged
Conversation
External API callers (e.g. a Hermes playbook step) that hit
crayo.run_autoclip without an AgentRun runId fall into the older,
synchronous fetchPageVideoToCrayoAsset() path instead of the tracked
background job. That path's `finally { sandbox.delete() }` never runs
if the caller's serverless function is killed on timeout rather than
throwing, so a failing/retrying external caller can leak one Daytona
sandbox per attempt well before its own 10-minute autoStopInterval
catches up — and a burst of these can exhaust the account's sandbox
concurrency limit, stalling unrelated, properly-tracked AutoClip runs
(including ones started from the Agent tab UI).
Add reapStaleMediaFetchSandboxes(), wired into the existing
/api/cron/ops backstop (runs every 15 min), to force-stop/delete
sandboxes labeled purpose=media-fetch older than 8 minutes. Scoped
only to that label so it never touches media-fetch-job sandboxes,
which are already tracked by tickMediaFetchJob/abortMediaFetchJob and
have their own 60-minute autoStopInterval.
Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01LYohqvCYqttwGukV8MDRYV
|
Deployment failed for project clippyos with the following error: Learn More: https://vercel.link/3Fpeeb1 |
SomeRandmGuyy
deleted the
devin/1789013883-media-fetch-sandbox-reaper
branch
September 10, 2026 04:21
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
Adds a periodic reaper for orphaned Daytona sandboxes leaked by the synchronous (non-job)
crayo.run_autoclippath, and wires it into the existing/api/cron/opsbackstop.Why
While investigating a stuck live AutoClip run, I found a separate, recurring automated caller (Hermes, via the API channel, actor label
crayo-autoclip) hittingcrayo.run_autoclipevery 2–50 minutes and failing every time ("Agent activity" log, clientm2deZvboUP1FpkCjjPaWIhoLHDxKEpXT).Root cause:
crayo.run_autoclipcalls that arrive without anAgentRunrunId(i.e. fromautonomy-actions.server.ts, used by external API/Hermes callers — as opposed to the Agent tab UI, which always has arunId) fall into the older, synchronousfetchPageVideoToCrayoAsset()path inmedia-fetch.server.ts, not the tracked background-job path (media-fetch-job.server.ts, fixed for hangs in #43).That synchronous path creates its own Daytona sandbox and does the whole yt-dlp install/probe/download/upload inline within one serverless function call, budgeted just under Vercel's function-execution ceiling. Its cleanup lives in a
finally { sandbox.delete() }— but if the function is killed on timeout instead of throwing, thatfinallynever runs, orphaning the sandbox until Daytona's own 10-minuteautoStopIntervalcatches up. With this external caller retrying every few minutes, multiple orphaned sandboxes can be alive concurrently, plausibly exhausting the Daytona account's concurrency limit — which is consistent with why a fresh, unrelated, manually-triggered AutoClip run (via the safe async path) was hanging indefinitely at "Sandbox started. Installing yt-dlp..." waiting for capacity.There was no sandbox reaper/cleanup job anywhere in the codebase before this change.
What this PR does
reapStaleMediaFetchSandboxes()inmedia-fetch-job.server.ts: lists Daytona sandboxes labeledpurpose: media-fetch(the untracked synchronous path only — notmedia-fetch-job, which is already tracked bytickMediaFetchJob/abortMediaFetchJoband has its own 60-minute auto-stop), force-stops/deletes any older than 8 minutes and not already stopped./api/cron/ops, which already runs every 15 minutes as a backstop, alongside the existing Linear queue sweep and media-fetch job ticker.This does not change the fast-path/hang-fix logic from #43 — it's a safety net that stops orphaned sandboxes from silently eating account-level capacity, regardless of exactly why an individual synchronous call failed.
Verification
npm run typecheck— clean.npm run lint— 0 errors; only pre-existing warnings (confirmed viagit stashdiff — none newly introduced by this change).autonomy-actions.server.ts→handleCrayoActionwith norunId→runAutoclip→fetchPageVideoToCrayoAsset) and confirmed thefinallycleanup gap and the Daytona SDK'sdaytona.list({ labels, limit })shape used elsewhere indaytona.server.ts(collectLabeled).media-fetch-labeled sandbox count stays low.Follow-up (not in this PR)
Pausing the failing Hermes
crayo.run_autoclipplaybook step itself should still happen separately (I couldn't cleanly identify which of the several "Hermes Agent" API keys in Settings → Automation & Hermes is the one calling in ascrayo-autoclip, since the UI's "Last used" timestamps didn't match the very recent activity-log entries) — this PR only stops the resulting sandbox leak, it doesn't stop the retries themselves.🤖 Generated with Claude Code
https://claude.ai/code/session_01LYohqvCYqttwGukV8MDRYV