Skip to content

fix(uploads): SkillForge text fallback + queue self-heal - #26

Merged
criptogus merged 1 commit into
mainfrom
claude/fix-mcp-oauth-callback-brZDc
May 28, 2026
Merged

criptogus merged 1 commit into
mainfrom
claude/fix-mcp-oauth-callback-brZDc

Conversation

@criptogus

Copy link
Copy Markdown
Owner

Summary

Diagnostic round 2 surfaced two production failures behind the freshly-shipped REST upload path. Same root areas, two different layers.

NEW-1: SkillForge author 0/1 on every inline upload

  • processBulkUpload was returning uploaded:0, failed:1 for every call. Both structured-output fallback models died:
    • openai/gpt-4o-mini → Bad Request (provider strict mode rejecting the JSON schema generated from PackageDraftSchema's open record<string, any> slots).
    • google/gemini-2.5-flash → No object generated: response did not match schema.
  • The catch swallowed the upstream body — only e.message reached logs, so the symptom was a generic "Bad Request" with no signal.

Fixes

  • describeAttemptError extracts e.text (NoObjectGeneratedError), e.responseBody (HTTP errors), and e.cause.message. Real provider body now lands in logs.
  • After all structured attempts exhaust, a text fallback runs: plain generateText asking for raw JSON, then parsed locally and validated with PackageDraftSchema. Tolerant extractor handles bare JSON, ```json fences, and brace-balanced first-object recovery. Loses provider-side schema enforcement but yields a draft instead of a fatal error.

NEW-2: queue zombies — uploads stuck in queued indefinitely

  • Worker hardcoded attempts: 1 on every claim, so MAX_ATTEMPTS gating could never trigger.
  • A job that crashed mid-flight (Vercel 60s budget, network flap) stayed in processing forever with no recovery path.

Fixes

  • Each drain pass scans processing jobs with started_at < now-2min:
    • attempts < 3 → flip back to queued, clear started_at.
    • attempts >= 3 → mark failed with an "abandoned" reason so the user sees a real outcome.
  • Claim increments attempts from the prior value instead of overwriting.
  • Return value adds requeued so cron logs surface the healing.

Out of scope (still TODO)

  • BUG-3 MCP-side Bearer sas_… rejection (suspected separate auth context inside the MCP tool, not the route). Will tackle next.
  • BUG-1 phantom super-agent CLI on npm — needs a publish or a snippet revision; not a code-in-this-repo fix.

Test plan

  • POST /api/packages/upload with a real file returns uploaded:1, failed:0 — text fallback should catch any remaining structured-output strictness issues.
  • When all models fail, response results[0].error now contains text=… or body=… snippets pointing at the upstream cause, not just "Bad Request".
  • Manually UPDATE package_upload_jobs SET status='processing', started_at=now()-interval '5 minutes' WHERE …; hit /api/jobs/drain-upload-queue?secret=…. Job flips back to queued (attempt 1) or failed with "abandoned" (attempt 3).
  • Drain response JSON now includes requeued counter.
  • Existing 14h-old zombie jobs reachable from /account/packages clear within one cron tick after this deploys.

https://claude.ai/code/session_019gMoupKKTVydpNwiiACQRd


Generated by Claude Code

Two related production failures observed in the 2nd-pass diagnostic:

NEW-1 — `processBulkUpload` returned `uploaded:0, failed:1` for every
inline call. Both fallback models died with provider-strict-mode
rejections of the structured-output JSON schema generated from
PackageDraftSchema (open `record<string, any>` slots → 400 on
gpt-4o-mini, "No object generated" on gemini-2.5-flash). The catch
swallowed the upstream body, so operators only saw "Bad Request".

  - Capture `e.text`, `e.responseBody`, `e.cause` per attempt so the
    real reason is in the logs instead of an opaque message.
  - After structured attempts exhaust, fall back to plain `generateText`
    asking for raw JSON, then parse + zod-validate locally. Loses
    provider-side schema enforcement but yields a draft instead of a
    fatal error. Tolerant extractor handles bare JSON, fenced JSON,
    and brace-balanced first-object recovery.

NEW-2 — uploads were stuck in `queued` indefinitely (>14h). Two bugs:
the worker hardcoded `attempts: 1` on each claim (so MAX_ATTEMPTS
gating could never trigger), and a job that crashed mid-flight stayed
in `processing` forever with no recovery path.

  - Each drain pass now scans for jobs in `processing` older than 2min
    and either requeues them (attempts < 3) or marks `failed` with an
    explicit "abandoned" reason (attempts >= 3). Returns `requeued`
    in the summary so cron logs make the healing visible.
  - Claim increments `attempts` from the prior value instead of
    overwriting to 1.

https://claude.ai/code/session_019gMoupKKTVydpNwiiACQRd
@criptogus
criptogus marked this pull request as ready for review May 28, 2026 02:31
@criptogus
criptogus merged commit 2033b8d into main May 28, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants