Skip to content

fix(ai-gateway): real root cause — env resolution + admin probe endpoint - #22

Merged
criptogus merged 1 commit into
mainfrom
claude/fix-ai-gateway-env-resolution
May 23, 2026
Merged

criptogus merged 1 commit into
mainfrom
claude/fix-ai-gateway-env-resolution

Conversation

@criptogus

Copy link
Copy Markdown
Owner

Why #21 didn't fix it

I couldn't reach the production MCP from the sandbox (network blocked), so I re-read the code instead and spotted the actual root cause: the model fallback chain from #21 never had a chance to run.

Every model attempt goes through getGatewayModel() — and that helper hard-coded the Lovable gateway URL and only read LOVABLE_API_KEY. The deployment is on the generic OpenAI-compatible gateway: .env.example documents AI_GATEWAY_BASE_URL + AI_GATEWAY_API_KEY + AI_GATEWAY_MODEL, not LOVABLE_API_KEY. So getGatewayModel threw "LOVABLE_API_KEY missing" 500 before the model id ever reached a provider. The fallback chain then dutifully retried the same throw four times and gave up.

Fix

src/lib/ai-gateway.ts

Resolve config in priority order:

  1. AI_GATEWAY_API_KEY (+ AI_GATEWAY_BASE_URL) — generic gateway, bearer auth.
  2. LOVABLE_API_KEY — legacy Lovable gateway, Lovable-API-Key header.
  3. OPENAI_API_KEY — direct OpenAI.

If none are set, the error names every env var so an operator sees the cause immediately in Vercel logs.

New getGatewayModel("default") resolves to the AI_GATEWAY_MODEL env var (default google/gemini-2.5-flash).

src/lib/admin/author.server.ts

  • Fallback chain now starts with "default" so a well-configured deployment doesn't pay the latency of a doomed initial attempt.
  • Dropped google/gemini-3-flash-preview (preview SKU that was the original failure point); added openai/gpt-4o-mini and openai/gpt-4o so the chain still recovers when only an OpenAI gateway is configured.
  • De-dups default against the explicit list so the same model isn't tried twice.
  • Logs describeGatewayConfig() once per author run so a future "what's wrong" investigation is one log-grep away.

src/routes/api/admin/ai-gateway-probe.ts (new)

GET /api/admin/ai-gateway-probe (admin auth required) returns:

{
  "config": {
    "configured": true,
    "source": "AI_GATEWAY_API_KEY",
    "baseURL": "https://api.openai.com/v1",
    "defaultModel": "openai/gpt-4o-mini",
    "authStyle": "bearer"
  },
  "probe": { "ok": true, "ms": 412, "text": "ok" }
}

On failure, probe.error carries the provider's response verbatim. This is the thing I wanted to hit from a browser tab when uploads first broke.

Verification path for the operator

After deploy:

  1. GET https://superagentskill.com/api/admin/ai-gateway-probe (with admin session) → confirm config.configured: true and probe.ok: true. If probe fails, the error field tells you exactly what the upstream gateway returned.
  2. Try a 1-file upload_packages via MCP → should now produce a private draft. If not, Vercel logs will show [skillforge.author] <model> failed: <upstream error> for each attempt — that's the new actionable signal.

Files

  • src/lib/ai-gateway.ts — rewritten
  • src/lib/admin/author.server.ts — fallback chain + config logging
  • src/routes/api/admin/ai-gateway-probe.ts (new)

https://claude.ai/code/session_019gMoupKKTVydpNwiiACQRd


Generated by Claude Code

…ay-probe

PR #21 added a model fallback chain inside generateDraft, but uploads
still fail. Tracing the code: every model in the chain goes through
getGatewayModel() — and that helper hard-coded the Lovable gateway
URL and only read LOVABLE_API_KEY.

The deployment is on the generic OpenAI-compatible gateway: .env.example
documents AI_GATEWAY_BASE_URL + AI_GATEWAY_API_KEY + AI_GATEWAY_MODEL,
not LOVABLE_API_KEY. So getGatewayModel threw "LOVABLE_API_KEY missing"
500 BEFORE the fallback chain ever called the model — the fallback
chain just retried the same throw, four times.

This change:

* getGatewayModel now resolves the gateway from env in priority order:
  AI_GATEWAY_API_KEY (+ AI_GATEWAY_BASE_URL) > LOVABLE_API_KEY (legacy
  Lovable URL) > OPENAI_API_KEY (direct OpenAI). The unresolved-config
  error names every env var so the operator sees the cause immediately.
* Auth-header style adapts to the gateway: bearer for the generic /
  OpenAI paths, Lovable-API-Key header for Lovable.
* New `getGatewayModel("default")` resolves to AI_GATEWAY_MODEL env
  (falls back to google/gemini-2.5-flash). The author fallback chain
  now tries `default` FIRST so a well-configured deployment doesn't
  pay the latency of a doomed initial attempt. Removes the dead
  "google/gemini-3-flash-preview" entry; adds gpt-4o-mini and gpt-4o
  so the chain still recovers when only an OpenAI gateway is
  available.
* describeGatewayConfig() — diagnostic helper returning which env
  supplied the key + base URL + default model. Logged once per
  author attempt so a future "what's wrong" investigation is one
  log-grep away.
* New /api/admin/ai-gateway-probe (admin only) returns the resolved
  config AND runs a 1-token round-trip against the default model,
  reporting latency or the provider's error verbatim. This is the
  thing I wanted to run from a browser tab when uploads broke.

https://claude.ai/code/session_019gMoupKKTVydpNwiiACQRd
@criptogus
criptogus marked this pull request as ready for review May 23, 2026 02:47
@criptogus
criptogus merged commit dbbd144 into main May 23, 2026
1 check passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants