fix(ai-gateway): real root cause — env resolution + admin probe endpoint - #22
Merged
Merged
Conversation
…ay-probe PR #21 added a model fallback chain inside generateDraft, but uploads still fail. Tracing the code: every model in the chain goes through getGatewayModel() — and that helper hard-coded the Lovable gateway URL and only read LOVABLE_API_KEY. The deployment is on the generic OpenAI-compatible gateway: .env.example documents AI_GATEWAY_BASE_URL + AI_GATEWAY_API_KEY + AI_GATEWAY_MODEL, not LOVABLE_API_KEY. So getGatewayModel threw "LOVABLE_API_KEY missing" 500 BEFORE the fallback chain ever called the model — the fallback chain just retried the same throw, four times. This change: * getGatewayModel now resolves the gateway from env in priority order: AI_GATEWAY_API_KEY (+ AI_GATEWAY_BASE_URL) > LOVABLE_API_KEY (legacy Lovable URL) > OPENAI_API_KEY (direct OpenAI). The unresolved-config error names every env var so the operator sees the cause immediately. * Auth-header style adapts to the gateway: bearer for the generic / OpenAI paths, Lovable-API-Key header for Lovable. * New `getGatewayModel("default")` resolves to AI_GATEWAY_MODEL env (falls back to google/gemini-2.5-flash). The author fallback chain now tries `default` FIRST so a well-configured deployment doesn't pay the latency of a doomed initial attempt. Removes the dead "google/gemini-3-flash-preview" entry; adds gpt-4o-mini and gpt-4o so the chain still recovers when only an OpenAI gateway is available. * describeGatewayConfig() — diagnostic helper returning which env supplied the key + base URL + default model. Logged once per author attempt so a future "what's wrong" investigation is one log-grep away. * New /api/admin/ai-gateway-probe (admin only) returns the resolved config AND runs a 1-token round-trip against the default model, reporting latency or the provider's error verbatim. This is the thing I wanted to run from a browser tab when uploads broke. https://claude.ai/code/session_019gMoupKKTVydpNwiiACQRd
criptogus
marked this pull request as ready for review
May 23, 2026 02:47
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why #21 didn't fix it
I couldn't reach the production MCP from the sandbox (network blocked), so I re-read the code instead and spotted the actual root cause: the model fallback chain from #21 never had a chance to run.
Every model attempt goes through
getGatewayModel()— and that helper hard-coded the Lovable gateway URL and only readLOVABLE_API_KEY. The deployment is on the generic OpenAI-compatible gateway:.env.exampledocumentsAI_GATEWAY_BASE_URL+AI_GATEWAY_API_KEY+AI_GATEWAY_MODEL, notLOVABLE_API_KEY. SogetGatewayModelthrew"LOVABLE_API_KEY missing"500 before the model id ever reached a provider. The fallback chain then dutifully retried the same throw four times and gave up.Fix
src/lib/ai-gateway.tsResolve config in priority order:
AI_GATEWAY_API_KEY(+AI_GATEWAY_BASE_URL) — generic gateway, bearer auth.LOVABLE_API_KEY— legacy Lovable gateway,Lovable-API-Keyheader.OPENAI_API_KEY— direct OpenAI.If none are set, the error names every env var so an operator sees the cause immediately in Vercel logs.
New
getGatewayModel("default")resolves to theAI_GATEWAY_MODELenv var (defaultgoogle/gemini-2.5-flash).src/lib/admin/author.server.ts"default"so a well-configured deployment doesn't pay the latency of a doomed initial attempt.google/gemini-3-flash-preview(preview SKU that was the original failure point); addedopenai/gpt-4o-miniandopenai/gpt-4oso the chain still recovers when only an OpenAI gateway is configured.defaultagainst the explicit list so the same model isn't tried twice.describeGatewayConfig()once per author run so a future "what's wrong" investigation is one log-grep away.src/routes/api/admin/ai-gateway-probe.ts(new)GET /api/admin/ai-gateway-probe(admin auth required) returns:{ "config": { "configured": true, "source": "AI_GATEWAY_API_KEY", "baseURL": "https://api.openai.com/v1", "defaultModel": "openai/gpt-4o-mini", "authStyle": "bearer" }, "probe": { "ok": true, "ms": 412, "text": "ok" } }On failure,
probe.errorcarries the provider's response verbatim. This is the thing I wanted to hit from a browser tab when uploads first broke.Verification path for the operator
After deploy:
GET https://superagentskill.com/api/admin/ai-gateway-probe(with admin session) → confirmconfig.configured: trueandprobe.ok: true. If probe fails, theerrorfield tells you exactly what the upstream gateway returned.upload_packagesvia MCP → should now produce a private draft. If not, Vercel logs will show[skillforge.author] <model> failed: <upstream error>for each attempt — that's the new actionable signal.Files
src/lib/ai-gateway.ts— rewrittensrc/lib/admin/author.server.ts— fallback chain + config loggingsrc/routes/api/admin/ai-gateway-probe.ts(new)https://claude.ai/code/session_019gMoupKKTVydpNwiiACQRd
Generated by Claude Code