feat(gateway): answer Claude Code's one-token model probes - #440
Merged
Merged
Conversation
Switching models in Claude Code failed against any model whose candidate
translates Messages to Responses:
API error: 400 {"error":{"message":"Invalid 'max_output_tokens': integer
below minimum value. Expected a value >= 16, but got 1 instead.",
"code":"invalid_request_body"}}
Claude Code validates a model by generating one token against it. `/model
<id>` issues a non-streaming `POST /v1/messages?beta=true` with
`max_tokens: 1` and a fixed throwaway prompt, and calls the model unusable
if the request throws. Decompiled from the 2.1.226 binary, the CLI's
`model_validation` side query sends a single user turn holding one
ephemeral `Hi` text block; the same helper serves the CLI's other
one-token probes with their own fixed prompts (`quota` preflight, `test`
credential verification, `.` reachability). None of them wants generated
text — the caller reads `usage.input_tokens` / `usage.output_tokens` and
treats "did not throw" as the verdict.
A one-token cap is not portable: OpenAI's Responses API floors
`max_output_tokens` at 16 and rejects anything lower outright, so the
probe fails and the CLI reports the model missing.
Add an unconditional Messages interceptor that recognizes those probes and
answers them itself with a turn that stopped at the one-token cap — no
content blocks, `stop_reason: max_tokens`, and a zero `usage` block, which
Claude Code dereferences unconditionally. Registered ahead of the rest of
the chain, since it replaces the turn rather than shaping it.
Interception sits after model resolution, so an id no upstream serves
still fails at the serve layer with a 404 and the CLI still reports it not
found; only the pointless generation behind the probe is suppressed. The
turn makes no upstream call, records a usage row at zero, and contributes
no latency sample.
Detection requires the `claude-cli` User-Agent, `max_tokens: 1`, no tools,
and a single user turn holding one of the CLI's fixed probe prompts. That
is narrower than CLIProxyAPI, which short-circuits on `max_tokens === 1`
alone after reverting a tighter predicate; the extra conditions fail safe
toward today's behavior, at the cost of tracking a Claude Code build
detail.
The Messages module mints Anthropic-shaped opaque ids in two places — the `request_id` on a gateway-synthesized error envelope, and the message id on a gateway-answered turn — from the same expression. Give the module one minter parameterized by prefix. The rationale comment it carried was also wrong about its own output: `crypto.randomUUID().replace(/-/g, '').slice(0, 24)` yields 24 hex characters, not base62, and the first 24 hex characters of a UUIDv4 carry ~90 bits rather than ~96 because the version nibble is fixed and the variant nibble constrained.
…truthfully
Review of the probe interceptor found the accepted-prompt set too wide and
its research notes wrong on three points.
`quota` leaves the set. It is the CLI's rate-limit preflight, and its
caller consumes the `anthropic-ratelimit-unified-*` response headers rather
than the body — a synthesized turn cannot carry them, so answering it would
replace a working quota reading, on every upstream that serves the probe
today, with a silent blank. `.` leaves too: it is issued by the Bedrock /
Vertex / Mantle SDK clients, which route through their own base URLs and
never reach an `ANTHROPIC_BASE_URL` gateway. That leaves the `/model`
validation probe, which is the reported failure, and the credential check,
which reads nothing off the response at all.
Prompt matching becomes exact. Every literal is read off a binary rather
than inferred, so lowercasing the set and normalizing the input widened the
predicate to shapes no evidence attributes to the CLI, in a predicate the
file argues at length is deliberately narrow. An unobserved shape should
reach the upstream rather than be answered from a guess.
Corrections to the research notes, each verified against the 2.1.226 binary
or the GitHub API:
- Only the `/model` validation probe goes through the `eie` side-query
helper; `quota`, `test`, and `.` are independent call sites. The
"reads usage and treats no-throw as the verdict" claim holds for the
validation path only.
- CLIProxyAPI has no `max_tokens === 1` short-circuit; PR 3571 is an open
proposal to add one, and the revert it describes is internal to its own
branch history. It is not shipped behavior and not a baseline.
- The bare-string content form is what the current client sends for the
credential check, not a legacy shape. Labeling it as one read as a
compatibility shim for old clients, which it is not.
- `eie` accepts a `tools` argument; it is the one-token probes
specifically, not side queries in general, that never carry tools.
- Ordering the entry first needs no web-search-shim hazard to justify it;
the shim is inactive for a request with no tools regardless.
`isClaudeCodeProbe` goes back to module-private, matching its siblings, and
the prompt matrix now drives the interceptor rather than the predicate.
Two shape guards — a sole turn with more than one block, and a sole turn
that is not a user turn — had no test pinning them; both are now covered.
Round-2 review caught the accepted-prompt set asserting an evidence standard one of its members does not meet. `hello` came from a third-party bug report against v2.1.220, and re-reading that report shows its request body is a hand-written reproduction rather than a capture. No build on disk contains it: 2.1.221 through 2.1.226 all send `Hi` in ephemeral-block form at the `model_validation` call site, and `hello` appears in those binaries only in computer-use tool-schema examples and an IPC handshake frame. Matching became exact precisely so that an unobserved shape reaches the upstream instead of being answered from a guess, so the guessed entry goes. The litellm report stays as the reference for what it does support — that the literal is a build detail worth re-reading per release. Two comment corrections alongside it. The response-field inventory said `usage.input_tokens` / `usage.output_tokens` and `stop_reason` were the only fields read; the telemetry event also reads `_request_id` and the two cache counters, all behind `??`, so the load-bearing property is the unguarded dereference rather than the size of the inventory. And the test fixture's User-Agent note credited the header to the Anthropic SDK and presented one fixed string: Claude Code composes it itself as `claude-cli/<version> (external, <entrypoint>…)` with up to three optional trailing segments, which is why the predicate anchors on the prefix alone.
The shape of an Anthropic id — `msg_` for a message, `req_` for a request, followed by an opaque 24-character token — is a property of the Messages wire contract, not of the gateway that happens to synthesize bodies carrying one. Move the generator into `@floway-dev/protocols/messages` alongside `createRandomResponsesItemId`, its exact counterpart on the Responses side, and name it `generateAnthropicId` to match. Generating from `crypto.getRandomValues` rather than by slicing a UUID drops the version and variant nibbles the slice inherited, so the token carries 96 bits of entropy across its 24 characters instead of 90, and the comment no longer needs the caveat. The wire shape is unchanged.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Switching models in Claude Code fails against any model whose candidate translates Messages to Responses:
Claude Code validates a model by generating one token against it.
/model <id>issues a non-streamingPOST /v1/messages?beta=truewithmax_tokens: 1and a fixed throwaway prompt, and calls the model unusable if that request throws.Upstream research
Decompiled from the Claude Code 2.1.226 binary — the
model_validationside query:The verdict is "did not throw". The path dereferences
usage.input_tokens/usage.output_tokensunguarded to populate the CLI's own telemetry event, and reads_request_id,stop_reasonand the cache counters behind??. A missingusagethrows inside the CLI and reads as a failed probe.The CLI runs several other one-token probes, each from an independent call site rather than through
eie:quota(rate-limit preflight),test(credential verification),.(Bedrock / Vertex / Mantle reachability).A one-token cap is not portable: OpenAI's Responses API floors
max_output_tokensat 16 and rejects anything lower outright.Prior art: CLIProxyAPI#3571 proposes short-circuiting on
max_tokens === 1alone — it is open and unmerged, and CLIProxyAPI ships no such short-circuit today. LiteLLM#35061 reports this exact symptom and fixes it from the other side, by rewriting the upstream's output-limit error.Change
An unconditional
answerClaudeCodeProbeMessages interceptor, registered ahead of the rest of the chain because it replaces the turn rather than shaping it. On a match it answers with a turn that stopped at the one-token cap: no content blocks,stop_reason: max_tokens, zerousage.Detection requires all of: the
claude-cli/<semver>User-Agent prefix,max_tokens: 1, no tools, and a single user turn whose text matches one ofHi/testexactly.The synthesized turn needs an Anthropic-shaped
msg_id, which the Messages error envelopes already minted privately. That generator moves to@floway-dev/protocols/messagesasgenerateAnthropicId, alongsidecreateRandomResponsesItemId, its counterpart on the Responses side — the id shape is a property of the wire contract, not of the gateway that happens to synthesize bodies carrying one.Decisions worth a look
quotais excluded because its caller reads theanthropic-ratelimit-unified-*response headers, not the body — answering it would replace a working quota reading, on every upstream that serves the probe today, with a silent blank..is excluded because it only ever reaches a Bedrock / Vertex / Mantle base URL.requests: 1and no metrics — visible in the dashboard, billed at zero. It attaches noperformancecontext, so it contributes no latency sample: there was no upstream call to measure.max_tokensto the target's floor and passing through — keeps real upstream validation but bills a real turn. Out of scope here, and the two are not mutually exclusive.Test Plan
Driven against a local Node instance with a fake OpenAI-compatible upstream that reproduces the
max_output_tokens >= 16floor.claudeCLI:/model gpt-5.6-sol[1m]reportsSet model to gpt-5.6-sol[1m] for this session onlyinstead of the 400usage_requestsrow, zerousagetoken-metric rows, and zeroperformance_summaryrowsThere's an issue with the selected modelmessage_start→message_delta→message_stopSSE sequencepnpm run typegen,pnpm run lint,pnpm run typecheckpnpm run test— 534 files / 5531 testspnpm run test:installers,pnpm run check:agents-md,pnpm run check:generated-assets,pnpm run check:verify-parity,pnpm run build:web