Skip to content

feat(gateway): answer Claude Code's one-token model probes - #440

Merged
Menci merged 8 commits into
mainfrom
claude-code-model-probe
Aug 8, 2026
Merged

Menci merged 8 commits into
mainfrom
claude-code-model-probe

Conversation

@Menci

@Menci Menci commented Aug 8, 2026 •

Copy link
Copy Markdown
Owner

Problem

Switching models in Claude Code fails against any model whose candidate translates Messages to Responses:

❯ /model gpt-5.6-sol[1m]
  ⎿  API error: 400 {"error":{"message":"Invalid 'max_output_tokens': integer below minimum value. Expected a value >= 16, but got 1 instead.","code":"invalid_request_body"}}

Claude Code validates a model by generating one token against it. /model <id> issues a non-streaming POST /v1/messages?beta=true with max_tokens: 1 and a fixed throwaway prompt, and calls the model unusable if that request throws.

Upstream research

Decompiled from the Claude Code 2.1.226 binary — the model_validation side query:

await eie({model:r, max_tokens:1, maxRetries:0, querySource:"model_validation",
  messages:[{role:"user",content:[{type:"text",text:"Hi",cache_control:{type:"ephemeral"}}]}]}),
  lra.set(r,!0), {valid:!0}

The verdict is "did not throw". The path dereferences usage.input_tokens / usage.output_tokens unguarded to populate the CLI's own telemetry event, and reads _request_id, stop_reason and the cache counters behind ??. A missing usage throws inside the CLI and reads as a failed probe.

The CLI runs several other one-token probes, each from an independent call site rather than through eie: quota (rate-limit preflight), test (credential verification), . (Bedrock / Vertex / Mantle reachability).

A one-token cap is not portable: OpenAI's Responses API floors max_output_tokens at 16 and rejects anything lower outright.

Prior art: CLIProxyAPI#3571 proposes short-circuiting on max_tokens === 1 alone — it is open and unmerged, and CLIProxyAPI ships no such short-circuit today. LiteLLM#35061 reports this exact symptom and fixes it from the other side, by rewriting the upstream's output-limit error.

Change

An unconditional answerClaudeCodeProbe Messages interceptor, registered ahead of the rest of the chain because it replaces the turn rather than shaping it. On a match it answers with a turn that stopped at the one-token cap: no content blocks, stop_reason: max_tokens, zero usage.

Detection requires all of: the claude-cli/<semver> User-Agent prefix, max_tokens: 1, no tools, and a single user turn whose text matches one of Hi / test exactly.

The synthesized turn needs an Anthropic-shaped msg_ id, which the Messages error envelopes already minted privately. That generator moves to @floway-dev/protocols/messages as generateAnthropicId, alongside createRandomResponsesItemId, its counterpart on the Responses side — the id shape is a property of the wire contract, not of the gateway that happens to synthesize bodies carrying one.

Decisions worth a look

  • Which probes we answer. Only the ones a synthesized turn can satisfy honestly. quota is excluded because its caller reads the anthropic-ratelimit-unified-* response headers, not the body — answering it would replace a working quota reading, on every upstream that serves the probe today, with a silent blank. . is excluded because it only ever reaches a Bedrock / Vertex / Mantle base URL.
  • Exact matching. Every accepted literal is read off a binary rather than inferred, so an unobserved shape reaches the upstream instead of being answered from a guess. The cost is that these literals are a Claude Code build detail: a release that renames one re-exposes the 400 until the new literal is read off that build and added.
  • Interception sits after model resolution, inside the per-candidate interceptor chain. An id no upstream serves therefore still fails with a 404 and the CLI still reports it not found — real model validation is preserved; only the pointless generation behind the probe is suppressed.
  • Telemetry. The intercepted turn makes no upstream call and records a usage row with requests: 1 and no metrics — visible in the dashboard, billed at zero. It attaches no performance context, so it contributes no latency sample: there was no upstream call to measure.
  • Not done: clamping. LiteLLM's route — raising max_tokens to the target's floor and passing through — keeps real upstream validation but bills a real turn. Out of scope here, and the two are not mutually exclusive.

Test Plan

Driven against a local Node instance with a fake OpenAI-compatible upstream that reproduces the max_output_tokens >= 16 floor.

  • Real claude CLI: /model gpt-5.6-sol[1m] reports Set model to gpt-5.6-sol[1m] for this session only instead of the 400
  • The switch produces zero upstream requests, one usage_requests row, zero usage token-metric rows, and zero performance_summary rows
  • A real inference turn on the switched model still proxies and bills normally — 7 input / 3 output, one upstream request
  • A model no upstream serves is still rejected — the CLI reports There's an issue with the selected model
  • Streaming probe variant emits a well-formed message_start → message_delta → message_stop SSE sequence
  • pnpm run typegen, pnpm run lint, pnpm run typecheck
  • pnpm run test — 534 files / 5531 tests
  • pnpm run test:installers, pnpm run check:agents-md, pnpm run check:generated-assets, pnpm run check:verify-parity, pnpm run build:web

Menci added 8 commits August 8, 2026 17:50
Switching models in Claude Code failed against any model whose candidate
translates Messages to Responses:

  API error: 400 {"error":{"message":"Invalid 'max_output_tokens': integer
  below minimum value. Expected a value >= 16, but got 1 instead.",
  "code":"invalid_request_body"}}

Claude Code validates a model by generating one token against it. `/model
<id>` issues a non-streaming `POST /v1/messages?beta=true` with
`max_tokens: 1` and a fixed throwaway prompt, and calls the model unusable
if the request throws. Decompiled from the 2.1.226 binary, the CLI's
`model_validation` side query sends a single user turn holding one
ephemeral `Hi` text block; the same helper serves the CLI's other
one-token probes with their own fixed prompts (`quota` preflight, `test`
credential verification, `.` reachability). None of them wants generated
text — the caller reads `usage.input_tokens` / `usage.output_tokens` and
treats "did not throw" as the verdict.

A one-token cap is not portable: OpenAI's Responses API floors
`max_output_tokens` at 16 and rejects anything lower outright, so the
probe fails and the CLI reports the model missing.

Add an unconditional Messages interceptor that recognizes those probes and
answers them itself with a turn that stopped at the one-token cap — no
content blocks, `stop_reason: max_tokens`, and a zero `usage` block, which
Claude Code dereferences unconditionally. Registered ahead of the rest of
the chain, since it replaces the turn rather than shaping it.

Interception sits after model resolution, so an id no upstream serves
still fails at the serve layer with a 404 and the CLI still reports it not
found; only the pointless generation behind the probe is suppressed. The
turn makes no upstream call, records a usage row at zero, and contributes
no latency sample.

Detection requires the `claude-cli` User-Agent, `max_tokens: 1`, no tools,
and a single user turn holding one of the CLI's fixed probe prompts. That
is narrower than CLIProxyAPI, which short-circuits on `max_tokens === 1`
alone after reverting a tighter predicate; the extra conditions fail safe
toward today's behavior, at the cost of tracking a Claude Code build
detail.
The Messages module mints Anthropic-shaped opaque ids in two places — the
`request_id` on a gateway-synthesized error envelope, and the message id on
a gateway-answered turn — from the same expression. Give the module one
minter parameterized by prefix.

The rationale comment it carried was also wrong about its own output:
`crypto.randomUUID().replace(/-/g, '').slice(0, 24)` yields 24 hex
characters, not base62, and the first 24 hex characters of a UUIDv4 carry
~90 bits rather than ~96 because the version nibble is fixed and the
variant nibble constrained.
…truthfully

Review of the probe interceptor found the accepted-prompt set too wide and
its research notes wrong on three points.

`quota` leaves the set. It is the CLI's rate-limit preflight, and its
caller consumes the `anthropic-ratelimit-unified-*` response headers rather
than the body — a synthesized turn cannot carry them, so answering it would
replace a working quota reading, on every upstream that serves the probe
today, with a silent blank. `.` leaves too: it is issued by the Bedrock /
Vertex / Mantle SDK clients, which route through their own base URLs and
never reach an `ANTHROPIC_BASE_URL` gateway. That leaves the `/model`
validation probe, which is the reported failure, and the credential check,
which reads nothing off the response at all.

Prompt matching becomes exact. Every literal is read off a binary rather
than inferred, so lowercasing the set and normalizing the input widened the
predicate to shapes no evidence attributes to the CLI, in a predicate the
file argues at length is deliberately narrow. An unobserved shape should
reach the upstream rather than be answered from a guess.

Corrections to the research notes, each verified against the 2.1.226 binary
or the GitHub API:

  - Only the `/model` validation probe goes through the `eie` side-query
    helper; `quota`, `test`, and `.` are independent call sites. The
    "reads usage and treats no-throw as the verdict" claim holds for the
    validation path only.
  - CLIProxyAPI has no `max_tokens === 1` short-circuit; PR 3571 is an open
    proposal to add one, and the revert it describes is internal to its own
    branch history. It is not shipped behavior and not a baseline.
  - The bare-string content form is what the current client sends for the
    credential check, not a legacy shape. Labeling it as one read as a
    compatibility shim for old clients, which it is not.
  - `eie` accepts a `tools` argument; it is the one-token probes
    specifically, not side queries in general, that never carry tools.
  - Ordering the entry first needs no web-search-shim hazard to justify it;
    the shim is inactive for a request with no tools regardless.

`isClaudeCodeProbe` goes back to module-private, matching its siblings, and
the prompt matrix now drives the interceptor rather than the predicate.
Two shape guards — a sole turn with more than one block, and a sole turn
that is not a user turn — had no test pinning them; both are now covered.
Round-2 review caught the accepted-prompt set asserting an evidence
standard one of its members does not meet. `hello` came from a third-party
bug report against v2.1.220, and re-reading that report shows its request
body is a hand-written reproduction rather than a capture. No build on disk
contains it: 2.1.221 through 2.1.226 all send `Hi` in ephemeral-block form
at the `model_validation` call site, and `hello` appears in those binaries
only in computer-use tool-schema examples and an IPC handshake frame.

Matching became exact precisely so that an unobserved shape reaches the
upstream instead of being answered from a guess, so the guessed entry goes.
The litellm report stays as the reference for what it does support — that
the literal is a build detail worth re-reading per release.

Two comment corrections alongside it. The response-field inventory said
`usage.input_tokens` / `usage.output_tokens` and `stop_reason` were the
only fields read; the telemetry event also reads `_request_id` and the two
cache counters, all behind `??`, so the load-bearing property is the
unguarded dereference rather than the size of the inventory. And the test
fixture's User-Agent note credited the header to the Anthropic SDK and
presented one fixed string: Claude Code composes it itself as
`claude-cli/<version> (external, <entrypoint>…)` with up to three optional
trailing segments, which is why the predicate anchors on the prefix alone.
The shape of an Anthropic id — `msg_` for a message, `req_` for a request,
followed by an opaque 24-character token — is a property of the Messages
wire contract, not of the gateway that happens to synthesize bodies
carrying one. Move the generator into `@floway-dev/protocols/messages`
alongside `createRandomResponsesItemId`, its exact counterpart on the
Responses side, and name it `generateAnthropicId` to match.

Generating from `crypto.getRandomValues` rather than by slicing a UUID
drops the version and variant nibbles the slice inherited, so the token
carries 96 bits of entropy across its 24 characters instead of 90, and the
comment no longer needs the caveat. The wire shape is unchanged.
@Menci Menci changed the title feat(gateway): answer Claude Code's one-token probes at the gateway feat(gateway): answer Claude Code's one-token model probes Aug 8, 2026
@Menci
Menci merged commit 5fef0c1 into main Aug 8, 2026
8 checks passed
@Menci
Menci deleted the claude-code-model-probe branch August 8, 2026 16:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant