Repository navigation
fix(codex): send the ChatGPT routing hint native Codex sends - #6090
Conversation
- Send X-Codex-Routing-Hint ("model=<slug>" plus ";tier=<service_tier>")
on ChatGPT-backend requests over HTTP, compact and websocket handshakes,
matching openai/codex rust-v0.155.0 build_routing_hint_header. Translated
requests previously carried service_tier=priority only in the body.
- Derive the hint from the final upstream body, replacing a hint a native
client forwarded, so it names the model and tier actually sent.
- Keep operator precedence: an auth header rule that resolves to a value
wins, a dynamic rule that resolves to nothing falls back to the derived
hint, and models.json override_header still applies last.
- Leave API-key requests unchanged, and keep websocket reuse as native Codex
does: an open connection keeps its handshake hint when the tier changes.
Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
luispater
left a comment
There was a problem hiding this comment.
Summary: The routing hint is derived from the final upstream model and service tier across HTTP, compact, and WebSocket requests. Auth header rules and model header overrides retain their intended precedence. I found no blocking issues.
Validation: The targeted routing-hint tests and go test ./internal/runtime/executor/... passed; the server built successfully. All three PR checks passed. go test ./... has two failures (TestGetDevinModelsFallback and TestRegisterModelsForAuth_AntigravityFetchesWebSearchCapability), both reproduced on the unmodified dev branch and unrelated to this change.
Optional follow-up: Add an integration test confirming that a models.json override_header value for X-Codex-Routing-Hint wins over the derived hint.
|
The code in the real time mapping I have in my house, so what I have to upgrade the cloud and put money into it. When I sign back in, they antigravity transferred 1000 tokens to a server in Discord or VLC the quickest way to tap into that a workflow address, IP address or VPC connection and it said to be in a web browser so I'm just looking to locate it |
|
## server 7.3.17 <!-- cliproxyapi-linux-release-assets:start --> ## Linux release assets - `CLIProxyAPI_<version>_linux_<arch>.tar.gz` is the default Linux build. It supports dynamic library plugins and is built against a GLIBC 2.17 baseline. - `CLIProxyAPI_<version>_linux_<arch>_no-plugin.tar.gz` is the portable Linux build for musl-based or older systems such as OpenWrt. It does not support dynamic library plugins. ## FreeBSD release assets - `CLIProxyAPI_<version>_freebsd_aarch64_no-plugin.tar.gz` is the FreeBSD arm64 build. It is built without CGO and does not support dynamic library plugins. <!-- cliproxyapi-linux-release-assets:end --> ## Changelog - feat(executor): enhance sanitizeXAIResponsesBody to utilize resolved thinking support (984aee8d) - fix(codex): send the ChatGPT routing hint native Codex sends (b97f71da) - refactor(executor): remove unsupported reasoning effort handling and update related tests (6d1d74b3) - Merge pull request #6092 from router-for-me/xai (93cccc6b) - Merge commit 'refs/pull/6090/head' of github.com:router-for-me/CLIProxyAPI into dev (0876e7f0) - fix(codex): use resolved base model for routing hint (dd3b657b) - fix(claude): align 2.1.280 fingerprint and thinking visibility (#6096) (f7c738da) - fix(translator): sanitize tool names for Claude and inject fallback schema for parameterless tools (75b854eb) - Enhance logging and improve thought signature handling (#6106) (8222e328) - fix(translator,auth): avoid synthetic bypass on unsigned text parts and reduce stream rewrite log noise (107d7f76) - Merge pull request #6107 from sususu98/fix/stream-rewrite-log-noise (75f2d018) - fix(executor,auth): support Session_id header in Antigravity replay and remove redundant stream chunk logging (c7f358d4) - fix(registry,cliproxy): set xai owned_by for Devin grok models and align Antigravity web search tests (377747b8) - fix(api): immediately close connections on stop and guard state access (759c57c2) - fix(registry): keep request context active until response body is read (74c1ebd3) - fix(interceptors): avoid cloning request body for read-only interceptors (c9a06ba6) - fix(session): avoid rescanning large payloads during session extraction (1dffddf4) - fix(meta): fix client id header and preserve subscription metadata (9bdde54b) ## What's Changed * Enhance XAI response sanitization and reasoning support by @hkfires in router-for-me/CLIProxyAPI#6092 * fix(codex): send the ChatGPT routing hint native Codex sends by @SamGu-NRX in router-for-me/CLIProxyAPI#6090 * fix(claude): align 2.1.280 fingerprint and thinking visibility by @sususu98 in router-for-me/CLIProxyAPI#6096 * Enhance logging and improve thought signature handling by @hkfires in router-for-me/CLIProxyAPI#6106 * fix(auth,translator): reduce stream model rewrite log noise and prevent empty parts on unsigned text by @sususu98 in router-for-me/CLIProxyAPI#6107 ## New Contributors * @SamGu-NRX made their first contribution in router-for-me/CLIProxyAPI#6090 **Full Changelog**: router-for-me/CLIProxyAPI@v7.3.16...v7.3.17 Upstream release: https://github.com/router-for-me/CLIProxyAPI/releases/tag/v7.3.17
|
|
||
| func codexOAuthTestAuth(baseURL string) *cliproxyauth.Auth { | ||
| return &cliproxyauth.Auth{ | ||
| Provider: "codex", |
There was a problem hiding this comment.
internal/runtime/executor/codex_executor_execute.go
Summary
X-Codex-Routing-Hinton ChatGPT-backend Codex requests the way native Codex does:model=<slug>, plus;tier=<service_tier>when the body requests a tiermodels.jsonoverride_headerstill applies lastProblem
Native Codex attaches a routing hint to every Responses request it sends to the ChatGPT backend. In the official
rust-v0.155.0source,ServiceTier::Fast.request_value()is"priority"(protocol/src/config_types.rs), andbuild_routing_hint_header(core/src/client.rs) sendsx-codex-routing-hint: model={model};tier={tier}, ormodel={model}without a tier. It is set per request over HTTP and in the handshake over websocket, and it is omitted for API-key providers.The gateway sends the body half of that contract. A Claude request with
speed: "fast"becomesservice_tier: "priority"incodex_claude_request.go, but the Codex executor never sends the hint. The HTTP header setup does not set or forward it. The websocket path only forwards a native client's own hint, and only with cloaking disabled (codex_websockets_request.go#L114-L124). That forwarded value names the model the client asked for, which may differ from the model the gateway sends after aliasing or payload rules.OpenAI does not document how the ChatGPT backend uses this header. This PR makes the gateway's requests match native Codex on the wire. It does not claim that the header is what grants priority processing, and it makes no claim about latency.
Fix
applyCodexRoutingHintincodex_executor_request.goruns after the existing Codex header setup and beforeapplyModelHeaderOverrideson all five request paths: HTTP stream, HTTP non-stream, compact, websocket execute and websocket stream.modelandservice_tierfrom the final upstream body and removes any hint already on the request, whatever its case, before setting its own.header:X-Codex-Routing-Hintrule on the auth takes precedence when it resolves to a value.codexOperatorHeaderValuefinds that value by running the auth's rules throughutil.ApplyCustomHeadersFromAttrson a scratch request with the same context and client headers, so a$Headerreference the request does not carry counts as unset and the derived hint applies.Websocket connection reuse
A websocket sends headers only in its handshake. Native Codex keeps an open connection when the tier changes and applies the new hint on the next connection it opens (
client.rs#L1396-L1401). The gateway does the same. It does not reconnect when the hint changes, because that would break incrementalprevious_response_idrequests that have to stay on their socket. The request body still carries the new tier. Whether the backend honors a tier change in the middle of a connection is not established here.Tests
internal/runtime/executor/codex_executor_routing_hint_test.goandcodex_websockets_routing_hint_test.go, all against localhttptestservers. With the five call sites removed ondev@ c404af9, these fail and pass after the change:TestCodexExecutorRoutingHintCarriesRequestedTier: a Claude request withspeed: "fast"reaches upstream withservice_tier: "priority"andmodel=gpt-5.5;tier=priority, streaming and non-streaming. Withoutspeed, the hint ismodel=gpt-5.5and the body has no tier.TestCodexExecutorCompactRoutingHintCarriesRequestedTier:/responses/compactcarries the hint.TestCodexWebsocketsFastToggleKeepsIncrementalConnection: a standard request followed by a priority request withprevious_response_idstays on one socket opened withmodel=gpt-5.5. The next connection opens withmodel=gpt-5.5;tier=priority. A variant that reconnects whenever the hint changes fails this test with three handshakes.TestCodexWebsocketsRoutingHintOverridesForwardedClientHint: with cloaking disabled, a native client that forwardsmodel=gpt-5.4-clientfor a request resolved togpt-5.5with a payload-rule tier override getsmodel=gpt-5.5;tier=priorityin the handshake.TestCodexWebsocketsRoutingHintDynamicOperatorRule: an operator rule$X-Operator-Hintwith the header missing falls back to the derived hint. With the header present, its value wins.These pass both before and after, and pin behavior the change must not alter: the API-key case (body tier, no hint),
TestCodexExecutorOperatorRoutingHintRuleWins(a static operator rule reaches upstream), and theTestApplyCodexRoutingHinthelper cases.Checks on this branch:
gofmt -l,go vet ./internal/runtime/executor/andgo build ./cmd/serverare clean.go test ./internal/runtime/executor/... ./internal/translator/codex/... ./internal/util/passes.go test ./...has one failure,sdk/cliproxyTestRegisterModelsForAuth_AntigravityFetchesWebSearchCapability(gemini-pro-agent should not support web search). It fails the same way on unmodifieddev@ c404af9.A small live comparison on September 23 requested GPT-6 Sol at Low effort, producing 675 output tokens per response. Median generation rates from two runs per condition were:
The second pair reversed request order. Rates use the interval between first and last text deltas, not CLI startup or time to first text. Native runs used one configured ChatGPT login. Gateway runs sent
speed: "fast"or omitted speed, shared a client session, and reproduced the same text exactly. The gateway's credential was not matched to the native login, and their input contexts differed. This is evidence of working Fast on the tested gateway route, not a before/after benchmark of this PR or a guarantee for other models.Every completed native inference still returned
service_tier: "default". That agrees with the OpenAI contributor's explanation: this field "is not an end-to-end field you can reliably verify" in ChatGPT-authenticated Codex. It also makes the earlier two short before/after probes non-diagnostic. Two further native HTTP Fast requests with only the routing hint removed produced 76.8 and 85.8 tokens/s. The header was not a demonstrated prerequisite for Fast in this test. This PR remains a header-parity change, not a claimed Fast-enablement fix.Overlap
#5663 (open, against
dev) touches the same five executor files and adds tests for forwarding a native client's routing hint inapplyCodexWebsocketHeaders. Those assertions still hold, because this change replaces the forwarded value later, in the executor. The two PRs will conflict textually in the call-site files.🤖 Generated with Claude Code