Skip to content

fix(codex): send the ChatGPT routing hint native Codex sends - #6090

Merged
luispater merged 1 commit into
router-for-me:devfrom
SamGu-NRX:fix/codex-routing-hint
Sep 24, 2026
Merged

luispater merged 1 commit into
router-for-me:devfrom
SamGu-NRX:fix/codex-routing-hint

Conversation

@SamGu-NRX

@SamGu-NRX SamGu-NRX commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • send X-Codex-Routing-Hint on ChatGPT-backend Codex requests the way native Codex does: model=<slug>, plus ;tier=<service_tier> when the body requests a tier
  • build the hint from the final upstream body, so it names the model and tier the gateway actually sends, and replace a hint a native client forwarded
  • keep operator configuration in charge: an auth header rule that resolves to a value wins, and models.json override_header still applies last

Problem

Native Codex attaches a routing hint to every Responses request it sends to the ChatGPT backend. In the official rust-v0.155.0 source, ServiceTier::Fast.request_value() is "priority" (protocol/src/config_types.rs), and build_routing_hint_header (core/src/client.rs) sends x-codex-routing-hint: model={model};tier={tier}, or model={model} without a tier. It is set per request over HTTP and in the handshake over websocket, and it is omitted for API-key providers.

The gateway sends the body half of that contract. A Claude request with speed: "fast" becomes service_tier: "priority" in codex_claude_request.go, but the Codex executor never sends the hint. The HTTP header setup does not set or forward it. The websocket path only forwards a native client's own hint, and only with cloaking disabled (codex_websockets_request.go#L114-L124). That forwarded value names the model the client asked for, which may differ from the model the gateway sends after aliasing or payload rules.

OpenAI does not document how the ChatGPT backend uses this header. This PR makes the gateway's requests match native Codex on the wire. It does not claim that the header is what grants priority processing, and it makes no claim about latency.

Fix

applyCodexRoutingHint in codex_executor_request.go runs after the existing Codex header setup and before applyModelHeaderOverrides on all five request paths: HTTP stream, HTTP non-stream, compact, websocket execute and websocket stream.

  • It reads model and service_tier from the final upstream body and removes any hint already on the request, whatever its case, before setting its own.
  • An operator header:X-Codex-Routing-Hint rule on the auth takes precedence when it resolves to a value. codexOperatorHeaderValue finds that value by running the auth's rules through util.ApplyCustomHeadersFromAttrs on a scratch request with the same context and client headers, so a $Header reference the request does not carry counts as unset and the derived hint applies.
  • API-key requests are left untouched, matching native Codex.

Websocket connection reuse

A websocket sends headers only in its handshake. Native Codex keeps an open connection when the tier changes and applies the new hint on the next connection it opens (client.rs#L1396-L1401). The gateway does the same. It does not reconnect when the hint changes, because that would break incremental previous_response_id requests that have to stay on their socket. The request body still carries the new tier. Whether the backend honors a tier change in the middle of a connection is not established here.

Tests

internal/runtime/executor/codex_executor_routing_hint_test.go and codex_websockets_routing_hint_test.go, all against local httptest servers. With the five call sites removed on dev @ c404af9, these fail and pass after the change:

  • TestCodexExecutorRoutingHintCarriesRequestedTier: a Claude request with speed: "fast" reaches upstream with service_tier: "priority" and model=gpt-5.5;tier=priority, streaming and non-streaming. Without speed, the hint is model=gpt-5.5 and the body has no tier.
  • TestCodexExecutorCompactRoutingHintCarriesRequestedTier: /responses/compact carries the hint.
  • TestCodexWebsocketsFastToggleKeepsIncrementalConnection: a standard request followed by a priority request with previous_response_id stays on one socket opened with model=gpt-5.5. The next connection opens with model=gpt-5.5;tier=priority. A variant that reconnects whenever the hint changes fails this test with three handshakes.
  • TestCodexWebsocketsRoutingHintOverridesForwardedClientHint: with cloaking disabled, a native client that forwards model=gpt-5.4-client for a request resolved to gpt-5.5 with a payload-rule tier override gets model=gpt-5.5;tier=priority in the handshake.
  • TestCodexWebsocketsRoutingHintDynamicOperatorRule: an operator rule $X-Operator-Hint with the header missing falls back to the derived hint. With the header present, its value wins.

These pass both before and after, and pin behavior the change must not alter: the API-key case (body tier, no hint), TestCodexExecutorOperatorRoutingHintRuleWins (a static operator rule reaches upstream), and the TestApplyCodexRoutingHint helper cases.

Checks on this branch:

  • gofmt -l, go vet ./internal/runtime/executor/ and go build ./cmd/server are clean.
  • go test ./internal/runtime/executor/... ./internal/translator/codex/... ./internal/util/ passes.
  • go test ./... has one failure, sdk/cliproxy TestRegisterModelsForAuth_AntigravityFetchesWebSearchCapability (gemini-pro-agent should not support web search). It fails the same way on unmodified dev @ c404af9.

A small live comparison on September 23 requested GPT-6 Sol at Low effort, producing 675 output tokens per response. Median generation rates from two runs per condition were:

Route Standard tokens/s Fast tokens/s
Native Codex 0.156.1, WebSocket 56.1 84.2
Native Codex 0.156.1, HTTP fallback 56.4 84.2
Patched v7.3.16 gateway, Claude-compatible HTTP 55.5 84.1

The second pair reversed request order. Rates use the interval between first and last text deltas, not CLI startup or time to first text. Native runs used one configured ChatGPT login. Gateway runs sent speed: "fast" or omitted speed, shared a client session, and reproduced the same text exactly. The gateway's credential was not matched to the native login, and their input contexts differed. This is evidence of working Fast on the tested gateway route, not a before/after benchmark of this PR or a guarantee for other models.

Every completed native inference still returned service_tier: "default". That agrees with the OpenAI contributor's explanation: this field "is not an end-to-end field you can reliably verify" in ChatGPT-authenticated Codex. It also makes the earlier two short before/after probes non-diagnostic. Two further native HTTP Fast requests with only the routing hint removed produced 76.8 and 85.8 tokens/s. The header was not a demonstrated prerequisite for Fast in this test. This PR remains a header-parity change, not a claimed Fast-enablement fix.

Overlap

#5663 (open, against dev) touches the same five executor files and adds tests for forwarding a native client's routing hint in applyCodexWebsocketHeaders. Those assertions still hold, because this change replaces the forwarded value later, in the executor. The two PRs will conflict textually in the call-site files.

🤖 Generated with Claude Code

- Send X-Codex-Routing-Hint ("model=<slug>" plus ";tier=<service_tier>")
  on ChatGPT-backend requests over HTTP, compact and websocket handshakes,
  matching openai/codex rust-v0.155.0 build_routing_hint_header. Translated
  requests previously carried service_tier=priority only in the body.
- Derive the hint from the final upstream body, replacing a hint a native
  client forwarded, so it names the model and tier actually sent.
- Keep operator precedence: an auth header rule that resolves to a value
  wins, a dynamic rule that resolves to nothing falls back to the derived
  hint, and models.json override_header still applies last.
- Leave API-key requests unchanged, and keep websocket reuse as native Codex
  does: an open connection keeps its handshake hint when the tier changes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

@luispater luispater left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Summary: The routing hint is derived from the final upstream model and service tier across HTTP, compact, and WebSocket requests. Auth header rules and model header overrides retain their intended precedence. I found no blocking issues.

Validation: The targeted routing-hint tests and go test ./internal/runtime/executor/... passed; the server built successfully. All three PR checks passed. go test ./... has two failures (TestGetDevinModelsFallback and TestRegisterModelsForAuth_AntigravityFetchesWebSearchCapability), both reproduced on the unmodified dev branch and unrelated to this change.

Optional follow-up: Add an integration test confirming that a models.json override_header value for X-Codex-Routing-Hint wins over the derived hint.

@luispater
luispater merged commit 0876e7f into router-for-me:dev Sep 24, 2026
3 checks passed
@philly55T-hash

Copy link
Copy Markdown

The code in the real time mapping I have in my house, so what I have to upgrade the cloud and put money into it. When I sign back in, they antigravity transferred 1000 tokens to a server in Discord or VLC the quickest way to tap into that a workflow address, IP address or VPC connection and it said to be in a web browser so I'm just looking to locate it

@philly55T-hash

Copy link
Copy Markdown

Summary: The routing hint is derived from the final upstream model and service tier across HTTP, compact, and WebSocket requests. Auth header rules and model header overrides retain their intended precedence. I found no blocking issues.

Validation: The targeted routing-hint tests and go test ./internal/runtime/executor/... passed; the server built successfully. All three PR checks passed. go test ./... has two failures (TestGetDevinModelsFallback and TestRegisterModelsForAuth_AntigravityFetchesWebSearchCapability), both reproduced on the unmodified dev branch and unrelated to this change.

Optional follow-up: Add an integration test confirming that a models.json override_header value for X-Codex-Routing-Hint wins over the derived hint.
philly55T-hash

github-actions Bot added a commit to omarcresp/cliproxyapi-flake that referenced this pull request Sep 24, 2026
## server 7.3.17

<!-- cliproxyapi-linux-release-assets:start -->
## Linux release assets

- `CLIProxyAPI_<version>_linux_<arch>.tar.gz` is the default Linux build. It supports dynamic library plugins and is built against a GLIBC 2.17 baseline.
- `CLIProxyAPI_<version>_linux_<arch>_no-plugin.tar.gz` is the portable Linux build for musl-based or older systems such as OpenWrt. It does not support dynamic library plugins.

## FreeBSD release assets

- `CLIProxyAPI_<version>_freebsd_aarch64_no-plugin.tar.gz` is the FreeBSD arm64 build. It is built without CGO and does not support dynamic library plugins.

<!-- cliproxyapi-linux-release-assets:end -->

## Changelog

- feat(executor): enhance sanitizeXAIResponsesBody to utilize resolved thinking support (984aee8d)
- fix(codex): send the ChatGPT routing hint native Codex sends (b97f71da)
- refactor(executor): remove unsupported reasoning effort handling and update related tests (6d1d74b3)
- Merge pull request #6092 from router-for-me/xai (93cccc6b)
- Merge commit 'refs/pull/6090/head' of github.com:router-for-me/CLIProxyAPI into dev (0876e7f0)
- fix(codex): use resolved base model for routing hint (dd3b657b)
- fix(claude): align 2.1.280 fingerprint and thinking visibility (#6096) (f7c738da)
- fix(translator): sanitize tool names for Claude and inject fallback schema for parameterless tools (75b854eb)
- Enhance logging and improve thought signature handling (#6106) (8222e328)
- fix(translator,auth): avoid synthetic bypass on unsigned text parts and reduce stream rewrite log noise (107d7f76)
- Merge pull request #6107 from sususu98/fix/stream-rewrite-log-noise (75f2d018)
- fix(executor,auth): support Session_id header in Antigravity replay and remove redundant stream chunk logging (c7f358d4)
- fix(registry,cliproxy): set xai owned_by for Devin grok models and align Antigravity web search tests (377747b8)
- fix(api): immediately close connections on stop and guard state access (759c57c2)
- fix(registry): keep request context active until response body is read (74c1ebd3)
- fix(interceptors): avoid cloning request body for read-only interceptors (c9a06ba6)
- fix(session): avoid rescanning large payloads during session extraction (1dffddf4)
- fix(meta): fix client id header and preserve subscription metadata (9bdde54b)


## What's Changed
* Enhance XAI response sanitization and reasoning support by @hkfires in router-for-me/CLIProxyAPI#6092
* fix(codex): send the ChatGPT routing hint native Codex sends by @SamGu-NRX in router-for-me/CLIProxyAPI#6090
* fix(claude): align 2.1.280 fingerprint and thinking visibility by @sususu98 in router-for-me/CLIProxyAPI#6096
* Enhance logging and improve thought signature handling by @hkfires in router-for-me/CLIProxyAPI#6106
* fix(auth,translator): reduce stream model rewrite log noise and prevent empty parts on unsigned text by @sususu98 in router-for-me/CLIProxyAPI#6107

## New Contributors
* @SamGu-NRX made their first contribution in router-for-me/CLIProxyAPI#6090

**Full Changelog**: router-for-me/CLIProxyAPI@v7.3.16...v7.3.17


Upstream release: https://github.com/router-for-me/CLIProxyAPI/releases/tag/v7.3.17

func codexOAuthTestAuth(baseURL string) *cliproxyauth.Auth {
return &cliproxyauth.Auth{
Provider: "codex",

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.


internal/runtime/executor/codex_executor_execute.go

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants