fix(upstream): report a chat challenge as a retryable overload and stop hammering Qwen during one - #181
Merged
Conversation
maxff77
force-pushed
the
fix/chat-challenge
branch
from
September 27, 2026 02:32
00f082f to
a4ae88f
Compare
…p hammering Qwen during one Since 2026-09-23 11:35Z Qwen answers most chat generations with a single frame `ret: ['FAIL_SYS_USER_VALIDATE', 'RGV587_ERROR::SM::哎哟喂,被挤爆啦,请稍后重试']`, even for a bare "你好" (Rfym21#180). Measured on a production deployment, 2026-09-23..26: 477 of 502 chat sends challenged, every one between 07:00Z and 22:00Z; the 25 that passed were all at night (UTC), and the same accounts passed by night and were challenged by day. Creating the chat and uploading the history file kept succeeding on the same account and egress seconds before each challenge. So it is neither the context size nor the account. What clients got instead: 502 upstream_error (OpenAI) or 500 api_error (Anthropic) with no Retry-After, and the message "Agent 上下文可能过大或账号需要 验证". Agentic clients retry within a second, each retry re-uploads and re-parses the history, and the parse rate limiter answers 529 on top. - A chat challenge now maps to 529 overloaded_error / 503 with Retry-After: 30, and says what Qwen said: upstream busy, retry later. - After 3 challenges in a row, sendChatRequest answers 529/503 for 60 s before creating a chat or uploading anything; the first response with `choices` closes the breaker. One breaker for the process: every account shares the same egress in the deployments this was measured on. - No account cooldown or warning for a challenge: it cooled healthy accounts for 5 minutes. The single agent-mode failover stays. - /v1/messages no longer forgets a reused history prefix on a chat challenge (Qwen rejected before reading it), so the client's retry does not re-parse. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…mages and failover Review follow-ups on the chat-challenge breaker: - Half-open state: once the cooldown ends, the first request goes out as the only probe and re-arms the window for the rest. The probe's answer closes the breaker; a challenge reopens it. Before, every waiting client came back in the same second. - An answer from a stream that was already running when the breaker opened now only clears strikes. It no longer closes an open breaker, since it says nothing about whether Qwen accepts new requests. - CHAT_CHALLENGE_BREAKER_SECONDS (default 60, 0 disables), an injectable clock for tests, a warn log on open and an info log on close. This mirrors the parse breaker. - A slider captcha (`/punish?`, no 被挤爆) gets its own message instead of "upstream busy". The breaker-open refusal says requests are paused. - The plain /v1/chat/completions handlers go through upstreamErrorShape like agent mode, so a chat challenge is 503 `upstream_unavailable` there too. - Image and video generation post to the same chat endpoint. They now check the breaker before creating a chat, recognise the challenge frame, and answer 503 with Retry-After instead of a generic 500. - A chat-challenge account switch in agent mode keeps the history prefix instead of re-uploading and re-parsing the whole history. Qwen refused before reading it, and normal rotation already reuses the prefix across accounts. A quota switch still re-uploads. - Tests reset the process-wide breaker in beforeEach. New tests cover: the half-open probe, an in-flight answer, the captcha message, OpenAI non-stream 503, /v1/messages keeping the prefix on a challenge (and forgetting it on other errors), the image path, and the failover prefix key. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…allenge is a real 529/503 /v1/messages and the OpenAI agent stream wrote HTTP 200 plus message_start (or the role chunk) before reading anything from Qwen. So the breaker's probe request, the one that actually reaches Qwen, got its chat challenge as an in-stream error event. Claude Code shows that as "API error · Retrying in 5s" and ignores the retry_after inside it. Pre-commit refusals already produced a real 529 that it honors (~60 s spacing, observed live). Both paths now commit lazily. The SSE headers and first protocol event are written on the first upstream frame that passed assertNoUpstreamFailure (usually response.created, so thinking turns still start at once), on the first ping or heartbeat (the cap), or when an attempt ends. A chat challenge or quota on the first frame therefore arrives with nothing committed and goes out as a real 529/503/429 with Retry-After. Any other failure before the first frame commits and goes out in the stream exactly as before, so it does not become a 5xx that SDKs retry against upload/parse. There is still one reader of the stream, so no bytes are buffered or replayed and a challenge costs one breaker strike. The long silence the early pings fixed was the correction retry, which still runs committed. - anthropic.js: ensureMessageStart() on the first onDelta, before every ping (runWithAnthropicPing beforePing), after each attempt, and in the attempt catch for other failures. - chat.js: commitStream() through a new on_upstream_frame runtime hook, the SSE heartbeat (beforeBeat), writeDelta, and after the runtime returns. - openai-agent-runtime.js: a challenge switch whose replay cannot start rethrows the challenge instead of an opaque 502. - chat.image.video.js: a challenge on a stream request that is still uncommitted is a 503 as JSON (t2v presets text/event-stream). - upstream-error.js: Retry-After for a chat challenge is capped at 59 s, because the Anthropic and OpenAI SDKs ignore values of 60 or more. The reachability matrix is rewritten. - Tests: first-frame quota/challenge now pins a real status. Mid-stream variants keep the in-event retry_after coverage. New tests cover the commit points (first frame, ping/heartbeat cap, empty attempt, other failures still in-stream, t2v/t2i) and the agent challenge switch. tools/bun-smoke.js checks 529/503 + JSON + Retry-After over real Bun HTTP. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…error The challenge case already rethrew; quota fell through to an opaque 502 upstream_retry_failed. Rethrowing it lets the agent stream answer a real 429 with Retry-After when nothing was committed yet, like a first-frame quota. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… too Qwen also answers /api/v2/chat/completions with 200 text/html: Aliyun's captcha page (seen 2026-09-19, the HTML form Rfym21#179 fixed on the image path). The text path never looked at the body: with no `data:` frame no detector ran, the breaker counted nothing, and every surface retried 2-3 times against the WAF before reporting an empty answer (502 upstream_empty_output / an in-stream api_error) with no Retry-After. sendChatRequest now reads a text/html 200 (capped at 256 KiB) and runs it through the shared detector: the captcha page is one strike and the same 529/503 + Retry-After as the JSON challenge, thrown before any stream is handed on. Any other HTML body is handed on byte for byte, so its path is unchanged. Covered in chat-challenge.test.js (captcha page -> 529 + one strike; other HTML -> replayed, no strike; fails with the screen removed) and over real Bun HTTP in bun-smoke (one send, not a retry loop). Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
… per review Review of the rebase onto Rfym21#178/Rfym21#179 found, besides the new HTML screen: - t2v asks for responseType 'json', so a non-JSON body (the captcha page, or a `data:`-wrapped challenge frame) reaches the controller as a string that parseUpstreamImageError dropped on its failed JSON.parse. The captcha page's first URL came back as the "video" with a 200, no strike, no Retry-After. On origin/main too; Rfym21#179's raw-text detection only ran on the streaming branch. String bodies now go through the same raw-text parser (parseUpstreamBody) on both the inline check and the reader. - A t2v challenge with a non-2xx status was parsed, and counted, twice (inner catch, then outer catch). The inner catch now throws it. - Image/video never reached noteChatAnswer: an answered image request could take the half-open probe and keep text blocked for another full window, and image answers did not break a run of strikes. Both answer paths now call it; a request without a prompt is rejected before it can take the probe. - A connection cut while reading a text/html body was handled as a transport failure of the POST (re-POST, account penalty, prefix invalidation) even when the captcha markers had already arrived. The bytes read are screened first. - The busy-vs-captcha wording looked only at the detector's filtered signals, so "被挤爆" without "RGV587" read as a captcha. It looks at the whole payload again, as before the rebase. - The breaker comment claimed only the probe's own answer closes a half-open breaker; any answer does. Documented, with its ceiling. Each fix has a test in chat-challenge.test.js that fails with the fix removed. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
Second review pass. The half-open probe re-arms the window for everyone else; a probe that failed before Qwen answered or challenged it (no account, no chat_id, upload failure, an image_edit without a text part) left it probing, and every request got 529/503 for another full window. - assertChatChallengeBreakerClosed says whether the caller is the probe; sendChatRequest (now a thin guard around postChatRequest) and the image/video flow release it (releaseChatProbe) on any failure that is not a chat challenge, so the next request probes at once. A probe that reaches Qwen and ends with neither an answer nor a challenge still costs one window; documented. - t2v closes the half-open breaker as soon as Qwen accepts the task (its whole body was already screened), instead of after the video is polled to completion, minutes later. Left out, never seen from Qwen (every challenge measured arrived with HTTP 200, whole): a non-2xx challenge on the streamed image path, and a captcha page cut mid-body on it. Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
maxff77
force-pushed
the
fix/chat-challenge
branch
from
September 27, 2026 02:35
a4ae88f to
452911c
Compare
Rfym21
added a commit
that referenced
this pull request
Sep 27, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Refs #180.
What #180 is
Since 2026-09-23 11:35Z Qwen answers most chat generations with one frame and nothing else:
It happens on a bare
你好: no tools, no upload, one POST to/api/v2/chat/completions, answered in <1 s. Measured on a production deployment (209 accounts, one egress), 2026-09-23 11:35Z → 09-26 15:28Z:chats/newand the history upload/parse succeed on the same account and egress seconds before each challenge, so it is not the context size either.data.urlis Aliyun's x5sec slider captcha (…/_____tmd_____/punish?x5secdata=…), and a logged-in browser passes at peak.你好, breaker off), headers aligned with the current web client (0.3.x: noversion, browsertimezone,source: h5, with and withoutAuthorization: Bearer) and a direct egress without the proxy were answered 2–3 times in 30 each, the same as the current headers.Rebased onto current
main(#176, #178, #179): the challenge is recognised by #179's shareddetectWafChallenge(JSON frame and HTML page); this PR only adds the strike count and the wording on top of it.This PR does not make Qwen accept those requests. It makes the proxy report the refusal truthfully and stop making it worse.
What clients got before
/v1/chat/completions(plain and agent mode, stream and non-stream)upstream_error, no Retry-After / 200 + error frameupstream_unavailable+ Retry-After/v1/messages(stream and non-stream)api_error/ 200 +api_erroreventoverloaded_error+ Retry-After/v1/images/*,/v1/videostext/htmlupstream_empty_output/ in-stream "empty answer"Retry-After is 30 s for a single challenge and up to 59 s while the breaker is open (the Anthropic and OpenAI SDKs ignore a Retry-After of 60 s or more).
Streams commit on Qwen's first valid frame.
/v1/messagesand the OpenAI agent stream used to write HTTP 200 +message_start(or the role chunk) before reading Qwen, so a challenge could only go out as an in-stream error event, which Claude Code shows as "API error · Retrying in 5s" while ignoring itsretry_after. Both now commit on the first upstream frame that passed validation (usuallyresponse.created, so thinking still starts at once), on the first ping/heartbeat (cap), or when an attempt ends. A challenge or quota on the first frame is therefore a real status with Retry-After; any other early failure still commits and goes out in the stream as before. One reader, nothing buffered or replayed, one breaker strike per challenge. A challenge after the commit (e.g. on a correction retry) still goes out as an event/frame carryingretry_after.The message was
Agent 上下文可能过大或账号需要验证, which sends users after the wrong cause. It now says what Qwen said:被挤爆啦→Qwen 上游繁忙,触发风控验证(被挤爆啦),请稍后重试 / Qwen chat challenge: upstream busy, retry later/punish?, Captcha #157) →… captcha required, retry later… repeated upstream challenges, requests paused, retry laterAgentic clients (Claude Code) retried the 500 within a second. Each retry re-uploaded and re-parsed the history, so the parse rate limiter started answering 529 on top, and the parse budget of the egress was burnt on requests Qwen was going to refuse anyway.
Changes
describeUpstreamFailure: a chat challenge isoverloaded(529 / 503). The plain/v1/chat/completionshandlers now build their error throughupstreamErrorShape, like agent mode.upstream-error.js), process-global on purpose (in the measured deployment every account leaves through one egress):choices) clears strikes;sendChatRequestand the image/video path answer 529 / 503 before creating a chat or uploading history, forCHAT_CHALLENGE_BREAKER_SECONDS(default 60,0disables). An answer from a stream that was already running clears strikes but does not close it;responseType: 'json', so a non-JSON body (the captcha page, or adata:-wrapped frame) arrived as a string that the JSON-only check dropped, and the page's first URL came back as the "video"; string bodies now go through fix(images): repair the image/video path (TDZ crash) and stop reporting WAF/captcha as a successful generation #179's raw-text parser. A challenge with a non-2xx status is counted once, not twice.sendChatRequestreads a200 text/htmlbody (capped at 256 KiB) and runs it through the same detector before handing on any stream, so it is one strike and a real 529/503 instead of an empty answer after 2–3 sends. A connection cut mid-page is still the challenge if the markers already arrived. Any other HTML body is handed on byte for byte.recordChallenge/challengeCooldownUntil): it took healthy accounts out of rotation for 5 minutes. The single agent-mode account switch from fix(agent): handle mid-stream quota_limit and upstream_waf_challenge with auto account rotation in runOpenAIAgentTurn #166 stays, but on a chat challenge it keeps the history prefix instead of re-uploading it (Qwen refused before reading it; normal rotation already reuses the prefix across accounts). A quota switch still re-uploads./v1/messagesno longer forgets a reused history prefix on a chat challenge, so the client's retry does not re-parse the history.Not covered:
sendChatRequest, so a request with images still uploads them while the breaker is open;t2i/image_edit), or a captcha page cut mid-body there. Never seen: every challenge measured arrived whole, with HTTP 200.Tests
npm testgate: PASS, 1225 tests / 137 suites / 0 fail;bun run lintclean;bun run test:bun(real Bun HTTP against a fake upstream) now also asserts a first-frame challenge gives529/503+application/json+retry-afteron/v1/messagesand agent streams, and that the captcha page astext/htmlis a 529 after one send.tests/chat-challenge.test.js(new, 32): wire shape and both messages (also被挤爆withoutRGV587); captcha page astext/html(and cut mid-page) → one strike, no re-POST, no account penalty, other HTML replayed; t2v captcha page /data:frame / non-2xx → 503 with exactly one strike; an answered image closes a half-open breaker and breaks a run of strikes; a probe that fails before asking Qwen (text, or an image_edit without text) hands the probe on; t2v closes a half-open breaker when the task is accepted; every one of these fails with its fix removed; opens on the 3rd strike; an in-flight answer does not close it; exactly one probe after the cooldown, closed by its answer or reopened by its challenge;sendChatRequestrefuses without touching the network; OpenAI non-stream 503upstream_unavailable;/v1/messageskeeps the prefix on a challenge and forgets it on other errors; image path 503.Updated:
agent-account-failover.test.js(no cooldown; challenge switch keeps the prefix key, quota switch drops it),agent-protocol.test.js(503upstream_unavailable+ Retry-After instead of 502). Both reset the process-wide breaker inbeforeEach.Live check (2026-09-26, 17:04–17:28Z, challenge window, one egress)
你好: 503 +Retry-After: 30, 30, then 60 when the breaker opened; the next request was refused in 6 ms with no upstream send (4 requests, 3 sends in the log)./v1/messagesstream while open: real 529overloaded_error(refused before headers are committed). With the breaker closed: 200 +overloaded_errorevent withretry_after.Retry-After: 45); after the cooldown, of two concurrent requests one went out as the probe (challenged, breaker reopened for 60 s) and the other was refused in 11 ms./v1/messagesstream:true→529application/json+Retry-After: 30, noevent:lines; agent stream →503 upstream_unavailable+Retry-After: 59; Claude Code waits the Retry-After between retries.text/htmlform on the text path (Qwen sent only the JSON frame during these checks); covered by unit tests and the Bun smoke.🤖 Generated with Claude Code