Repository navigation
noema-review fails closed on provider transport errors (502/429) after gateway failover; blocks unchanged consumer PRs #2165
Description
Activity
Run 34758483435 (four-pillars#38 @d4750b1) also failed:
HTTP Error 429: Too Many Requests; caller attempts=1, duration=574.7s, served_model=dots-studio/dots-3-note-preview:free. All three re-dispatched runs on 2026-09-13 ~13:00Z ended in provider transport errors (1×502, 2×429); none produced a verdict. Not re-dispatching again in this window to avoid adding 429 pressure.seonghobae commented
on Sep 13, 2026 ContributorAuthorMore actionsAdditional consumer evidence inherited from
ContextualWisdomLab/noema#711(2026-09-13 UTC/KST observation), preserving the owner boundary here rather than adding provider/routing recovery logic to Noema:consumer PR exact/observed head job/run evidence terminal transport evidence fast-mlsirm#18600fc9dfbfjob 103700705680HTTP 502 Bad Gateway;caller attempts=1;duration=1250.7s;phase=response_error;served_model=google/gemma-4-31b-itat ~09:24 UTCfast-mlsirm#164270f1a4c8noema-review run observed ~12:50 UTC HTTP 502 Bad Gateway;caller attempts=1;duration=758.0s;phase=response_error;served_model=deepseek-ai/deepseek-v4-flash-0731fast-mlsirm#16429867fed9job 100377077377, 2026-09-03 08:11 UTCHTTP 502;duration=498.2s;phase=connecting;served_model=unknownThese extend #2165 beyond four-pillars and show the same required-review availability class on fast-mlsirm, including two materially different served models after 12–21 minute waits. The consumer-side issue also records that
gh run rerunfrom the leaf cannot rerun the central workflow (404), and no automatic re-dispatch was observed.Interpretation remains bounded:
caller attempts=1is the Noema gate's one gateway request and is not proof that contextual-orchestrator attempted only one provider; current central source intentionally delegates provider retry/failover to contextual-orchestrator. The repair should therefore stay in the canonical central/orchestrator owner paths: typed transport-unavailable evidence/re-dispatch policy here, and provider routing/retry/failover only in contextual-orchestrator. Noema/fast-mlsirm should not add provider/model fallback logic.This comment fully transfers the valid factual delta from noema#711 into this existing owner issue so the leaf issue can close as a duplicate without losing evidence.
Cross-link so this does not fork into parallel tracks: #2148 tracks the same free-pool 429 exhaustion for private targets (ZDR-only pool = 3 OpenRouter
:freeroutes on one account; proposes honoring providerRetry-Afterrather than an arbitrary retry budget) and #1915 tracks the orchestrator/free provider-family SPOF. Additional public-target occurrences today on.githubPR #2155 head3452ed560: noema-review job 103714113314 (9× HTTP 429, no verdict) and strix job 103714672313 (6×HTTPErrorfrom the gateway, no verdict). Your item 2 (per-run count of providers/models attempted before the terminal error) is already partly emitted in the sidecar preflight artifactcontextual-orchestrator-preflight.json(candidate_count/probed_count/rejected_count, see #2148) — the gap is surfacing it in the check summary, not producing it.Same provider-capacity class in the strix required workflow: four-pillars#41 @d263e66 run 34758864273 job 103728389760 fails at
Provision contextual-orchestrator Strix sidecar— preflight evidence showsaccount_skip_after_429: 2and thensidecar exited before healthz (status 1); stderr: discovery_diagnostics_complete. No consumer code path is involved.Additional occurrence: four-pillars#41 @d263e66 noema-review (pull_request_target run, 2026-09-13) —
HTTP Error 502: Bad Gateway; caller attempts=1, duration=1434.8s, served_model=deepseek-ai/deepseek-v4-flash-0731. Fourth consecutive transport-class failure with attempts=1 on this consumer; no verdict produced.Recurrence 2026-09-14 on contextual-orchestrator#1177 (noema run 34772771262, job 103769999194, sidecar pin
767e67fb): preflightready_count: 1, review call failed after 562.4 s withHTTP Error 429served bygoogle/gemma-4-31b-it:free, caller attempts=1. Gateway-side fix under review: contextual-orchestrator#1179 (parse Retry-After, skip cooling routes, bounded wait, honest 429+Retry-After, readinessearliest_ready_seconds; #1180 exposes readiness at inference scope for the sidecar). Will post re-validation here once the sidecar pin advances past those PRs.New failure shape on four-pillars#44 @2ac1a92 (job 103763057524, 2026-09-13T18:17Z → 2026-09-14T00:17Z): step
Prepare Noema model verdictran for the full 6h Actions ceiling and was cancelled (The operation was canceled.) without any verdict or provider error. Same one-shot caller path; this time the gateway call never returned instead of failing fast with 502/429.New occurrence with a materially different shape from the ones already recorded here — worth adding because it bears on item 1 (honoring
Retry-After) and on runner occupancy..githubPR #2191, head86b49cb5b,noema-reviewjob103865304457:Noema gateway transport failed: HTTPError: HTTP Error 429: Too Many Requests; caller attempts=1, duration=4181.8s, phase=response_error, served_model=inclusionai/ling-3.0-flash-sante:free Noema gateway attempt outcome=failed phase=response_error duration=4181.8s ... (gateway owns repair/failover)The distinguishing detail is
caller attempts=1withduration=4181.8s: one attempt that occupied a runner for 70 minutes and then returned 429. That is not the rapid-burst pattern recorded earlier on this issue and on #2148 (9 × 429 within a single job); here the gateway held a single request for over an hour before failing it.Two consequences worth separating:
- Capacity.
inclusionai/ling-3.0-flash-sante:freeis one of the three OpenRouter free ZDR routes enumerated in Private-target review pool is three OpenRouter free ZDR routes; a single 429 burst fails every review closed #2148. A 429 arriving 70 minutes in means the exhaustion was not visible at admission time, so aRetry-After-aware admission check (item 1) would have to be re-evaluated during the call, not only before it. - Runner occupancy. This failure mode consumes ~70 minutes of a required-workflow runner and produces no verdict. That is a different cost from a fast failure and is relevant to the queue-saturation issues (ops: diagnose and bound organization GitHub Actions queue starvation #712/Actions queue saturation: 120 open PRs + self-amplifying scheduler block all org merges (pg-erd-cloud: 0 merges since 2026-08-20) #1531): slow failures are far more expensive than the failure count alone suggests.
Context for the rate this class occurs at: in a census of the last 100 completed
opencode-review-dispatchruns (posted on #1931), model-pool/gateway exhaustion accounted for 14 of 83 failures (17%); the window produced 3 successful runs out of 100. No action requested — recording the timing evidence since this issue asked for per-run attempt data.- Capacity.
A failure mode the title does not yet cover: a ready gateway, a 400, and no failover
From
ContextualWisdomLab/four-pillarsPR #48 at headadb90c2, run 34825522104, job 104002823130, 2026-09-14.The terminal error is not 502 or 429:
##[error]Noema gateway transport failed: HTTPError: HTTP Error 400: Bad Request; caller attempts=1, duration=971.8s, phase=response_error, served_model=meta/llama-3.2-90b-vision-instructThree things about it are worth separating.
The gateway was healthy, not exhausted. The uploaded
contextual-orchestrator-preflight.jsonfrom the same job reports:field value gateway.status ready gateway.finish_reason stop candidate_count 24 probed_count 16 ready_count 8 rejected_count 6 deferred_count 2 So this is not the provider-exhaustion path. Preflight finished with eight routes marked ready and the gateway answering.
A vision model was chosen to review a text pull request.
served_modelismeta/llama-3.2-90b-vision-instruct. The sidecar log shows it probed successfully at 19:09:14 and entered the ready set. The changed files on that PR are a Python test module and a changelog entry; there is no image in the request.One 400 ended the review, with seven other ready routes unused.
attempts=1. A 400 is not transient, so not retrying the same route is right, but the run had eight ready routes and fell back to none of them. The call also occupied 971.8 seconds, roughly sixteen minutes, before returning.Why this matters to a consumer
The consumer repository cannot influence any of it. It cannot choose the route, cannot see the catalog, and cannot rerun the org-owned workflow. What it gets is a required check that fails with a transport error, which then makes the OpenCode reviewer refuse to approve, which holds
CHANGES_REQUESTEDon PRs whose own checks are green. That cascade is now visible on six of the seven four-pillars PRs that have any failure at all.The distinction I would ask the issue to carry
The existing title describes failing closed on 502 and 429 after gateway failover. This run failed closed on a non-transient 400 from a ready route while other ready routes existed. The remedy is different: 502 and 429 want backoff and capacity, whereas this wants the next ready route to be tried, and wants a text review not to be served by a vision-instruct model in the first place.
Consumer-side evidence only. I have no visibility into the routing policy and am not proposing a change to it.
Correction to my own framing: this is intermittent, not an outage
My earlier comment reported a specific failing run accurately, but a reader could take it as the service being down. It is not, and the distinction changes what the fix is.
cwl-noema-reviewapprovedContextualWisdomLab/four-pillars#46at 2026-09-14T17:52Z, hours after the failure I described on #48 at 19:11Z. Surveying every open PR in that repository for the latest verdict from either agent:PR latest agent verdict four-pillars#46 cwl-noema-reviewAPPROVED, 2026-09-14 17:52four-pillars#44 opencode-agentCHANGES_REQUESTED, 2026-09-14 09:58four-pillars#41 opencode-agentCHANGES_REQUESTED, 2026-09-13 14:07four-pillars#31, #37, #38 opencode-agentCHANGES_REQUESTED, 2026-08-31 to 09-01four-pillars#35, #39, #48, #52 through #57, #59 none yet Both agents do reach verdicts. They just do it unevenly, and a large set of newer heads has been waiting without one.
Why the intermittency supports the diagnosis rather than weakening it
A service that is simply down fails uniformly. This does not. It succeeded on one head and, on another head hours later, selected
meta/llama-3.2-90b-vision-instructfor a text-only diff, spent 971.8 seconds, returned a non-transient 400, and stopped atattempts=1while its own preflight reported eight ready routes.That is the signature of per-run route selection, not capacity loss. Which run gets a usable route decides whether a consumer PR moves. So the remedy is still the one I asked for, and now with a clearer reason: try the next ready route after a non-transient failure, and keep a vision-instruct model out of a text review in the first place.
Consumer-side observation only. I have no visibility into the routing policy.
- added a commit that references this issue
on Sep 17, 2026 - addedbugSomething isn't workingSomething isn't workingpriority: criticalImmediate blocker, P0, urgent deadlock, or critical incidentImmediate blocker, P0, urgent deadlock, or critical incident
on Sep 19, 2026 seonghobae commented
on Sep 23, 2026 ContributorAuthorMore actionsFresh OriginWeave canary shows the bounded same-head continuation is now attempted after provider-capacity failure, but the continuation path itself cannot dispatch with the selected reviewer credential.
Consumer:
ContextualWisdomLab/OriginWeave#37
Exact head:219b43bfa87ab4fd90a77bb962df2797925bf661
Required Noema run/job:35816714301/107075636896
Protected central workflow source:.github@e6334e229581a918e2f22de18733b76fa65d7e71Observed sequence:
- exact-head/live-PR validation, repository-scoped
cwl-noema-reviewapp-token mint, CO sidecar bootstrap, HWP reader provisioning all succeeded; orchestrator/freepreflight became ready with five usable routes and a successful gateway chat/completions probe;Prepare Noema model verdictthen ended after 1217.9 s with HTTP 429 /provider_capacity_unavailableoninclusionai/ling-3.0-flash-sante:free;- the workflow correctly classified that transport result as continuation-eligible, selected bounded delay 157 s, re-read the live PR/head, and attempted same-head
repository_dispatchattempt 1/2; - that dispatch failed deterministically with
HTTP 403 Resource not accessible by integration.
The selected app token is minted with target-repository Actions read, Checks read, Contents read, Metadata read, Pull requests write, Security events read, Statuses read, Vulnerability alerts read. GitHub's create-repository-dispatch endpoint requires repository Contents write permission for fine-grained/App tokens, so this token cannot perform the continuation it is asked to perform. No verdict was prepared or published; fail-closed behavior is intact.
Please extend this issue's acceptance to cover the retry transport itself, not only 429/5xx classification:
- prove same-head continuation can actually be materialized with a least-privileged authorized dispatch path;
- do not broaden the reviewer publication credential merely to make retry dispatch work if a separate dispatcher capability or an in-run bounded retry can preserve tighter separation;
- contract-test the 429/5xx → bounded continuation path including credential capability, live-head revalidation, attempt counter, and no duplicate/stale dispatch;
- prove an unchanged consumer head reaches a fresh Noema review attempt after the first capacity failure.
OriginWeave will not alter the central workflow, app permissions, or blind-rerun this required job. This is canonical
.githubowner work.- exact-head/live-PR validation, repository-scoped
OriginWeave exact-consumer RCA addition, 2026-10-02 KST:
PR337 head
23e4ca5c78b5930064b91ff28383c7df89843b44, required Noema run36881970827attempt2 / job110463974265, immutable central workflow source37b10243cec3d160ecc9c1be75c71428b160a703.Downloaded raw job log and artifact11176574625 (
noema-sidecar-evidence) and verified archive SHA256 against GitHub digest4cb7b04dc49627383ca1e121803be955cb29c8fcfc51a07c2fb0c1c2dc43b047. The sidecar artifact changes the interpretation of the terminal check summary:- Preflight reports gateway ready (plain chat), candidate_count24/probed16/ready6; that is not execution-shape or sustained-review success.
- Actual review request
27af8ead556d49c8a79d2aa499e4f36abegins onnvidia_nim_google_gemma_4_31b_itat16:16:10UTC. The primary route ends RemoteDisconnected at16:25:50, followed by the same model's sub route ending RemoteDisconnected at16:30:21. - The gateway next attempts llama3.2-11b primary/sub routes, then llama3.2-90b. The last route returns non-transient provider400 at16:31:41; gateway terminal response is
invalid_request_error, status400, total latency930676ms. - Therefore
caller attempts=1means one gateway request, not one provider call. The final served_model does not account for the preceding fifteen-minute route chain, and no-failover is not an accurate description of this particular run. - The sanitized provider error body is omitted. Exact payload/schema/context cause is not established from these artifacts; a same-looking historical single-tool-call rejection must not be substituted for this request's missing diagnostics.
PR338 head
4f8e2da905d9a0e66b8fab49b38e35ee36072eadhas a genuine same-head Noema review and success via Gemma4-31b at209.8s. That comparative result is not proof that changing/pinning models fixes PR337. Existing policy remains: canonical contextual-orchestrator owns request-shaped admission, provider failure classification and safe failover; the leaf must not widen400 retry, choose a provider/model, or weaken required review.No caller rerun requested after recovery of this evidence. Required current-head OpenCode and CodeQL producers are separately queued at their actual single-member runner groups, not a Noema source finding. Native coverage artifacts for both consumers were downloaded, digest checked and verifier passed; these do not substitute for required review/scan terminal evidence. Please use this exact specimen in existing gateway-owner diagnosis, preserving explicit unknown/error outcomes and secret-safe diagnostics.
Fresh OriginWeave#339 exact-head specimen for existing issue2165, 2026-10-02 05:10 KST. No central workflow or credential changes requested from the consumer.
Head08d6a0adc43179f6331b3021b729cb5fa22d1e48, base87c4daa1830bac5a5228b6036752ad5633232085, immutable central source37b10243cec3d160ecc9c1be75c71428b160a703. Required Noema run36916933233, review job110553189840, continuation job110561054845.
Downloaded sidecar artifact11190394252; ZIP digest1754d75214213b4695398a9264046d3a2207254a31928ac47503989665e5551e verified against API digest. Plain-chat preflight ready5/16probed of24candidates, but actual review request4edf5a8a3b0349c4999e5fafd8cce0a6 lasted1087202.9ms. Gemma primary/sub and llama90b failed with RemoteDisconnected; llama11b primary/sub routing also occurred. The final OpenRouter free pool routes all returned429. Gateway returned429/rate_limit_exceeded/provider_capacity_unavailable. Caller attempts1 is one gateway call, not one provider attempt. No model verdict was produced.
Bounded continuation selected attempt1 and67second delay, then failed at20:09:08UTC with403 Resource not accessible by integration; error payload identifies create-repository-dispatch-event endpoint. Actual review log shows tokens scoped to repositories:OriginWeave with permission-contents:read; continuation attempts repository_dispatch on ContextualWisdomLab/.github. This reproduces the existing target-scoped-reviewer versus central-dispatch capability mismatch documented here for OriginWeave#37, on a fresh exact consumer head. It is distinct from my actor REST API rate-limit403 and from PR337's final provider400.
Existing canonical owner acceptance still applies: preserve reviewer/dispatcher privilege separation; prove authorized bounded same-head continuation including live head/base, counter and no stale/duplicate dispatch. Consumer cannot solve this by source edits, changing model, widening credentials, author approval or repeated blind reruns. Required review remains failed/nonpassing; local/native current-head Rust, coverage and other scan success do not replace it.
Fresh OriginWeave#330 current-head transport and recovery specimen for existing issue2165; 2026-10-02 09:20 KST. This is new run evidence, not a request to weaken review or change consumer credentials.
Head
54bbea576158da50187d7d2c0a256d6d8a41f853, protected base87c4daa1830bac5a5228b6036752ad5633232085, pinned central source37b10243cec3d160ecc9c1be75c71428b160a703. Required Noema run36940579678, review job110630927623, continuation job110640107237, all bound to this same head. Run completed FAILURE; no formal current-head review was published.Actual raw sequence:
- Plain-chat preflight ready3/probed16/candidates24 is not review-shape or sustained-execution success. The preflight request returned200 separately.
- Actual review request
b250e643314f4161b5818900a5611a1atraversed primary/sub DeepSeek provider-response failures and Gemma attempts. Gemma NIM ended504; the subsequent seven OpenRouter free-route calls ended429. Gateway terminal POST returned429/rate_limit_exceeded after1608766ms. Caller reports one gateway attempt,1608.8s,phase=response_error,outcome=provider_capacity_unavailable. This differs from PR337's final400 and PR340's HTTP200 followed by semantic-validation rejection. - The workflow classified this capacity failure as continuation-eligible, selected bounded91-second delay and attempt1, then performed live head/base checks and attempted central repository_dispatch.
- At
2026-10-01T23:58:39Zcontinuation failedResource not accessible by integration (HTTP403); its response identifies the create-repository-dispatch-event endpoint. No successful continuation notice or materialized retry is established by this job.
Downloaded exact sidecar artifact
11201036678; ZIP SHA25657b37f4816a51fa1c991db4753f9728e935aaa3602717678357e52fd57c6ec3amatches API digest; contained only sanitized sidecar stderr and preflight JSON. No provider response body, raw model response or secret was extracted.Pinned continuation workflow lines999–1072 selects
PR_REVIEW_MERGE_TOKEN || github.tokenand posts to.github/dispatches. The actual selected token identity/capabilities are masked and cannot be distinguished from this log. The review job separately minted an OriginWeave-only app token with Contents read; do NOT infer that this same token was used by the continuation. Exact proven cause at the recovery boundary is central dispatch authorization rejection, not the leaf actor's unrelated REST rate-limit403. The underlying provider-response error bodies remain unknown.An offline import of unchanged stdlib-only
noema_transport_redispatch.pypinned SHA256affabe216f1a426330869196424d04368b72f397d0ced635f14c66dfaec9170creproduces delay91 for this head/attempt0 and checks spent-budget, malformed-counter and Retry-After controls. This confirms bounded scheduling only; it cannot repair provider capacity or grant dispatch capability.Candidate corrective actions remain canonical-owner work: diagnose sanitized provider-response failures; verify the existing authorized least-privileged dispatch route separately from reviewer publication; then exercise bounded same-head continuation with current base/head, counter and duplicate/stale rejection. Verify actual retry run and terminal review publication, not just a dispatch command or comment. The consumer has not altered credentials, workflows, app permissions, routing/model selection, protected rules or required checks, and has not triggered a blind rerun/duplicate dispatch. Native current-head CI/coverage remains valid but does not substitute for this missing required review.
Fresh OriginWeave#339 current-head specimen, 2026-10-02 14:23 KST; existing issue2165, no duplicate issuer/reviewer/producer or retry requested.
Current head65ac1132d91e6561fec7277298eb671a61bb5a88; base87c4daa1830bac5a5228b6036752ad5633232085; exact event pull_request_target; central materialized source37b10243cec3d160ecc9c1be75c71428b160a703. Required run36956118364, review job110679385174, continuation110696465567. The current failure is not inherited from predecessor08d6.
Verified raw sequence:
- Substantive request2bfeb3d83f014702927cf127351810ba terminated HTTP429/rate_limit_exceeded after3115659ms, caller attempts1/duration3115.7s/phase=response_error. Exact sidecar shows10 RemoteDisconnected failures and7 OpenRouter HTTP429 failures inside this one gateway request;23 logged route starts are not23 caller attempts.
- Local typed verdict validation and publication were not reached. This is not PR340's HTTP200-plus-semantic-validation rejection. Preflight200 belongs to a different request and is not sustained review acceptance.
- Existing bounded continuation selected attempt1, waited177s, then central POST repos/ContextualWisdomLab/.github/dispatches failed at2026-10-02T03:49:58Z with403 Resource not accessible by integration. No successful continuation or formal verdict was produced.
- Event-pinned workflow selects PR_REVIEW_MERGE_TOKEN || github.token for continuation. Exact masked credential selection/ACL is unknown; current local gh actor central-admin capability does not prove that workflow credential's authority. Provider429 body is omitted, so quota versus throttling versus other capacity subcause is unknown.
Exact artifact11208196151 ZIP SHA256a76c331f231885c89f89a27e1b8966688b25fe3a7ee5b474f03f56bd47e6f17e matched the API digest. Parent independently verified39 retained source-file hashes, actual final errors, correlated10+7 failures and12 offline binding/classifier assertions. Sanitized evidence preserved outside scratch; reportSHA2562d6bb655c44eab0a35e816ebffa29fcc66bcbafe99b81309929e0739e6c1a11a.
Existing owner corrective path remains: verify credential provenance/central dispatch authorization without exposing values, preserve reviewer-versus-dispatcher capability separation, diagnose this correlated provider route chain, then prove actual live-head/base bounded continuation and terminal verdict through the existing governed route. Do not widen credentials, select a consumer provider, synthesize approval or blindly rerun the failed consumer. Consumer-native/current-head local tests and100%coverage do not replace required Noema review. No recovery/approval/merge is claimed.
2026-10-02: actual Noema 504 and personal-gateway adoption audit
Consumer
korean-writing-skills#4@e59806d7abf3c7f13478e5997b4d0cbdb20f3ba6, canonical Noema producer36967223479, model job110713563147.The sanitized direct job log at
2026-10-02T05:27:01Zrecords HTTP 504 Gateway Timeout, caller attempts=1, duration=975.9s, served_model=google/gemma-4-31b-it, outcome=provider_capacity_unavailable. It records bounded continuation eligibility in 162s (attempt1/2). The overall producer later reads CANCELLED, not SUCCESS; no qualifying current-head approval exists. Eligibility alone is not evidence that continuation was dispatched. Strix producer36967224264advanced to actualRun Strix (quick)at05:22:11Z and remains active; no final report yet.Separate implementation audit of current protected central source37b10243 and pinned contextual-orchestrator01bf92a3 confirms the original request to connect an external authenticated personal gateway has not reached central adoption. All three consumers retain the fixed loopback sidecar and default provider discovery. The owner generic
configured_gateway_sourcecontract exists, but the central launcher invokes defaultdiscover_all_models()and never connects it. Repo5/org14 Secret metadata contains noLLM_GATEWAY_API_KEY; no inference Secret timestamp in the inspected scope reflects the October1 request. This is not a claim that no key exists on the server or that unseen environments were audited.Do not treat a new key, CLI smoke, existing provider success, or a skipped workflow as completed Noema/OpenCode/Strix normalization. The actual remaining delivery is secure request-specific key custody + adopted provider wiring/policy + authenticated same-head preflight and final published scan/review receipt. Preserve ZDR, free-pool and exact-head authority; no synthetic approval or duplicate active dispatch. Existing personal-system owner has received the verified audit and a request for the actual sole writer/receipt. No secret value, server mutation or new issuer was used in this audit.
Fresh current-head Noema owner specimen: three published OriginWeave regression PRs independently reproduced the same provider-capacity/continuation boundary. This is evidence for the existing issue, not a request to rerun consumers, change credentials, or weaken review.
- PR343 exact head
1f11ac98f74fa761dd194c7883313c936b64db59, protected base87c4daa1830bac5a5228b6036752ad5633232085, required run36979365948. Model job110750194532failed after one gateway request returned HTTP 429 /provider_capacity_unavailableat 1,404,868.9 ms. Its digest-verified sidecar artifact11216386904correlates request280caad20a9a4487a71136c29d0c4069: 19 route starts, including 14 logged 429s, two 504s, two RemoteDisconnected failures and one ProviderResponseError. - PR344 exact head
a73925abb803ebd0f9fdd5199576d672e1ae2197, same base, required run36981798235. Model job110757804924failed after one gateway request returned HTTP 429 at 3,060,768.3 ms. Digest-verified artifact11218083452correlates requestef17bd92759d4a0f99370a8845f21030: 22 route starts, seven logged 429s, five RemoteDisconnected and three ProviderResponseError failures. - PR345 exact head
4618c76e6b029838dcb1b8437d11b5f3872e5a69, same base, required run36985657837. Model job110770047099failed after one gateway request returned HTTP 429 at 3,026,667.7 ms. Digest-verified artifact11219758834correlates request2f1c8fa838634cd5b8d1d8a749c99634: 20 route starts, seven logged 429s, six RemoteDisconnected and one ProviderResponseError failure.
For all three, the exact pinned Noema workflow was
37b10243cec3d160ecc9c1be75c71428b160a703. Preflight found three ready routes from 24 candidates after probing 16; this is not review success. Each existing continuation job (110760388540,110777494134,110790580019) waited its bounded 70/106/67 seconds, then failed with HTTP 403Resource not accessible by integrationat the repository-dispatch endpoint. No successful continuation or formal verdict publication is evidenced. The pinned workflow source selectsPR_REVIEW_MERGE_TOKEN || github.token; the exact selected masked credential identity is unknown, so do not conflate it with the separately minted reviewer token.This follows the earlier PR330/PR339 provider-capacity and continuation specimens in this issue. PR337's distinct final HTTP400 failure remains separately classified. These three new exact-head outcomes corroborate the existing dispatch-authorization rejection; they do not establish hidden upstream error bodies or the exact selected token principal. Please have the existing canonical owner/operator diagnose the correlated provider routes and prove the authorized least-privilege continuation path, keeping dispatch and review-publication capabilities separate. The consumer lane made no rerun, dispatch, credential/workflow/protection change or merge. Full sanitized evidence manifest:
/Users/seonghobae/orca/reports/hermes-rolling-migration/fleet/originweave-noema-three-head-failure-specimens-20261002/manifest.json(SHA-256d79092c77b60d7118d1639ccbea2755083f0cb9840e3c094e6cc195e28909309).- PR343 exact head
Additional exact-head specimen for existing issue2165, observed October 2, 2026, KST. This adds the now-terminal OriginWeave#346 outcome; it is not a duplicate consumer retry request.
Head e6cade752b9f1959bef27ce4e916bb94f19787dc, base87c4daa1830bac5a5228b6036752ad5633232085; required Noema run36991050500, review job110787150585. One caller request ended HTTP429/provider_capacity_unavailable after3169.6s. Digest-verified artifact11222610929 (ZIP SHA256 d56e27f4ff8e92c15b5c00d7579a8a16febeebce60be342aa224a5cd4dc616c0) correlates request5554ce142cc14bd29f51a29d99c42a1b and gateway latency3169628.8ms. Its23 logged internal route starts are not23 caller attempts. Logged failures include7 HTTP429,7 RemoteDisconnected and2 ProviderResponseError. Omitted route-terminal details and hidden upstream error bodies remain unknown.
Existing continuation job110807346314 waited115 seconds (attempt1), checked the exact live PR identity, and failed HTTP403/Resource not accessible by integration at the central repository-dispatch endpoint at2026-10-02T10:49:57Z. The pinned continuation selects PR_REVIEW_MERGE_TOKEN || github.token; actual masked selection/ACL is unknown. It must not be conflated with the separately minted repository-scoped review token. No successful retry or formal current-head verdict is evidenced.
Distinct scope clarification: Strix run36991050600/job110787160346 returned SUCCESS because its selector found no scannable changed files. Exact artifact11220221536 ZIP SHA256301062c87f689cebeedba1dd54086dd6389f3f726bea6de064b18a9d57ca0ed8 matches the API digest, and gate-console records the exclusion. This is a workflow scope-admission result, not a substantive vulnerability scan. Preflight200 is not scan or review acceptance.
The dependency-order corrective route remains the existing authorized owner/operator: identify the workflow dispatcher principal without exposing values; prove its least-privilege central dispatch authorization separately from reviewer publication; diagnose the correlated provider capacity; then prove bounded same-head continuation and actual terminal published verdict. The consumer made no provider/credential/workflow/protection changes, manual dispatch or rerun. Existing CodeQL36991253724 and OpenCode36991246340 remain queued; retain those producers rather than replacing them. Review approval and protected merge remain blocked.
seonghobae commented
on Oct 3, 2026 ContributorAuthorMore actions2026-10-03 exact-head recurrence: bounded capacity budget and unclassified 400
Two current
.githubReady heads add distinct owner evidence without changing either consumer branch:- PR #2563 exact
050075e5e764be43faca28506dc38754b5aa2b8f: parent Noema run37096111077failed with HTTP 502 after 149.4s and classifiedprovider_capacity_unavailable. Its authenticated same-head continuation advanced through the bounded budget; final repository-dispatch run37097033391, job111129001699, ran withNOEMA_TRANSPORT_RETRY_ATTEMPT=2and failed after 198.2s with HTTP 429 fromgoogle/gemma-4-31b-it:free, explicitly reporting that the automatic re-dispatch budget was exhausted. No verdict was published. - PR #2564 exact
d2f1f1530a0e37ac6d17c23fceac8a9a861dbbe1: Noema run37095593470, job111124893064, admitted the exact live head and then failed after 1,647.2s with HTTP 400 fromapodex/apodex-1.1-mini:free,caller attempts=1. It was not classified as capacity, socontinue-noema-transportcorrectly remained skipped; no verdict was published. This repeats the existing non-transient-400/unused-route subtype already documented on this issue.
Both jobs also failed their non-Draft evidence upload with the repository artifact-storage quota exhausted. That is deliberately fail-closed for Ready PRs; #2563 only repairs the already-model-skipped Draft path. No blind rerun, wake commit, paid fallback, or gate weakening is justified. The consumer PRs remain merge HOLD pending executable exact-head review evidence and qualifying approval.
- PR #2563 exact
seonghobae commented
on Oct 3, 2026 ContributorAuthorMore actionsFresh exact-head provider-capacity specimen from #1725 at
229027280e8bce6cc4f13e722b2b0c280a1ac17d:- Initial Noema run
37105208116, job111152370524: HTTP 429 after 1299.6s, served modelgoogle/gemma-4-26b-a4b-it:free, outcomeprovider_capacity_unavailable; bounded continuation was eligible. - Final continuation run
37107023081, job111157492058: attempt 2 again returned HTTP 429 after 413.2s, served modelgoogle/gemma-4-31b-it:free; the workflow correctly reported that the automatic re-dispatch budget was exhausted and review remains required. - Both evidence uploads also hit the GitHub artifact-storage quota.
This is terminal exact-head provider/quota evidence, not a #1725 source defect and not approval. No blind rerun, paid fallback, bypass or merge was used.
- Initial Noema run
Billing PR182: executed provider failure and separate artifact quota — October 4, 2026
Current consumer head remains
04df0fa5c9b9ccd644732a8b9627917ecacee482in metering-billing-platform#182. Noema repository-dispatch run37186250757 used trusted workflow source37b10243cec3d160ecc9c1be75c71428b160a703, admitted the exact live target, provisioned the review sidecar, then failed its actual Prepare Noema model verdict step in job111388712286:- HTTP429 after559.3s, phase=response_error, served_model=google/gemma-4-31b-it:free, caller attempts1, outcome=provider_capacity_unavailable;
- provider_attempt_count remains unknown; caller count1 must not be described as one gateway/provider attempt;
- artifact upload separately failed: Artifact storage quota has been hit. No evidence upload or verdict publication succeeded;
- continue-noema-transport job111390510063 succeeded and dispatched same-head attempt1. The parent run's terminal conclusion is cancelled after continuation, not a source-review rejection;
- native successor run37186870944/job111390596055 is already executing Prepare Noema model verdict oncwlab-s1-02. No manual duplicate dispatch or paid fallback was submitted.
The consumer's independently reviewed22-path source and722test/100% committed-head evidence remain valid local receipts, not hosted approval. A source fix or a longer caller timeout is not established by this429. Existing gateway/review capacity and artifact-retention owners must recover these separate operating gates through their existing contracts. Do not remove review/retention safeguards, delete historical artifacts indiscriminately, or mark the PR approved while evidence publication failed.
Safe causal logs and exact run/job/source binding are retained under /Users/seonghobae/orca/reports/metering-billing-platform/autoplan-20261004/noema37186250757-result.json and noema37186250757-executed-errors.log. Literal shell guard echoes were excluded from executed-error evidence. This adds a current consumer specimen to the existing incident, not a competing workflow repair or new scanner run.
2026-10-04 BandScope causal repair owner continuation: central#2126 current76ab3eb5515afdc3f1d61560ad951e34d7776e3f Ready, source tests5233pass/4skip/40subtests plus scoped471×2/100%/independentreview. ReadyNoema37195475102/job111416368239 actualprovider429/provider_capacity_unavailable caller1/306.7s; separateCreateArtifactquota. Built-in37195949589 nowcompleted/cancelled; existing successor37196887825 created10:54:18Z target2126 exacthead in runname, currentlyinprogress. No manualretry/dispatch or modelpool/payment/permissionchanges. Actor/header/runheadsha37b10243 is trustedcontrolbase, not sourcePRhead. Assigned existingissue seonghobae; preserve exactheadbound continuationand terminalevidence, failclosed review. Primarycapacitystorage/billingcoordination belongs#2356/#712; no sourcefinding from provider429 or missingartifact. Existingbounded successor must settle before anynewrequest. Localverification receipt remains distinct from genuineformalapproval.
Terminal unchangedhead continuation evidence, 2026-10-04 20:18 KST: last existing Noema37196887825 for central#212676ab3eb is completedFAILURE. PrepareNoema20:08:04KST HTTP429/provider_capacity_unavailable caller1 duration673.9s; retryattempt2, automaticredispatchbudget exhausted, review remains required. Separate sidecarartifactCreateArtifactquota persists. Actualselectedcwl-noema-review installation actor is genuine; no sourceCHANGES_REQUESTES synthesized. Original37195475102 and successor37195949589 were automaticallycancelled by continuation, no manualretryperformed. Parentfull5233pass/4skip/40subtests is localsourceevidence only. Needcanonicalgateway/providercapacity andstoragepayerresolution before genuinelyfreshreview; donotexpandfreepool into paidfallback, weaken review, inventstatus or repeatedlyrerun. ContinueindependentBandScope#827 causallyreproducedstale jobpollrepair meanwhile.
Current exact-head terminal correction — 2026-10-05 KST
Fresh bounded read corrected the prior-head specimen: existing central #2126 current head is 7f378a2, OPEN/Ready/unmerged, CHANGES_REQUESTED; no qualifying current-head OpenCode verdict. Original Noema run37202697311 and successor37203136028 are completed/cancelled. Final built-in successor37203661799 is completed/FAILURE and bound by its display title to this current source head; control head37b10243cec3d160ecc9c1be75c71428b160a703 is the trusted dispatch source, not the PR source.
Actual final log: at 2026-10-04 22:13:18 KST the gateway returned HTTP429/provider_capacity_unavailable, caller attempts1, duration795.4s, served_model google/gemma-4-31b-it:free. Automatic re-dispatch budget exhausted; review remains required. At22:13:19KST CreateArtifact independently failed with storage quota. Caller attempts1 is not a gateway-provider-attempt count. Fresh current-head GraphQL annotations separately show hosted startup refusal due to billing lock. REST /user403 was API rate exhaustion, while GraphQL viewer seonghobae and target reads succeeded; no credential replacement is justified.
No manual rerun/dispatch, paid fallback, provider/permission/billing change, synthetic status, wake commit or artifact deletion. Preserve current-head evidence, security/SBOM/release/forensic receipts and replay ledgers. Existing #2356/#712 own payer/operator read-only lock/payment authorization/plan/storage/budget verification; repo ADMIN is not billing authority. Please keep this exact-current-head terminal specimen on the existing capacity/storage recovery lane. Consumer continues safe scratch-only dependency compatibility acceptance; source tests do not substitute for hosted review or mandatory checks.
Evidence: /Users/seonghobae/.hermes/cache/scratch/bandscope-cron-20261005-0900/noema-specific-runs.json and noema-currenthead-terminal-log.json. No fresh model request or formal approval claimed.
2026-10-08 16:33 KST Health central producer PR2600 current head076a71caa3947b1ea4e1af099ea06ec3f3ea7eee is affected by this existing incident. Required Noema run37728671202/job113153229724 failed with gateway504 capacity; automatic continuation37733646970/job113168410206 failed15:00KST with HTTP429/provider_capacity_unavailable after907.4s (served_modelgoogle/gemma-4-31b-it:free), automatic redispatch budget exhausted. Both also report artifact storage quota; artifacts endpoints for these runs200/count0, but repository-wide artifacts endpoint500/emptybody(requestD9AA:50B32:1CA75:219D0:6AC7461C07:28:36UTC), so no deletion or quota mitigation claimed. Noema remains required. Law-CI peer has accepted read-only gateway RCA; free-pool/ZDR/modelbudget unchanged. OpenCode pending37729324305 is a separate offline dedicated runner incident; runner owner is recovering that existing listener/guest. No blind retry or fake approval.
Symptom
Required Noema Reviewfails on unchanged four-pillars heads with provider transport errors after the gateway exhausted its failover chain. The check is required by the org ruleset, so PRs with clean code checks stay BLOCKED.Confirmed facts (2026-09-13)
Prepare Noema model verdict→Noema gateway transport failed: HTTPError: HTTP Error 502: Bad Gateway; caller attempts=1, duration=231.8s, phase=response_error, served_model=meta/muse-glimmer-30b. Sidecar log shows manyprovider_attempt_failed … transient=True(HTTPError/TimeoutError) before the final 502.HTTP Error 429: Too Many Requests; caller attempts=1, duration=315.3s, served_model=dots-studio/dots-3-note-preview:free.TimeoutError: timed outfrom a fixedopener.open(request, timeout=120)inscripts/ci/noema_review_gate.py; that fixed timeout is no longer in HEAD 64f483d, so this is a distinct, current failure mode.Hypothesis (unconfirmed)
The noema caller performs exactly one gateway request by design (docstring near
noema_review_gate.py:1520delegates failover to the orchestrator). When every free-pool provider is transiently unavailable or rate-limited, the orchestrator surfaces the last provider status (502/429) and the gate fails closed with no bounded retry or explicit "provider capacity unavailable, retry scheduled" outcome. Free-pool 429 pressure may correlate with concurrent org-wide review sweeps.Impact
Every consumer repository PR is blocked at merge regardless of code quality until a manual
repository_dispatch: noema-reviewhappens to land on a healthy provider window.Next actions (owner: ContextualWisdomLab/.github; orchestrator behaviour: contextual-orchestrator)
tests/test_noema_review_gate.py.Internal target for a decision on (1): 2026-09-20. Next review: 2026-09-16.
Related: four-pillars#37, four-pillars#38, four-pillars#35.