fix(bridge): cap reconcile_pr_fixes same-tier retries by CODER_AGENT_SLOTS - #346
Conversation
…nbounded coder concurrency) Three coder-creating paths were uncapped, so a batch of work could recreate coders concurrently and thrash the single-slot local 27B (observed: 6 at once, pipeline to a crawl). #344 capped the issue-claim loop, reconcile_failures, and drain_pr_fixes; this completes the set so NO path can spawn a coder past the cap: - reconcile_pr_fixes same-tier retry (retry / retry-progress): gate the recreate on the current coder's free_slots; defer (leave the tombstone) when full. - reconcile_pr_fixes escalation: gate the next-tier recreate on THAT tier's free_slots (coder-frontier has its own cap), so a burst of escalations cannot exceed the stronger coder either. - redrive_infra: gate the infra-recovery recreate on free_slots; return False to defer (reconcile_infra_parked keeps the marker and retries next tick). coder_load is threaded from run_tick through reconcile_failures -> infra redrive -> reconcile_pr_fixes -> drain, mutated in place, so all paths draw from one shared per-tick pool. Every defer is transient (retry next tick), never a silent BLOCK. Escalation to frontier is still reached -- gated by frontier's own slots, not the local coder's.
b4b9c17 to
df3417e
Compare
There was a problem hiding this comment.
AI Automated Review
Full PR review.
Analysis engine: MiniMax-M2.7@https://litellm.jory.dev/v1 (anthropic) — primary route
Recommendation: Approve
This PR completes the work started in PR 344 by gating the three remaining uncapped create_workload paths with CODER_AGENT_SLOTS.
Change-by-Change Findings
bridge/main.py — redrive_infra
- Added
free_slotsimport. - Before recreating a coder after infra recovery, checks
free_slots(redrive_coder, coder_load, cfg.coder_slots) <= 0and returnsFalse(defer) when full. - Correctly increments
coder_load[redrive_coder]aftercreate_workload(manifest)to account for the new in-flight load. - Passes
agent_load=coder_load, agent_slots=cfg.coder_slotstoreconcile_pr_fixesso the shared per-tick pool is consistent across all paths.
bridge/prfix.py — reconcile_pr_fixes
- Added
agent_load=None, agent_slots=Noneparameters; local aliasesload = agent_load if agent_load is not None else {}andslots = agent_slots or {}. - Same-tier retry gate (lines 443–451): Before recreating onto
_prfix_current_coder(wl), checksfree_slots(retry_coder, load, slots) <= 0. Defers withretry-deferred:coder-busywhen full — leaves the tombstone (nodelete), so it retries next tick. Escalation intentionally not gated here. - Escalation gate (lines 519–524): Checks
free_slots(next_coder, load, slots) <= 0for the next tier's coder (coder-frontier), not the current one. Defers withescalate-deferred:coder-busywhen that tier is full. - Correctly increments
load[retry_coder]andload[next_coder]after eachcreate_workloadcall.
tests/test_prfix.py
test_reconcile_retry_deferred_when_coder_at_capacity: verifies retry defers at capacity; created/deleted both empty, result containsretry-deferred:coder-busy:coder.test_reconcile_retry_proceeds_and_draws_load_when_slot_free: verifies retry proceeds when free;loadmutated in place to{"coder": 1}for the drain that follows.test_reconcile_escalate_not_gated_by_busy_qwen: verifies escalation is not blocked by a busy same-tier coder (coder:1/1), only deferred when frontier itself is full.test_reconcile_retry_uncapped_when_no_slots_configured: verifiesagent_slots={}means no gate, proceeds with{"coder": 9}load.test_reconcile_escalate_deferred_when_frontier_at_capacity: verifies escalation defers whencoder-frontieris at4/4; no delete on defer (asserts AssertionError on delete).
Tool Harness Findings
No tool calls were issued; reviewing the corpus directly. No findings.
Unknowns or Needs Verification
No linked issue context was provided, so the "Linked Issue Fit" section is omitted. The PR body references PR 344 as the predecessor fix; the repository history confirms PR 344 capped three paths (issue-claim loop, reconcile_failures, drain_pr_fixes) and this PR addresses the three remaining paths (redrive_infra, same-tier retry, escalation) — the scope is consistent with the described audit.
The CI check results show test: success and docker: success, confirming the full test suite passes and the container builds.
Root cause
A batch of work could recreate coders concurrently on the single-slot local 27B and prefill-thrash it (observed: 6 at once, pipeline to a crawl). The bridge has several
create_workloadpaths that spawn a coder, and not all checkedCODER_AGENT_SLOTS. #344 capped three (issue-claim loop,reconcile_failures,drain_pr_fixes); three more were uncapped.This PR — caps the remaining three, so no path spawns a coder past the cap
reconcile_pr_fixessame-tier retry (retry/retry-progress) — the path that caused the incident: it recreated a failed pr-fix's coder unconditionally. Now gated on the current coder'sfree_slots; defers (leaves the tombstone) when full.reconcile_pr_fixesescalation — recreates on the next tier (coder-frontier). Now gated on that tier'sfree_slots, so a burst of escalations can't exceed frontier's cap either. Escalation is still reached — it's gated by frontier's slots, never blocked by a busy qwen.redrive_infra— the infra-recovery path recreated coders when a model went healthy again, uncapped. Now returnsFalse(defer) when the coder's full; the recovery driver keeps the infra marker and retries next tick.coder_loadis threaded fromrun_tickthroughreconcile_failures → infra redrive → reconcile_pr_fixes → drain, mutated in place, so all paths draw from one shared per-tick pool.Every defer is transient (
*-deferred:coder-busy, retried next tick) — never a silent BLOCK, which is the "nothing silently parked" requirement.Audit
All
create_workloadsites confirmed gated:main.pyissue-claim (#344) + infra-redrive (here);retry.pyissue-retry (#344);prfix.pydrain (#344) + pr-fix retry/escalate (here).Tests
test_prfix.py: retry defers at capacity (load-bearing), retry proceeds + draws load when free, escalation not gated by a busy qwen, escalation defers when frontier itself is full, uncapped when no slots configured. Full suite: 654 passed.