fix(bridge): cap concurrent pr-fix coders by CODER_AGENT_SLOTS - #340
Conversation
drain_pr_fixes dispatched a Workload for every actionable queued item with no capacity check, and _load_by_coder_agent only counted issue Workloads (created-by=dispatch-bridge), never pr-fix ones (dispatch-bridge-prfix). With pr-fixes routed to a cloud API that was harmless; routing NORMAL pr-fixes to the single-slot local coder surfaced it as three concurrent coders thrashing the model. - _load_by_coder_agent now counts pr-fix Workloads too (one set-based list call; pr-fix Workloads are pipeline-shaped, resolved by _coder_agent_name's pipeline fallback) so issue and pr-fix coders share one accounting. - drain_pr_fixes takes agent_load + agent_slots, skips an item whose coder is at capacity (leaving it QUEUED), and draws the load down as it dispatches. - run_tick threads the coder load + slots into the drain.
There was a problem hiding this comment.
AI Automated Review
Full PR review.
Analysis engine: MiniMax-M2.7@https://litellm.jory.dev/v1 (anthropic) — primary route
Recommendation: Approve
This PR correctly addresses the concurrency issue where PR-fix coders were not being bounded by CODER_AGENT_SLOTS, causing pileup on single-slot local models.
Change-by-Change Findings
bridge/main.py
- Added
_active_coder_pool_workloads()(lines 639-661): New helper that queries both issue (dispatch-bridge) and pr-fix (dispatch-bridge-prfix) Workloads in one set-based list call. The docstring explicitly explains why this matters: pr-fix Workloads share the same coder Agent (and GPU) as issue coders, so they must be counted together to prevent same-tick collisions. - Modified
_load_by_coder_agent()(line 717): Now calls_active_coder_pool_workloadsinstead of_active_workloadswhen workloads is None, ensuring the load accounting includes pr-fix Workloads. - Threaded capacity into
drain_pr_fixes()(line 1287): Passescoder_loadandcfg.coder_slotsto enforce slot-based dispatch limits.
bridge/prfix.py
- Added parameters (line 198):
agent_loadandagent_slotsoptional parameters with clear docstring explaining the capacity-bound behavior. - Added slot guard (lines 227-229): Before dispatching, checks
free_slots(coder_agent, load, slots) <= 0and skips withskip:coder-busy:{coder_agent}log line. - Increment load counter (line 242): After successful dispatch, increments
load[coder_agent]so subsequent items in the same drain pass are aware of earlier dispatches.
tests/test_bridge_runtime.py
test_counts_prfix_workload_toward_shared_pool: Verifies_load_by_coder_agentcounts a pr-fix Workload (created-by=dispatch-bridge-prfix) toward the shared pool.
tests/test_prfix.py
test_drain_caps_concurrent_coders_at_slot_capacity: Two items on same coder, slot=1 → first dispatches, second skipscoder-busy.test_drain_counts_inflight_load_against_cap: Pre-existing in-flight coder (agent_load) occupies the slot, preventing fresh dispatch.test_drain_uncapped_when_no_slots_configured: Emptyagent_slotspreserves legacy dispatch-all behavior.test_drain_higher_capacity_lane_dispatches_more: Lane with 4 slots dispatches all 3 items.
Tool Harness Findings
No tool calls were issued; reviewing corpus directly.
Standards Compliance
No repository standards file (AGENTS.md) applies to the code changes (it governs issue filing conventions, not implementation patterns). The implementation follows existing repository patterns for agent_load/agent_slots usage (evidenced by bridge/workload.py and bridge/main.py pre-existing patterns).
CI Check Results
Both checks passed:
docker: successtest: success (646 passed, including the 5 new load-bearing tests)
Known Bound
The PR body correctly documents the known same-tick sharing bound: issue coders assigned earlier in the same tick are not reflected in coder_load passed to the drain, so a single tick could place one issue coder and one pr-fix coder on the same coder. This is bounded and self-corrects next tick; full same-tick accounting would require changing run_once's copy semantics and was deferred.
Unknowns / Needs Verification
None. The diff is complete, tests are present and passing, and the implementation is internally consistent with established patterns in the codebase.
…ad across the tick (#344) reconcile_failures (the Failed-Workload retry recreation) was the third workload-creation path and the only one still uncapped: #340 capped the issue-claim loop and the pr-fix drain, but a retried coder was recreated regardless of coder capacity. Observed in-cluster: wl-...-616 retried (attempt 3/3) onto the local coder while a pr-fix coder was already running -- two coders on the single-slot 27b. - reconcile_failures now resolves the coder before the delete, checks free_slots, and defers a retry whose coder is at capacity (leaves the Failed tombstone; list_failed re-offers it next tick). - coder_load is computed once before the retry pass and threaded through all three passes (reconcile -> issue-claim -> pr-fix drain). reconcile mutates it in place, so a retried coder is visible to the later passes: the three paths draw from one pool instead of each independently filling the same slot.
What
Make the pr-fix dispatch path honor
CODER_AGENT_SLOTS. It previously had no concurrency cap.The incident
PR_FIX_LANE_AGENTS.NORMALwas moved fromcoder-frontier(cloud API, tolerates concurrency) tocoder(local qwen3.8-27b, single 3090 slot). Three NORMAL pr-fixes then dispatched at once and thrashed the model with prefill churn.Two gaps combined:
drain_pr_fixescreated a Workload for every actionable queued item — nocoders_saturated/slot check._load_by_coder_agentcounted only issue Workloads (created-by=dispatch-bridge), never pr-fix ones (created-by=dispatch-bridge-prfix), so pr-fix coders were invisible to the cap even where it was applied.Harmless while pr-fixes ran on a cloud API; a pileup on a single-slot local model is not.
Fix (fail-forward — keeps NORMAL pr-fix on the local coder)
_load_by_coder_agentcounts both issue and pr-fix Workloads (one set-basedcreated-by in (dispatch-bridge,dispatch-bridge-prfix)list). pr-fix Workloads are pipeline-shaped, so their coder resolves via_coder_agent_name's pipeline fallback (fix(bridge): count pipeline-shaped coders toward CODER_AGENT_SLOTS #338). Issue and pr-fix coders now share one accounting.drain_pr_fixestakesagent_load+agent_slots, skips an item whose coder is at capacity (leaves it QUEUED withskip:coder-busy), and draws the load down per dispatch.run_tickthreads the coder load + slots into the drain.Tests
Five added (all load-bearing — verified they fail without the change): drain caps at slot capacity, counts in-flight load, stays uncapped when no slots configured, dispatches more on a higher-capacity lane, and
_load_by_coder_agentcounts a pr-fix Workload. Full suite: 646 passed.Known bound
coder_loadpassed to the drain is the tick-start load (in-flight issue + pr-fix). Issue coders assigned earlier in the same tick aren't reflected, so a single tick can still place one issue coder and one pr-fix coder on the same coder. Bounded and self-corrects next tick — a large improvement over the previous unbounded pr-fix dispatch. Full same-tick sharing would require threadingrun_once's in-tick assignments into the drain; deferred to avoid changingrun_once's copy semantics.