Skip to content
This repository was archived by the owner on Sep 20, 2026. It is now read-only.

fix(bridge): cap concurrent pr-fix coders by CODER_AGENT_SLOTS - #340

Merged
joryirving merged 1 commit into
mainfrom
koji/prfix-coder-cap
Sep 17, 2026
Merged

joryirving merged 1 commit into
mainfrom
koji/prfix-coder-cap

Conversation

@joryirving

Copy link
Copy Markdown
Contributor

What

Make the pr-fix dispatch path honor CODER_AGENT_SLOTS. It previously had no concurrency cap.

The incident

PR_FIX_LANE_AGENTS.NORMAL was moved from coder-frontier (cloud API, tolerates concurrency) to coder (local qwen3.8-27b, single 3090 slot). Three NORMAL pr-fixes then dispatched at once and thrashed the model with prefill churn.

Two gaps combined:

  1. drain_pr_fixes created a Workload for every actionable queued item — no coders_saturated/slot check.
  2. _load_by_coder_agent counted only issue Workloads (created-by=dispatch-bridge), never pr-fix ones (created-by=dispatch-bridge-prfix), so pr-fix coders were invisible to the cap even where it was applied.

Harmless while pr-fixes ran on a cloud API; a pileup on a single-slot local model is not.

Fix (fail-forward — keeps NORMAL pr-fix on the local coder)

  • _load_by_coder_agent counts both issue and pr-fix Workloads (one set-based created-by in (dispatch-bridge,dispatch-bridge-prfix) list). pr-fix Workloads are pipeline-shaped, so their coder resolves via _coder_agent_name's pipeline fallback (fix(bridge): count pipeline-shaped coders toward CODER_AGENT_SLOTS #338). Issue and pr-fix coders now share one accounting.
  • drain_pr_fixes takes agent_load + agent_slots, skips an item whose coder is at capacity (leaves it QUEUED with skip:coder-busy), and draws the load down per dispatch.
  • run_tick threads the coder load + slots into the drain.

Tests

Five added (all load-bearing — verified they fail without the change): drain caps at slot capacity, counts in-flight load, stays uncapped when no slots configured, dispatches more on a higher-capacity lane, and _load_by_coder_agent counts a pr-fix Workload. Full suite: 646 passed.

Known bound

coder_load passed to the drain is the tick-start load (in-flight issue + pr-fix). Issue coders assigned earlier in the same tick aren't reflected, so a single tick can still place one issue coder and one pr-fix coder on the same coder. Bounded and self-corrects next tick — a large improvement over the previous unbounded pr-fix dispatch. Full same-tick sharing would require threading run_once's in-tick assignments into the drain; deferred to avoid changing run_once's copy semantics.

drain_pr_fixes dispatched a Workload for every actionable queued item with no
capacity check, and _load_by_coder_agent only counted issue Workloads
(created-by=dispatch-bridge), never pr-fix ones (dispatch-bridge-prfix). With
pr-fixes routed to a cloud API that was harmless; routing NORMAL pr-fixes to
the single-slot local coder surfaced it as three concurrent coders thrashing
the model.

- _load_by_coder_agent now counts pr-fix Workloads too (one set-based list
  call; pr-fix Workloads are pipeline-shaped, resolved by _coder_agent_name's
  pipeline fallback) so issue and pr-fix coders share one accounting.
- drain_pr_fixes takes agent_load + agent_slots, skips an item whose coder is
  at capacity (leaving it QUEUED), and draws the load down as it dispatches.
- run_tick threads the coder load + slots into the drain.

@its-saffron its-saffron Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI Automated Review

Full PR review.

Analysis engine: MiniMax-M2.7@https://litellm.jory.dev/v1 (anthropic) — primary route

Recommendation: Approve

This PR correctly addresses the concurrency issue where PR-fix coders were not being bounded by CODER_AGENT_SLOTS, causing pileup on single-slot local models.

Change-by-Change Findings

bridge/main.py

  • Added _active_coder_pool_workloads() (lines 639-661): New helper that queries both issue (dispatch-bridge) and pr-fix (dispatch-bridge-prfix) Workloads in one set-based list call. The docstring explicitly explains why this matters: pr-fix Workloads share the same coder Agent (and GPU) as issue coders, so they must be counted together to prevent same-tick collisions.
  • Modified _load_by_coder_agent() (line 717): Now calls _active_coder_pool_workloads instead of _active_workloads when workloads is None, ensuring the load accounting includes pr-fix Workloads.
  • Threaded capacity into drain_pr_fixes() (line 1287): Passes coder_load and cfg.coder_slots to enforce slot-based dispatch limits.

bridge/prfix.py

  • Added parameters (line 198): agent_load and agent_slots optional parameters with clear docstring explaining the capacity-bound behavior.
  • Added slot guard (lines 227-229): Before dispatching, checks free_slots(coder_agent, load, slots) <= 0 and skips with skip:coder-busy:{coder_agent} log line.
  • Increment load counter (line 242): After successful dispatch, increments load[coder_agent] so subsequent items in the same drain pass are aware of earlier dispatches.

tests/test_bridge_runtime.py

  • test_counts_prfix_workload_toward_shared_pool: Verifies _load_by_coder_agent counts a pr-fix Workload (created-by=dispatch-bridge-prfix) toward the shared pool.

tests/test_prfix.py

  • test_drain_caps_concurrent_coders_at_slot_capacity: Two items on same coder, slot=1 → first dispatches, second skips coder-busy.
  • test_drain_counts_inflight_load_against_cap: Pre-existing in-flight coder (agent_load) occupies the slot, preventing fresh dispatch.
  • test_drain_uncapped_when_no_slots_configured: Empty agent_slots preserves legacy dispatch-all behavior.
  • test_drain_higher_capacity_lane_dispatches_more: Lane with 4 slots dispatches all 3 items.

Tool Harness Findings

No tool calls were issued; reviewing corpus directly.

Standards Compliance

No repository standards file (AGENTS.md) applies to the code changes (it governs issue filing conventions, not implementation patterns). The implementation follows existing repository patterns for agent_load/agent_slots usage (evidenced by bridge/workload.py and bridge/main.py pre-existing patterns).

CI Check Results

Both checks passed:

  • docker: success
  • test: success (646 passed, including the 5 new load-bearing tests)

Known Bound

The PR body correctly documents the known same-tick sharing bound: issue coders assigned earlier in the same tick are not reflected in coder_load passed to the drain, so a single tick could place one issue coder and one pr-fix coder on the same coder. This is bounded and self-corrects next tick; full same-tick accounting would require changing run_once's copy semantics and was deferred.

Unknowns / Needs Verification

None. The diff is complete, tests are present and passing, and the implementation is internally consistent with established patterns in the codebase.

@joryirving
joryirving merged commit a136630 into main Sep 17, 2026
3 checks passed
@joryirving
joryirving deleted the koji/prfix-coder-cap branch September 17, 2026 15:41
@its-miso its-miso Bot mentioned this pull request Sep 17, 2026
joryirving added a commit that referenced this pull request Sep 17, 2026
…ad across the tick (#344)

reconcile_failures (the Failed-Workload retry recreation) was the third
workload-creation path and the only one still uncapped: #340 capped the
issue-claim loop and the pr-fix drain, but a retried coder was recreated
regardless of coder capacity. Observed in-cluster: wl-...-616 retried (attempt
3/3) onto the local coder while a pr-fix coder was already running -- two
coders on the single-slot 27b.

- reconcile_failures now resolves the coder before the delete, checks
  free_slots, and defers a retry whose coder is at capacity (leaves the Failed
  tombstone; list_failed re-offers it next tick).
- coder_load is computed once before the retry pass and threaded through all
  three passes (reconcile -> issue-claim -> pr-fix drain). reconcile mutates it
  in place, so a retried coder is visible to the later passes: the three paths
  draw from one pool instead of each independently filling the same slot.
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant