Skip to content
This repository was archived by the owner on Sep 20, 2026. It is now read-only.

fix(bridge): cap reconcile_pr_fixes same-tier retries by CODER_AGENT_SLOTS - #346

Merged
joryirving merged 1 commit into
mainfrom
koji/prfix-reconcile-coder-cap
Sep 18, 2026
Merged

joryirving merged 1 commit into
mainfrom
koji/prfix-reconcile-coder-cap

Conversation

@joryirving

@joryirving joryirving commented Sep 18, 2026 •

Copy link
Copy Markdown
Contributor

Root cause

A batch of work could recreate coders concurrently on the single-slot local 27B and prefill-thrash it (observed: 6 at once, pipeline to a crawl). The bridge has several create_workload paths that spawn a coder, and not all checked CODER_AGENT_SLOTS. #344 capped three (issue-claim loop, reconcile_failures, drain_pr_fixes); three more were uncapped.

This PR — caps the remaining three, so no path spawns a coder past the cap

  • reconcile_pr_fixes same-tier retry (retry / retry-progress) — the path that caused the incident: it recreated a failed pr-fix's coder unconditionally. Now gated on the current coder's free_slots; defers (leaves the tombstone) when full.
  • reconcile_pr_fixes escalation — recreates on the next tier (coder-frontier). Now gated on that tier's free_slots, so a burst of escalations can't exceed frontier's cap either. Escalation is still reached — it's gated by frontier's slots, never blocked by a busy qwen.
  • redrive_infra — the infra-recovery path recreated coders when a model went healthy again, uncapped. Now returns False (defer) when the coder's full; the recovery driver keeps the infra marker and retries next tick.

coder_load is threaded from run_tick through reconcile_failures → infra redrive → reconcile_pr_fixes → drain, mutated in place, so all paths draw from one shared per-tick pool.

Every defer is transient (*-deferred:coder-busy, retried next tick) — never a silent BLOCK, which is the "nothing silently parked" requirement.

Audit

All create_workload sites confirmed gated: main.py issue-claim (#344) + infra-redrive (here); retry.py issue-retry (#344); prfix.py drain (#344) + pr-fix retry/escalate (here).

Tests

test_prfix.py: retry defers at capacity (load-bearing), retry proceeds + draws load when free, escalation not gated by a busy qwen, escalation defers when frontier itself is full, uncapped when no slots configured. Full suite: 654 passed.

its-saffron[bot]

This comment was marked as outdated.

…nbounded coder concurrency)

Three coder-creating paths were uncapped, so a batch of work could recreate
coders concurrently and thrash the single-slot local 27B (observed: 6 at once,
pipeline to a crawl). #344 capped the issue-claim loop, reconcile_failures, and
drain_pr_fixes; this completes the set so NO path can spawn a coder past the cap:

- reconcile_pr_fixes same-tier retry (retry / retry-progress): gate the recreate
  on the current coder's free_slots; defer (leave the tombstone) when full.
- reconcile_pr_fixes escalation: gate the next-tier recreate on THAT tier's
  free_slots (coder-frontier has its own cap), so a burst of escalations cannot
  exceed the stronger coder either.
- redrive_infra: gate the infra-recovery recreate on free_slots; return False to
  defer (reconcile_infra_parked keeps the marker and retries next tick).

coder_load is threaded from run_tick through reconcile_failures -> infra redrive
-> reconcile_pr_fixes -> drain, mutated in place, so all paths draw from one
shared per-tick pool. Every defer is transient (retry next tick), never a silent
BLOCK. Escalation to frontier is still reached -- gated by frontier's own slots,
not the local coder's.

@its-saffron its-saffron Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

AI Automated Review

Full PR review.

Analysis engine: MiniMax-M2.7@https://litellm.jory.dev/v1 (anthropic) — primary route

Recommendation: Approve

This PR completes the work started in PR 344 by gating the three remaining uncapped create_workload paths with CODER_AGENT_SLOTS.

Change-by-Change Findings

bridge/main.py — redrive_infra

  • Added free_slots import.
  • Before recreating a coder after infra recovery, checks free_slots(redrive_coder, coder_load, cfg.coder_slots) <= 0 and returns False (defer) when full.
  • Correctly increments coder_load[redrive_coder] after create_workload(manifest) to account for the new in-flight load.
  • Passes agent_load=coder_load, agent_slots=cfg.coder_slots to reconcile_pr_fixes so the shared per-tick pool is consistent across all paths.

bridge/prfix.py — reconcile_pr_fixes

  • Added agent_load=None, agent_slots=None parameters; local aliases load = agent_load if agent_load is not None else {} and slots = agent_slots or {}.
  • Same-tier retry gate (lines 443–451): Before recreating onto _prfix_current_coder(wl), checks free_slots(retry_coder, load, slots) <= 0. Defers with retry-deferred:coder-busy when full — leaves the tombstone (no delete), so it retries next tick. Escalation intentionally not gated here.
  • Escalation gate (lines 519–524): Checks free_slots(next_coder, load, slots) <= 0 for the next tier's coder (coder-frontier), not the current one. Defers with escalate-deferred:coder-busy when that tier is full.
  • Correctly increments load[retry_coder] and load[next_coder] after each create_workload call.

tests/test_prfix.py

  • test_reconcile_retry_deferred_when_coder_at_capacity: verifies retry defers at capacity; created/deleted both empty, result contains retry-deferred:coder-busy:coder.
  • test_reconcile_retry_proceeds_and_draws_load_when_slot_free: verifies retry proceeds when free; load mutated in place to {"coder": 1} for the drain that follows.
  • test_reconcile_escalate_not_gated_by_busy_qwen: verifies escalation is not blocked by a busy same-tier coder (coder:1/1), only deferred when frontier itself is full.
  • test_reconcile_retry_uncapped_when_no_slots_configured: verifies agent_slots={} means no gate, proceeds with {"coder": 9} load.
  • test_reconcile_escalate_deferred_when_frontier_at_capacity: verifies escalation defers when coder-frontier is at 4/4; no delete on defer (asserts AssertionError on delete).

Tool Harness Findings

No tool calls were issued; reviewing the corpus directly. No findings.

Unknowns or Needs Verification

No linked issue context was provided, so the "Linked Issue Fit" section is omitted. The PR body references PR 344 as the predecessor fix; the repository history confirms PR 344 capped three paths (issue-claim loop, reconcile_failures, drain_pr_fixes) and this PR addresses the three remaining paths (redrive_infra, same-tier retry, escalation) — the scope is consistent with the described audit.

The CI check results show test: success and docker: success, confirming the full test suite passes and the container builds.

@joryirving
joryirving merged commit 6760ad4 into main Sep 18, 2026
3 checks passed
@joryirving
joryirving deleted the koji/prfix-reconcile-coder-cap branch September 18, 2026 18:04
@its-miso its-miso Bot mentioned this pull request Sep 18, 2026
Sign up for free to subscribe to this conversation on GitHub. Already have an account? Sign in.

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant