Per work item:
- maxAttempts: 3 (implementor runs before escalating to human)
- maxCostPerWorkItem: $5 (cumulative across all runs for this item)
Per rework cycle (reviewer rejects → implementor re-runs):
- maxReworkCycles: 2 (before escalating to human)
Global:
- maxDailyCost: $50 (hard stop across all agents)
- maxConcurrentRuns: 3 (you already have per-work-item guards)
The key insight: retry and rework are different things. Retry is "same input, try again" (almost never useful for LLMs — they'll make the same mistake). Rework is "new input (review
feedback), try again" — this is valuable because the reviewer's comments are new context.
- No auto-retry on failure. If an agent errors out, park it for human review. LLMs don't benefit from retry the way network calls do.
- Auto-rework with budget caps. Reviewer rejects → feed comments back to implementor → re-run. But max N cycles and max $X per work item. Exceed either → escalate to human.
This gives you the automation leverage without the runaway cost risk. The budget is a policy — which connects to your policy system