Follow-up to rockymadden's review on #3369 (finding 2).
tests/experiments/nightly.yaml hands every task turn_timeout: 900 / task_timeout: 1200, sized for the codex arm. The claude-code arm streams the whole session as one turn, so turn_timeout is its effective wall budget — and 20+ BPMN tasks (plus other suites) still inherit the default. Result: per-task budget PRs keep recurring (#3150 → #3369) each time a task's claude duration drifts across the codex-sized cap.
Proposal: a claude-specific experiment config (e.g. nightly-claude.yaml) or per-agent run_limits support, defaulting the claude arm to ~1800/2400, so the next slow task is a config line instead of a third PR of this shape.
Evidence for the sizing: #3369's analysis — tool execution is 16–50s of 900–1800s walls, the rest is model generation; three-day nightly history shows tasks passing with 3–37s of margin under the current defaults.
🤖 Generated with Claude Code
Follow-up to rockymadden's review on #3369 (finding 2).
tests/experiments/nightly.yamlhands every taskturn_timeout: 900/task_timeout: 1200, sized for the codex arm. The claude-code arm streams the whole session as one turn, soturn_timeoutis its effective wall budget — and 20+ BPMN tasks (plus other suites) still inherit the default. Result: per-task budget PRs keep recurring (#3150 → #3369) each time a task's claude duration drifts across the codex-sized cap.Proposal: a claude-specific experiment config (e.g.
nightly-claude.yaml) or per-agentrun_limitssupport, defaulting the claude arm to ~1800/2400, so the next slow task is a config line instead of a third PR of this shape.Evidence for the sizing: #3369's analysis — tool execution is 16–50s of 900–1800s walls, the rest is model generation; three-day nightly history shows tasks passing with 3–37s of margin under the current defaults.
🤖 Generated with Claude Code