Goal
Collect enough comparable delivery evidence to evaluate Cursor, Muse, Antigravity, and Devin as Code Mower builders without manufacturing low-value production changes.
Sampling target
Complete five bounded, real issues per provider. Balance each provider's five samples across the same task classes where practical:
- focused test or fixture hardening;
- small CLI text/JSON consistency change;
- documentation or adoption-flow correction;
- privacy/redaction guardrail;
- provider/configuration validation.
Each issue should normally fit below 300 changed lines and finish within one working session. Existing real backlog is preferred; do not create duplicate implementations merely to complete a quota.
Controlled protocol
- Freeze the task contract and acceptance criteria before selecting the builder.
- Use a fresh branch/worktree and explicit builder identity.
- Do not reveal peer-review findings until the builder declares its first implementation complete.
- Require established Codex and Claude review, excluding the author lane.
- Record all audit findings with one disposition: accepted-fixed, owner-decided, false-positive, duplicate, or infrastructure.
- Record start-to-PR, PR-to-merge, total elapsed time, interventions, CI retries, audit/fix rounds, accepted severity, reported tokens/cost, and post-merge health.
- Treat unavailable cost as unavailable, never zero.
- Upload metadata-only evidence after dry-run inspection. Never upload source, diffs, transcripts, issue bodies, raw stdout/stderr, auth output, local paths, or secrets.
Builder targets
Completed samples
Infrastructure evidence that does not count as a completed sample
Remaining samples must come from real backlog. Do not create low-value or duplicate changes solely to complete the target.
Antigravity and Devin currently require explicit manual adapter sessions. Their dispatch/setup time must be recorded as orchestration overhead rather than omitted. The experiment should inform whether first-class adapters are worth implementing.
Decision output
After 20 completed samples, publish a provider scorecard with confidence intervals and a promotion recommendation for each role. Do not rank providers solely by raw BLOCKED rate because reviewer assignment and task mix can confound it.
Goal
Collect enough comparable delivery evidence to evaluate Cursor, Muse, Antigravity, and Devin as Code Mower builders without manufacturing low-value production changes.
Sampling target
Complete five bounded, real issues per provider. Balance each provider's five samples across the same task classes where practical:
Each issue should normally fit below 300 changed lines and finish within one working session. Existing real backlog is preferred; do not create duplicate implementations merely to complete a quota.
Controlled protocol
Builder targets
Completed samples
Infrastructure evidence that does not count as a completed sample
Remaining samples must come from real backlog. Do not create low-value or duplicate changes solely to complete the target.
Antigravity and Devin currently require explicit manual adapter sessions. Their dispatch/setup time must be recorded as orchestration overhead rather than omitted. The experiment should inform whether first-class adapters are worth implementing.
Decision output
After 20 completed samples, publish a provider scorecard with confidence intervals and a promotion recommendation for each role. Do not rank providers solely by raw BLOCKED rate because reviewer assignment and task mix can confound it.