Skip to content

Builder calibration: five bounded deliveries each for Cursor, Muse, Antigravity, and Devin #659

Description

@jeffhuber

Goal

Collect enough comparable delivery evidence to evaluate Cursor, Muse, Antigravity, and Devin as Code Mower builders without manufacturing low-value production changes.

Sampling target

Complete five bounded, real issues per provider. Balance each provider's five samples across the same task classes where practical:

  1. focused test or fixture hardening;
  2. small CLI text/JSON consistency change;
  3. documentation or adoption-flow correction;
  4. privacy/redaction guardrail;
  5. provider/configuration validation.

Each issue should normally fit below 300 changed lines and finish within one working session. Existing real backlog is preferred; do not create duplicate implementations merely to complete a quota.

Controlled protocol

  • Freeze the task contract and acceptance criteria before selecting the builder.
  • Use a fresh branch/worktree and explicit builder identity.
  • Do not reveal peer-review findings until the builder declares its first implementation complete.
  • Require established Codex and Claude review, excluding the author lane.
  • Record all audit findings with one disposition: accepted-fixed, owner-decided, false-positive, duplicate, or infrastructure.
  • Record start-to-PR, PR-to-merge, total elapsed time, interventions, CI retries, audit/fix rounds, accepted severity, reported tokens/cost, and post-merge health.
  • Treat unavailable cost as unavailable, never zero.
  • Upload metadata-only evidence after dry-run inspection. Never upload source, diffs, transcripts, issue bodies, raw stdout/stderr, auth output, local paths, or secrets.

Builder targets

Completed samples

Infrastructure evidence that does not count as a completed sample

Remaining samples must come from real backlog. Do not create low-value or duplicate changes solely to complete the target.

Antigravity and Devin currently require explicit manual adapter sessions. Their dispatch/setup time must be recorded as orchestration overhead rather than omitted. The experiment should inform whether first-class adapters are worth implementing.

Decision output

After 20 completed samples, publish a provider scorecard with confidence intervals and a promotion recommendation for each role. Do not rank providers solely by raw BLOCKED rate because reviewer assignment and task mix can confound it.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    builder-calibrationControlled builder comparison and calibration workepicEpic tracking issue

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions