Skip to content

perf(repo): plan shards on the measured cost of every verdict - #225

Merged
kiro-systemf[bot] merged 2 commits into
mainfrom
stryker/shard-balance-actual
Oct 7, 2026
Merged

kiro-systemf[bot] merged 2 commits into
mainfrom
stryker/shard-balance-actual

Conversation

@systemfsoftware-maker

@systemfsoftware-maker systemfsoftware-maker commented Oct 7, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

stryker plan now splits shards by the time each verdict measurably cost. Before, a verdict decided without a test (CompileError, NoCoverage, Ignored) recorded no cost, so the plan priced it as if every covering test would run for it.

Before / after (main run 37602076567, 11m08s, published CLI 17.0.2)

In that run, stryker-js had 3547 CompileError mutants. Each was priced at about 68 s of covering tests, about 241,914 s in total. All 1607 Killed and Survived mutants together measured 3,216 s. The plan therefore balanced CompileError counts (177–178 per shard), while the test time each shard actually ran ranged from 88 to 237 s. Shard wall time tracks that test time: stryker-js part = 99 s + 1.51 × recorded test seconds, r = 0.91.

today (37602076567) recorded costs (replayed LPT)
test seconds per shard 88–237 145–179
slowest Mutation step 498 s ~452 s (model)
run wall-clock 11m08s ~10m22s (model)
CompileErrors priced per shard ~68 s each their share of the check call (milliseconds)

The right-hand column is a projection: I replayed the planner's LPT using the previous run's recorded actualMs. Those recorded costs predict the next run's per-mutant cost with r = 0.81, against 0.51 for predictedMs. A main run on a release containing this change replaces the model with measurements.

Design decisions

  • Check time is the measured cost. The checker pool times each check call and divides the elapsed time evenly across the call's mutants. A mutant seen by several checkers sums its shares. Ignored mutants, and NoCoverage with no checker, record 0: nothing ran.
  • Budget unchanged. The run budget still prices only verdicts that ran a test. budget.predictedSeconds and the budget gate do not move.
  • StaticVerdict changes. The StaticVerdict event's costMs now includes check time for static mutants a checker rejected, and the changeset says so.
  • Verdict-Semantics: unchanged: no status changes.

Borrowed: cargo-mutants and Bazel test sharding by recorded per-test durations: plan on the last run's measured cost, and use predictions only as a cold-start fallback. Rejected: pricing test-free verdicts at 0, because LPT would put every zero-cost item in the same least-loaded bin and all 3547 checks would land in one shard. Also rejected: capping CompileErrors per shard, which is a second heuristic rather than a measurement.

Validation

  • Failing before. With src reverted to main in a scratch worktree, all 3 new scenarios fail:
    • The plan scenario runs the built CLI run, then plan --full, and asserts on the written ShardPlan. It fails with everyScheduledMutantCarriesAMeasuredCost: false, thePlanPricesTheRecordedCosts: false, thePlanPricesNoWholeSuitePrediction: false.
    • The report scenario fails with rejectedMutantsAreCompileErrorsChargedTheCheckTime: false, ignoredRecordsZeroCosts: false, noTestVerdictsPublishTheirMeasuredCostOnTheStream: false.
    • The no-checker scenario fails with everyUncoveredMutantIsPricedAtZero: false.
  • Passing after. All three pass. Format, lint, typecheck and check:ci exit 0, and the changeset gate passes.
  • Review. Correctness review: no blocking findings. Its two test-strength findings were fixed: rejected CompileErrors must cost > 0, and the plan test runs the real plan path. A simplify pass merged a duplicate zero-cost helper.

Effect on main

Main's plan job runs the published CLI, so this takes effect only after a release. The first plan after that release still reads cost records written by 17.0.2, where these costs are null. The run after that is the first one planned on these recorded costs.

A verdict the engine decided without running a test recorded no cost, so
`stryker plan` priced it at a whole-suite prediction. On run 37602076567 that
covered 3547 CompileError, 276 NoCoverage and 741 Ignored mutants at about
68 s each: 241,914 predicted seconds against 3,216 measured test seconds for
every Killed and Survived mutant. LPT then balanced CompileError counts (every
shard got 177-178) rather than real work, and shard wall time tracked the
measured test time (r=0.91) at 221-415 s per shard.

A verdict a checker decided now records its share of the check call that
decided it, an ignored verdict or one left uncovered with no checker records
zero, and the run budget still counts only the verdicts that ran a test, so a
shard plan balances on recorded times

Verdict-Semantics: unchanged
Untested verdicts recorded no cost, so the planner priced them at a whole-suite prediction; the doc names the regression gate

Verdict-Semantics: unchanged
@systemfsoftware-maker systemfsoftware-maker changed the title stryker/shard balance actual perf(repo): plan shards on the measured cost of every verdict Oct 7, 2026

@kiro-systemf kiro-systemf Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Planner weights shards by each verdict's measured cost; CompileError records share of its check call, Ignored records 0; no status changes. Failing-before scenarios included. 8/8 green.

@kiro-systemf
kiro-systemf Bot merged commit 81a869e into main Oct 7, 2026
8 checks passed
@kiro-systemf
kiro-systemf Bot deleted the stryker/shard-balance-actual branch October 7, 2026 12:06
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant