Repository navigation
perf(repo): plan shards on the measured cost of every verdict - #225
Merged
Merged
Conversation
A verdict the engine decided without running a test recorded no cost, so `stryker plan` priced it at a whole-suite prediction. On run 37602076567 that covered 3547 CompileError, 276 NoCoverage and 741 Ignored mutants at about 68 s each: 241,914 predicted seconds against 3,216 measured test seconds for every Killed and Survived mutant. LPT then balanced CompileError counts (every shard got 177-178) rather than real work, and shard wall time tracked the measured test time (r=0.91) at 221-415 s per shard. A verdict a checker decided now records its share of the check call that decided it, an ignored verdict or one left uncovered with no checker records zero, and the run budget still counts only the verdicts that ran a test, so a shard plan balances on recorded times Verdict-Semantics: unchanged
Untested verdicts recorded no cost, so the planner priced them at a whole-suite prediction; the doc names the regression gate Verdict-Semantics: unchanged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
stryker plannow splits shards by the time each verdict measurably cost. Before, a verdict decided without a test (CompileError, NoCoverage, Ignored) recorded no cost, so the plan priced it as if every covering test would run for it.Before / after (main run 37602076567, 11m08s, published CLI 17.0.2)
In that run, stryker-js had 3547 CompileError mutants. Each was priced at about 68 s of covering tests, about 241,914 s in total. All 1607 Killed and Survived mutants together measured 3,216 s. The plan therefore balanced CompileError counts (177–178 per shard), while the test time each shard actually ran ranged from 88 to 237 s. Shard wall time tracks that test time: stryker-js part = 99 s + 1.51 × recorded test seconds, r = 0.91.
The right-hand column is a projection: I replayed the planner's LPT using the previous run's recorded
actualMs. Those recorded costs predict the next run's per-mutant cost with r = 0.81, against 0.51 forpredictedMs. A main run on a release containing this change replaces the model with measurements.Design decisions
checkcall and divides the elapsed time evenly across the call's mutants. A mutant seen by several checkers sums its shares. Ignored mutants, and NoCoverage with no checker, record 0: nothing ran.budget.predictedSecondsand the budget gate do not move.costMsnow includes check time for static mutants a checker rejected, and the changeset says so.Verdict-Semantics: unchanged: no status changes.Borrowed: cargo-mutants and Bazel test sharding by recorded per-test durations: plan on the last run's measured cost, and use predictions only as a cold-start fallback. Rejected: pricing test-free verdicts at 0, because LPT would put every zero-cost item in the same least-loaded bin and all 3547 checks would land in one shard. Also rejected: capping CompileErrors per shard, which is a second heuristic rather than a measurement.
Validation
run, thenplan --full, and asserts on the written ShardPlan. It fails witheveryScheduledMutantCarriesAMeasuredCost: false,thePlanPricesTheRecordedCosts: false,thePlanPricesNoWholeSuitePrediction: false.rejectedMutantsAreCompileErrorsChargedTheCheckTime: false,ignoredRecordsZeroCosts: false,noTestVerdictsPublishTheirMeasuredCostOnTheStream: false.everyUncoveredMutantIsPricedAtZero: false.check:ciexit 0, and the changeset gate passes.Effect on main
Main's plan job runs the published CLI, so this takes effect only after a release. The first plan after that release still reads cost records written by 17.0.2, where these costs are null. The run after that is the first one planned on these recorded costs.