feat(repo): agent-ready failure diagnostics (stack #146-#156) - #159
Closed
ryanleecode wants to merge 10 commits into
Closed
ryanleecode wants to merge 10 commits into
ryanleecode wants to merge 10 commits into
Conversation
…operty (#146) * docs(repo): plan agent-ready failure diagnostics Verdict-Semantics: unchanged * fix(repo): draw every deciding timeout class in the timeout-fields property Verdict-Semantics: unchanged
* feat(repo): end every failed run in one typed failure record A failed run now emits a FailureRecord (catalog code, stage, structured evidence, cause chain, next action, replay capsule, trace id) on stderr, as the schema-3.0 terminal stream event, and in reports/mutation/failure.json. StageError and PrepareError are replaced by RunFailure, exit codes come from the failure catalog, and failed baseline tests carry project-relative file, line and stack from the Vitest runner Verdict-Semantics: unchanged * chore(repo): guard the failure catalog as a versioned contract document Verdict-Semantics: unchanged * test(repo): assert the frozen-runner dry run through its failure record Trunk's frozen-runner scenario read the removed StageError stage and log capture; it now reads the RunFailure evidence stage Verdict-Semantics: unchanged * fix(repo): declare stryker.failure.code optional on the run span A successful run carries no failure code, so the e2e trace contract refused every passing run span with a missing key Verdict-Semantics: unchanged * refactor(repo): declare the optional span attribute schema in its schema module Verdict-Semantics: unchanged * test(repo): pin summary annotations to the record code and assert the failing-run record once The summary-annotation law accepted any single annotation, so a draw with no located evidence let a constant renderer pass and the fork refused the file as vacuous on macOS. The failing-run journey checked the terminal record in two consecutive steps, which the fork refuses as a second check on the same state Verdict-Semantics: unchanged * test(repo): show the run's failure record when a framework-run scenario expects success A CI rerun of the claimed-type scenario stopped after file discovery and the Then reported only that the incremental state was not JSON, so the failure that stopped the run was lost Verdict-Semantics: unchanged
…properties (#148) mutant-set-policy and interpret-vitest-mutant-run now declare cover classes for every branch that decides their subject, and their generators reach each class constructively under the mutation-worker budget Verdict-Semantics: unchanged
* feat(repo): give every failure record a one-step reproduction The Vitest runner reports a reproduce argv per failed test, and a pure workflow decides each record's capsule: the runner command under the mutation-worker budget, the dry-run fallback, the --mutant rerun for new survivors, or DoesNotReplay with its stand-in Verdict-Semantics: unchanged * test(repo): expect the runner's vitest replay in the failing-run capsule The capsule layer replays a failed baseline test through the runner's vitest run <file> -t <test> argv; the e2e journey still expected the stryker run fallback Verdict-Semantics: unchanged
) MSP answers engine failures with the record in error.data and keeps serving; MCP refuses reruns with the record and gains get_failure; a failed run with the sarif reporter writes a failed invocation per record; trace tests pin stryker.failure.code and the record's traceId; the run log drops the orphan possible-causes block and the bundled WASI warning Verdict-Semantics: unchanged
* ci(repo): render mutation job failures from the failure record The mutation job decodes the stream's terminal record, renders the summary and annotations through the contract renderers, builds RecordMissing, JobTimedOut and BinaryMissing records itself, reports a run that evaluated no mutants as exactly that, and no longer calls a failed run an infrastructure failure Verdict-Semantics: unchanged * ci(repo): keep a failed outcome visible when a run reused every verdict The evaluated-none label replaced the stryker outcome, so a verdict-failed run served entirely from cache read as evaluated no mutants instead of failure Verdict-Semantics: unchanged * docs(solutions): a failure record is the only rendering of a failure
…155) A preflight job dry-runs every planned package and publishes the coverage that shards reuse through incrementalSources. A dry-run-only run now writes that coverage to its incremental file, keeping any verdicts there, and the run-inputs digest no longer covers dryRunOnly. Each job exports its traces, a failed shard's parts are kept 30 days, the report job uploads one merged SARIF run and publishes main's survivor baseline Verdict-Semantics: unchanged
…rs (#156) * ci(repo): run the mutation lane on pull requests and gate new survivors A pull request runs the same preflight and sharded incremental run against main's restored verdict cache, saves no cache and records no timings, then gates its merged report against the survivor baseline main published. A run that evaluated no mutants says so in a notice instead of reading as a pass Verdict-Semantics: unchanged * ci(repo): plan mutation shards from the full mutant set's cost This stack's own pull request ran 7465 invalidated stryker-js mutants in one job planned at 9m43s and hit the 1800s cap: the timing record held main's reuse-shrunk duration. Each entry now carries the reuse line's ran and reused counts, the record scales a run to the full mutant set and keeps its previous value when nothing ran, and shards download the preflight coverage only when the preflight published some Verdict-Semantics: unchanged
Brings in #153 (cleanTempDir follows the run outcome) and #154 (worker boot timeout). Conflicts resolved: - worker-client.blueprint.ts: keep main's boot timeout (timeoutOrElse + WORKER_BOOT_TIMEOUT, no connect retry) and the stack's workerKind on WorkerBootTimeoutError. - substituted-worker.fixture.ts: bootPingWorker (moved here by #154) passes workerKind: 'testRunner'. - dry-run-failure.integration.test.ts: main's runWithRunnerPlugin and deaf-runner scenario against the stack's RunFailure (stage read from evidence.stage; no log capture, since the stack asserts the failure record instead of logs). - worker-launcher.integration.test.ts: take main's version (helpers moved to the fixture).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Lands the agent-ready failure diagnostics stack (plan
docs/plans/2026-10-01-2306-feat-agent-ready-failure-diagnostics-plan.md) onmain. The eight stacked PRs were merged into this feature branch; this PR brings the branch intomainwith a merge commit.Merge from main
maingained #153 and #154 after the stack was cut, so the branch carries one merge commit frommainwith these conflicts resolved:worker-client.blueprint.ts: main's boot timeout (timeoutOrElsewithWORKER_BOOT_TIMEOUT, no connect retry), plus the stack'sworkerKindonWorkerBootTimeoutError.substituted-worker.fixture.ts:bootPingWorker, which fix(repo): fail the run instead of hanging when a plugin worker never accepts its socket #154 moved here, now passesworkerKind: 'testRunner'.dry-run-failure.integration.test.ts: main'srunWithRunnerPluginand deaf-runner scenario, written against the stack'sRunFailure. The stage is read fromevidence.stage, and there is no log capture because the stack asserts the failure record instead of logs.worker-launcher.integration.test.ts: main's version, since the helpers moved into the fixture.Testing
@systemfsoftware/stryker-jsand its dependencies pass.gate:taskspasses 111 of 112 tasks. The one failure wasgit-diff.integration.test.ts, which failed only because my local global git config signs commits with gpg; with a cleanHOMEit passes 2/2. The conflict-area suites (dry-run-failure, worker-launcher, worker-boot-timeout, clean-temp-dir) pass 14/14.test/e2e/testResources/failing-fixture:BaselineTestsFailedrecord namessrc/thing.test.ts:7:21, with next actionfixCode, otherwisefixTest;reports/mutation/failure.jsonis written;reproduce:command reruns exactly the failing test.LeakedStatefailure on feat(repo): give every failure record a one-step reproduction #149 was a flake.Follow-ups
mutation.ymlusescancel-in-progress: truekeyed ongithub.ref, which was already true on main. A second merge tomaincancels the full main run that the PR lane now depends on for its survivor baseline and timings. Main should probably queue rather than cancel.Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.