Skip to content

feat(repo): agent-ready failure diagnostics (stack #146-#156) - #159

Closed
ryanleecode wants to merge 10 commits into
mainfrom
feat/agent-ready-failure-diagnostics
Closed

ryanleecode wants to merge 10 commits into
mainfrom
feat/agent-ready-failure-diagnostics

Conversation

@ryanleecode

@ryanleecode ryanleecode commented Oct 2, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Lands the agent-ready failure diagnostics stack (plan docs/plans/2026-10-01-2306-feat-agent-ready-failure-diagnostics-plan.md) on main. The eight stacked PRs were merged into this feature branch; this PR brings the branch into main with a merge commit.

Merge from main

main gained #153 and #154 after the stack was cut, so the branch carries one merge commit from main with these conflicts resolved:

  • worker-client.blueprint.ts: main's boot timeout (timeoutOrElse with WORKER_BOOT_TIMEOUT, no connect retry), plus the stack's workerKind on WorkerBootTimeoutError.
  • substituted-worker.fixture.ts: bootPingWorker, which fix(repo): fail the run instead of hanging when a plugin worker never accepts its socket #154 moved here, now passes workerKind: 'testRunner'.
  • dry-run-failure.integration.test.ts: main's runWithRunnerPlugin and deaf-runner scenario, written against the stack's RunFailure. The stage is read from evidence.stage, and there is no log capture because the stack asserts the failure record instead of logs.
  • worker-launcher.integration.test.ts: main's version, since the helpers moved into the fixture.

Testing

Follow-ups

  • mutation.yml uses cancel-in-progress: true keyed on github.ref, which was already true on main. A second merge to main cancels the full main run that the PR lane now depends on for its survivor baseline and timings. Main should probably queue rather than cancel.

View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

systemfsoftware-maker and others added 10 commits October 2, 2026 14:15
…operty (#146)

* docs(repo): plan agent-ready failure diagnostics

Verdict-Semantics: unchanged

* fix(repo): draw every deciding timeout class in the timeout-fields property

Verdict-Semantics: unchanged
* feat(repo): end every failed run in one typed failure record

A failed run now emits a FailureRecord (catalog code, stage, structured evidence, cause chain, next action, replay capsule, trace id) on stderr, as the schema-3.0 terminal stream event, and in reports/mutation/failure.json. StageError and PrepareError are replaced by RunFailure, exit codes come from the failure catalog, and failed baseline tests carry project-relative file, line and stack from the Vitest runner

Verdict-Semantics: unchanged

* chore(repo): guard the failure catalog as a versioned contract document

Verdict-Semantics: unchanged

* test(repo): assert the frozen-runner dry run through its failure record

Trunk's frozen-runner scenario read the removed StageError stage and log capture; it now reads the RunFailure evidence stage

Verdict-Semantics: unchanged

* fix(repo): declare stryker.failure.code optional on the run span

A successful run carries no failure code, so the e2e trace contract refused every passing run span with a missing key

Verdict-Semantics: unchanged

* refactor(repo): declare the optional span attribute schema in its schema module

Verdict-Semantics: unchanged

* test(repo): pin summary annotations to the record code and assert the failing-run record once

The summary-annotation law accepted any single annotation, so a draw with no located evidence let a constant renderer pass and the fork refused the file as vacuous on macOS. The failing-run journey checked the terminal record in two consecutive steps, which the fork refuses as a second check on the same state

Verdict-Semantics: unchanged

* test(repo): show the run's failure record when a framework-run scenario expects success

A CI rerun of the claimed-type scenario stopped after file discovery and the Then reported only that the incremental state was not JSON, so the failure that stopped the run was lost

Verdict-Semantics: unchanged
…properties (#148)

mutant-set-policy and interpret-vitest-mutant-run now declare cover classes for every branch that decides their subject, and their generators reach each class constructively under the mutation-worker budget

Verdict-Semantics: unchanged
* feat(repo): give every failure record a one-step reproduction

The Vitest runner reports a reproduce argv per failed test, and a pure workflow decides each record's capsule: the runner command under the mutation-worker budget, the dry-run fallback, the --mutant rerun for new survivors, or DoesNotReplay with its stand-in

Verdict-Semantics: unchanged

* test(repo): expect the runner's vitest replay in the failing-run capsule

The capsule layer replays a failed baseline test through the runner's vitest run <file> -t <test> argv; the e2e journey still expected the stryker run fallback

Verdict-Semantics: unchanged
)

MSP answers engine failures with the record in error.data and keeps serving; MCP refuses reruns with the record and gains get_failure; a failed run with the sarif reporter writes a failed invocation per record; trace tests pin stryker.failure.code and the record's traceId; the run log drops the orphan possible-causes block and the bundled WASI warning

Verdict-Semantics: unchanged
* ci(repo): render mutation job failures from the failure record

The mutation job decodes the stream's terminal record, renders the summary and annotations through the contract renderers, builds RecordMissing, JobTimedOut and BinaryMissing records itself, reports a run that evaluated no mutants as exactly that, and no longer calls a failed run an infrastructure failure

Verdict-Semantics: unchanged

* ci(repo): keep a failed outcome visible when a run reused every verdict

The evaluated-none label replaced the stryker outcome, so a verdict-failed run served entirely from cache read as evaluated no mutants instead of failure

Verdict-Semantics: unchanged

* docs(solutions): a failure record is the only rendering of a failure
…155)

A preflight job dry-runs every planned package and publishes the coverage that shards reuse through incrementalSources. A dry-run-only run now writes that coverage to its incremental file, keeping any verdicts there, and the run-inputs digest no longer covers dryRunOnly. Each job exports its traces, a failed shard's parts are kept 30 days, the report job uploads one merged SARIF run and publishes main's survivor baseline

Verdict-Semantics: unchanged
…rs (#156)

* ci(repo): run the mutation lane on pull requests and gate new survivors

A pull request runs the same preflight and sharded incremental run against main's restored verdict cache, saves no cache and records no timings, then gates its merged report against the survivor baseline main published. A run that evaluated no mutants says so in a notice instead of reading as a pass

Verdict-Semantics: unchanged

* ci(repo): plan mutation shards from the full mutant set's cost

This stack's own pull request ran 7465 invalidated stryker-js mutants in one job planned at 9m43s and hit the 1800s cap: the timing record held main's reuse-shrunk duration. Each entry now carries the reuse line's ran and reused counts, the record scales a run to the full mutant set and keeps its previous value when nothing ran, and shards download the preflight coverage only when the preflight published some

Verdict-Semantics: unchanged
Brings in #153 (cleanTempDir follows the run outcome) and #154 (worker boot timeout). Conflicts resolved:
- worker-client.blueprint.ts: keep main's boot timeout (timeoutOrElse + WORKER_BOOT_TIMEOUT, no connect retry) and the stack's workerKind on WorkerBootTimeoutError.
- substituted-worker.fixture.ts: bootPingWorker (moved here by #154) passes workerKind: 'testRunner'.
- dry-run-failure.integration.test.ts: main's runWithRunnerPlugin and deaf-runner scenario against the stack's RunFailure (stage read from evidence.stage; no log capture, since the stack asserts the failure record instead of logs).
- worker-launcher.integration.test.ts: take main's version (helpers moved to the fixture).
@ryanleecode ryanleecode closed this Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants