Skip to content

Crash site reporting - #13

Merged
richardcocks merged 8 commits into
mainfrom
crash-site-reporting-reland
Aug 22, 2026
Merged

richardcocks merged 8 commits into
mainfrom
crash-site-reporting-reland

Conversation

@richardcocks

Copy link
Copy Markdown
Owner

Attempt to fix the issue with crashes that happen inside erlang processes.

Currently a crash can be reported in stderr / the error log, but the exit code is 0, and a pass is reported.

This attempts to fix this.

richardcocks and others added 8 commits August 21, 2026 23:26
background_job_test starts an unlinked process that panics and then
passes — the case lpil asked about: with crash reports swallowed, how do
you know? The e2e proves the answer end to end: a default run shows
nothing of the crash, and --show-crash-reports prints the BEAM report
after the summary. run_command_merged folds stderr into the captured
stream for that one test; the stdout-only harness keeps proving the
reports never reach stdout.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
A process a test starts that dies without taking the test down (an
unlinked worker, a fire-and-forget job) left the test passing and the
run green; the BEAM's crash report was the only trace, and the capture
commits on this branch had hidden even that. gleeunit behaves the same
way (pass, exit 0, report printed mid-run).

Each test process now runs under its own group leader, an io-forwarding
proxy that every process it starts inherits and every crash report
records, so a captured report is charged to exactly the test whose
process died, under --parallel too. After each test the runner claims its
reports and folds them into the outcome: a passing test with a crash
fails as BackgroundCrashDetail (or is a Todo when the process died of a
todo); a test's own failure wins. Reports nobody claims - arrived after
their test finished, or from outside any test - are printed at the end
and fail the run. Only crash-shaped logger events are captured; other
logger output still goes to stderr. --show-crash-reports now additionally
prints the raw BEAM reports.

Emulator "Error in process" events are handled in logger_proxy, not the
logger server, so that is the process the pre-claim flush round-trips.

The playground's background_job_test now fails on Erlang (2 pass, 5
fail) and still passes on JavaScript, where nothing is spawned.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
The previous commit charged a background process's death to its test by
capturing the BEAM's crash report and matching the dying process's group
leader. That report is asynchronous: measured against the monitor DOWN a
test observes, it is only ~81% enqueued in time even after flushing
logger_proxy, so a real crash could still slip through as a pass - the very
bug this is meant to close.

Pass/fail now comes from tracing instead. run_test traces each test's
process with erlang:trace(procs, set_on_spawn), so every process it starts,
transitively, reports its exit to a per-test collector. A trace exit is
delivered before the test can send its result, and trace_delivered/1 flushes
any still in transit before the collector is read, so attribution is exact
and never a race - under --parallel as much as sequentially. The collector
keeps tracing after the test is claimed, so a worker that outlives its test
and crashes later is recorded as unattributed and fails the run. Nested runs
(vouch's own suite runs tests that run tests) clear the trace inherited from
the outer test before setting their own, since a process may have one tracer.

The logger capture is kept, demoted to display only: it holds the raw BEAM
reports off the output streams and hands them to --show-crash-reports, and
no longer drives any outcome. CrashReport carries just the exit reason now;
the report text is gone from the outcome path.

Cost is ~100us per test (trace setup plus the trace_delivered barrier),
doubling a 2000-empty-test run from 86ms to ~200ms - negligible for a real
suite and still far faster than gleeunit, and the deterministic guarantee is
worth it for a test runner.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Process-per-test isolated the test function, not the tree it spawned.
Nothing in the BEAM tears that tree down when a test ends, and a `normal`
exit is ignored by a non-trapping link, so a worker a test started ran on
into later tests and could crash long after its own test was reported —
which is what the unattributed path existed to catch after the fact.

The per-test trace already had the information to prevent it: `procs`
delivers `spawn` and `spawned` alongside the exits, so the collector can
keep a live set as cheaply as it discards those messages today. At the
claim it hands that set back; the runner kills it in spawn order, so an
ancestor is gone before anything it would restart, and records each one as
a leak named by what started it. A Gleam closure reaches the trace as
erlang:apply over a fun, so origins go through fun_info to recover the
function the closure was written in.

Ordering carries the correctness. The trace flush comes first, so every
death the test caused on its own is accounted for before the runner causes
any. Then the collector switches to discarding, so vouch's own kills are
not reported back as the test's crashes — the trap `killed` sets, since it
is not a clean exit. Then the kill, then a second round only if the first
killed anything. With nothing left alive the collector stops, retiring the
one lingering process per test as well.

Leaks do not fail the run: the process was cleaned up before it could reach
another test, and the report is what stops that being silent.
--keep-leaked-processes reports without killing, for a suite that shares a
process across tests deliberately, or where a leaked process is linked to
something outside its tree — a kill propagates along that link. Processes
started through an already-running supervisor or application:start are
spawned outside the test's tree and are untouched either way.

Also fixes two bugs found reviewing the branch:

- forward_crashes/2 was not exported. erlang:hibernate/3 resumes through an
  exported call, so hibernating into a local function succeeds and fails
  with undef on wake-up: every late crash more than one idle period after
  the claim was silently lost, and the collector died. The existing test
  passed only because its worker crashed inside the 200ms window.

- A background `todo` could mark a passing test skipped. Gleam attributes a
  closure's panic to the enclosing named function, so a worker written
  inline in a test body carries the test's own module and function — the
  pair classify_panic reads as "this test body is pending". Skipped
  contributes 0 to the exit code, so the crash disappeared entirely.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
kill_leaked ran on every test, so a test that cleaned up after itself
still paid two collector round-trips and a second node-wide
trace_delivered/1. Measured on a 2000-trivial-test suite that leaks
nothing, that was 215ms -> 283ms, about 35us per test spent proving there
was nothing to do.

The claim already computes the live set, so both sides can gate on it: the
collector stops immediately when nothing outlived the test, and
claim_crashes skips the kill rounds. Nothing can appear afterwards — the
test sends its result as its last act, so with no live descendant there is
no one left to spawn. Both sides ask through live_pids/leaks_of, which
filter alike, so they cannot disagree about emptiness.

Same suite after: 184ms sequential, 196ms parallel, against 198ms and
200ms for the branch without any of the leak work — within noise. The
sweep and kill only cost anything when there is something to sweep.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@richardcocks
richardcocks merged commit 406c3bd into main Aug 22, 2026
3 checks passed
@richardcocks
richardcocks deleted the crash-site-reporting-reland branch August 22, 2026 13:44
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant