Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
11 changes: 11 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,16 @@
# Changelog

## Unreleased

A process that crashes now fails the test:

playground_test.background_job_test
Background process crashed at src/playground.gleam:26
background job crashed: queue is full

There is also a process leak detector. This doesn't fail the
run, and `--keep-leaked-processes` leaves them running instead.

## v1.2.0

Watch mode added for the JavaScript target, supporting both Node and Deno
Expand Down
24 changes: 21 additions & 3 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -52,6 +52,8 @@ gleam test -- --format=teamcity
gleam test -- --junit=report.xml
gleam test -- --timeout=1000 # Erlang target only
gleam test -- --parallel # Erlang target only
gleam test -- --show-crash-reports # Erlang target only, see below
gleam test -- --keep-leaked-processes # Erlang target only, see below
gleam test -- --color=never # console colour: auto | always | never
gleam run -m vouch -- watch # rerun the suite on file change
```
Expand Down Expand Up @@ -83,13 +85,29 @@ On Deno, watch mode also needs `allow_run = ["gleam"]` to spawn the inner runs (
| Outcome | Meaning | Exit code contribution |
| --- | --- | --- |
| pass | ran without panic | 0 |
| fail | assert/panic/crash/timeout | 1 |
| todo | hit `todo` in code under test | 1 — unimplemented is still not done |
| fail | assert/panic/crash/timeout, or a process the test started crashed | 1 |
| todo | hit `todo` in code under test (directly or in a process it started) | 1 — unimplemented is still not done |
| skip | `todo` within test function | 0 |

## Crash reports

A process that crashes fails the test:

playground_test.background_job_test
Background process crashed at src/playground.gleam:26
background job crashed: queue is full


## Leaked processes

There is a process leak detector. This does not fail the run.
You can keep leaked processes alive by passing `--keep-leaked-processes` and they will be left running.
They will still be reported in the summary.

## Target differences

Erlang target supports `--parallel` and `--timeout=n`
Erlang target supports `--parallel`, `--timeout=n`, `--show-crash-reports` and
`--keep-leaked-processes`

Deno needs permissions in your project's `gleam.toml`.
You need `allow_read` for discovery and source quoting.
Expand Down
7 changes: 4 additions & 3 deletions docs/chapters/architecture.typ
Original file line number Diff line number Diff line change
Expand Up @@ -63,9 +63,10 @@ The FFI contract, in full:
[Terminal facts (TTY, env vars)],
[`io:columns/0`, `os:getenv/1`],
[`isTTY` / `Deno.stdout.isTerminal`, env lookup],
[Route BEAM diagnostics to stderr],
[re-add the default logger handler],
[not needed],
[Catch a crash behind a test, kill what it leaked / keep BEAM reports off
the streams],
[`erlang:trace` a per-test process tree; capture logger handler],
[not needed (no processes)],
[Halt with exit code],
[`erlang:halt/1`],
[`process.exit` / `Deno.exit`],
Expand Down
64 changes: 53 additions & 11 deletions docs/chapters/execution.typ
Original file line number Diff line number Diff line change
Expand Up @@ -23,17 +23,59 @@ fidelity for ordinary failures. The runner waits on three possibilities:
killed and reported as a timeout failure. A hung test costs its timeout,
not the whole run.

BEAM diagnostics (crash reports for processes that tests spawned) are routed
to stderr at run start by re-adding the default logger handler with
`type: standard_error` — they stay visible in a terminal but can never
corrupt piped stdout. (`logger:update_handler_config` silently ignores a
runtime type change; remove-and-re-add is required. The reports are
asynchronous and are often lost to `erlang:halt` entirely.)

Known limitation, accepted for v1: a test that spawns unlinked long-lived
processes, registers global names, or mutates shared ETS tables can still
leak state between tests. Process-per-test is the 90% solution; full
sandboxing is out of scope.
A process that dies *behind* a test — an unlinked worker, a fire-and-forget
job nothing is linked to or monitoring — would leave the test passing, its
death known only to the BEAM's crash report. vouch traces each test's
process tree: `erlang:trace(TestPid, true, [procs, set_on_spawn, {tracer,
Collector}])` gives every process the test starts, transitively, a per-test
collector that receives a trace message for each abnormal exit. This is the
pass/fail signal, and it is race-free where the crash report is not: a trace
exit is delivered before the test can report its result, and
`erlang:trace_delivered/1` flushes any still in transit before the collector
is read. After each test the runner folds the collected crash reasons into
the outcome — a passing test whose worker crashed fails
(`BackgroundCrashDetail`), or is a todo if the worker died of a `todo`, the
test's own failure winning a tie. A crash a collector sees after its test
was claimed (a worker outliving its test) is recorded as unattributed;
those are reported at the end and fail the run. Nesting — vouch's own suite
runs tests that run tests — is handled by clearing the inherited trace
before setting each test's own, since a process may have only one tracer.

The BEAM's own crash reports (the emulator's "Error in process", proc_lib
crash reports, gen_\* terminate and supervisor child reports) are separately
diverted from the default logger handler into an ETS table by a capture
handler whose filter admits only crash-shaped events, purely to keep them
off the output streams and hand their raw text to `--show-crash-reports`;
they no longer drive any outcome. The default handler is re-added with
`type: standard_error` for everything else routed through logger, so it
stays visible but can never corrupt piped stdout.
(`logger:update_handler_config` silently ignores a runtime type change;
remove-and-re-add is required.)

The same trace closes the process half of that leak. A test's tree is not
torn down when the test ends — nothing in the BEAM does that, and a `normal`
exit is ignored by a non-trapping link — so a worker a test started outlives
it by default. At the claim the collector already knows which descendants
are still alive, from the `spawn` and `spawned` messages `procs` delivers
alongside the exits. Those are killed in spawn order, so an ancestor dies
before anything it would restart, and recorded as leaks for the end of the
run. Order matters throughout: the flush comes first, so every death the
test caused on its own is accounted for before the runner causes any; then
the collector is switched to discarding, so vouch's own kills are not
reported back as the test's crashes (the trap `killed` sets, since it is not
a clean exit); then the kill; then a second round only if the first killed
anything, in case a survivor restarted a child. With nothing left alive the
collector stops, which also retires the one lingering process per test.
`--keep-leaked-processes` reports without killing, for a suite that shares a
process across tests deliberately or where a leaked process is linked
outside its tree, since a kill propagates along that link.

Known limitation, accepted for v1: a test that registers global names or
mutates shared ETS tables can still leak state between tests, and a process
started through an already-running supervisor or `application:start` is
spawned outside the test's tree, so it is neither traced, reported, nor
killed. Process-per-test plus killing what a test spawned is the 90%
solution; full sandboxing is out of scope.

== JavaScript target: sequential in-process

Expand Down
7 changes: 5 additions & 2 deletions docs/chapters/test-model.typ
Original file line number Diff line number Diff line change
Expand Up @@ -82,9 +82,12 @@ pub type TestOutcome {

pub type FailureDetail {
PanicDetail(GleamPanic)
UnknownDetail(raw: Dynamic)
UnknownDetail(raw: Dynamic, site: Option(CrashSite))
TimeoutDetail(after_ms: Int)
ExitDetail(raw: Dynamic)
ExitDetail(raw: Dynamic, site: Option(CrashSite))
// A process the test started died behind its back; `cause` is the
// PanicDetail or UnknownDetail that killed it (Erlang target only).
BackgroundCrashDetail(cause: FailureDetail)
}
```

Expand Down
30 changes: 30 additions & 0 deletions examples/playground/src/playground.gleam
Original file line number Diff line number Diff line change
Expand Up @@ -16,3 +16,33 @@ pub fn rate_limit(_requests: Int) -> Bool {
@external(erlang, "playground_ffi", "sleep")
@external(javascript, "./playground_ffi.mjs", "sleep")
pub fn sleep(ms: Int) -> Nil

/// A fire-and-forget worker with a bug: the caller gets Nil back and
/// carries on, while the unlinked process it started dies. Nothing is
/// linked to or monitoring the worker, so the calling test would pass —
/// the BEAM's crash report is the only trace of the death, and vouch
/// charges it to the test as a "Background process crashed" failure.
pub fn start_background_job() -> Nil {
spawn(fn() { panic as "background job crashed: queue is full" })
}

/// A worker with no owner: the caller returns while it is still running.
/// Nothing crashes — this is the leak, not the crash — and on the BEAM vouch
/// kills it when the test that started it ends, so it cannot run on into a
/// later test, and reports it after the summary.
pub fn start_long_worker() -> Nil {
spawn(fn() { sleep(60_000) })
}

@external(erlang, "playground_ffi", "spawn")
@external(javascript, "./playground_ffi.mjs", "spawn")
fn spawn(job: fn() -> Nil) -> Nil

/// Stands in for a dependency whose .beam has gone stale or missing: no
/// config_parser module exists, so on the BEAM the call raises undef and
/// the report names config_parser:parse/1 from the stacktrace. On
/// JavaScript the FFI throws a plain TypeError instead — a raw crash with
/// no site, which is all a JS stacktrace offers.
@external(erlang, "config_parser", "parse")
@external(javascript, "./playground_ffi.mjs", "parse")
pub fn parse_config(path: String) -> String
7 changes: 6 additions & 1 deletion examples/playground/src/playground_ffi.erl
Original file line number Diff line number Diff line change
@@ -1,6 +1,11 @@
-module(playground_ffi).
-export([sleep/1]).
-export([sleep/1, spawn/1]).

sleep(Ms) ->
timer:sleep(Ms),
nil.

%% Unlinked on purpose: the crash must not take the caller down.
spawn(Job) ->
erlang:spawn(Job),
nil.
12 changes: 12 additions & 0 deletions examples/playground/src/playground_ffi.mjs
Original file line number Diff line number Diff line change
Expand Up @@ -5,3 +5,15 @@ export function sleep(ms) {
// spin
}
}

// JavaScript has no process to leak: a thrown error in a deferred callback
// would take the whole runtime down, not a sibling. The job is simply not
// run, so background_job_test passes on JavaScript and fails on the BEAM,
// where the worker really does crash — a documented target difference.
export function spawn(_job) {}

// The JavaScript half of parse_config: a raw runtime error, standing in for
// a broken FFI dependency. On Erlang the module itself is missing instead.
export function parse(path) {
throw new TypeError(`config_parser is not loaded (parsing ${path})`);
}
42 changes: 41 additions & 1 deletion examples/playground/test/playground_test.gleam
Original file line number Diff line number Diff line change
@@ -1,10 +1,19 @@
//// One test per outcome flavour, so a plain run shows vouch's full range:
////
//// gleam test - 2 pass, 3 fail, 2 todo, 1 skip
//// gleam test - 3 pass, 5 fail, 2 todo, 1 skip
//// (JavaScript: 4 pass, 4 fail — no
//// process crashes behind
//// background_job_test there)
//// gleam test -- --timeout=100 - slow_test also fails as a timeout
//// gleam test -- --filter=slow - just the slow test
//// gleam test -- --format=json - the same as a JSONL stream
//// gleam test -- --junit=report.xml - plus a CI-ready XML report
//// gleam test -- --show-crash-reports
//// - plus the full BEAM crash report
//// from background_job_test's worker
//// gleam test -- --keep-leaked-processes
//// - leave leaky_worker_test's worker
//// running instead of killing it

import playground
import vouch
Expand All @@ -17,6 +26,29 @@ pub fn fast_pass_test() {
assert playground.add(1, 2) == 3
}

/// The test body passes — but the unlinked worker it starts crashes behind
/// its back, and nothing is linked to or monitoring it. The BEAM's crash
/// report is the only trace; vouch charges it to this test, which fails as
/// "Background process crashed". `--show-crash-reports` prints the raw
/// report too. (On JavaScript there is no worker, so this passes.)
pub fn background_job_test() {
playground.start_background_job()
// Give the job a moment, as a test of fire-and-forget code typically
// does; it also makes the worker's death land before the test ends, so
// the outcome is the same on a one-scheduler CI box.
playground.sleep(20)
assert playground.add(0, 0) == 0
}

/// Starts a worker and returns while it is still running. The test passes —
/// leaking a process is not a failure — but on the BEAM vouch kills the
/// worker at the end of this test and lists it after the summary, so it
/// cannot pollute a later test. (On JavaScript there is no worker at all.)
pub fn leaky_worker_test() {
playground.start_long_worker()
assert playground.add(2, 2) == 4
}

/// Passes with the default 5000ms timeout; fails with --timeout=100.
pub fn slow_test() {
playground.sleep(300)
Expand All @@ -35,6 +67,14 @@ pub fn failing_predicate_test() {
assert playground.within_budget(spend)
}

/// A crash that is not a Gleam panic: the code under test calls into a
/// missing module — the stale-.beam problem. The report names what was
/// being called ("Crashed: Undef calling config_parser:parse/1") instead of
/// an information-free "Crashed: Undef".
pub fn crashing_config_test() {
playground.parse_config("gleam.toml")
}

/// A failing pattern match: renders the unmatched value.
pub fn failing_let_assert_test() {
let assert Ok(n) = Error("the port was closed")
Expand Down
19 changes: 18 additions & 1 deletion src/vouch.gleam
Original file line number Diff line number Diff line change
Expand Up @@ -54,6 +54,8 @@ fn run_tests(args: List(String)) -> Nil {
cfg.filter,
cfg.timeout_ms,
cfg.parallel,
cfg.show_crash_reports,
cfg.kill_leaked_processes,
)
config.Console, Some(path) ->
runner.run(
Expand All @@ -64,29 +66,44 @@ fn run_tests(args: List(String)) -> Nil {
cfg.filter,
cfg.timeout_ms,
cfg.parallel,
cfg.show_crash_reports,
cfg.kill_leaked_processes,
)
config.Json, None ->
runner.run(jsonl.reporter(), cfg.filter, cfg.timeout_ms, cfg.parallel)
runner.run(
jsonl.reporter(),
cfg.filter,
cfg.timeout_ms,
cfg.parallel,
cfg.show_crash_reports,
cfg.kill_leaked_processes,
)
config.Json, Some(path) ->
runner.run(
reporter.pair(jsonl.reporter(), junit.reporter(path)),
cfg.filter,
cfg.timeout_ms,
cfg.parallel,
cfg.show_crash_reports,
cfg.kill_leaked_processes,
)
config.TeamCity, None ->
runner.run(
teamcity.reporter(),
cfg.filter,
cfg.timeout_ms,
cfg.parallel,
cfg.show_crash_reports,
cfg.kill_leaked_processes,
)
config.TeamCity, Some(path) ->
runner.run(
reporter.pair(teamcity.reporter(), junit.reporter(path)),
cfg.filter,
cfg.timeout_ms,
cfg.parallel,
cfg.show_crash_reports,
cfg.kill_leaked_processes,
)
}
}
Expand Down
25 changes: 25 additions & 0 deletions src/vouch/internal/config.gleam
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,18 @@ pub type Config {
timeout_ms: Int,
color: ColorChoice,
parallel: Parallelism,
/// Reprint, in full, the BEAM crash reports captured during the run
/// (processes the tests started dying) after the summary. Off by
/// default: a death under a test already shows in the test's outcome,
/// as its own failure or a "Background process crashed" one.
show_crash_reports: Bool,
/// Kill the processes a test leaves running when it ends, on the Erlang
/// target. On by default: process-per-test is only a real boundary if
/// nothing survives the test that started it. Turn it off when a suite
/// deliberately shares a process across tests, or when a leaked process
/// is linked to something outside the test's tree — killing it would
/// propagate.
kill_leaked_processes: Bool,
)
}

Expand All @@ -49,6 +61,8 @@ pub fn from_args(args: List(String)) -> Result(Config, String) {
timeout_ms: default_timeout_ms,
color: Auto,
parallel: Sequential,
show_crash_reports: False,
kill_leaked_processes: True,
),
)
}
Expand Down Expand Up @@ -90,6 +104,10 @@ fn parse(args: List(String), config: Config) -> Result(Config, String) {
"--parallel expects a worker count of at least 1, got: " <> value,
))
}
["--show-crash-reports", ..rest] ->
parse(rest, Config(..config, show_crash_reports: True))
["--keep-leaked-processes", ..rest] ->
parse(rest, Config(..config, kill_leaked_processes: False))
[arg, ..] ->
case string.starts_with(arg, "-") {
True -> Error(usage("unknown option: " <> arg))
Expand Down Expand Up @@ -131,6 +149,13 @@ fn usage(problem: String) -> String {
<> " --timeout=ms per-test timeout on the Erlang target (default 5000)\n"
<> " --parallel[=n] run tests concurrently on the Erlang target with n\n"
<> " workers (default: one per scheduler)\n"
<> " --show-crash-reports\n"
<> " also print the full BEAM crash reports of processes\n"
<> " that died during the run, after the summary (Erlang)\n"
<> " --keep-leaked-processes\n"
<> " do not kill the processes a test leaves running when\n"
<> " it ends; they are killed and reported by default\n"
<> " (Erlang)\n"
<> " --color=mode console colour: auto (default), always, never;\n"
<> " auto respects NO_COLOR and non-TTY output\n"
<> "\n"
Expand Down
Loading
Loading