Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
24 changes: 24 additions & 0 deletions docs/ml-worker-evaluation.md
Original file line number Diff line number Diff line change
Expand Up @@ -102,6 +102,30 @@ Wiring Tier 2 — the BETH rail and disposition-corpus calibration — is tracke
[#1974](https://github.com/Xore/APIARY/issues/1974). It is not tracked by this
paragraph.

### 2026-09-25 — disposition export/census landed; Tier 2 calibration remains gated

`ml-worker/benchmarks/disposition_corpus.py` now exports the closed
operator-disposition population and writes a hashed census outside the
repository. The report carries a canonical-content SHA-256 (the digest field
is excluded from its own hash), so the saved artifact can be independently
verified. It is read-only: open and legacy unlabelled alerts remain in the
full alert denominator but are excluded from the labelled snapshot, and the
command never updates Elasticsearch or synthesises a verdict.

The census is a gate, not a calibration result. A zero-label or single-class
labelled subset is reported as non-calibratable. Even after labels accumulate,
the corpus is **precision-only**: `write_anomaly()` returns before persistence
below `ML_ALERT_THRESHOLD`, so it can describe precision within alerts and
within-alert calibration, but it can never measure deployment recall or
ordinary below-threshold calibration. Any Tier 2 report that consumes this
corpus must repeat that limitation and must not treat unlabelled or absent
events as negatives.

The decision record should receive a result only after a concrete snapshot is
attached to the run and its class/model-state diversity is sufficient. The
live census is therefore not copied into this file as a durable result; rerun
the command against the deployment and retain the generated report and hash.

## Findings carried in from the research phase

Recorded so they are not rediscovered, each with the reason it matters here.
Expand Down
55 changes: 55 additions & 0 deletions ml-worker/benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -8,6 +8,61 @@ governance model. The decision record is

**This produces the ruler, not a detector.**

## Operator-disposition corpus

`disposition_corpus.py` is a separate read-only export/census for the
`ml-anomalies` index. It does not train, calibrate, mutate, re-label, or
auto-disposition an alert.

```bash
python3 ml-worker/benchmarks/disposition_corpus.py \
--es-host "$ES_HOST" \
--output "$HOME/ml-worker-qualification/dispositions.ndjson" \
--report "$HOME/ml-worker-qualification/dispositions-census.json"
```

The endpoint may come from `--es-host`, `ES_HOST`, or `ELASTICSEARCH_URL`; no
deployment endpoint is hardcoded. Credentials may come from `ES_API_KEY` or
`ELASTICSEARCH_API_KEY`, or from the `ES_USERNAME`/`ES_PASSWORD` pair. The
tool issues search and PIT lifecycle requests only. It first opens one PIT,
uses it for the all-alert status/time census and the closed-label export, then
closes it. The exported rows and hashed report are written atomically outside
the repository. The report records a canonical-content SHA-256 (computed
without the digest field) and prints the same digest after the census.

Open and legacy documents without a disposition are counted in the full alert
denominator but excluded from the labelled export. Every closed row contains
the production anomaly timestamp, score, detector scores/contributors,
threshold, model state, sensor, disposition, actor, and disposal time.
`disposition_reason` is free text and is replaced with `[REDACTED]` by
default; the explicit `--include-reason` override is available only when an
operator has a separate review and storage boundary in place.
The census reports:

- total alerts and every status count (including `<missing>`);
- closed-label count, class balance, and labelled/all-alert denominator;
- all-alert and labelled time ranges;
- labelled sensor, scoring model, threshold, distinct-model-state, and
missing-field counts, including per-status breakdowns;
- the SHA-256 of the immutable NDJSON snapshot and the canonical-content
SHA-256 of the JSON report; and
- an explicit zero/single-class calibration gate.

The output begins with a non-negotiable warning:

> **PRECISION-ONLY:** only above-threshold alerts are persisted, so this
> corpus cannot measure deployment recall or calibrate ordinary
> below-threshold traffic.

An unlabelled alert is not a negative, and an absent below-threshold event is
not evidence that it was correctly rejected. #2986's calibration work may use
the snapshot only after the census shows enough labels and model-state
diversity; this command never performs that calibration.

For an offline verification, pass a complete local JSON array of ES-shaped
hits to `--fixture`. The fixture includes open/legacy rows as well as closed
rows so the two denominators can be tested without a network.

## Safety properties

- Fixtures are the per-sensor documents from `ml-worker/tests/fixtures.py` —
Expand Down
Loading
Loading