Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 4 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -33,3 +33,7 @@ analysis/ghidra/revdeck/

.grit
.orchestrator/

# External BETH benchmark data. The harness never downloads or writes it.
ml-worker/benchmarks/data/
ml-worker/benchmarks/beth-data/
28 changes: 24 additions & 4 deletions docs/ml-worker-evaluation.md
Original file line number Diff line number Diff line change
Expand Up @@ -47,16 +47,36 @@ checks against a candidate over the per-sensor fixture corpus and emits a
hashed JSON report. It is the first reusable acceptance bar `ml-worker` has
had, and it makes "evaluate this candidate offline" answerable at all.

**Tier 2 (accuracy) is blocked on [#1797](https://github.com/Xore/APIARY/issues/1797).**
There is no labelled corpus. Until BETH is confirmed usable — or ruled out —
this benchmark cannot rank candidates, only reject misbehaving ones. Recorded
plainly rather than shipping a harness that looks like it ranks.
**Tier 2 (accuracy) was blocked on [#1797](https://github.com/Xore/APIARY/issues/1797).**
There was no labelled corpus at that point. The date is retained as the
historical Tier 1 landing record; the later BETH verdict below supersedes the
blocked status.

**No candidate has been promoted or rejected on this evidence yet.** The two
deployed detectors have not been run through it as a qualification; doing so is
the next step, along with the live-threshold measurement
(#1794-b) that the alert-budget metric and the promotion gate both need.

### 2026-09-25 — BETH Tier 2 sanity rail is available

`ml-worker/benchmarks/evaluate_accuracy.py beth` now provides the bounded
parallel-corpus rail. It consumes the three published BETH process CSVs as-is,
uses the authors' encoding from `BETH_Dataset_Analysis/dataset.py`, fits on
`train`, uses `val` only for the held-out diagnostic split, and scores `test`
once per declared seed without row re-splitting. It reports AUROC as the
headline, AUPRC alongside it, mean ± standard deviation over at least five
seeds, and each split's `sus` base rate beside every printed metric. A missing
local data directory is a clear CLI error; the harness never downloads data
or fabricates results.

The command is importable and its parser/validation can be tested without the
39.8 MB corpus. Runtime reports record CSV and optional archive MD5s outside the
repository. The report is evidence about the detector architecture/pipeline
only: it is **not** ground truth for deployed traffic and cannot justify a
composite-weight or `ML_ALERT_THRESHOLD` change. The disposition corpus tracked
by [#3295](https://github.com/Xore/APIARY/issues/3295) remains a separate
deployment-labelled source and is not imported here.

### 2026-09-05 — #1797 has ruled; Tier 2 is no longer blocked on it

The paragraph above is superseded. Both [#1794](https://github.com/Xore/APIARY/issues/1794)
Expand Down
71 changes: 50 additions & 21 deletions ml-worker/benchmarks/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,8 +13,9 @@ governance model. The decision record is
- Fixtures are the per-sensor documents from `ml-worker/tests/fixtures.py` —
already reviewed, already TEST-NET addresses and reserved names. A second
fixture set would mean two things to keep in step with the sensors.
- The benchmark never trains, downloads, or deploys anything. It calls the
candidate's scoring path and nothing else.
- The Tier 1 benchmark never trains, downloads, or deploys anything. It calls
the candidate's scoring path and nothing else. The separate BETH Tier 2 rail
is offline and also performs no deployment writes.
- No Elasticsearch, no network. Everything is in-process.

## Run
Expand Down Expand Up @@ -45,27 +46,54 @@ and serious in production:
**A skipped check is not a passed check.** Skips are counted and reported
separately so a partially-exercised candidate cannot read as a clean one.

## Tier 2 — accuracy, and why it is not here
## Tier 2 — BETH architecture sanity rail

[`evaluate_accuracy.py`](evaluate_accuracy.py) is the bounded, labelled-corpus
rail. It is intentionally a **parallel-corpus sanity check**, not deployment
qualification: BETH is eBPF process telemetry, while `ml-worker` scores
honeypot/network events (`dst_port`, payload entropy, GeoIP, and credentials).
There is no honest feature mapping in either direction. The DNS half is not a
usable alternative: its six published files are byte-identical.

The BETH source is the anonymous Kaggle dataset
[`katehighnam/beth-dataset`](https://www.kaggle.com/datasets/katehighnam/beth-dataset),
licensed **CC0 1.0 Public Domain** and frozen at version 3. Prepare it outside
this repository; this command never downloads it. The expected directory layout
is direct, not nested:

```text
$HOME/beth/
labelled_training_data.csv
labelled_validation_data.csv
labelled_testing_data.csv
```

Run at least five fixed seeds. The harness fits PCA and IsolationForest on
`train`, uses `val` only as the held-out diagnostic split (IsolationForest has
no early-stopping or calibration step here), and scores `test` once per seed.
The report records runtime MD5s for all three CSVs, optionally the prepared
archive with `--archive`, split row counts and base rates, package versions,
elapsed time, and the SHA-256 of the JSON report. The report path must be
outside the repository. The published split row/label/host fingerprint is
checked before any model is fitted; a local extraction must not be re-split.

A labelled ranking corpus, **blocked on
[#1797](https://github.com/Xore/APIARY/issues/1797)**. If BETH does not map onto
our feature extractors, this benchmark ships as Tier 1 only — said out loud
here rather than shipping a harness that quietly cannot rank.
```bash
python3 ml-worker/benchmarks/evaluate_accuracy.py beth \
--data-dir "$HOME/beth" \
--output "$HOME/ml-worker-qualification/beth.json"
```

When it lands:
AUROC is the headline and AUPRC is reported alongside it, as mean ± standard
deviation over the selected seeds. Every printed metric is accompanied by the
published per-split `sus` base rates (approximately 0.17% / 0.42% / 90.7%).
The paper's iForest reference is about 0.850 AUROC. The harness uses a fixed
±0.05 screening band only to flag a pipeline miss; it is not a promotion gate
or a reason to retune a deployed threshold.

- **AUPRC is the headline.** Anomalies are rare and AUROC flatters under skew.
- AUROC secondary, for comparability with BETH's own published baselines.
- **Alert-budget precision**: precision *and* the absolute alert count per day
at the operating threshold. That number decides whether the dashboard is
usable.
- **Seed variance**: mean ± std over ≥ 3 fixed seeds. A delta smaller than the
seed spread is not a result. Non-negotiable — the LLM side measured a
±1-point run-to-run spread on a fixed model and prompt, and single-run
deltas smaller than that had been read as rankings for months.
- Calibration (ECE, Brier, reliability diagrams) only after Platt scaling on a
held-out labelled split. An ECDF rank transform makes scores commensurable
but is **not** calibration; never report ECE on one and call it calibrated.
**No-promotion rule:** this corpus is never evidence for changing the deployed
composite weights or `ML_ALERT_THRESHOLD`. The disposition corpus in
[#3295](https://github.com/Xore/APIARY/issues/3295) is the separate,
deployment-labelled source and is not imported by this BETH command.

## Banned metrics

Expand All @@ -76,7 +104,8 @@ When it lands:
trigger-happy detectors. Build the obvious thing and you get a harness that
ranks a coin flip above the LSTM-AE while looking rigorous. Use **PA%K** with
K stated if a segment metric is wanted.
- Any metric computed on a random or time-shuffled split.
- Any metric computed on a random or time-shuffled split. The BETH harness
consumes the three published files in place and never re-splits rows.

## Split rule — non-negotiable, and it applies to existing code

Expand Down
Loading
Loading