From 9dc38790efd6bed2c2688b6b24c9df20513837d5 Mon Sep 17 00:00:00 2001 From: "Ioannis L." <44038245+blitzcrieg1@users.noreply.github.com> Date: Sun, 23 Aug 2026 12:31:08 +0300 Subject: [PATCH] docs: order the roadmap by what reaches somebody else's SIEM The roadmap listed the dashboard once, as a bullet in "Later", with no reason attached. That reads as an oversight rather than a decision, and an oversight is something the next reader helpfully corrects. Three separate strategy documents this month each arrived at "demote the dashboard" as a finding, which is what happens when a decision is only implied by ordering. So the ordering is now stated. `Next` is sorted by one test: does this make the sensor land in a SIEM somebody else already runs. That puts the Splunk TA listing and per-host identity above the OTel ingest work, because the TA works and has tests and is still a directory you copy onto a search head by hand. Both of those needed an issue to satisfy this file's own second rule. Per-host identity had one. Splunk certification had none anywhere, so it has one now: enterprise #8, filed after checking that nothing in that repo has been near AppInspect. The dashboard gets a named section at the bottom instead of a stray bullet, with the reason and the limit. It is a local inspection surface for one machine, the SIEM is the console, and a triage queue here would be the second-best version of something the customer already bought. It also says the part that stops this reading as abandonment: it builds in CI, and the Next 16 migration went in because dependency alerts had to close, not because a screen needed adding. The Sigma pack is not on the roadmap because it shipped in #98. It is in CHANGELOG.md, per the first rule at the top of the file. The Dependabot row is gone for the same reason, the queue being empty. Status table refreshed: 1125 tests, and the detection row now says the rules are exported as Sigma, which is a claim a reader can check against docs/integrations/sigma/ rather than take. Co-Authored-By: Claude Opus 5 --- CHANGELOG.md | 39 +++++++++++++++++++++++++++++++++++++++ ROADMAP.md | 37 ++++++++++++++++++++++++++++++------- 2 files changed, 69 insertions(+), 7 deletions(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 5335066..2297e3f 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,6 +9,45 @@ separately (currently `1.2.0`) and changes additively. ## [Unreleased] +### Added + +- **The fifteen sequence detections ship as Sigma rules.** The Sigma pack + already in `docs/integrations/sigma/` covered the recorder's own health: + heartbeat silence, degraded coverage, MCP schema drift, denial bursts, + untriaged criticals. All useful, and none of it the product. A Splunk or + Sentinel team that wanted to alert on `credential-exfil` had to write the + search themselves from a schema document, which is a strange gap for a project + whose claim is that your SIEM stays the console. If the console is theirs, the + content has to be portable. + + The rules are generated by `tools/generate_sigma_pack.py`, which replays the + benchmark corpus through the real engine and reads the `Detection` objects it + emits. Hand-writing fifteen would put rule ids, severities and MITRE ids in a + second place that drifts from the first, which is the failure this repository + keeps finding in its own documents. A severity that changes in the rules + changes here on the next run, and a test fails if nobody reruns it. + + Sixteen files for fifteen rule ids, and the extra one is a real finding. + `encoded-command-download` emits two severities deliberately: critical for + remote code fetched and executed, low for local content piped into an + interpreter. Keying the pack on rule id alone produced one rule at whichever + severity the corpus yielded first, claiming critical while some firings are + low. Each severity now gets its own rule, pinned on `action.outcome`. One + static level would either page on the quiet variant or stay silent on the loud + one. + + Five tests hold it: every built-in rule has a Sigma rule, no Sigma rule + outlives a deleted one, severities match what the engine emits, the documents + are valid Sigma, and the UUIDs are derived from the rule id rather than random + so regenerating does not churn an id a consumer pinned. Closes + [#6](https://github.com/blitzcrieg1/agentmetry/issues/6). + + The two rules with no corpus case, `host-subagent-swarm-burst` and + `off-hours-activity`, carry an explicit entry in the generator with the reason. + That table is checked against `BUILTIN_RULE_IDS` rather than trusted: a + sixteenth rule with neither a corpus case nor an entry fails the script instead + of silently shipping a pack that claims more than it has. + ## [0.5.0] - 2026-08-21 ### Added diff --git a/ROADMAP.md b/ROADMAP.md index 5a4dff8..a93faef 100644 --- a/ROADMAP.md +++ b/ROADMAP.md @@ -1,6 +1,6 @@ # Roadmap -**Last refreshed: 2026-08-22.** State, not aspiration. Dates are absolute, +**Last refreshed: 2026-08-23.** State, not aspiration. Dates are absolute, because the previous version phased everything in "weeks 3 to 6" from a July start and every window had elapsed while the file still read as current. @@ -30,15 +30,15 @@ Two rules, so it does not happen again. --- -## Where this actually is, on 2026-08-22 +## Where this actually is, on 2026-08-23 | | | |---|---| | Version | 0.5.0 on PyPI, released 2026-08-21 | | Canonical event schema | 1.2.0, additive | -| Detection rules | 15 sequence rules, ATT&CK on every event, ATLAS on the AI-specific subset | +| Detection rules | 15 sequence rules, ATT&CK on every event, ATLAS on the AI-specific subset. All 15 exported as Sigma, generated from the engine | | Benchmark | 50 recorded sessions, 26 attack and 24 benign, 0 missed and 0 false positives | -| Tests | 1113 passing, 81% coverage, floor enforced in CI | +| Tests | 1125 passing, 81% coverage, floor enforced in CI | | Dogfood gate | **2 of 4 consecutive green weeks**, clock started 2026-08-08 | | Own trail | 27,504 hash-chained lines, 24 external anchors | | Adoption | Public alpha. No design partner tenant yet, no reference customer | @@ -94,7 +94,6 @@ the most important one on the page. | **First design partner contact** | none, tracked in the enterprise repo | Zero messages sent against ten researched accounts. Nothing else on this list matters as much | | Detection precision pass | [#44](https://github.com/blitzcrieg1/agentmetry/issues/44) [#49](https://github.com/blitzcrieg1/agentmetry/issues/49) [#50](https://github.com/blitzcrieg1/agentmetry/issues/50) [#51](https://github.com/blitzcrieg1/agentmetry/issues/51) [#55](https://github.com/blitzcrieg1/agentmetry/issues/55) | Five false-positive sources in frozen files. Land as one pass with #55 first, after the dogfood gate closes, so the ruleset fingerprint moves once | | Evidence pack integrity covers `meta` | [#75](https://github.com/blitzcrieg1/agentmetry/issues/75) | The date range on an evidence pack can currently be rewritten without breaking the hash. Needs a schema bump so existing packs keep verifying | -| Work the Dependabot queue | n/a | 15 open. The four GitHub Actions bumps touch `release.yml` and want the dry run before merging | **Held deliberately until 2026-09-05:** anything that edits `detection/rules.py`, `detection/traits.py`, `detection/engine.py` or @@ -108,14 +107,21 @@ rather than promised. ## Next (September to October 2026) +Reordered 2026-08-23 around one test: **does this item make the sensor land in +somebody else's SIEM?** The three at the top are what a security team touches +before they touch anything this project renders itself. Everything the dashboard +wants has moved to the bottom of the page. + | Item | Issue | Note | |---|---|---| +| **Splunk TA through AppInspect and a Splunkbase listing** | [enterprise #8](https://github.com/blitzcrieg1/agentmetry-enterprise/issues/8) | The TA works and has tests. It is a directory you copy onto a search head, which is a different conversation from a listing a security team can find. No public copy claims certification until it exists | +| **Per-host identity on a fleet trail** | [enterprise #1](https://github.com/blitzcrieg1/agentmetry-enterprise/issues/1) | A fleet trail is tamper-evident and not yet attributable. Ed25519 per host, so a forwarded event says which machine signed it | | **Ingest Claude Code OpenTelemetry as a third capture tier** | [#56](https://github.com/blitzcrieg1/agentmetry/issues/56) | Anthropic ships observed approval decisions we currently infer. Ingesting beats competing, and it closes [#45](https://github.com/blitzcrieg1/agentmetry/issues/45) for free. Reasoning in [the notes](https://agentmetry.ai/blog/why-not-just-opentelemetry) | +| OTLP **export** | none yet | Distinct from #56, which is ingest. Table stakes for teams standardised on a collector | | Tool response sizes in the trail | [#45](https://github.com/blitzcrieg1/agentmetry/issues/45) | The trail cannot tell a config lookup from a database dump | | Per-project scoping | [#37](https://github.com/blitzcrieg1/agentmetry/issues/37) | One trail currently mixes every repo on a machine | | Benchmark coverage for the six uncovered rules | [#36](https://github.com/blitzcrieg1/agentmetry/issues/36) [#25](https://github.com/blitzcrieg1/agentmetry/issues/25) | 13 of 15 rules have corpus coverage. Benign sessions harvested from the real trail beat invented ones | | Agent-directed technique taxonomy | [#47](https://github.com/blitzcrieg1/agentmetry/issues/47) | Partly addressed by the ATLAS layer in 0.5.0. Reassess what is genuinely still unlabelled | -| OTLP **export** | none yet | Distinct from #56, which is ingest. Table stakes for teams standardised on a collector | --- @@ -125,11 +131,28 @@ rather than promised. ([#35](https://github.com/blitzcrieg1/agentmetry/issues/35)) - Windsurf and VS Code Copilot hook installers ([#7](https://github.com/blitzcrieg1/agentmetry/issues/7)) -- Dashboard triage queue ([#27](https://github.com/blitzcrieg1/agentmetry/issues/27)) - MCP audit proxy over SSE and streamable HTTP, not only stdio - STIX/TAXII export of detections - DLP beyond regex. Real, and it waits for revenue +### The dashboard, last on purpose + +It is a **local inspection surface for the machine the sensor runs on**, and it +stays one. The SIEM is the console, which is the claim on every page of the +site, and a triage queue built here would be the second-best version of a +feature the customer already bought. + +It is not being deleted and it is not unmaintained: it builds in CI, and the +Next 16 migration went in because dependency alerts had to close, not because a +screen needed adding. That is the level of attention it gets. + +- Keyboard triage queue for the Detections tab + ([#27](https://github.com/blitzcrieg1/agentmetry/issues/27)) +- Remove the removed product's state model + ([#26](https://github.com/blitzcrieg1/agentmetry/issues/26)) +- Work through the 12 `set-state-in-effect` sites and let the rule error again + ([#95](https://github.com/blitzcrieg1/agentmetry/issues/95)) + --- ## Where this sits against what else exists