diff --git a/.github/workflows/quality.yml b/.github/workflows/quality.yml
index 047de130d..7ac6dac1c 100644
--- a/.github/workflows/quality.yml
+++ b/.github/workflows/quality.yml
@@ -1589,6 +1589,32 @@ jobs:
home: true
run: python scripts/check-docs-reachable.py
+ # #3395: #2458 scans only README.md and docs/**, so the component
+ # trees under arcane/**, analysis/**, sandbox/** and branding/** were
+ # never link-checked -- the agent-intrusion-corpus README's dead
+ # `../../docs/...` hop (two levels short of a five-deep path) lived
+ # there unnoticed. Whole-tree relative-link resolution. Fenced blocks
+ # are skipped: their paths belong to the reader's project, not ours.
+ - name: Every tracked doc's relative links resolve (#3395)
+ home: true
+ run: python scripts/check-doc-links.py
+
+ # #3395: two of the forty mermaid diagrams in the tree did not parse
+ # at all -- AI_TRIAGE.md named a node `call` (a reserved mermaid
+ # token) and ghidra/README.md left a colon unquoted inside an edge
+ # label. Both rendered as an error box on GitHub and nothing noticed,
+ # so mermaid parsing is now a gate rather than a review habit. Renders
+ # every block in one headless browser; mermaid is resolved from the
+ # npx cache, so the step is self-bootstrapping (no pinned version to
+ # drift out of date with the docs). The warm-up below is what makes
+ # that true: a fresh runner's npx cache is empty, so the lookup inside
+ # check-mermaid.mjs finds nothing and the gate cannot run at all.
+ - name: Mermaid diagrams parse (#3395)
+ home: true
+ run: |
+ npx -y @mermaid-js/mermaid-cli --version
+ node scripts/check-mermaid.mjs
+
# #2576: the CAPE implementation-plan doc claimed the detail page
# (/cape/{sha256}) was admin-gated. Nothing in backend-service
# enforces that -- require_service_token is the actual gate on
diff --git a/README.md b/README.md
index b76c87a8c..7e82aae5d 100644
--- a/README.md
+++ b/README.md
@@ -25,9 +25,12 @@ flowchart LR
wg --> home["home APIARY stacks @ 10.8.0.2"]
```
-**All core sensors run without compose profiles.** The only profile is the
-optional on-demand `geoip-update` maintenance job. 38 deployment pieces —
-32 independent Arcane-managed stacks under `arcane/home/` plus 6 more at
+**All core sensors run without compose profiles.** The only profiles are the
+optional on-demand maintenance jobs `geoip-update` and `threat-intel` (both in
+`honeypot-init`); the four `["legacy"]` worker stacks are defined for rollback
+but do not run, and `sandbox/ghosts`'s `["test"]` client is not a sensor.
+39 deployment pieces —
+33 independent Arcane-managed stacks under `arcane/home/` plus 6 more at
their own repository-root paths, all at home, plus the VPS (see
[docs/ARCANE-GIT-SYNC.md](docs/ARCANE-GIT-SYNC.md) for how a repo commit
reaches the live host, and [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) for
@@ -38,16 +41,17 @@ why the home side split into this many Compose stacks):
| `honeypot-keycloak` ([arcane/home/honeypot-keycloak/compose.yml](arcane/home/honeypot-keycloak/compose.yml)) | **home** | Arcane-managed Keycloak/PostgreSQL identity stack; only Keycloak is reachable from VPS Traefik over WireGuard |
| `honeypot-init` ([arcane/home/honeypot-init/compose.yml](arcane/home/honeypot-init/compose.yml)) | **home** | one-shot bootstrap jobs: log paths, Elasticsearch templates, Arkime schema, persona validation |
| `honeypot-cowrie`, `honeypot-dionaea`, `honeypot-conpot`, `honeypot-dnp3`, `honeypot-http`, `honeypot-multipot` (`arcane/home/honeypot-/compose.yml`, one directory each) | **home** | the sensors: Cowrie, Dionaea (+ TFTP relay), Conpot personas, DNP3, HTTP/API honeypots, multipot |
-| `honeypot-dicompot`, `honeypot-dns-honeypot`, `honeypot-citrix`, `honeypot-cisco-asa`, `honeypot-rdp`, `honeypot-endlessh`, `honeypot-beelzebub`, `honeypot-hellpot`, `honeypot-elasticpot`, `honeypot-galah`, `honeypot-sentrypeer`, `honeypot-mailoney` (`arcane/home/honeypot-/compose.yml`, one directory each) | **home** | more sensors: DICOM medical-imaging decoy, DNS UDP reflection bait (response-capped, never a real amplification vector), Citrix ADC/NetScaler Gateway decoy (CVE-2019-19781), Cisco ASA WebVPN+IKE decoy (CVE-2018-0101), RDP decoy, SSH pre-auth tarpit, vendored multi-protocol deception runtime (SSH/LDAP/MCP/HTTP, #1418), vendored HTTP bot tarpit (#1419), vendored Elasticsearch decoy distinct from multipot's own (#1423), vendored LLM-powered HTTP honeypot behind its own broker-guarded bridge onto the shared Ollama instance (#1420), vendored SIP/VoIP fraud-detection honeypot (#1424), vendored SMTP honeypot taking over port 25 from multipot's own retired handler (#1422) — the row's `honeypot-wordpot` / WordPress/CMS decoy slot was removed when wordpot retired (#2381) |
+| `honeypot-dicompot`, `honeypot-dns-honeypot`, `honeypot-citrix-honeypot`, `honeypot-cisco-asa-honeypot`, `honeypot-sonicwall-sma`, `honeypot-rdp-honeypot`, `honeypot-endlessh`, `honeypot-beelzebub`, `honeypot-hellpot`, `honeypot-elasticpot`, `honeypot-galah`, `honeypot-sentrypeer`, `honeypot-mailoney` (`arcane/home/honeypot-/compose.yml`, one directory each) | **home** | more sensors: DICOM medical-imaging decoy, DNS UDP reflection bait (response-capped, never a real amplification vector), Citrix ADC/NetScaler Gateway decoy (CVE-2019-19781), Cisco ASA WebVPN+IKE decoy (CVE-2018-0101), SonicWall SMA1000 Work Place/AMC decoy (CVE-2026-83548 Work Place SSRF), RDP decoy, SSH pre-auth tarpit, vendored multi-protocol deception runtime (SSH/LDAP/MCP/HTTP, #1418), vendored HTTP bot tarpit (#1419), vendored Elasticsearch decoy distinct from multipot's own (#1423), vendored LLM-powered HTTP honeypot behind its own broker-guarded bridge onto the shared Ollama instance (#1420), vendored SIP/VoIP fraud-detection honeypot (#1424), vendored SMTP honeypot taking over port 25 from multipot's own retired handler (#1422) — the row's `honeypot-wordpot` / WordPress/CMS decoy slot was removed when wordpot retired (#2381) |
| `honeypot-canarytokens` ([arcane/home/honeypot-canarytokens/compose.yml](arcane/home/honeypot-canarytokens/compose.yml)) | **home** | self-hosted honeytoken platform (#1426) -- planted-artifact deception, not a listening protocol decoy; `canarytokens-adapter` translates its webhook alerts into this repo's shared JSON event shape. The dashboard's Settings > Canarytokens pane (#1487) creates PDF/Word/Excel/custom-image/Windows-Folder/QR tokens on demand for use *outside* this honeypot (#1662 dropped the stale design doc that described the pre-cutover plan; the shipped pane is authoritative) |
| `honeypot-tanner` ([arcane/home/honeypot-tanner/compose.yml](arcane/home/honeypot-tanner/compose.yml)) | **home** | SNARE + TANNER application-emulation boundary |
| `honeypot-elk` ([arcane/home/honeypot-elk/compose.yml](arcane/home/honeypot-elk/compose.yml)) | **home** | Filebeat, Elasticsearch, Kibana, EveBox, Arkime |
-| `honeypot-agent-intrusion-worker` ([arcane/home/honeypot-agent-intrusion-worker/compose.yml](arcane/home/honeypot-agent-intrusion-worker/compose.yml)) | **home** | correlates sensor/Suricata events into campaigns, scores them against deterministic criticality rules, writes the `agent-intrusion-campaigns` index the dashboard's `/agent-campaigns` route reads |
-| `honeypot-attacker-identity-worker`, `honeypot-correlator-worker`, `honeypot-payload-inventory-worker` (`arcane/home/honeypot-/compose.yml`, one directory each) | **home** | three more workers that had their own top-level compose file but had drifted out of the deploy/installer inventory before #1502's audit caught it (same class of gap #560 and #891 each fixed once before) -- attacker-identity correlation, cross-sensor campaign correlation, and payload inventory tracking |
+| `honeypot-agent-intrusion-worker` ([arcane/home/honeypot-agent-intrusion-worker/compose.yml](arcane/home/honeypot-agent-intrusion-worker/compose.yml)) | **home** | the labelled corpus plus the Tier 1 contract benchmark. The worker itself was ported to Rust in #1610 and now runs as `WORKER_LOOPS=agent-intrusion` inside `honeypot-dashboard`'s `backend-worker`; the Python stack is retained under the `legacy` profile for rollback only, and the live `agent-intrusion-campaigns` index is written by the Rust loop |
+| `honeypot-attacker-identity-worker`, `honeypot-correlator-worker`, `honeypot-payload-inventory-worker` (`arcane/home/honeypot-/compose.yml`, one directory each) | **home** | three more workers that had their own top-level compose file but had drifted out of the deploy/installer inventory before #1502's audit caught it (same class of gap #560 and #891 each fixed once before) -- attacker-identity correlation, cross-sensor campaign correlation, and payload inventory tracking. All three were retired by the same #1649 pass as the agent-intrusion worker: #1610 ported them to Rust, where they now run as `WORKER_LOOPS=attacker-identity` and `WORKER_LOOPS=correlator` on `honeypot-dashboard`'s `backend-worker` and `WORKER_LOOPS=payload-inventory` on `backend-worker-payload-inventory`. Like row above, the Python stacks are kept under the `legacy` profile for rollback only and are not the live writers |
| `honeypot-dashboard` ([arcane/home/honeypot-dashboard/compose.yml](arcane/home/honeypot-dashboard/compose.yml)) | **home** | the live investigation dashboard: TanStack Start frontend (`dashboard-next`) in front of the Rust axum `backend-service` API tier, plus its worker containers (importer, networkless enrichment, the `WORKER_LOOPS` aggregation loops) and the services-adapter Docker control surface. Live since #1628's cutover completed 2026-08-22 -- the Go dashboard is deleted; see [docs/DASHBOARD-CUTOVER.md](docs/DASHBOARD-CUTOVER.md) and [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) |
-| `honeypot-dashboard-backend` ([arcane/home/honeypot-dashboard-backend/compose.yml](arcane/home/honeypot-dashboard-backend/compose.yml)) | **home** | the write-capable, host-mounted `backend-service` instance (:8082), split out from `honeypot-dashboard` by #1622 -- same route table plus the analysis request-spool mounts; only this instance can dispatch `analysis/ghidra`/sandbox jobs |
+| `honeypot-dashboard-backend` ([arcane/home/honeypot-dashboard-backend/compose.yml](arcane/home/honeypot-dashboard-backend/compose.yml)) | **home** | the unprivileged read-only `backend-service` API tier (:8081), split out from `honeypot-dashboard` by #1622 so Arcane can redeploy the API tier without touching `dashboard-next`; the write-capable, host-spool-mounted instance is `backend-service-mounted` (:8082), which stayed in `honeypot-dashboard` |
| `honeypot-payload-analysis` ([arcane/home/honeypot-payload-analysis/compose.yml](arcane/home/honeypot-payload-analysis/compose.yml)) | **home** | payload dedup + YARA scanning |
| `honeypot-utilities` ([arcane/home/honeypot-utilities/compose.yml](arcane/home/honeypot-utilities/compose.yml)) | **home** | autoheal, log rotation, disk-space monitoring, reporting |
+| `unsloth` ([arcane/home/unsloth/compose.yml](arcane/home/unsloth/compose.yml)) | **home** | Unsloth Studio, the interactive leg of the round-7 training work area (#3080); operator starts and stops it in Arcane so it can release VRAM between cold-benchmark legs |
| [`vps/`](vps/) | **VPS** | Traefik, portbridge raw tunnels, Suricata, WireGuard HTTP bridges, and isolated Keycloak OIDC gateways |
Every stack above is a directory-aware Arcane Git sync driven by
@@ -90,7 +94,7 @@ for where those fit.
| [docs/BACKUP-ESSENTIALS.md](docs/BACKUP-ESSENTIALS.md) | What is backed up so the stack can be rebuilt, where the three copies go, and the full restore procedure |
| [scripts/install.sh](scripts/install.sh) | Single entry point for host provisioning — `sudo ./scripts/install.sh --profile home\|vps`, which dispatches to the installer below (or [scripts/install-vps.sh](scripts/install-vps.sh)) with that profile's answers file. Both share their retry/step/logging framework via [scripts/lib/install-common.sh](scripts/lib/install-common.sh) ([#1609](https://github.com/Xore/APIARY/issues/1609)) |
| [scripts/install-homeserver.sh](scripts/install-homeserver.sh) | Unattended provisioning script (Docker, GPU/NVIDIA, WireGuard, Arcane, the stacks themselves) for a manually-installed base Ubuntu system — fill in [scripts/install-homeserver.conf.example](scripts/install-homeserver.conf.example) first, same idea as a Windows `autounattend.xml` answer file. First cut, see [#518](https://github.com/Xore/APIARY/issues/518) |
-| [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) | System architecture and data flow — trust boundaries, container map, event ingestion, correlation/enrichment (p0f, HASSH/JA3/JA4, GeoIP), payload lifecycle, sandbox detonation, evidence types (6 diagrams) |
+| [docs/ARCHITECTURE.md](docs/ARCHITECTURE.md) | System architecture and data flow — trust boundaries, container map, event ingestion, correlation/enrichment (p0f, HASSH/JA3/JA4, GeoIP), payload lifecycle, sandbox detonation, evidence types (4 diagrams) |
| [docs/SENSORS.md](docs/SENSORS.md) | The sensor table, resource budgets, investigation UIs, SNARE+TANNER, Suricata, Arkime, and how real attacker IPs survive the tunnel |
| [docs/OPERATIONS.md](docs/OPERATIONS.md) | Persona inventory, the seeded cowrie filesystem, GeoIP, and how to actually read the data (dashboard, Kibana, Arkime, backups) |
| [docs/ip-reporting-plan.md](docs/ip-reporting-plan.md) | Defensive IP-blocklist reporting (AbuseIPDB/Blocklist.de), dry-run by default |
@@ -103,7 +107,7 @@ for where those fit.
| [docs/CONTAINER-UPDATES.md](docs/CONTAINER-UPDATES.md) | How to check pinned images for updates, assess compatibility, verify empirically, and pin by digest |
| [docs/TESTING.md](docs/TESTING.md) | The three testing tiers -- CI, live feature smoke tests, and the full clean-reinstall release gate -- and how to repeat each one |
| [docs/STACK-REBUILD.md](docs/STACK-REBUILD.md) | Runbook for a full deliberate reset — stop order, what's preserved vs wiped, and the ordering/permission pitfalls to avoid |
-| [deploy-profiles/](deploy-profiles/) | Named deployment shapes (full / ICS-only / web-only) — which of the 20 split home stacks run for a given deployment, plus a validator catching cross-stack drift before deploy |
+| [deploy-profiles/](deploy-profiles/) | Named deployment shapes (full / ICS-only / web-only) — which of the 26 split home stacks run for a given deployment, plus a validator catching cross-stack drift before deploy |
| [docs/RECOVERY.md](docs/RECOVERY.md) | `factory-reset.sh` — one entry point for "back up, optionally wipe/reset, restart" on the same host |
| [docs/ROADMAP.md](docs/ROADMAP.md) / [docs/WORK-LEDGER.md](docs/WORK-LEDGER.md) | What order work happens in, and how issues are claimed/reviewed |
| [docs/ml-worker-plan.md](docs/ml-worker-plan.md), [docs/gpu-llm-analysis-worker.md](docs/gpu-llm-analysis-worker.md), [docs/gpu-ml-worker-acceleration.md](docs/gpu-ml-worker-acceleration.md) | The homeserver's NVIDIA GPU running local LLM log/payload analysis and CUDA-accelerated anomaly detection — no data leaves the machine |
diff --git a/SECURITY.md b/SECURITY.md
index b15c38084..410168456 100644
--- a/SECURITY.md
+++ b/SECURITY.md
@@ -13,3 +13,12 @@ live malware, private keys, production `.env` files, packet captures containing
private traffic, or unredacted logs to a public issue.
Supported security fixes target the current `main` branch.
+
+`scripts/check-public-leaks.py` enforces the "leak a real secret" half of this
+policy on every change, and fails CI on private keys, GitHub/AWS/Slack tokens,
+literal credential assignments, credentials embedded in URLs, deployment
+`.env` files, private-key and packet-capture binaries, and the
+deployment-specific addresses and default password this repository must never
+name. Exactly one tracked `.env` is exempt, and it is the decoy honeyfs file
+under `arcane/home/honeypot-cowrie/` that exists for attackers to find — not a
+credential, and not something to "fix" by deleting.
diff --git a/arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/README.md b/arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/README.md
index e78e53c60..9196316e3 100644
--- a/arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/README.md
+++ b/arcane/home/honeypot-agent-intrusion-worker/analysis/agent-intrusion-corpus/README.md
@@ -15,7 +15,7 @@ original consumer and is retired) that renders its output — all proven
against that one corpus rather than hand-built fixtures alone. The
prerequisite research (mapping the campaign to APIARY's actual trust
boundaries) is
-[`docs/agent-intrusion-threat-model.md`](../../docs/agent-intrusion-threat-model.md);
+[`docs/agent-intrusion-threat-model.md`](../../../../../docs/agent-intrusion-threat-model.md);
phase 4's preventive-control audits (Dockerfile digest pinning, an
assessed ARKIME secret-delivery finding) are documented there, not here,
since they touch the wider repo rather than this directory. Phases 1-5
@@ -259,7 +259,8 @@ change reviewed content. To extend the corpus, add events directly to
`corpus.jsonl` is not itself schema-versioned (no top-level `version`
field) — the file's own git history is the version record, matching how
-this repo treats `analysis/yara/` and other reviewed-fixture directories.
+this repo treats `arcane/home/honeypot-payload-analysis/analysis/yara/` and
+other reviewed-fixture directories.
A future breaking change to `schema.json` (a required field added/removed,
an enum value changed) should bump `schema.json`'s own `$id` and add a
note here, not silently reinterpret old corpus rows under a new meaning.
diff --git a/docs/ARCANE-GIT-SYNC.md b/docs/ARCANE-GIT-SYNC.md
index 43fc3c151..a2909baf0 100644
--- a/docs/ARCANE-GIT-SYNC.md
+++ b/docs/ARCANE-GIT-SYNC.md
@@ -1,6 +1,6 @@
# Arcane Git sync
-How the 37 home-hosted stacks (31 that migrated under `arcane/home/` plus 6
+How the 39 home-hosted stacks (33 that live under `arcane/home/` plus 6
that were already self-contained and stayed at their existing path) get to
the live host, replacing the old model of copying or symlinking top-level
`docker-compose.*.yml` files into place. Everything here was confirmed live
@@ -9,7 +9,10 @@ taken from Arcane's docs — see the risk that motivated that in "Version/API
compatibility" below. Census numbers and the live sync-store state were
re-verified 2026-08-27 against the pinned `v2.9.0` image and its own sqlite
store (#2549): 31 in-tree directories + 6 self-contained = 37, exactly the
-manifest's entry count.
+manifest's entry count *at that date*. The manifest has since grown to 39 —
+#2911 swapped the `pihole` entry for `technitium` and #3092 added `unsloth`
+directly under `arcane/home/` — re-counted 2026-09-27 from
+`arcane/manifests/home-production.json`: 33 in-tree + 6 self-contained = 39.
## The model
@@ -18,19 +21,21 @@ Each stack gets its own **directory-aware Git sync**: Arcane clones the
selected `compose.yml` (not just that one file) under
`/var/dockge/stacks//`, and deploys it. The manifest at
[`arcane/manifests/home-production.json`](../arcane/manifests/home-production.json)
-is the single source of truth for which 37 stacks exist, what branch/path
+is the single source of truth for which 39 stacks exist, what branch/path
each syncs from, and any per-stack sync limits — `scripts/install-homeserver.sh`,
CI, and this doc all read from it rather than maintaining separate lists.
- The 32 `honeypot-*` stacks live under `arcane/home//`: their build
context and git-tracked config were moved there from repository root
(see each compose file's own `#1502` comment for what moved and why).
+ `unsloth` is the 33rd directory there — #3092 added it in-tree rather
+ than at a root path, so it never went through the #1502 move.
- The 6 other stacks (`auth-events-worker`, `llm-worker`, `ml-worker`,
- `analysis/ghidra`, `sandbox/ghosts`, `pihole`) were already self-contained
- and stayed at their existing path — moving them would have broken real
- references from `scripts/install-homeserver.sh`, CI workflows, and
- `deploy.yml`'s own ghidra-worker resync step. See each one's own compose
- file header for the specifics.
+ `analysis/ghidra`, `sandbox/ghosts`, `technitium`) were already
+ self-contained and stayed at their existing path — moving them would
+ have broken real references from `scripts/install-homeserver.sh`, CI
+ workflows, and `deploy.yml`'s own ghidra-worker resync step. See each
+ one's own compose file header for the specifics.
- `honeypot-arcane` itself is **not** in the manifest and never will be —
syncing the thing that has to already be running before any sync can
happen is a bootstrap loop, not a simplification. It stays
@@ -58,15 +63,27 @@ been provisioned once.
`scripts/install-homeserver.sh`'s `step_arcane_import_stacks` reads the
manifest and creates one `POST /environments/0/gitops-syncs` per matching
entry (environment `0` is Arcane's single "Local Docker" environment on a
-one-host deployment). Its selection filter matches every `honeypot-*`
-entry — and, since #1505, three of the six non-`honeypot-*` stacks too:
-`auth-events-worker`, `llm-worker` and `ml-worker` are imported by that
-step as well (each confirmed to have no host-local state beyond `.env`).
-The other three keep their dedicated installer steps for reasons specific
-to each: `pihole`'s non-`.env` host state, `analysis/ghidra`'s conditional
-GPU compose overlay, and `sandbox/ghosts`'s Arcane build-context
-limitation (#1506) — see the script's own Phase 8 header comment for the
-reasoning behind each.
+one-host deployment). Its selection filter matches all 32 `honeypot-*`
+entries — and, since #1505, three of the seven non-`honeypot-*` entries
+too: `auth-events-worker`, `llm-worker` and `ml-worker` are imported by
+that step as well (each confirmed to have no host-local state beyond
+`.env`). That is 35 of the manifest's 39 entries. The other three the
+installer provisions itself keep their dedicated steps for reasons specific
+to each: `technitium`'s non-`.env` host state (its `config/` directory needs
+non-root ownership, #2911 — this is the step `pihole` used to have),
+`analysis/ghidra`'s conditional GPU compose overlay, and `sandbox/ghosts`'
+s Arcane build-context limitation (#1506) — see the script's own Phase 8
+header comment for the reasoning behind each.
+
+The 39th entry, `unsloth` (#3092), is the one the installer does not reach
+at all: it matches neither arm of the filter, and the script has no
+`step_unsloth_*` of its own. That is not a defect — `arcane/home/unsloth/compose.yml`'s
+own header documents it as an operator-started stack ("Deployed through
+Arcane like every other homeserver stack … Never `docker compose up` by
+hand"), deliberately carrying no `restart:` policy so the cold-benchmark legs
+in `analysis/ghidra/training/` can have the card to themselves. It is
+covered where fleet-wide operations are concerned — `docs/STACK-REBUILD.md`
+lists it in both reset loops — just not by the from-scratch import path.
To import (or re-import) by hand instead, `POST` the manifest's entries to
`/environments/0/gitops-syncs/import` — the bulk-import shape matches the
@@ -106,16 +123,19 @@ letting Arcane's own directory check block the whole import.
not track** — not just `.env` and `secrets/`. Confirmed live: `pihole`
also keeps its DNSCrypt resolver config and its own Pi-hole database
directly under its own top-level directory (`dnscrypt-proxy/`,
- `etc-pihole/`, `etc-dnsmasq.d/`), a shape none of the other 37 stacks
- have. Back up the *actual* bind-mount sources a stack's compose file
+ `etc-pihole/`, `etc-dnsmasq.d/`), a shape only a handful of stacks have
+ — the live one now is `technitium`, whose `config/` (zones, settings,
+ blocklists) needs non-root ownership for its distroless image, which is
+ exactly why the installer still provisions it by hand (#2911). Back up the *actual* bind-mount sources a stack's compose file
declares, not an assumed `.env`/`secrets/` checklist — read the compose
file if in doubt.
2. Back up everything found in step 1.
3. Remove the stack's current directory.
4. Create the Arcane sync (`syncDirectory: true`) — this deploys
immediately, and will legitimately fail-closed if a required secret
- isn't present yet (expected, not a bug — see canarytokens/ghosts/
- keycloak/dashboard's own `:?required` variables).
+ isn't present yet (expected, not a bug — see the four stacks that do
+ declare `:?required` variables: `honeypot-canarytokens`, `ghosts`,
+ `technitium` and `unsloth`).
5. Restore everything backed up in step 1, **preserving original ownership
and permissions, not just content**. Confirmed live: restoring a secret
file as `root:root` when the container expects the previous owning
@@ -173,7 +193,14 @@ a stack means:
These are platform behaviors, not something fixable from a compose file
alone. Re-verify against whatever Arcane version is pinned in
-`docker-compose.arcane.yml` if it's ever upgraded.
+`docker-compose.arcane.yml` if it's ever upgraded. **That upgrade has
+happened and the re-verification has not:** every item below was confirmed
+against `v2.8.0`–`v2.9.0` (including the no-`profiles:`-support finding,
+which cites the upstream issue as still open at `v2.8.1`), while
+`docker-compose.arcane.yml` now pins `ghcr.io/getarcaneapp/manager:v2.11.1`
+by digest. Nothing here is claimed to be false of `v2.11.1` — it is claimed
+to be unconfirmed against it, and three minor versions of upstream is
+exactly the gap that re-verification is for.
- **`"project directory is not inside a mounted directory"` is a false
warning against every stack with a relative bind-mount source, under
@@ -194,18 +221,24 @@ alone. Re-verify against whatever Arcane version is pinned in
is no per-message log-suppression Arcane exposes, and `hp-arcane`'s
container logs aren't ingested into this repo's ELK pipeline (it tails
application log files, not `docker logs` streams), so there is no
- repo-side filter either. Fires for exactly the 9 stacks with at least one
- relative-source bind mount under this identity mount (`honeypot-cowrie`,
- `honeypot-dionaea`, `honeypot-elk`, `honeypot-init`, `honeypot-keycloak`,
- `honeypot-payload-analysis`, `honeypot-tanner`, `honeypot-utilities`,
- `pihole`) — every relative mount on the 7 of those 9 with running
- containers was confirmed to resolve to its real, existing host path, zero
- missing-on-host mounts. (`pihole` and `honeypot-keycloak` were the two
- without running containers, so their mounts were reasoned from the same
- identity-mount arithmetic rather than observed — #2853, #2764. `pihole`
- has since been replaced by `technitium` in the manifest, #2911; the
- warning's mechanism is per-relative-mount and unchanged by the swap, but
- the stack name in this list is the pre-#2911 one.) **Decision: live with it — this repo has no fix
+ repo-side filter either. Fires for exactly the 10 stacks with at least one
+ relative-source bind mount under this identity mount (`ghidra`,
+ `honeypot-cowrie`, `honeypot-dionaea`, `honeypot-elk`, `honeypot-init`,
+ `honeypot-keycloak`, `honeypot-payload-analysis`, `honeypot-tanner`,
+ `honeypot-utilities`, `technitium`; the 10 is re-derived 2026-09-27 from
+ each of the manifest's 39 entries' `volumes:` sources) — every relative
+ mount on the 7 of the 9 that existed at the 2026-09-03 check and had
+ running containers was confirmed to resolve to its real, existing host
+ path, zero missing-on-host mounts. (`pihole` and `honeypot-keycloak` were
+ the two without running containers, so their mounts were reasoned from the
+ same identity-mount arithmetic rather than observed — #2853, #2764.
+ `pihole` has since been replaced by `technitium` in the manifest, #2911 —
+ hence the post-#2911 name above; the warning's mechanism is
+ per-relative-mount and unchanged by the swap. The tenth stack, `ghidra`,
+ is not new: its `./revdeck-proxyfix/*` mounts date from #1173, so the
+ 2026-09-03 live pass simply never reached it, and the "9" it recorded was
+ a nine-name subset of the real set rather than a different count.)
+ **Decision: live with it — this repo has no fix
available, and the warning is confirmed harmless.** Re-verified live
2026-09-03: still firing at the same 9 stacks, ~1 warning per relative
mount per sync. Not reported upstream (`getarcaneapp/arcane`) as of this
@@ -247,7 +280,7 @@ alone. Re-verify against whatever Arcane version is pinned in
(`dashboard/vendor/`, exactly 650 tracked files at the #1502 move,
re-counted from git history), which exceeded the default back then. That
tree is gone with the Go tier itself (#1659) and the synced directory is
- down to 268 tracked files (re-counted 2026-08-27,
+ down to 279 tracked files (re-counted 2026-09-27,
`git ls-files arcane/home/honeypot-dashboard`) — under the default
again, so the bump is dormant headroom, kept because raising it is an
Arcane-side change rather than a repo one. The failure mode itself
@@ -262,9 +295,11 @@ alone. Re-verify against whatever Arcane version is pinned in
`ghosts`'s sync (`sandbox/ghosts/compose.yml`) always failed — either a
~40s-then-500 with no detail, or (once `maxSyncTotalSize` alone was
raised) a fast `file count limit exceeded` — because
- `sandbox/ghosts/vendor/ghosts-src/` (963 tracked files, ~132 MB) sits in
- the same directory tree as the compose file, so the sync's whole-directory
- walk of `sandbox/ghosts/` (989 files, 135,789,139 bytes in total) blows
+ `sandbox/ghosts/vendor/ghosts-src/` (935 tracked files, 135,601,433 bytes)
+ sits in the same directory tree as the compose file, so the sync's
+ whole-directory walk of `sandbox/ghosts/` (962 files, 135,748,394 bytes —
+ both re-counted 2026-09-27; the 2026-08-31 live walk quoted below measured
+ 989/135,789,139) blows
past *both* defaults (`maxSyncFiles: 500`, `maxSyncTotalSize: 50MB`), not
just the one either failure message names. Fixed the same way as the
`honeypot-dashboard` case above — both limits raised on the sync record
@@ -347,8 +382,11 @@ alone. Re-verify against whatever Arcane version is pinned in
## Local environment overrides
Compose's own `.env`-in-project-directory interpolation already covers
-every `${VAR}` reference in these 37 stacks — none of them use `env_file:`,
-and none needed it added. Arcane's effective environment merge
+every `${VAR}` reference in these 39 stacks — none of them *need* `env_file:`
+and only one declares it: `analysis/ghidra/docker-compose.ghidra.yml`'s
+`revdeck`-profile service carries `env_file: [{path: .env, required: false}]`
+(#110), which duplicates the interpolation it sits beside rather than
+carrying anything. None needed it added. Arcane's effective environment merge
(`project.env` + `.env.git` → `.env`) feeds that same mechanism
transparently, so a local override set through Arcane's own UI for a synced
project works exactly like editing `.env` by hand always did; no stack-file
@@ -407,12 +445,28 @@ Two things worth knowing when setting an override:
## Promotion workflow and change control
Decided in #1507: **release/tag promotion**, with `autoSync` enabled for
-exactly the three stacks where a sync is the whole deploy. As of
-2026-08-27 that policy has **never been put into effect** — every sync
-still tracks `main` and nothing auto-deploys (#2549 re-derived the live
-state). What actually runs is the manual model below; #2577 holds the
-one-time activation steps and the dangling-sync cleanup if that ever
-changes.
+exactly the three stacks where a sync is the whole deploy. **The manifest
+half of that policy is in effect: every one of the manifest's 39 entries
+declares `branch: "production"`, and #1943 (2026-08-25) is the commit that
+changed it from `main` — re-derived 2026-09-27, and unchanged since except
+for the `pihole`→`technitium` and `unsloth` entries that inherited it.**
+The live-store half is not. As of the last reads (2026-08-27 #2549, 2026-09-03)
+every live sync still tracked `main` and nothing auto-deployed, so the
+manifest and the host disagree about what a sync follows. #2577 holds the
+one-time activation steps and the dangling-sync cleanup.
+
+**The pointer the manifest names does not exist yet.** `git ls-remote
+--heads origin` on 2026-09-27 returned `main` and assorted agent/dependabot/
+design branches and no `refs/heads/production`; `scripts/promote-release.sh
+--list` reports it as "branch does not exist yet". A sync pointed at
+`production` therefore fails the same way a tag does — Arcane prefixes
+`refs/heads/` onto whatever it is given (see "Arcane cannot track a tag"
+below) — so this is the one place where the manifest is currently ahead of
+reality in a way that bites: the documented bulk-import path,
+`POST /environments/0/gitops-syncs/import`, would create 39 syncs that all
+fail on their first run, on a branch that nothing creates until someone runs
+`scripts/promote-release.sh `. What actually runs today is the manual
+model below.
### Two facts (still true today)
@@ -457,27 +511,36 @@ treat every sync as a per-project restart and bound the blast radius
accordingly: scope each run to one stack at a time and verify the result — rather than firing
a fleet-wide sync pass and racing the timeout. (A purpose-built
single-project script is a natural follow-up; ship it separately so the
- doc's claim of 'every sync is a restart' is reviewable on its own.)**34 of the 37 stacks build an image.** Only `honeypot-elk`,
-`honeypot-keycloak` and `pihole` pull — re-derived 2026-08-27 from the 37
-manifest compose paths (34 carry `build:`; the same three pullers the
-#1502-era text named, which said "35 of the 38" before #2381 retired
-wordpot and f139fe24 retired the Go ip-enrichment-worker). For any
+ doc's claim of 'every sync is a restart' is reviewable on its own.)**34 of the 39 stacks build an image.** Five pull —
+re-derived 2026-09-27 from the 39 manifest compose paths (34 carry
+`build:`): `honeypot-elk`, `honeypot-keycloak`, `llm-worker`, `technitium`
+and `unsloth`. `llm-worker` is in that set only because the manifest
+deploys `llm-worker/docker-compose.captured-data-deploy.yml`, an overlay
+with no `build:` of its own — `llm-worker/docker-compose.yml` still has
+one. The #1502-era text named "35 of the 38" and the same three pullers,
+before #2381 retired wordpot and f139fe24 retired the Go
+ip-enrichment-worker. For any
building stack, `autoSync: true` would mean every merge produces a
deployment that looks successful and changes nothing — worse than a
manual process, because it is unattended.
### What actually runs
-- **Every sync tracks `branch: "main"`** — all 38 live `gitops_syncs`
- rows (the 37 manifest stacks plus #2577's dangling `honeypot-wordpot`
- orphan). The rows predate #1507's decision and nothing re-pointed
- them; no `production` pointer exists.
+- **Every sync tracks `branch: "main"`** on the live host, while the
+ manifest has said `branch: "production"` for all 39 entries since
+ #1943 — so the rows predate #1507's decision and nothing re-pointed
+ them. Every live `gitops_syncs` row is one of 39 manifest stacks plus
+ #2577's dangling `honeypot-wordpot` orphan, so 40; the 2026-08-27 read
+ saw 38 of them, before #2911's `technitium` and #3092's `unsloth`. The
+ gap between the two is the thing to resolve before the next import, and
+ the branch to resolve it against does not exist yet (see "Promotion
+ workflow" above).
- **`autoSync` is 0 everywhere, including the three the manifest flags.**
- `honeypot-elk`, `honeypot-keycloak` and `pihole` carry `autoSync: true`
- in `arcane/manifests/home-production.json`, but the live store has
- `auto_sync = 0` on all 38 rows — the elk/keycloak/pihole auto-follow
- policy is silently inert: a promotion, or any push, will not deploy
- them.
+ `honeypot-elk`, `honeypot-keycloak` and `technitium` carry
+ `autoSync: true` in `arcane/manifests/home-production.json`, but the live
+ store has `auto_sync = 0` on every row — the elk/keycloak/technitium
+ auto-follow policy is silently inert: a promotion, or any push, will not
+ deploy them.
- **Every deploy is manual**: sync → build → redeploy per stack. The
order matters and is not arbitrary: `honeypot-dashboard` must sync
before `honeypot-dashboard-backend` builds, because the Rust source
@@ -485,8 +548,8 @@ manual process, because it is unattended.
- **Promotion is CI-only.** `scripts/promote-release.sh v0.1.0` exists
and still refuses a ref that is not a tag, and a tag that is not an
ancestor of `main` — so if a pointer ever exists, what reaches it has
- always been through CI. Today it moves nothing, because there is no
- pointer to move.
+ always been through CI. As of 2026-09-27 it still moves nothing: there
+ is no `production` branch on `origin` to move.
### The #1507 design, for the record
@@ -501,8 +564,14 @@ manual process, because it is unattended.
single-sync PATCH work.
#2577 closed (PR #2704) having done only the wordpot-orphan half of its own
-scope — the #1507 activation itself was never done, and #2858 (below)
-decided explicitly not to do it in this round either.
+scope, and #2858 (below) decided explicitly not to do the activation in this
+round either. Two halves of "the #1507 activation" are worth keeping
+distinct: the manifest half **was** done, in #1943 on 2026-08-25, which is
+what put `branch: "production"` and the three `autoSync: true` flags into
+`arcane/manifests/home-production.json`; the live-store half — re-pointing
+the rows already in Arcane at that branch and setting the
+`pullImageAfterSync`/`redeployAfterSync` fields the manifest cannot carry —
+is what is still outstanding.
## Proving which revision is deployed (#3315)
@@ -642,7 +711,7 @@ investigated "a merged PR never reached the host" issues.
**Decision: leave `autoSync: false` fleet-wide. Do not activate #1507's
three-puller policy in this round either.** Reasoning:
-- Turning `autoSync` on for all 37 projects means every merge to `main`
+- Turning `autoSync` on for all 39 projects means every merge to `main`
redeploys the fleet unattended — a real increase in blast radius, and
exactly the kind of change that should not happen as a side effect of
fixing a visibility gap. (This round's own brief calls this out
@@ -676,26 +745,31 @@ three-puller policy in this round either.** Reasoning:
ops-triage pass. `autoSync` therefore remains `false` fleet-wide for now.
- #2854's one-shot-abort hazard (a `restart: no` job's clean `exit(0)`
making Arcane report `failed` on a deploy that actually completed) is
- fully scoped to `honeypot-init`. It is **not** the only file in the repo
+ scoped to `honeypot-init` **and, since #3128 (2026-09-08), `honeypot-elk`.**
+ It is **not** the only file in the repo
with that shape, and a grep alone does not establish the claim — YAML
writes it three ways, so the check has to be
`git grep -nE "restart:[[:space:]]*[\"']?no[\"']?"`, which on `origin/main`
- returns fourteen hits — twelve real declarations across seven files (all
- seven are in the table below), plus two prose matches, in this document and
- in `scripts/compose-drift-watch.py`'s header. What scopes the hazard is that
- only one of the seven is a service Arcane actually starts:
+ returns twenty hits — thirteen real declarations across eight files (all
+ eight are in the table below), plus seven prose matches: four in this
+ document, two in `scripts/arcane-sync-drift-report.py`, one in
+ `scripts/compose-drift-watch.py`'s header. (This bullet's own totals were
+ fourteen / twelve / seven / two / one at the 2026-09-04 round-6 pass.)
+ What scopes the hazard is that
+ only two of the eight are services Arcane actually starts:
| hit | why it cannot trip the hazard |
|---|---|
- | `arcane/home/honeypot-init/compose.yml` (6×) | **this is the exposed one** |
+ | `arcane/home/honeypot-init/compose.yml` (6×) | **this is an exposed one** |
+ | `arcane/home/honeypot-elk/compose.yml:540` | **the other exposed one** — `arkime-pcap-init`, a `chown`/`chmod` one-shot added by #3128, not profile-gated, in a manifest entry Arcane syncs and starts like any other |
| `sandbox/ghosts/compose.yml:162` | `ghosts` *is* an Arcane project, but the service is `ghosts-client-test`, gated behind `profiles: ["test"]`, so a default `up` never creates it |
- | `llm-worker/docker-compose.{production-session-canary,synthetic-canary}.yml` | canary overlays; the manifest deploys `llm-worker/docker-compose.yml`, which has no `restart: no` |
+ | `llm-worker/docker-compose.{production-session-canary,synthetic-canary}.yml` | canary overlays; the manifest deploys `llm-worker/docker-compose.captured-data-deploy.yml`, which has no `restart: no` |
| `docker-compose.sandbox.yml:48` | per-detonation sandbox lifecycle, not an Arcane manifest entry |
- | `vps/docker-compose.yml:956` | the VPS stack, deployed by a different mechanism entirely |
+ | `vps/docker-compose.yml:966` | the VPS stack, deployed by a different mechanism entirely |
| `sandbox/ghosts/vendor/ghosts-src/Ghosts.Api/docker-compose.yml:85` | vendored upstream source, never deployed |
- All six excluded files were checked against
- `arcane/manifests/home-production.json`'s 37 entries and their
+ All seven excluded files were checked against
+ `arcane/manifests/home-production.json`'s 39 entries and their
`dockerComposePath` values, not inferred from the file paths.
`honeypot-init` is not one of #1507's three auto-sync candidates, so this
decision doesn't change its exposure either way: it keeps deploying
@@ -707,6 +781,18 @@ three-puller policy in this round either.** Reasoning:
0755, owned by `github-deploy-runner`, dated 2026-09-03 23:07. #2908 is
closed.
+ **`honeypot-elk` is a #1507 auto-sync candidate, so this one is not
+ hypothetical the way `honeypot-init` is.** Two consequences follow, and
+ neither is established here: whether the live `honeypot-elk` record
+ actually reports `failed` (the #2854 shape is a strong reason to expect
+ it does, but this has not been read off the live API), and, if it does,
+ that `scripts/arcane-sync-drift-report.py`'s `KNOWN_STRUCTURAL_FAILURES`
+ still names only `honeypot-init` — so the report would exit non-zero on a
+ perfectly healthy fleet, which is the exact "permanently red and therefore
+ ignored" failure mode its own comment says the exemption exists to
+ prevent. Worth one `GET /environments/0/gitops-syncs` before #1507's
+ activation is revisited.
+
**What makes Arcane gitops-sync drift visible, since nothing did before:**
`scripts/arcane-sync-drift-report.py` — read-only, on-demand (not wired into
a scheduled workflow: that would need an Arcane API key available to a CI
@@ -734,7 +820,20 @@ permanently-red-and-therefore-ignored failure mode
`scripts/isolation-audit.sh`'s own tiering comment was written against.
Exempted projects are still **printed**, with the reason, under an `EXEMPT`
heading; they are not silenced. Any project that reports `failed` without
-being named there still fails the run.
+being named there still fails the run. `KNOWN_STRUCTURAL_FAILURES` names
+exactly one project, `honeypot-init`; if the `honeypot-elk` exposure
+described under #2858 above turns out to be real, that table needs a second
+entry for the same reason, or a healthy fleet goes permanently red.
+
+**The baseline is hardcoded to `origin/main`.** `scripts/arcane-sync-drift-report.py`
+fetches `origin main` and measures `..origin/main`
+unconditionally — it does not read each record's own `branch`. That is
+correct while every live sync tracks `main` (it still does, per the reads
+above), but it becomes the wrong measure the moment #1507's live-store
+activation re-points rows at `production`: a sync correctly caught up to
+the promoted release would then read as N commits behind `main` for every
+merge since the promotion. Same shape as the `honeypot-elk` gap above —
+worth resolving in the same pass.
Exit-code behaviour was demonstrated in both directions (2026-09-03) by
driving the shipped `main()` with a synthetic record set: a fleet whose only
@@ -752,7 +851,7 @@ owed and worth one look once a current key is to hand.
The fleet has a second one — `deploy.yml`'s rsync into `/opt/stacks/apiary`
(the same inode as `/var/dockge/stacks/apiary`) — that no sync record covers.
That channel had gone 18 days without a successful run, which is what made a
-fleet reading `lastSyncCommit == main` on all 37 records still run a
+fleet reading `lastSyncCommit == main` on all 39 records still run a
weeks-old copy of every rsynced script: `diagnostics.yml`'s
isolation-invariants step executes the *deployed* `isolation-audit.sh` from
that path. Tracked and fixed as **#2908**, now closed — re-derived
diff --git a/docs/ARCHITECTURE.md b/docs/ARCHITECTURE.md
index 478063967..7d69aa789 100644
--- a/docs/ARCHITECTURE.md
+++ b/docs/ARCHITECTURE.md
@@ -25,8 +25,8 @@ rotting again; as of this writing that turns up eight distinct groups
A public VPS terminates attacker traffic — Suricata sniffs it, Traefik
routes HTTP through Keycloak-backed auth, portbridge relays raw protocol
ports — and forwards everything over a home-initiated WireGuard tunnel to
-a homeserver running **31 Arcane-managed sensor/worker/utility stacks**
-(plus 6 more at repository-root paths; 37 sync entries in
+a homeserver running **33 Arcane-managed sensor/worker/utility stacks**
+(plus 6 more at repository-root paths; 39 sync entries in
[`arcane/manifests/home-production.json`](../arcane/manifests/home-production.json),
which is authoritative — not `.github/workflows/deploy.yml`). Sensors
write JSON logs to shared host directories; Filebeat ships them into
@@ -56,7 +56,7 @@ flowchart LR
direction TB
kc["honeypot-keycloak Keycloak + private PostgreSQL"]
init["honeypot-init bootstrap jobs → *.done markers"]
- sensors["Sensor stacks ×22 each its own single-member network"]
+ sensors["Sensor stacks ×20 each its own single-member network"]
tanner["honeypot-tanner SNARE+TANNER+nested Docker"]
elk["honeypot-elk Filebeat · Elasticsearch · Kibana EveBox · Arkime · zeek-proxy"]
dash["honeypot-dashboard (+ -backend) frontend-next · backend-service ×2 worker loops · services-adapter"]
@@ -119,7 +119,6 @@ flowchart TB
subgraph stack["honeypot-dashboard"]
fe["frontend-next :19090 TanStack Start (Node cluster) server functions · SSE hub · BFF cookie"]
- bs["backend-service :8081 Rust axum — the API surface 100+ routes under /api/v1"]
bsm["backend-service-mounted :8082 same route table + host spool mounts write-capable instance"]
loops["backend-worker loops role picked by WORKER_LOOPS: alert-notifier · attacker-identity · agent-intrusion · correlator · dashboard-rollups · threat-intel · zeek-proxy-attribution"]
imp["backend-worker importer es-results-importer, shard-partitionable"]
@@ -129,6 +128,10 @@ flowchart TB
sock[("/var/run/docker.sock")]
end
+ subgraph stackbe["honeypot-dashboard-backend (#1622)"]
+ bs["backend-service :8081 Rust axum — the API surface 100+ routes under /api/v1 + user-retention-sweep · reports-scheduler"]
+ end
+
es[("Elasticsearch")]
analyst -->|"HTTPS"| t --> fe
@@ -147,14 +150,20 @@ Division of labor:
backend access flows through typed server functions — the browser never
speaks to Elasticsearch or sees service tokens. Live updates ride one
shared SSE stream whose frames match the Rust emitter's `event` naming.
-- **backend-service (:8081)** is the unprivileged API tier: constant-time
- service-token middleware, 30s ES timeouts, PIT + `search_after`
- pagination everywhere, CAS writes. It also hosts the embedded worker
- loops — the same image plays each role selected by `WORKER_LOOPS`.
+- **backend-service (:8081)** — the unprivileged API tier, and the only
+ service in the sibling `honeypot-dashboard-backend` stack (#1622 split it
+ out so Arcane can redeploy the API tier without touching `dashboard-next`).
+ Constant-time service-token middleware, 30s ES timeouts, PIT +
+ `search_after` pagination everywhere, CAS writes. It also hosts two
+ embedded worker loops of its own (`user-retention-sweep`,
+ `reports-scheduler`) — the same image plays each role selected by
+ `WORKER_LOOPS`.
- **backend-service-mounted (:8082)** is the same code with the host-side
request-spool mounts (CAPE/Ghidra/GitHub-analysis/GHOSTS/sandbox/
- Windows-sandbox/Rev·Deck). Only this instance can dispatch analysis
- jobs; frontend callers resolve it explicitly via `{mounted: true}`, so
+ Windows-sandbox/Rev·Deck), and it lives in `honeypot-dashboard` itself,
+ not in the sibling stack — the name that says "mounted" is the one that
+ carries the mounts. Only this instance can dispatch analysis jobs;
+ frontend callers resolve it explicitly via `{mounted: true}`, so
capability follows configuration, not URL guessing.
- **Worker containers**: importer mirrors root-owned result spools into
`*-analysis-v1` indices (read-only, never writes back — local JSON stays
@@ -182,9 +191,9 @@ flowchart TB
markers[("state/init-markers/*.done")]
loginit & esinit & arkinit & snareclone --> markers
- subgraph sg["Sensor stacks ×21 (isolated networks)"]
+ subgraph sg["Sensor stacks ×20 (isolated networks)"]
direction LR
- cow["cowrie"] & dion["dionaea+tftp"] & conp["conpot ×6"] & rest["dnp3 · dicompot · dns · citrix cisco-asa · rdp · endlessh · http/api multipot · mailoney · beelzebub · hellpot elasticpot · galah · sentrypeer canarytokens"]
+ cow["cowrie"] & dion["dionaea+tftp"] & conp["conpot ×6"] & rest["dnp3 · dicompot · dns · citrix cisco-asa · sonicwall-sma · rdp · endlessh http/api · multipot · mailoney · beelzebub hellpot · elasticpot · galah · sentrypeer canarytokens"]
end
logsT[("logs/<sensor>")]
@@ -216,7 +225,7 @@ entrypoint instead — the cross-stack readiness contract documented in
## Event ingestion (summary)
The full pipeline — PROXY-aware vs tunnel-blind sensor split, the
-ingest-time `via_port` join, the 12-step `geoip-honeypot` processor chain,
+ingest-time `via_port` join, the 14-processor `geoip-honeypot` chain,
and the dashboard's four read paths — is
[PIPELINES.md §1](PIPELINES.md#1-event-ingestion). Facts that shape
everything else:
diff --git a/docs/BACKUP-ESSENTIALS.md b/docs/BACKUP-ESSENTIALS.md
index 02ca18445..ffb32d735 100644
--- a/docs/BACKUP-ESSENTIALS.md
+++ b/docs/BACKUP-ESSENTIALS.md
@@ -17,14 +17,14 @@ for restoring onto a replacement host see
| | |
|---|---|
-| `homeserver/env/*.env` | all 41 Arcane/Dockge stack `.env` files |
+| `homeserver/env/*.env` | one file per Arcane/Dockge stack (40 under `/var/dockge/stacks/` as of 2026-09-27) |
| `homeserver/secrets/` | secret files kept beside a stack rather than in its `.env` |
| `homeserver/wireguard/` | `wg0.conf` including the private key |
| `homeserver/installer/` | `install-homeserver.conf` — the installer's answers file, which exists only on the root filesystem a reinstall wipes |
| `homeserver/technitium/` | hand-maintained Technitium DNS config |
| `homeserver/keycloak/keycloak.sql.gz` | `pg_dump` of the identity DB — realm, clients, client secrets, users |
| `homeserver/es-operator-state/` | the dashboard's operator-authored Elasticsearch documents, as mapping + NDJSON — see [Operator state](#operator-state) |
-| `homeserver/volumes/` | `dashboard-state`, `arcane-data`, `evebox-config`, `canarytokens-redis-data`, `es-importer-state` |
+| `homeserver/volumes/` | `dashboard-state`, `honeypot-arcane_arcane-data`, `honeypot-elk_evebox-config`, `honeypot-canarytokens_canarytokens-redis-data`, `honeypot-dashboard_es-importer-state` — the Arcane-prefixed names are the real volume names |
| `vps/env/vps.env`, `vps/secrets/`, `vps/traefik/`, `vps/wireguard/` | the VPS's entire config surface, including the Traefik origin certificates |
| `*/manifest/` | host reference notes — disks, volumes, containers, WireGuard, nftables |
| `repo/docs/`, `repo/scripts/`, `repo/analysis/` | this repository's runbooks and operational scripts |
@@ -133,7 +133,7 @@ Three locations, all written by the workstation, which is the backup host:
|---|---|---|---|
| 1 | `/run/media/xore//apiary-backups` | ext4 (Crucial X8 USB) | udisks auto-mount — only present while plugged in |
| 2 | `~/apiary-backups` | XFS (internal) | always available |
-| 3 | `homeserver:/mnt/usb-recovery/apiary-backups` | ext4 (Samsung T7, label `APIARY-BACKUP`) | mounted from fstab by UUID with `nofail` |
+| 3 | `homeserver:/mnt/usb-recovery/apiary-backups` | ext4 (Samsung PSSD T7, label `APIARY-BACKUP`) | mounted from fstab by UUID with `nofail` |
Location 3 was a Ventoy stick formatted exfat until 2026-08-23, mounted
read-only and absent from `/etc/fstab` — so every write to it failed and it
@@ -279,7 +279,12 @@ repository — `install-homeserver.conf.example` carries only placeholders.
`vps/secrets/oidc/`.
5. **Volumes.** For each `homeserver/volumes/.tar.gz`, with the stack
stopped, create the volume and unpack into it through a networkless
- container:
+ container. `` is the **full real volume name** — the archive is
+ written as `$volume.tar.gz` by `backup-essentials.sh`, so four of the
+ five carry their Arcane project prefix
+ (`honeypot-arcane_arcane-data.tar.gz`, and so on). Creating a
+ short-named `arcane-data` volume instead would restore into a volume
+ no stack is mounted against.
```bash
docker volume create
docker run --rm --network none -v :/dst -v "$PWD/homeserver/volumes:/src:ro" \
@@ -376,7 +381,9 @@ gone, for two reasons that happen to point the same way:
Also found and worth knowing: `honeypot-keycloak/.env` carries a full set of
`RESTIC_*` variables pointing at `/mnt-2/apiary-keycloak`, but that repository
-directory does not exist, its password file (`secrets/restic-password`) does
+directory does not exist — and as of 2026-09-27 neither does `/mnt-2` itself,
+which has been decommissioned, so the path cannot start working by accident.
+Its password file (`secrets/restic-password`) does
not exist, `restic` is not installed on the homeserver and no unit references
it. It is dead configuration — no Keycloak restic backup has ever run. The
`keycloak.sql.gz` dump in both scripts here covers that gap.
diff --git a/docs/CGNAT-DEPLOYMENT.md b/docs/CGNAT-DEPLOYMENT.md
index 948f079b8..ce81c3f86 100644
--- a/docs/CGNAT-DEPLOYMENT.md
+++ b/docs/CGNAT-DEPLOYMENT.md
@@ -22,13 +22,19 @@ flowchart TD
validation), `honeypot-elk`, `honeypot-cowrie`, `honeypot-dionaea`,
`honeypot-conpot`, `honeypot-dnp3`, `honeypot-http`, `honeypot-multipot`,
`honeypot-payload-analysis`, `honeypot-tanner`, `honeypot-dashboard`,
- `honeypot-utilities`, the standalone honeypots (cisco-asa, citrix, rdp,
- dicompot, dns-honeypot, endlessh, beelzebub, hellpot, elasticpot, galah,
+ `honeypot-dashboard-backend`, `honeypot-utilities`, the standalone
+ honeypots (cisco-asa, citrix, rdp, sonicwall-sma, dicompot, dns-honeypot,
+ endlessh, beelzebub, hellpot, elasticpot, galah,
sentrypeer, mailoney, canarytokens), `honeypot-keycloak`, and the
- workers (ip-enrichment, agent-intrusion, attacker-identity, correlator,
- payload-inventory) — 31 stacks total, one Arcane-managed directory each
+ workers (agent-intrusion, attacker-identity, correlator,
+ payload-inventory) — 32 stacks total, one Arcane-managed directory each
under `arcane/home//` (`honeypot-wordpot` sat here until #2381
- retired it). `honeypot-init` still deploys first; every
+ retired it; `ip-enrichment-worker` was in that worker list until
+ f139fe24 retired the Go service, `honeypot-dashboard-backend` joined when
+ #1622 split it out of `honeypot-dashboard`, and `honeypot-sonicwall-sma`
+ when #3131 added it — re-counted 2026-09-27 against
+ `arcane/manifests/home-production.json`, which holds 33 in-tree entries:
+ these 32 plus `unsloth`). `honeypot-init` still deploys first; every
sensor stack waits on its completion markers at its own entrypoint rather
than a Compose-level dependency, same reasoning as before, just across
more projects now. See `docs/STACK-REBUILD.md` for the full current list
@@ -39,7 +45,7 @@ flowchart TD
- VPS: plain Docker Compose manages `/root/vps/docker-compose.yml`.
Unchanged by #1502 — VPS deployment stays outside Arcane entirely, as
that issue's own scope decision.
-- Each of the 31 migrated stacks' Compose source (build context, git-tracked config,
+- Each of the 33 in-tree stacks' Compose source (build context, git-tracked config,
`compose.yml` with an explicit top-level `name:` pinned to its live
project name) lives self-contained under `arcane/home//` in this
repository. Arcane clones the repo and materializes the *entire
@@ -52,25 +58,31 @@ flowchart TD
these syncs on a from-scratch install, driven by the single source of
truth at `arcane/manifests/home-production.json`. Six more home-hosted
stacks (`auth-events-worker`, `llm-worker`, `ml-worker`,
- `analysis/ghidra`, `sandbox/ghosts`, `pihole`) are Arcane-managed too but
+ `analysis/ghidra`, `sandbox/ghosts`, `technitium`) are Arcane-managed too but
were already self-contained, so they kept their existing repository-root
path instead of moving. Three of those six (`auth-events-worker`,
`llm-worker`, `ml-worker`) are also imported by
`step_arcane_import_stacks` itself now (#1505 — confirmed to have no
host-local state beyond `.env`); the other three keep their own dedicated
- installer steps for reasons specific to each (`pihole`'s non-`.env` host
- state, `analysis/ghidra`'s conditional GPU compose overlay, and
+ installer steps for reasons specific to each (`technitium`'s non-`.env` host
+ state — the step `pihole` had until #2911 swapped the two — `analysis/ghidra`'s conditional GPU compose overlay, and
`sandbox/ghosts`'s confirmed Arcane build-context limitation, #1506) —
see `scripts/install-homeserver.sh`'s own Phase 8 header comment for the
- full reasoning behind each.
+ full reasoning behind each. That filter reaches 35 of the manifest's 39
+ entries; `unsloth` is the one the installer reaches neither way (no
+ `step_unsloth_*`, and not a `honeypot-*` name), by design — see
+ `docs/ARCANE-GIT-SYNC.md`'s "Manifest import".
- The public gateway source is under `vps/`.
Arcane is used only on the home server. The VPS uses `docker compose` directly.
See `docs/ARCANE-GIT-SYNC.md` for the sync model, cutover procedure, and
-confirmed Arcane v2.8.0 platform limitations (a required compose variable
+confirmed Arcane platform limitations (a required compose variable
in a port-binding position, remote build contexts pinned to a Git tag, the
sync file-count limit, and stale project records after a `destroy` call
-all have confirmed workarounds documented there).
+all have confirmed workarounds documented there). Those were each confirmed
+against `v2.8.0`–`v2.9.0`; `docker-compose.arcane.yml` now pins
+`manager:v2.11.1`, and none of them has been re-confirmed against that
+image — its own section header says to re-verify on upgrade.
## WireGuard addressing
@@ -135,12 +147,15 @@ the only internet-facing component.
8. Run `python3 analysis/verify-stack.py` (with `DASHBOARD_SERVICE_TOKEN`
from `honeypot-dashboard/.env`) and inspect `/source-health`.
-Each stack is a folder under your Arcane stacks dir (default `/opt/stacks/`).
-Upload the whole home folder via SFTP — compose **and** the build
-sub-folders (`cowrie/`, `multipot/`, `http-honeypot/`, `dashboard/`, …) —
-since Arcane's own editor only edits the compose file. After editing Go
-source or honeyfs content, rebuild from the `APIARY` stack's Arcane
-**terminal**: `docker compose -f compose.yml up -d --build`.
+Each stack is a folder under your Arcane stacks dir (`/var/dockge/stacks`;
+`/opt/stacks` is a symlink to it, #1185). **Since #1502 nothing is uploaded
+by SFTP** — Arcane materializes each stack's whole directory from its Git
+sync, and `honeypot-wordpot` aside the source of truth is the repository, not
+a hand-copied folder. The SFTP-upload and "edit then rebuild from Arcane's
+terminal" instructions this paragraph used to give are part of the pre-#258
+model the callout above already flags; what replaces them is a commit plus a
+sync, and a separate `POST /projects/{id}/build` for the stacks that have a
+`build:` service — see `docs/ARCANE-GIT-SYNC.md`.
### Boot-safe home networking and VPS log mounts
@@ -263,11 +278,12 @@ template is in [`vps/traefik/dynamic.yml`](../vps/traefik/dynamic.yml):
`honeypot-http` (`decoy.`) + `honeypot-web` (catch-all) → fake nginx,
`honeypot-snare` (`www-portal.` and `snare.`) → SNARE, one
native-OIDC route for the dashboard (no gateway, since #1026), one native-OIDC
-route for Arcane (no gateway, #1185), and six forward-auth-protected
+route for Arcane (no gateway, #1185), and six gateway-fronted
investigation routes sitting behind their own Keycloak-backed `oauth2-proxy`
gateway: Kibana, TANNER, EveBox, Arkime, Rev·Deck, and the Traefik dashboard
-itself. Each has a matching
-`socat-hp-*` bridge in [`vps/docker-compose.yml`](../vps/docker-compose.yml).
+itself. Five of the six have a matching
+`socat-hp-*` bridge in [`vps/docker-compose.yml`](../vps/docker-compose.yml);
+the Traefik dashboard's gateway terminates on the VPS, not through one.
Traefik is an HTTP(S) reverse proxy — it adds TLS, per-subdomain routing and
auth to the web honeypots and dashboards. The other protocols (SSH, SMB,
@@ -325,14 +341,22 @@ Cloudflare answers every proxied hostname with **526**.
`deploy.yml` never overwrites `traefik/certs/`. On a normal deploy it only
checks that `origin.pem` still parses.
-### The forward-auth bridge, generically
+### The gateway-fronted chain, generically
Six investigation UIs (Kibana, TANNER, EveBox, Arkime, Rev·Deck,
the Traefik dashboard) reach home through the identical chain — one pattern,
-six routers in `vps/traefik/dynamic.yml`, six `socat-hp-*` bridges, each
-fronted by its own Keycloak-backed `oauth2-proxy` gateway container, not
-six different mechanisms. The honeypot dashboard and Arcane are the two
-exceptions — both speak native OIDC directly, no gateway — see the note below.
+six routers in `vps/traefik/dynamic.yml`, six `oauth2-proxy` gateway
+containers, five `socat-hp-*` bridges, not six different mechanisms. The
+gateway *is* the router's upstream rather than a `forwardAuth` middleware:
+`honeypot-kibana`'s `service:` is a loadBalancer at `http://oidc-kibana:4180`,
+and that container's `OAUTH2_PROXY_UPSTREAMS` is the socat bridge
+(`http://socat-hp-kibana:5601`). `grep -c forwardAuth vps/traefik/dynamic.yml`
+is 0, so nothing in the config uses the forward-auth middleware form. The
+sixth gateway, `oidc-traefik`, has no socat hop at all — its upstream is
+Traefik's own dashboard (`http://traefik:8081`) on the VPS, which is why the
+bridge count is five and not six. The honeypot dashboard and Arcane are the
+two exceptions — both speak native OIDC directly, no gateway — see the note
+below.
```mermaid
sequenceDiagram
@@ -340,7 +364,7 @@ sequenceDiagram
actor Op as operator's browser
participant CF as Cloudflare (proxied DNS)
participant TR as Traefik (TLS termination + routing)
- participant OA as oauth2-proxy (forward-auth, one per service)
+ participant OA as oauth2-proxy (gateway, one per service)
participant KC as Keycloak (honeypot-keycloak, at home)
participant SOC as socat-hp-* (VPS container)
participant WG as WireGuard tunnel
@@ -348,14 +372,13 @@ sequenceDiagram
Op->>CF: HTTPS request, e.g. kibana.
CF->>TR: proxied, real client IP in X-Forwarded-For
- TR->>OA: forward-auth check
+ TR->>OA: routed to the app's own gateway (oidc-kibana:4180)
alt no valid session
OA-->>Op: redirect to Keycloak login (auth.)
Op->>KC: authenticate (password + mandatory TOTP)
KC-->>OA: OIDC callback, session established
end
- OA-->>TR: identity headers
- TR->>SOC: request, security-headers applied
+ OA->>SOC: proxied request, identity headers added
SOC->>WG: raw TCP, VPS listen port → 10.8.0.2:home-exposed-port
WG->>APP: delivered to the app's own internal port
APP-->>Op: response, relayed back through the same chain
diff --git a/docs/CI-CD.md b/docs/CI-CD.md
index ab87058ef..6c725f946 100644
--- a/docs/CI-CD.md
+++ b/docs/CI-CD.md
@@ -97,15 +97,18 @@ flowchart TB
containersHome["containers.yml — all image builds"]
securityHome["security.yml — all CodeQL languages"]
pagesHome["pages.yml artifact build"]
+ imageScanHome["image-security-scan.yml"]
end
prPush --> quality
prPush -->|"PR: build only, never published"| containerBuild
+ prPush --> codeql
+ prPush --> pages
mainPush --> quality
mainPush --> containerBuild
mainPush --> codeql
mainPush --> pages
- mainPush -.->|"every workflow's compute jobs — only after passing the ci-router trust gate + heartbeat; pull_request needs repo variable CI_HOMESERVER_PRS, and forks can never qualify"| ciSelfHosted
+ mainPush -.->|"those five workflows' ci-target jobs — only after passing the ci-router trust gate + heartbeat; pull_request needs repo variable CI_HOMESERVER_PRS, and forks can never qualify"| ciSelfHosted
```
**`honeypot-ci` does not see `pull_request` by default, by design.** A
@@ -114,9 +117,11 @@ job; a self-hosted runner's job runs as a real process on real
home-network infrastructure. A malicious test file in an unreviewed PR
(`os.system(...)`, a crafted Go `TestMain`) would execute wherever that
runner has access — the same reasoning `production-home`'s own deployment
-runner (below) already applies. Every workflow's executor routing (each
-caller's own `ci-target` job, which since #2571 always calls the shared
-`.github/workflows/ci-router.yml`) trusts
+runner (below) already applies. Executor routing is a caller-side job named
+`ci-target`, which since #2571 calls the shared
+`.github/workflows/ci-router.yml`. Five of the repo's twenty workflows have
+one: `quality.yml`, `containers.yml`, `security.yml`, `pages.yml` and
+`image-security-scan.yml`. It trusts
push-to-main (already reviewed and merged), the `schedule` and
`workflow_dispatch` (an operator's own machinery); same-repo pull
requests need the repository variable `CI_HOMESERVER_PRS=true`, and fork
@@ -558,10 +563,14 @@ stack on the host lives entirely in Arcane's Git-sync machinery — see
[ARCANE-GIT-SYNC.md](ARCANE-GIT-SYNC.md) for the full contract (its
non-obvious cornerstones: creating a sync *is* an initial deploy, a sync
materializes files without redeploying — live `redeploy_after_sync`
-defaults to 0, though the manifest schema cannot express it — and every
-synced stack runs `autoSync: false`: the #1507 tag-promotion /
-`production`-pointer policy was decided but never deployed, so all syncs
-track `main` and deploys are manual; ARCANE-GIT-SYNC.md's promotion
+defaults to 0, though the manifest schema cannot express it — and deploys
+are manual. #1507's tag-promotion / `production`-pointer policy was only
+half activated: #1943 (2026-08-25) put `branch: "production"` and three
+`autoSync: true` flags into `arcane/manifests/home-production.json`, so all
+39 entries now name that branch rather than `main` — but the live store
+still reads `auto_sync = 0` on every row, and `origin` has no
+`refs/heads/production` for the pointer to name, so nothing follows a
+promotion and deploys stay manual; ARCANE-GIT-SYNC.md's promotion
section carries the live-state evidence).
This workflow deliberately stopped touching those directories entirely:
running an rsync/build loop alongside Arcane's own sync would put two
@@ -718,10 +727,9 @@ alert/intelligence history in the old one is gone.
Everything else that was still monolithic as of the earlier revision of
this section (`dionaea`, `payload-dedupe`, `yara-scanner`, and the Tanner
-group) has since split out too -- see the `honeypot-dionaea` and
-`honeypot-payload-analysis` section below; only the Tanner group remains in
-`APIARY`, as part of its own internal `depends_on` chain not yet
-worth splitting.
+group) has since split out too -- see the `honeypot-dionaea`,
+`honeypot-payload-analysis` and `honeypot-tanner` sections below. Nothing
+remains in `APIARY`: the root `docker-compose.yml` is `services: {}`.
#### Dashboard redeploy (single replica; #266 rolling pair retired, #1659 legacy `dashboard` removed)
@@ -872,13 +880,26 @@ open handles into the log directories this script wipes for this target.
### honeypot-elk (#258)
`arcane/home/honeypot-elk/compose.yml` bundles the ELK/analysis plane (`elasticsearch`,
-`kibana`, `filebeat`, `evebox`, `arkime-capture`, `arkime-viewer`,
-`pcap-sync`) into one stack at `/opt/stacks/honeypot-elk` -- the last group
+`kibana`, `filebeat`, `evebox`, `pcap-sync`, `arkime-pcap-init`,
+`arkime-capture`, `arkime-viewer`, `extracted-file-importer`, `zeek-proxy`)
+into one stack at `/opt/stacks/honeypot-elk` -- the last group
that was still in the monolithic file. Kept together, not split further:
-all seven sit on the shared `honeynet` network and either read from or
-write to the one Elasticsearch instance, so splitting them apart would
-turn every one of those relationships into a cross-stack shared resource
-for services that only ever make sense running together.
+they share the one Elasticsearch instance, the `arkime-pcap` volume and the
+host's `logs/` bind-mount tree, so splitting them apart would turn every one
+of those relationships into a cross-stack shared resource for services that
+only ever make sense running together.
+
+The shared-network story is narrower than it looks, and the count above is
+not "ten on `honeynet`". Seven of the ten declare `honeynet` explicitly
+(`elasticsearch`, `kibana`, `filebeat`, `evebox`, `arkime-capture`,
+`arkime-viewer`, `extracted-file-importer`); `elasticsearch` is additionally
+on `llm-data`. `pcap-sync` and `arkime-pcap-init` declare no `networks:` at
+all and so ride the project's implicit default network -- `pcap-sync` moves
+rotated pcaps through host bind-mounts and a marker file, and
+`arkime-pcap-init` is a one-shot `chown` of the `arkime-pcap` volume, so
+neither needs the shared network. `zeek-proxy` sets `network_mode: host`
+outright and reaches the sensor plane through host-published ports, which is
+the point of it.
`honeynet` and `llm-data` get the usual explicit shared `name:` treatment.
`es-data` does **not**, despite appearances: `honeypot-init`'s
@@ -1164,8 +1185,9 @@ The install also drops two things next to the unit:
The leading `+` runs that line as root even though the unit's own `User=`
is the unprivileged runner account, so no new sudoers grant was needed
-(unlike `compose-project-state.py` above, this runs as part of the unit's
-own privileged startup rather than from inside a workflow step).
+(unlike `scripts/compose-project-state.py`, the narrow root helper for
+`compose-drift-watch.py`, this runs as part of the unit's own privileged
+startup rather than from inside a workflow step).
**Why it exists.** A root process that writes into a runner's `_work`
checkout leaves files the runner user can never delete, and
@@ -1368,7 +1390,7 @@ The `Pick cache backend` step therefore chooses per executor:
the runner can actually write it, so a rebuild replay (#1609) recreates it
rather than leaving a hand-made directory nobody records. If the step has
not run on a given box, `Pick cache backend` emits a workflow warning and
-falls back to `type=gha` -- a slow build, not eighteen failed matrix rows.
+falls back to `type=gha` -- a slow build, not nineteen failed matrix rows.
**Bounding it.** `type=local` has *no* eviction: every export leaves
unreferenced blobs behind in `blobs/sha256/` forever.
@@ -1592,9 +1614,11 @@ Honest limitations:
`docker/login-action` targets `ghcr.io` and is gated
`if: github.event_name != 'pull_request'`, so on a PR every base-image pull
went out anonymous -- and Docker Hub meters anonymous pulls **per source
-IP**, at roughly 100 per 6h. With `CI_HOMESERVER_PRS=true` all 18 matrix
-rows leave this box through one address, and the tree carries **74
-non-`scratch` Hub `FROM` lines**. One cold run spends most of the budget;
+IP**, at roughly 100 per 6h. With `CI_HOMESERVER_PRS=true` all 19 matrix
+rows leave this box through one address, and the tree carries **75
+non-`scratch` `FROM` lines** across 56 tracked Dockerfiles, 68 of them
+resolving to Docker Hub (the other 7 are `mcr.microsoft.com`). One cold run
+spends most of the budget;
the run after it fails with `toomanyrequests` on whichever rows happen to
ask last. #2771's per-image `type=gha` scopes do not help: that cache holds
*our* layers, never the base image, so every run re-resolves every `FROM`
@@ -1709,9 +1733,9 @@ flowchart TB
checkout["actions/checkout"]
key["VPS_SSH_KEY written to a temp file, mode 0600"]
backup[("Snapshot: /root/vps-backups/ pre-deploy-<timestamp>.tar.gz, 10 most recent kept")]
- rsync["rsync vps/ -> /root/vps/ over SSH, excluding .env, traefik/certs/, traefik/dynamic.yml (VPS-owned, see table below)"]
+ rsync["rsync vps/ -> /root/vps/ over SSH, excluding .env, traefik/certs/, traefik/dynamic.yml, secrets/ (VPS-owned, see table below)"]
validate["SSH: docker compose config validates /root/vps/docker-compose.yml"]
- up["SSH: docker compose up -d --build"]
+ up["SSH: docker compose up -d --build --remove-orphans (#2813)"]
dynGen["Separate step: substitute DOMAIN into the committed *.honeypot.example placeholders, validate as YAML, no leftover placeholders -- all BEFORE touching the VPS"]
dynWrite["Copy to a temp path on the VPS, then write in place with cat -- never copy-then-rename (see below: Traefik's bind mount tracks the inode, not the path)"]
verify["Verify step: fail the job if certs or dynamic.yml are missing, empty, unparseable, or still placeholder"]
@@ -1739,10 +1763,14 @@ The VPS job runs on a short-lived GitHub-hosted Ubuntu runner:
`/root/vps-backups/pre-deploy-.tar.gz`, keeping the ten most
recent archives.
5. `rsync` sends only the repository's `vps/` directory over SSH to
- `/root/vps/`, **excluding** `traefik/dynamic.yml` (see below).
+ `/root/vps/`, **excluding** the four VPS-owned paths in the table below
+ (`.env`, `traefik/certs/`, `traefik/dynamic.yml`, `secrets/`).
6. A second SSH command runs on the VPS, validates
`/root/vps/docker-compose.yml`, and executes
- `docker compose up -d --build`.
+ `docker compose up -d --build --remove-orphans`. The flag matters: a
+ service removed from `docker-compose.yml` otherwise leaves its container
+ running forever, which is how #2813 found `socat-hp-wordpot` still up a
+ week after #2469 retired wordpot and dropped its forwarder rule.
7. A dedicated step generates the deployable `traefik/dynamic.yml` --
substitutes `DOMAIN` for every `*.honeypot.example` placeholder in the
committed template -- and validates the result (parses as YAML, no
@@ -1757,7 +1785,7 @@ The VPS job runs on a short-lived GitHub-hosted Ubuntu runner:
### Files the VPS owns, not the repository
`--delete-delay` removes destination files that no longer exist under the
-repository's `vps/` directory, and overwrites the ones that do. Three paths are
+repository's `vps/` directory, and overwrites the ones that do. Four paths are
therefore excluded from the main `rsync` because the VPS copy is authoritative
(or, for `dynamic.yml`, because it needs different handling entirely):
@@ -1766,6 +1794,11 @@ therefore excluded from the main `rsync` because the VPS copy is authoritative
| `.env` | Secrets and host-specific values. |
| `traefik/certs/` | Issued TLS certificates. They do not exist in the repository, so an unexcluded `--delete-delay` deletes them, and the workflow cannot reissue them. |
| `traefik/dynamic.yml` | Carries the deployment's real domain. The committed copy is a `*.honeypot.example` placeholder -- Traefik's file provider has no `${VAR}`-style substitution the way docker-compose already gives every other host-specific value in this repo, so this file can't just be templated in place the normal way. Deployed by its own dedicated step instead (step 7 above), which substitutes `DOMAIN` and writes the result separately. |
+| `secrets/` | Per-gateway OIDC cookie-secret/client-secret files (`OIDC_SECRETS_DIR=./secrets/oidc` in `vps/.env.example`). Git-ignored, so they are never present in the checkout at all -- and `--delete-delay` reads "absent from the source" as "delete it". That happened once for real and took down every oauth2-proxy gateway at the time, which is why the exclude exists. A client-secret is never regenerated: it has to match what is already registered with Keycloak, so recovering it means `kcadm get clients//client-secret`, and losing it means re-registering the client. |
+
+The pre-deploy backup step ahead of the bulk `rsync` archives the same four
+paths (`.env`, `traefik/certs`, `traefik/dynamic.yml`, `secrets/`) that
+actually exist, keeping the last ten under `/root/vps-backups/`.
The certificates were lost once, in a single `target: both` run before that
exclusion existed: Traefik fell back to self-signed and every router silently
@@ -1822,8 +1855,8 @@ scratch twice, reaching the same blocker both times.
**No workflow edit is needed.** Every `secrets.VPS_*` / `secrets.DOMAIN`
reference already sits inside a job that declares
-`environment: production-vps` — `deploy.yml`'s `vps` job (`:233`, environment
-at `:236`), `diagnostics.yml`'s `vps` job (`:439`/`:442`), and
+`environment: production-vps` — `deploy.yml`'s `vps` job (`:231`, environment
+at `:234`), `diagnostics.yml`'s `vps` job (`:292`/`:295`), and
`vps-start-blackhole.yml`'s `start-blackhole-profile` job (`:22`/`:24`). The
`home` jobs (`deploy.yml:21`, `diagnostics.yml:112`) read none of the five.
Environment secrets also shadow repository secrets of the same name, so
@@ -1852,7 +1885,7 @@ source rather than the password manager, rotate it deliberately rather than
as a side effect of the move.
`DOCKERHUB_USERNAME` / `DOCKERHUB_TOKEN` (added 2026-09-01) are read only by
-`containers.yml:157-160`, which declares **no** `environment:` at all — so
+`containers.yml:157-162`, which declares **no** `environment:` at all — so
they have to stay repository-scoped until that workflow gains one, and they
are correctly out of this migration's scope rather than merely deferred. They
are still the reason the repository-secret set grew from five to seven, which
@@ -1998,7 +2031,10 @@ diagnostics workflow itself.
Home:
GitHub -> outbound-polling self-hosted runner on homeserver
-> local rsync /opt/stacks/apiary
- -> Arcane compose.yml -> docker compose up
+ -> compose config --quiet (validation only, no `up`;
+ the root file is services: {})
+ -> Arcane's own Git-sync machinery materializes and
+ deploys stacks (ARCANE-GIT-SYNC.md)
VPS:
GitHub-hosted runner -> rsync + SSH over VPS_PORT
@@ -2006,6 +2042,11 @@ GitHub-hosted runner -> rsync + SSH over VPS_PORT
-> docker compose up on VPS
```
+The home path has not run a deploy in this workflow since #1502 — see
+"What deploy.yml actually runs (since #1502)" above for the full list of
+what it does instead. A `home` run that succeeds has still changed nothing
+on the host.
+
Selecting `both` creates both jobs from the same workflow run. They share the
`honeypot-production` concurrency group, but the home and VPS jobs are
otherwise independent: one can fail while the other succeeds. Always inspect
diff --git a/docs/DASHBOARD-CUTOVER.md b/docs/DASHBOARD-CUTOVER.md
index fd65fbe64..bafd9b99c 100644
--- a/docs/DASHBOARD-CUTOVER.md
+++ b/docs/DASHBOARD-CUTOVER.md
@@ -17,20 +17,21 @@ Status at drafting: first runbook for #1628. Unlike
end-to-end — refine it against what actually happens the first time it's
run, the way that doc was.
-**Confirmed live against the homeserver (2026-08-20): the `next` profile
-has never been activated.** `honeypot-dashboard`'s Arcane project reports
+**Pre-cutover snapshot, kept for the record — 2026-08-20, two days before
+the completion banner above made this true: the `next` profile
+had never been activated.** `honeypot-dashboard`'s Arcane project reported
exactly 4 running services — `dashboard`, `oidc-sessions`,
`es-results-importer`, `services-adapter`, all legacy — and zero
-`hp-apiary-*` containers exist anywhere on the host. `dashboard-next`,
+`hp-apiary-*` containers existed anywhere on the host. `dashboard-next`,
`backend-service`, and every Rust worker (`backend-worker`,
`backend-worker-importer`, `backend-worker-enrichment`,
-`backend-service-mounted`) have never run in production. This means:
-no bake period has started for any worker, so none of #1628's
-worker-retirement decisions can move to "retire" yet regardless of how
-much parity testing has landed in CI — that testing proves the Rust
+`backend-service-mounted`) had never run in production. This meant:
+no bake period had started for any worker, so none of #1628's
+worker-retirement decisions could move to "retire" yet regardless of how
+much parity testing had landed in CI — that testing proves the Rust
implementations are *correct*, not that they've *run* against real
-production load. The actual next step, once #1628's remaining ops-
-blocker items are resolved, is step 3 below (`cutover-dashboard.sh
+production load. The next step at the time, once #1628's remaining ops-
+blocker items were resolved, was step 3 below (`cutover-dashboard.sh
preflight`) for the very first time — not any worker's retirement.
Tracking issue for everything this cutover depends on:
@@ -49,8 +50,9 @@ port 19090, fronted by VPS Traefik's `honeypot-dashboard` router
(`vps/traefik/dynamic.yml`). Background loops (`notifyLoop`,
`reportScheduleLoop`) run inside this same binary/service.
-**New:** three tiers, all currently gated behind the `next` Compose
-profile so nothing binds a host port or receives traffic until cutover:
+**New:** three tiers, all originally gated behind the `next` Compose
+profile so nothing bound a host port or received traffic until cutover
+(the profile is gone now — see the status banner):
- `dashboard-next` — TanStack Start frontend/BFF (Arcane stack
`honeypot-dashboard`, same stack as the old `dashboard` service, for now)
@@ -121,7 +123,9 @@ retirement calls for you.
cover `/healthz`; still manually confirm a handful of golden-path
pages SSR correctly, `/api/live` streams, and login redirects to
Keycloak and completes. Run `port-tests/{backend-api,frontend-ssr,
- auth-flow}.sh` against this live instance if not already fresh.
+ auth-flow}.sh` against this live instance if not already fresh. (The
+ whole `port-tests/` directory is gone from the tree today — this step
+ is history, and the current equivalent is `.github/workflows/quality.yml`.)
5. **Re-point Traefik** — only if this cutover is ever cross-host; in
the current single-host topology the VPS-side `socat-hp-dashboard`
forward already points at a fixed home address
diff --git a/docs/DECEPTION-EXTENSIONS.md b/docs/DECEPTION-EXTENSIONS.md
index 14255209f..57fcb72d2 100644
--- a/docs/DECEPTION-EXTENSIONS.md
+++ b/docs/DECEPTION-EXTENSIONS.md
@@ -109,6 +109,7 @@ traces back to at least one decision.
| RDP decoy (rdphoneypot lineage) | Integrated | `arcane/home/honeypot-rdp-honeypot/` | #238 batch, per-decoy plan #412 |
| Cisco ASA VPN gateway decoy | Integrated | `arcane/home/honeypot-cisco-asa-honeypot/` | #238 batch, CVE context #414 |
| Citrix ADC gateway decoy | Integrated | `arcane/home/honeypot-citrix-honeypot/` | #238 batch, CVE context #414 |
+| SonicWall SMA1000 Work Place/AMC decoy | Integrated | `arcane/home/honeypot-sonicwall-sma/` | #3033; CVE-2026-83548 SSRF / CVE-2026-83549 AMC command-injection chain |
| DNS amplification bait | Integrated | `arcane/home/honeypot-dns-honeypot/` | #238 batch, safety-sensitive design #415 |
| Mailoney (SMTP) | Integrated | `arcane/home/honeypot-mailoney/` | #1422 |
| SentryPeer (VoIP/SIP) | Integrated | `arcane/home/honeypot-sentrypeer/` | #1424 |
diff --git a/docs/ES-CONSUME-PATTERNS.md b/docs/ES-CONSUME-PATTERNS.md
index e52349a51..08d299bed 100644
--- a/docs/ES-CONSUME-PATTERNS.md
+++ b/docs/ES-CONSUME-PATTERNS.md
@@ -121,8 +121,9 @@ analysis/es-consume/
│ # same pages in -> same consumed set + checkpoint out
└── tests/test_es_consume.py # vendoring registry + parity + query-shape contracts
ml-worker/es_consume.py # vendored copy (byte-for-byte asserted)
-attacker-identity-worker/esconsume.go # Go reference engine, behaviourally
- # identical, tested against the same fixtures
+arcane/home/honeypot-attacker-identity-worker/attacker-identity-worker/
+└── esconsume.go # Go reference engine, behaviourally
+ # identical, tested against the same fixtures
```
Tests, run in CI (`.github/workflows/quality.yml`):
diff --git a/docs/GEOIP-THREAT-INTEL.md b/docs/GEOIP-THREAT-INTEL.md
index 406646944..43a42cb54 100644
--- a/docs/GEOIP-THREAT-INTEL.md
+++ b/docs/GEOIP-THREAT-INTEL.md
@@ -51,11 +51,15 @@ worker's own reload/run interval (`threat_intel.rs`), no restart needed.
The deployed stack already works with manually supplied MMDB files. For official
automatic MaxMind updates, set `MAXMIND_ACCOUNT_ID` and
-`MAXMIND_LICENSE_KEY` in Dockge's stack environment, then enable the optional
-profile:
+`MAXMIND_LICENSE_KEY` in the `honeypot-init` stack's `.env` (Arcane's stack
+environment; Dockge was replaced by Arcane per
+[#1185](https://github.com/Xore/APIARY/issues/1185)), then enable the optional
+profile. `geoipupdate` is a service of the `honeypot-init` stack, which syncs to
+`/opt/stacks/honeypot-init`, not to the `/opt/stacks/apiary` checkout the `.mmdb`
+files themselves land in:
```bash
-cd /opt/stacks/apiary
+cd /opt/stacks/honeypot-init
docker compose -f compose.yml --profile geoip-update up -d geoipupdate
```
@@ -70,10 +74,16 @@ The fallback does not provide city, coordinates, ASN, organization, or IPv6.
`country.csv` and downloaded `.mmdb` files are intentionally ignored by Git;
credentials and licensed/generated databases must not be committed.
-MMDB databases and a manually-edited `threat-cidrs.csv` are loaded when
-`hp-dashboard` starts -- restart that container after replacing a database or
-hand-editing the file directly. `threat-cidrs.csv` refreshed by
-`refresh-threat-cidrs.sh` is the one exception: the running dashboard picks
-that up on its own (see above), no restart needed. Geolocation is
+The `.mmdb` files are read by three containers, so replacing one needs all
+three restarted: `hp-elasticsearch` (the `ingest-geoip` mount the
+`geoip-honeypot` processors read), and both Arkime containers —
+`hp-arkime-capture` and `hp-arkime-viewer` (each mounts `/opt/arkime/geo`).
+A fourth container, `hp-geoipupdate` in the `honeypot-init` stack, is the
+writer, not a consumer.
+`threat-cidrs.csv` is mounted read-only into the dashboard's `backend-worker`
+container, so hand-editing it directly needs that one restarted. A
+`threat-cidrs.csv` refreshed by `refresh-threat-cidrs.sh` is the one exception:
+the running worker picks that up on its own reload interval (see above), no
+restart needed. Geolocation is
approximate and must not be treated as proof of an attacker's physical
location.
diff --git a/docs/HOMESERVER-DISK-LAYOUT.md b/docs/HOMESERVER-DISK-LAYOUT.md
index 4f66454ae..37636d1e5 100644
--- a/docs/HOMESERVER-DISK-LAYOUT.md
+++ b/docs/HOMESERVER-DISK-LAYOUT.md
@@ -3,13 +3,22 @@
This documents the physical disk layout of the honeypot homeserver
(`supermicro`) as it actually exists today, and a generated Ubuntu
**autoinstall** config (the Ubuntu/subiquity equivalent of Windows'
-`autounattend.xml`) to reproduce that layout on a reinstall or a second
-build server. Captured 2026-08-04 as part of the #518 smoke-test research.
-
-Ubuntu Server's installer (`subiquity`) is driven by `curtin` under the
-hood — the fstab comments on this box literally say "was on /dev/sdX
-during curtin installation", confirming this machine was already installed
-this way rather than by hand.
+`autounattend.xml`) that reproduces the layout the box had at the time of
+the #518 smoke-test research.
+
+> **The autoinstall config below no longer describes this host.** The
+> physical table and the provisioning steps were re-measured read-only on
+> 2026-09-27. `supermicro` has since been reinstalled as **Rocky Linux
+> 10.2** and no longer runs the Ubuntu/`curtin` layout: it uses LVM (the
+> original notes recorded "no LVM"), `/var` sits on a *partition* of the
+> RAID LUN rather than the whole disk, and swap is a 32G LVM logical
+> volume rather than a swapfile. The original capture was 2026-08-04, when
+> the fstab comments did literally say "was on /dev/sdX during curtin
+> installation" — that evidence was sound for the Ubuntu install, which
+> has since been replaced. Keep the autoinstall file as the record of the
+> Ubuntu layout; do not use it as a rebuild target for the current host.
+> See `docs/HOST-TUNING.md` for the tuning that *does* apply to the Rocky
+> install.
## Why this layout, not one big disk
@@ -22,24 +31,36 @@ reinstall of the OS disk alone doesn't touch captured evidence.
## Physical layout (as installed)
+Re-measured read-only on 2026-09-27 via `lsblk`/`lvs`/`findmnt`/`df`.
+
| Device | Model | Size | Partition table | Filesystem | Mount | Role |
|---|---|---|---|---|---|---|
-| `nvme0n1` | Samsung MZVLW256HEHP | 238.5G | GPT | vfat (p1) / ext4 (p2) | `/boot/efi`, `/` | OS + EFI, boot disk |
-| `sdb` | AVAGO MR9440-8i (RAID LUN) | — | whole-disk (no partition table) | xfs | `/var` | Docker root, Arcane-managed stacks, container state (`/var/lib/docker`, `/var/dockge`) — now also `benchmarks/`, `training/`, `hf-cache/`, `buildx-cache/`, `ci-registry-mirror/`, `image-sbom/`, the former `/mnt-1` workload |
-| `sda` | Intel SSDSC2KB480G8L | 447.1G | GPT, 1 partition | xfs | `/mnt-2` | Reserved bulk storage (currently empty) |
-| `sr0` | ATAPI optical | — | — | — | — | Unused |
-
-**`/mnt-1` is decommissioned.** Its RAID VD (formerly `sdc`) suffered a
-two-drive fault on 2026-09-09 (#3158) and no longer enumerates as a block
-device at all; the mount was unwired (#3159, PR #3159) and everything that
-lived under it moved to `/var`. `/mnt-1` itself survives on the host only as
-a directory of compatibility symlinks into `/var` (`benchmarks`, `training`,
-`hf-cache`, `buildx-cache`, `ci-registry-mirror`) so any script still hard-
-coding the old path keeps resolving — new code should target `/var/*`
-directly. See #3158/#3159 for the incident and decommission detail; this
-table's `sdb` size is left unstated above rather than guessed, since the
-volume backing `/var` changed as part of that recovery and hasn't been
-re-measured for this doc.
+| `nvme0n1` | PC401 NVMe SK hynix 1TB | 953.9G | GPT, 3 partitions | vfat (p1, 600M) / xfs (p2, 2G) / LVM2_member (p3, 951.3G) | `/boot/efi`, `/boot`, — | OS boot disk |
+| └ `rl-root` | (LVM on `nvme0n1p3`) | 70G | — | xfs | `/` | OS root |
+| └ `rl-swap` | (LVM on `nvme0n1p3`) | 32G | — | swap | `[SWAP]` | Swap |
+| └ `rl-home` | (LVM on `nvme0n1p3`) | 849.3G | — | xfs | `/home` | Home |
+| `sdb` | AVAGO MR9440-8i (RAID LUN) | 8.7T | GPT, 1 partition (`sdb1`, whole remaining size) | xfs | `/var` | Docker root, Arcane-managed stacks, container state (`/var/lib/docker`, `/var/dockge`) — now also `benchmarks/`, `training/`, `hf-cache/`, `buildx-cache/`, `ci-registry-mirror/`, the former `/mnt-1` workload |
+| `sda` | PSSD T7 (**USB-attached**) | 465.8G | GPT, 1 partition (`sda1`) | ext4 | `/mnt/usb-recovery` | USB recovery disk, not local bulk storage |
+
+Three of the four rows changed since the 2026-08-04 capture, and none of
+the change is cosmetic. The OS disk is a different, 4x-larger NVMe; the
+former `sda` bulk-storage disk (`/mnt-2`) is now a USB-attached portable
+SSD mounted at `/mnt/usb-recovery`; and the boot disk is now LVM-backed
+with a separate `/home`. The old `sr0` ATAPI optical drive is no longer
+enumerated at all.
+
+**`/mnt-1` and `/mnt-2` are both decommissioned.** `/mnt-1`'s RAID VD
+(formerly `sdc`) suffered a two-drive fault on 2026-09-09 (#3158) and no
+longer enumerates as a block device at all; the mount was unwired (#3159,
+PR #3159) and everything that lived under it moved to `/var`. `/mnt-1`
+itself survives on the host only as a directory of compatibility
+symlinks into `/var` (`benchmarks`, `training`, `hf-cache`,
+`buildx-cache`, `ci-registry-mirror`) so any script still hard-coding the
+old path keeps resolving — new code should target `/var/*` directly. See
+#3158/#3159 for the incident and decommission detail. `/mnt-2` is gone for
+a different reason: its disk is the USB `PSSD T7` above, remounted at
+`/mnt/usb-recovery`, so it is no longer local bulk storage and must not be
+relied on for a rebuild.
`sdb` sits behind an AVAGO/LSI MR9440-8i hardware RAID controller and
appears to the OS as a SCSI LUN, not a raw disk — the controller's own
@@ -49,17 +70,21 @@ the controller's own tooling (`storcli`/`perccli` or vendor equivalent) if
the RAID config itself needs to be reproducible, not just the OS
partitioning on top of it.
-`/var` on its own disk is the key decision: `/var/lib/docker` is 103G and
-`/var/dockge` (bind-mounted stack data for all 23 Arcane-managed stacks, including
+`/var` on its own disk is the key decision, and it has only become more
+load-bearing: `/var/lib/docker` is **2.9T** and `/var/dockge` (stack data
+for the 45 directories under `/var/dockge/stacks/`, including
Elasticsearch indices, Cowrie logs, payload captures, sandbox disks) is
-229G — 332G combined, well past what the 238G OS disk could hold even
-before accounting for the OS itself. Putting `/var` on the 1.7T `sdb`
-disk instead of growing the root filesystem was the right call and should
-be preserved on any rebuild.
-
-Swap is an **8G swapfile** at `/swap.img` on the root filesystem, not a
-dedicated partition — simpler to resize than a swap partition and fine at
-this scale (91G RAM, swap is a safety margin not a working set).
+**350G**. `/var` is 70% full (6.1T of 8.8T) with 2.7T free. The manifest
+still declares 39 sync entries, 33 of which name one of the 34 directories
+under `arcane/home/`; `rex86-eval` is present on disk but **not** in the
+manifest (the other 6 manifest entries are root-level stacks). Putting `/var` on the RAID LUN instead of growing the root
+filesystem remains the right call and should be preserved on any rebuild.
+
+Swap is a **32G LVM logical volume** (`rl-swap`) in the `rl` volume group,
+not a dedicated partition and not a swapfile — the 8G `/swap.img`
+swapfile described in the 2026-08-04 capture no longer exists. 92G of RAM
+means swap is a safety margin rather than a working set, though it was
+under real pressure at measurement time (14.6G in use, priority -2).
`/var` also carries the two CI-created directories, both of which the
workflows cannot create for themselves (`/var` is `root:root 0755`, so a
@@ -117,11 +142,16 @@ with `homeserver-user-data.yaml` renamed to `user-data` alongside an empty
- SSH (key-only, no password auth) and the `xfsprogs`/`nvme-cli` packages
the manual partitioning step below needs.
-**What has to be done by hand, at the storage screen, using the physical
-layout table above as the target:** 3-disk layout (NVMe boot/OS: GPT,
-EFI + ext4 root; `/var`: whole-disk xfs, no partition table; `/mnt-2`:
+**What has to be done by hand, at the storage screen.** For reproducing
+the **former Ubuntu layout** (the one this template was written against,
+and the one the 2026-08-04 capture recorded): 3-disk layout (NVMe boot/OS:
+GPT, EFI + ext4 root; `/var`: whole-disk xfs, no partition table; `/mnt-2`:
GPT + single xfs partition), no LVM, 8G swapfile instead of a swap
-partition. `/mnt-1` is no longer part of the target layout (decommissioned,
+partition. That is **not** the live layout any more — the box now uses LVM
+with a separate `/home`, `/var` on a partition, and a 32G swap LV, and its
+disks have all been replaced (see the table above). Do not use this
+paragraph as a partition plan for the current host; it is a record of what
+the autoinstall flow produced. `/mnt-1` is no longer part of the target layout (decommissioned,
see above) — do not recreate it on a rebuild. The template does **not**
attempt to reproduce the AVAGO RAID controller's own LUN configuration
either — that has to happen before the OS installer ever sees a block
diff --git a/docs/HOST-TUNING.md b/docs/HOST-TUNING.md
index afe407186..74ae01858 100644
--- a/docs/HOST-TUNING.md
+++ b/docs/HOST-TUNING.md
@@ -8,6 +8,39 @@ sudo ./scripts/tune-rocky10.sh --dry-run # show what would change
sudo ./scripts/tune-rocky10.sh # apply
```
+> **Status (2026-09-27): the script is not fully applied on the live homeserver.**
+> The host was re-provisioned onto Rocky Linux 10.2 on 2026-09-03, and the
+> install it now has differs from the one this script was written against. Read
+> the checklist in "Verifying afterwards" as *the intended end state*, not as a
+> description of the box — re-measured read-only over `ssh homeserver` on
+> 2026-09-27, it currently fails two of its six lines (the two `zramctl` /
+> `swapon --show` lines; the other four commands pass). The bullets below are
+> keyed to the **tuning** numbers in the table, not to those six command lines,
+> so they list three tunings rather than two: a tuning can be inapplicable as a
+> durable change while the one command that observes it still reads correct
+> today. That is exactly the case for #2 and #4.
+>
+> - **#1 (zram) is not in effect.** `zram-generator` is not installed, no zram
+> module is loaded, `zramctl` prints nothing, and `swapon --show` lists only
+> the 32G LVM swap. An `/etc/systemd/zram-generator.conf` *does* exist, but it
+> is not the file this script writes: it has no `# Managed by` header and
+> sizes the device with `min(ram / 2, 32768)` rather than the computed MB
+> value the script emits, so it was hand-written — and nothing consumes it
+> while the package is missing.
+> - **#2 (scheduler rule) is not in effect.**
+> `/etc/udev/rules.d/60-apiary-ioscheduler.rules` does not exist. The
+> schedulers that *are* set came from elsewhere: `nvme0n1` is on `none` and
+> the 8.7T rotational `sdb` is on `mq-deadline`, but the non-rotational `sda`
+> is also on `none` where this script's rule would put it on `kyber`.
+> - **#4 (noatime) is half done.** `/var` carries `noatime` in `/etc/fstab`;
+> `/home`, which is a separate mount on this install and did not exist in
+> `/etc/fstab` before the re-provision, is still `defaults`.
+> - **#5 (CPU governor) is in effect** — `tuned-adm active` reports
+> `throughput-performance` and cpu0's governor is `performance`.
+>
+> The 20 GB figure in the sizing note below is still right: the live
+> `NVIDIA RTX 4000 Ada Generation` reports 20475 MiB.
+
These are the Rocky/RHEL-family equivalents of the tunings, not a
transliteration of the Debian recipe. The differences are the point: the
Debian instructions do not work here, and two of them do not survive a reboot
diff --git a/docs/KEYCLOAK-CUTOVER.md b/docs/KEYCLOAK-CUTOVER.md
index a9a244d3f..c294a8a70 100644
--- a/docs/KEYCLOAK-CUTOVER.md
+++ b/docs/KEYCLOAK-CUTOVER.md
@@ -2,6 +2,39 @@
Status: accepted Phase 0 decision record for #976 and epic #986.
+> **Status, verified against the repository 2026-09-27: the cutover is
+> implemented, not pending.** The line above is the decision that was accepted
+> at Phase 0; the hard cutover it specifies has since been carried out, so read
+> this file as the contract that governs the live identity tier rather than as a
+> proposal. What the repository shows today:
+>
+> - `honeypot-keycloak` is a manifest entry
+> (`arcane/manifests/home-production.json`) and a deployed stack
+> (`arcane/home/honeypot-keycloak/compose.yml` — Keycloak plus its
+> PostgreSQL), which is the "Fixed architecture" section's `honeypot-keycloak`
+> stack.
+> - The six "Isolated gateway" rows in the matrix below correspond one-to-one to
+> six `oauth2-proxy` services in `vps/docker-compose.yml`: `oidc-kibana`,
+> `oidc-tanner`, `oidc-evebox`, `oidc-arkime`, `oidc-revdeck` and
+> `oidc-traefik` (one `quay.io/oauth2-proxy/oauth2-proxy` image definition,
+> six services sharing its `oidc-gateway-environment` anchor). `arcane.*` and
+> the `dashboard.*`/`honeypot.*` row correctly have no gateway.
+> - Every identifier on the **Hard-cutover removal list** below is gone from
+> live configuration. `auth-portal`, `strip-auth-identity`, `xore_sso`,
+> `AUTH_INTROSPECTION_URL` and `AUTH_TARGET_HOST` now appear only in
+> documentation (this file, `docs/TESTING.md` and the dated
+> `docs/research/518-smoke-test-research.md` record), never in `vps/` or
+> `arcane/`. No `forward-auth`/`forwardAuth` middleware reference remains
+> under `vps/` or `arcane/`.
+> - The `Xore/auth-backend` line below is accurate as written and should not be
+> read as a live dependency: the *runtime* is retired and only the
+> presentation-only theme is still sourced from that repository, which is
+> what `docs/SENSORS.md` and `docs/KEYCLOAK-OPERATIONS.md` already say.
+>
+> The contract text itself is unchanged and still accurate; this note records
+> only that it has shipped. Day-to-day administration lives in
+> [`KEYCLOAK-OPERATIONS.md`](KEYCLOAK-OPERATIONS.md).
+
Implementation and operations are documented in
[`KEYCLOAK-OPERATIONS.md`](KEYCLOAK-OPERATIONS.md). `Xore/auth-backend` owns
exactly the presentation-only `themes/apiary` Keycloak theme, checked out
@@ -37,17 +70,30 @@ pinned Keycloak runtime does not provide a recovery-code required-action factory
| Route | Consumer | Integration | Keycloak client | Required role | Trust boundary and special traffic |
|---|---|---|---|---|---|
| `dashboard.*`, `honeypot.*` | APIARY dashboard, APIs, SSE, exports, embedded settings | Native authorization-code OIDC with PKCE | `apiary-dashboard` | `user`; mutations require `admin` | Dashboard validates tokens and owns its session. No proxy identity headers. PDF routes use the same dashboard session. |
-| `kibana.*` | Kibana | Isolated gateway | `kibana` | `user` | Gateway is the only network peer allowed to reach Kibana; preserve WebSockets and base paths. |
-| `evebox.*` | EveBox | Isolated gateway | `evebox` | `user` | Gateway is the only upstream path; preserve API, stream, and download behavior. |
-| `arkime.*` | Arkime | Isolated gateway | `arkime` | `user` | Arkime trusts the gateway-injected `X-Forwarded-User` identity (`authMode=header+digest`, #979) -- safe only because the upstream is unreachable except through the gateway (verified) and oauth2-proxy strips any client-supplied copy of that header before injecting its own (`OAUTH2_PROXY_SKIP_AUTH_STRIP_HEADERS`, pinned). The pre-existing local "admin" digest account remains as a fallback. |
-| `tanner.*` | TANNER UI | Isolated gateway | `tanner` | `user` | Gateway-only upstream network; preserve API and static assets. |
-| `rev.*` | RevDeck/Ghidra UI | Isolated gateway | `revdeck` | `user` | Gateway-only upstream network; preserve long responses, downloads, and streams. |
+| `kibana.*` | Kibana | Isolated gateway | `kibana` | `access` | Gateway is the only network peer allowed to reach Kibana; preserve WebSockets and base paths. |
+| `evebox.*` | EveBox | Isolated gateway | `evebox` | `access` | Gateway is the only upstream path; preserve API, stream, and download behavior. |
+| `arkime.*` | Arkime | Isolated gateway | `arkime` | `access` | Arkime trusts the gateway-injected `X-Forwarded-User` identity (`authMode=header+digest`, #979) -- safe only because the upstream is unreachable except through the gateway (verified) and oauth2-proxy strips any client-supplied copy of that header before injecting its own (`OAUTH2_PROXY_SKIP_AUTH_STRIP_HEADERS`, pinned). The pre-existing local "admin" digest account remains as a fallback. |
+| `tanner.*` | TANNER UI | Isolated gateway | `tanner` | `access` | Gateway-only upstream network; preserve API and static assets. |
+| `rev.*` | RevDeck/Ghidra UI | Isolated gateway | `revdeck` | `access` | Gateway-only upstream network; preserve long responses, downloads, and streams. |
| `traefik.*` | Traefik read-only dashboard | Isolated gateway | `traefik-dashboard` | `admin` | Gateway fronts `api@internal`; callback is excluded from recursive auth. |
| `arcane.*` | Arcane administrator UI (#1185, Dockge's replacement -- Dockge decommissioned) | Native authorization-code OIDC | `arcane` | `admin` | No gateway: Arcane authenticates directly against Keycloak. Root-equivalent risk (`/var/run/docker.sock` mounted read-write) -- `admin` role is granted only to the `administrators` group. |
| `auth.example.invalid` | OIDC login, discovery, JWKS, account console | Direct Keycloak | n/a | public protocol endpoints; authenticated account actions | Rate-limited edge route to the Keycloak WireGuard bridge. |
| `auth.example.invalid/admin` | Keycloak administration | Direct Keycloak | n/a | Keycloak administrator + MFA | Same host as the issuer (#1028); a `PathPrefix(/admin)` router gives the SPA's bootstrap burst a larger rate limit. Keycloak owns authentication; HTTP Basic would conflict with the SPA's Bearer API calls. |
| decoy/static/API/status/file/blog hosts currently lacking `forward-auth` | Public honeypot or explicitly application-owned auth | Public / unchanged | none | none | Never attach operator SSO merely because the hostname exists. Public collection must remain independent of IdP availability. |
+Role names differ per row on purpose. The realm's low-privilege client role is
+`access`, and every gateway enforces exactly `:access` through
+`OAUTH2_PROXY_ALLOWED_ROLES` in `vps/docker-compose.yml`;
+`traefik-dashboard` and `arcane` deliberately use `admin` instead
+(#1014/#1185, root-equivalent). The dashboard row's `user` is
+*not* a Keycloak role name: `resource_access.apiary-dashboard.roles` carries
+`access` and `admin`, and the dashboard collapses them to its own
+`user`/`admin` session role. Which human carries which role is a realm
+provisioning decision, not a repository fact: the committed
+`keycloak/realm/apiary-realm.json` defines the `users` and `administrators`
+groups with empty `roleMappings`, so treat the group membership the matrix
+implies as an operator-side grant to be verified in the live realm.
+
The deployment validator must fail when a protected router has neither native
OIDC ownership nor its named gateway. A redirect alone is not evidence: each
row must retain an authorized real-page/API check and an unauthorized denial
diff --git a/docs/KEYCLOAK-OPERATIONS.md b/docs/KEYCLOAK-OPERATIONS.md
index f1e6cb79a..a24223d2d 100644
--- a/docs/KEYCLOAK-OPERATIONS.md
+++ b/docs/KEYCLOAK-OPERATIONS.md
@@ -8,7 +8,19 @@ addresses, passwords, client secrets, cookies, and realm users out of Git.
## Resulting topology
- `honeypot-keycloak` is an Arcane-managed stack on the homeserver.
-- Keycloak and PostgreSQL use pinned upstream images; no local image is built.
+- No local image is built, and the two upstream images are pinned
+ *differently on purpose*. PostgreSQL is a digest pin
+ (`postgres:18.6-bookworm@sha256:…`) because its version is chosen, not
+ chased. Keycloak is deliberately **unpinned** —
+ `quay.io/keycloak/keycloak:latest` with `pull_policy: always` — because it
+ is the identity provider for the whole stack: it was digest-pinned to
+ 26.7.1 when CVE-2026-18963 (unauthenticated account takeover via
+ reset-credentials, CVSS 9.1) landed, and a digest pin is exactly what keeps
+ a known-vulnerable build running until someone edits a file. `pull_policy`
+ is load-bearing here, not decoration: without it a redeploy silently reuses
+ whatever `:latest` first resolved to, which is a pin again with none of the
+ honesty of one. Do not "tidy" the tag into a digest without reading the
+ comment above it in the compose file first.
- PostgreSQL is reachable only on the internal `keycloak-data` network.
- Keycloak publishes HTTP only on the homeserver WireGuard address.
- VPS Traefik terminates TLS and forwards the issuer and administrator hosts
@@ -353,7 +365,12 @@ sudo KEYCLOAK_RESTORE_CONFIRM=restore-keycloak-database \
## 7. Upgrade and rebuild procedure
1. Review Keycloak and PostgreSQL release notes.
-2. Update image tag and digest together in `arcane/home/honeypot-keycloak/compose.yml`.
+2. Update the PostgreSQL tag and digest together in
+ `arcane/home/honeypot-keycloak/compose.yml`. The Keycloak service needs no
+ edit — it is deliberately on `:latest` with `pull_policy: always` (see
+ "Resulting topology"), so a security release arrives by redeploying, and
+ accepting that trade is the decision this step is not asking you to
+ revisit.
3. Re-test the theme against the pinned Keycloak parent theme.
4. Validate Compose and the realm template.
5. Deploy to a disposable stack and exercise the acceptance tests.
diff --git a/docs/NETWORK.md b/docs/NETWORK.md
index 7f83c81f8..ce4bfa2be 100644
--- a/docs/NETWORK.md
+++ b/docs/NETWORK.md
@@ -63,7 +63,10 @@ flowchart LR
- **Home firewall**: none to reason about — the home server has no inbound
exposure at all. Every published container port binds `${HP_BIND}`
(normally `10.8.0.2`, the WireGuard address), never `0.0.0.0`. Verified
- repo-wide during the #1960 review: zero exceptions across all 32 stacks.
+ repo-wide during the #1960 review: no published port binds `0.0.0.0`
+ anywhere under `arcane/home/`. The one stack that spells the variable
+ differently is `unsloth`, whose two published ports use
+ `${UNSLOTH_BIND:-10.8.0.2}` — same tunnel-only default, different name.
## Ingress paths
diff --git a/docs/OPERATIONS.md b/docs/OPERATIONS.md
index 8f8803dff..e318d5507 100644
--- a/docs/OPERATIONS.md
+++ b/docs/OPERATIONS.md
@@ -42,9 +42,10 @@ Two independent geo integrations, same source files:
get country + ASN. (#2713: this used to point at a separate,
never-automated `arkime/geo/` directory populated by hand from db-ip.com —
retired in favor of the same files everything else already uses.)
-- **Elasticsearch** enriches every `suricata-*` and `honeypot-*` event through
- the `geoip-honeypot` ingest pipeline (set as `index.default_pipeline` on both
- index templates), writing ECS `source.geo` / `source.as` / `destination.geo`
+- **Elasticsearch** enriches every `suricata-*` and `honeypot-v2-*` event (and
+ the portbridge, zeek, extracted-files, huginn and traefik families) through
+ the `geoip-honeypot` ingest pipeline (set as `index.default_pipeline` on 8 of
+ the init stack's 33 index templates), writing ECS `source.geo` / `source.as` / `destination.geo`
with city-level lat/lon — this is what powers Kibana maps
(`source.geo.location` is mapped as `geo_point`), from
`GeoLite2-City.mmdb` mounted at
@@ -119,7 +120,7 @@ commands/credentials, payloads, enriched IDS alerts, and ingest failures.
produced their correlation score. The navbar alert badge shows unacknowledged
alert state, while source health uses neutral metric tiles for feeds,
Elasticsearch, Filebeat, and dead letters.
- `/api/campaigns` exposes the same correlation data. A balanced recent feed
+ `/api/v1/campaigns` exposes the same correlation data. A balanced recent feed
prevents one noisy sensor from hiding lower-volume sensors. The
portbridge connection log is used only to recover real source IPs; it is not
counted as a sensor or displayed as an event.
@@ -138,8 +139,9 @@ commands/credentials, payloads, enriched IDS alerts, and ingest failures.
OverviewPanels.tsx`), so there is no basemap env surface to set anymore.
The hourly activity chart also exposes exact counts on hover/focus. The 24-hour
KPI compares activity with the preceding 24 hours and labels large changes;
- source health reports dashboard heap, reserved and cgroup memory, uptime, and
- goroutine count through the same `/api/runtime` contract.
+ source health reports dashboard process uptime and memory (RSS + virtual)
+ as the runtime card on `/api/v1/source-health` — the Go heap/goroutine
+ figures that card used to show have no Rust equivalent and are gone.
Event metadata is directly pivotable: sessions, HASSH/JA3/JA4/User-Agent
fingerprints, exact commands and credentials, HTTP paths, IDS signatures and
categories, payload hashes, ASNs, organizations, and provider classes all
@@ -170,13 +172,17 @@ commands/credentials, payloads, enriched IDS alerts, and ingest failures.
SSE, and events pivot directly to Kibana, EveBox, Arkime, and VirusTotal.
Event tables support keyboard-accessible sorting, selectable columns, and an
expandable normalized-row JSON view; live events on investigation pages raise
- a transient notification. Browser API contracts live in `dashboard/frontend`
+ a transient notification. Browser API contracts live in
+ `arcane/home/honeypot-dashboard/frontend-next`
as strict TypeScript and compile to the committed, dependency-free production
bundle, so Node.js is only a development tool and never part of the container.
- **Operational APIs** — `/metrics` exposes Prometheus text metrics for event,
sensor, ingestion, Filebeat, runtime, dead-letter, and YARA health.
- `/dead-letters` investigates rejected Elasticsearch documents and
- `/api/intelligence/archive` exposes durable campaign/cluster snapshots.
+ `/dead-letters` investigates rejected Elasticsearch documents, and
+ durable campaign/cluster snapshots in `dashboard-intelligence-archive-v1`
+ are readable through the generic index-store route
+ `/api/v1/store/intelligence` (there is no dedicated intelligence route
+ in the Rust router).
Alert acknowledgements and captured-malware downloads require the
dashboard's own Keycloak-derived `admin` role.
- **Safe payload triage** — `yara-scanner` inventories all mounted Dionaea,
@@ -186,8 +192,9 @@ commands/credentials, payloads, enriched IDS alerts, and ingest failures.
snapshot API and other named volumes are archived separately. Test and restore
procedures are in [`docs/analysis/RECOVERY.md`](analysis/RECOVERY.md).
- **Kibana saved objects** (dashboards, visualizations, data views you build
- by hand) live only in Elasticsearch's `.kibana` index — an ES reset,
- migration, or upgrade loses them with no recovery path unless you've
+ by hand) live only in Elasticsearch's Kibana saved-objects index — on this
+ stack's Kibana 9.5.3 that is `.kibana_`, not a bare `.kibana` — so an ES
+ reset, migration, or upgrade loses them with no recovery path unless you've
exported first. Run `analysis/kibana-export.sh` before any ES-affecting
change (matching `KIBANA_URL` to how you reach Kibana — defaults to
`http://kibana:5601`, the in-cluster address); restore with
@@ -220,11 +227,12 @@ commands/credentials, payloads, enriched IDS alerts, and ingest failures.
python3 analysis/analyze.py /opt/stacks/apiary/logs --top 20
```
- **Kibana** → `https://kibana.` (Keycloak via the oauth2-proxy gateway). Data views already exist:
- `honeypot-*` and `suricata-*` (time field `@timestamp`) plus **Arkime
+ `honeypot-v2-*` and `suricata-*` (time field `@timestamp`) plus
+ `dead-letter-honeypot*`, alongside **Arkime
Sessions** (`arkime_sessions3-*`, time field `lastPacket`). All suricata and
honeypot events carry `source.geo` / `source.as` — build maps on
- `source.geo.location`. Arkime sessions have country + ASN only (the db-ip
- country database has no coordinates).
+ `source.geo.location`. Arkime sessions have country + ASN only (GeoLite2
+ Country has no coordinates).
- **Arkime** → `http://:19080` — full-packet session search over
everything Suricata captured on the VPS.
- **TANNER dashboard** → `https://tanner.` (Keycloak via the oauth2-proxy gateway) — web-attack analysis.
diff --git a/docs/PIPELINES.md b/docs/PIPELINES.md
index 9f17a29dd..58f330653 100644
--- a/docs/PIPELINES.md
+++ b/docs/PIPELINES.md
@@ -90,9 +90,11 @@ consume those files, never each other:
2. **The enrichment worker** rewrites watched sensors' files into
`logs/enriched/` before Filebeat sees them. That watch list began as
the five sensor families of #37/#38 (cowrie, dionaea, the conpot
- personas, dns-honeypot, cisco-asa-honeypot) and has grown to 16 named
+ personas, dns-honeypot, cisco-asa-honeypot) and has grown to 17 named
sources plus every conpot persona discovered on disk (six live,
- 2026-08-27) — 22 sources in all.
+ 2026-08-27) — 23 sources in all. The list is
+ `discover_sources` in
+ `arcane/home/honeypot-dashboard/backend-service/src/ip_enrichment/mod.rs`;
Which sensors the worker watches, and whether their files need rewriting
at all, follows one question per sensor: **does the sensor see the
@@ -100,8 +102,8 @@ attacker's real IP?**
| Group | Sensors | Why | Fix |
|---|---|---|---|
-| PROXY-aware | http, api-honeypot, multipot, tanner, dnp3, dicompot, citrix, rdp, endlessh, cisco-asa (WebVPN side), galah (proxied door, XFF), hellpot (proxied door, XFF) | the VPS-side portbridge (`vps/portbridge`) speaks HAProxy PROXY v1, or Traefik sets XFF in-band | none for attribution. Five of them (multipot, tanner, http-honeypot, citrix-honeypot, rdp-honeypot) are watched anyway, solely so canonical-field promotion (#1217) runs on their lines |
-| Tunnel-blind (joined) | cowrie, dionaea + its incident variant (#623), every conpot persona, dns-honeypot, cisco-asa (IKE side), elasticpot, mailoney, hellpot (raw door), beelzebub, sentrypeer, galah (raw door) | raw TCP relay; the log records the WireGuard peer (`10.8.0.1`, the VPS-side tunnel address) | `via_port` join against the portbridge connection log — the generic join for flat `src_ip`/`src_port` shapes (cowrie, dionaea, the conpot personas, dns-honeypot, cisco-asa IKE, elasticpot, mailoney); bespoke join paths for the rest (dionaea-incident's nested rewrite; beelzebub and sentrypeer derive their own address field, then join; hellpot and galah's raw door is joined and adjudicated against their forwarded-header claim) |
+| PROXY-aware | http, api-honeypot, multipot, tanner, dnp3, dicompot, citrix, rdp, endlessh, cisco-asa (WebVPN side), sonicwall-sma, conpot (its TCP personas — every one of the six sets `CONPOT_PROXY_PROTOCOL=1` and portbridge carries the `pp` flag on their TCP rules), galah (proxied door, XFF), hellpot (proxied door, XFF) | the VPS-side portbridge (`vps/portbridge`) speaks HAProxy PROXY v1, or Traefik sets XFF in-band | none for attribution. Six of them (multipot, tanner, http-honeypot, citrix-honeypot, rdp-honeypot, sonicwall-sma-honeypot) are watched anyway, solely so canonical-field promotion (#1217) runs on their lines |
+| Tunnel-blind (joined) | cowrie, dionaea + its incident variant (#623), conpot's UDP personas only (SNMP 161, BACnet 47808, IPMI 623 — PROXY v1 has no UDP form, so those three listeners get no prefix), dns-honeypot, cisco-asa (IKE side), elasticpot, mailoney, hellpot (raw door), beelzebub, sentrypeer, galah (raw door) | raw TCP relay; the log records the WireGuard peer (`10.8.0.1`, the VPS-side tunnel address) | `via_port` join against the portbridge connection log — the generic join for flat `src_ip`/`src_port` shapes (cowrie, dionaea, the conpot personas, dns-honeypot, cisco-asa IKE, elasticpot, mailoney); bespoke join paths for the rest (dionaea-incident's nested rewrite; beelzebub and sentrypeer derive their own address field, then join; hellpot and galah's raw door is joined and adjudicated against their forwarded-header claim) |
The join runs **at ingest time, not read time** (#37/#38): the networkless
`backend-worker-enrichment` container reads both files off disk and writes
@@ -133,11 +135,19 @@ Order (1:1 with `arcane/home/honeypot-init/analysis/elasticsearch-setup.sh`):
p0f OS guess. Stripped/empty sources write nothing — no empty-string
pollution — so ES-side terms aggs reproduce what the dashboard's
read-time classification produced without the dashboard running.
-3–5. GeoIP on suricata src/dst fields (`ignore_missing` no-ops elsewhere)
-6–9. GeoIP on honeypot/portbridge src fields
-10. Dionaea incident hash extraction (plain scan, no regex)
-11. Network-type classification from ASN org (scanner/cloud/hosting)
-12. Log4Shell deobfuscation flag (bounded depth/length)
+3. Traefik wire-tuple `community_id` (#1765) — hashes the tuple the request
+ was actually accepted on (`ClientAddr` → VPS address/entrypoint port), not
+ the client Traefik resolved after forwarded-header trust, so a Traefik
+ record and huginn's sidecar observation of the same TLS connection share
+ a key
+4. Generic `community_id` (#1742) — derives the key for any record that has a
+ 5-tuple but none of its own (Zeek's ~20 protocol logs carry `uid`; only
+ `conn.log` carries `community_id`), seed 0 to match `suricata.yaml`
+5–7. GeoIP on suricata src/dst fields (`ignore_missing` no-ops elsewhere)
+8–11. GeoIP on honeypot/portbridge src fields
+12. Dionaea incident hash extraction (plain scan, no regex)
+13. Network-type classification from ASN org (scanner/cloud/hosting)
+14. Log4Shell deobfuscation flag (bounded depth/length)
No processor makes a network call — GeoIP reads local `.mmdb` files.
@@ -170,6 +180,7 @@ flowchart TB
zpa["zeek-proxy-attribution every 120s · time-bounded join"]
alert["alert-notifier webhook fan-out, cooldown-gated"]
roll["dashboard-rollups every ROLLUP_RUN_INTERVAL_SECS (default 300s)"]
+ tint["threat-intel every 15m · 24h lookback"]
end
subgraph out["Durable entities"]
@@ -187,6 +198,8 @@ flowchart TB
raw --> zpa
atk & cmp & aic --> alert --> st
raw --> roll --> rll
+ raw --> tint
+ tint -.->|"rewrites source.as.type in place"| raw
```
| Loop | Reads | Writes | Cadence | Notes |
@@ -195,9 +208,10 @@ flowchart TB
| correlator | raw events | `campaigns-v1`, `attacker-clusters-v1` | every cycle | pure aggregations, recomputed from scratch; groups ≥2 IPs sharing fingerprint/hash/ASN/provider-class |
| agent-intrusion | raw events | `agent-intrusion-campaigns` | 300s | deterministic criticality rules escalate; LLM never gates escalation; deterministic sha256 campaign_id ⇒ upsert not duplicate |
| zeek-proxy-attribution | zeek flows + portbridge log | flow docs | 120s | attributes relayed flows to attackers; ordering rule above applies here too |
+| threat-intel | raw event indices | `source.as.type` in place | 15m run, 5m CIDR reload, 24h lookback | classifies source IPs against `threat-cidrs.csv`; intel labels win over the ingest pipeline's provider class, reproducing the retired Go dashboard's `firstNonEmpty(e.Intel, e.Provider)` precedence at the data layer |
| dashboard-rollups (#2046) | raw event indices (default pattern) | `overview-rollup-v1`, `geo-rollup-v1`, `attack-rollup-v1` | `ROLLUP_RUN_INTERVAL_SECS`, default 300s | pure-ES derived overviews the dashboard's overview/map/kill-chain reads slice cheaply instead of re-aggregating raw events per request |
-| ml-worker / llm-worker | payloads + events | anomaly scores + `dashboard-ml-anomaly-ack-v1` | continuous | scoring semantics tracked in #1969/#1974 |
-| payload-inventory | disk stores | `dashboard-payload-inventory-v1/-bytes-v1` | periodic scan | HEAD-exists fast path (#1221) |
+| ml-worker / llm-worker | payloads + events | `ml-anomalies` + `dashboard-ml-anomaly-ack-v1` | continuous | scoring semantics tracked in #1969/#1974 |
+| payload-inventory | disk stores | `dashboard-payload-inventory-v1`, `dashboard-payload-bytes-v1` | periodic scan | HEAD-exists fast path (#1221) |
| es-results-importer | root-owned result spools | `*-analysis-v1` | continuous | read-only mirror, shard-partitionable |
| vault-worker (#2290) | `*-analysis-v1`, `llm-analysis` | markdown notes under the knowledge-vault directory (#2289) | `VAULT_POLL_INTERVAL_SECONDS`, default 900s | one note per payload/session entity, sha256-keyed filename ⇒ upsert not duplicate; checkpointed via `knowledge-vault-state-v1`, batch-run so a capture flood can't swamp the vault |
@@ -223,7 +237,7 @@ flowchart LR
store & store2 & store3 --> yara["YARA scanner networkless · read-only"]
yara --> yout[("yara-results/results.json")]
- store & store2 & store3 & yout --> inv["inventory worker"] --> ix[("dashboard-payload-inventory-v1 + -bytes-v1")]
+ store & store2 & store3 & yout --> inv["inventory worker"] --> ix[("dashboard-payload-inventory-v1 dashboard-payload-bytes-v1")]
ix --> wb{"Analyst dispatch: payload workbench"}
wb -->|"hash-only .request markers"| spools["analysis spools: ghidra · linux sandbox · windows sandbox GHOSTS · revdeck · CAPE"]
@@ -262,16 +276,16 @@ index has exactly one writer):
| `attackers-v1` | attacker-identity-worker | backend-service (attackers, overview, graphs) |
| `campaigns-v1`, `attacker-clusters-v1` | correlator-worker | backend-service (clusters, kill-chain, investigate) |
| `agent-intrusion-campaigns` | backend-service agent_intrusion loop | agent-campaigns page |
-| `ghidra-analysis-v1`, `sandbox-analysis-v1`, `github-analysis-v1`, `cape-analysis-v1`, `revdeck-analysis-v1` | es-results-importer | identity worker, investigate/payload surfaces |
+| `ghidra-analysis-v1`, `sandbox-analysis-v1`, `github-analysis-v1`, `workbench-runs-v1`, `cape-analysis-v1`, `revdeck-analysis-v1` | es-results-importer | identity worker, investigate/payload surfaces |
| `yara-analysis-v1` | YARA join via inventory | backend-service charts + investigate |
-| `dashboard-payload-inventory-v1`, `-bytes-v1` | payload-inventory-worker | payloads page, charts |
+| `dashboard-payload-inventory-v1`, `dashboard-payload-bytes-v1` | payload-inventory-worker | payloads page, charts |
| `dashboard-canarytokens-v1` | canarytokens-adapter | canarytokens page + settings pane |
| `cowrie-ttylog-v1` | Filebeat | tty-replay, recordings |
| `mailoney-mail-v1` | Filebeat | sessions/mail views |
| `reporter-metrics-v1` | reporter | settings stats pane |
| `dashboard-alert-state-v1` | alert-notifier loop | alerts page |
| `overview-rollup-v1`, `geo-rollup-v1`, `attack-rollup-v1` | dashboard-rollups loop (#2046) | overview/map/kill-chain dashboard reads |
-| anomaly score + ack indices | ml/llm workers | ml-anomalies page, composite score |
+| `ml-anomalies`, `dashboard-ml-anomaly-ack-v1` | ml/llm workers | ml-anomalies page, composite score |
| `dashboard-users-v1`, `dashboard-workbench-runs-v1`, report/problem-report indices | backend-service itself | their pages |
Retention specifics (ILM, pcap ceilings, snapshots) live in
diff --git a/docs/README.md b/docs/README.md
index cea44a9eb..fb71dd31d 100644
--- a/docs/README.md
+++ b/docs/README.md
@@ -99,8 +99,8 @@ Analysis and sandbox components:
era references inside are historical)](dashboard-manual-ip-block-design.md)
Subdirectories (`analysis/`, `research/`, `sandbox/`, `vps/`, `autoinstall/`,
-`deploy-profiles/`, `archive/`) hold the same kinds of documents scoped to
+`deploy-profiles/`, `design-lab/`) hold the same kinds of documents scoped to
their component. Every doc must be reachable from this page through links;
dated record trees (`research/`, `benchmarks/`, the VM-detection results,
-`archive/`) are exempt. `scripts/check-docs-reachable.py` enforces this in CI
-(#3332).
+`design-lab/`, kept as a near-duplicate of `branding/design-lab/`) are
+exempt. `scripts/check-docs-reachable.py` enforces this in CI (#3332).
diff --git a/docs/RECOVERY.md b/docs/RECOVERY.md
index 12ab72c55..57bd95601 100644
--- a/docs/RECOVERY.md
+++ b/docs/RECOVERY.md
@@ -11,7 +11,7 @@ This repo already has the individual pieces T-Pot's single documented
"Factory Reset" sequence (stop, back up `data/`, wipe it, `git reset
--hard`, reinstall) covers -- [`analysis/backup-honeypot.sh`](../analysis/backup-honeypot.sh)
for the backup, [`docs/STACK-REBUILD.md`](STACK-REBUILD.md)'s live-verified
-runbook for the stop/wipe/restart sequence across the 32 independent
+runbook for the stop/wipe/restart sequence across the 33 independent
Arcane-managed stacks [#258](https://github.com/Xore/APIARY/issues/258)
split this into (Dockge originally, replaced by Arcane per
[#1185](https://github.com/Xore/APIARY/issues/1185); each now a
diff --git a/docs/ROADMAP.md b/docs/ROADMAP.md
index 54c936030..c32560049 100644
--- a/docs/ROADMAP.md
+++ b/docs/ROADMAP.md
@@ -13,6 +13,11 @@ Last audited: 2026-08-05
> Deleted by the 2026-08-30 bulk purge and restored verbatim by #2896/#2947
> on 2026-09-04. Nothing below was updated during that window — treat
> "last audited" above as still true, not as of the restore date.
+>
+> The *repository* did keep moving across that same window, though, so the
+> audit date is not a safe lower bound for "still true": reconcile any status
+> claim here against the code before relying on it. The CAPEv2 line above is
+> one already re-verified (2026-09-27) and found stale.
## Current baseline
@@ -35,21 +40,30 @@ Last audited: 2026-08-05
booting; the end-to-end submit-to-report path is verified for the
Linux/Wine sandbox and GitHub-analysis publishing. The Windows-11 golden
image epic ([#47](https://github.com/Xore/APIARY/issues/47)) is
- still open — see the Windows sandbox section below. CAPEv2 (#314-322)
- remains unbuilt and is post-0.1.0 backlog.
-- Documentation has been consolidated: every doc that used to be scattered
- next to its source now lives under `docs/`, mirroring the source tree
- ([#670](https://github.com/Xore/APIARY/issues/670), closed
- 2026-08-05).
-- A separate, not-yet-cut-over track: the dashboard is being rewritten as
- a TanStack Start frontend/BFF + Rust service tier
+ still open — see the Windows sandbox section below. CAPEv2 (#314-322) is
+ **authored, not deployed**: [#843](https://github.com/Xore/APIARY/issues/843)
+ (2026-09-01) landed `sandbox/cape/` — compose stack, Packer
+ `win11-cape.pkr.hcl`, CAPEv2 override units and a spool worker — but unlike
+ GHOSTS it has no entry in `arcane/manifests/home-production.json`, so nothing
+ Dockge-managed deploys it. Treat it as post-0.1.0 backlog whose build work
+ has already started, not as unbuilt.
+- Documentation has been consolidated: the subsystem docs that used to be
+ scattered next to their source now live under `docs/`, mirroring the source
+ tree ([#670](https://github.com/Xore/APIARY/issues/670), closed
+ 2026-08-05). The exceptions are deliberate and catalogued in
+ [`docs/README.md`](README.md) — per-stack and vendored `README.md` files stay
+ next to their code, and a few dated record trees are exempt from the
+ reachability gate (#3332). Root-level dated task records (`DIFF.md`,
+ `EVIDENCE.md`, `HANDOFF-3097.md`) are the residue of that sweep and are not
+ covered by it.
+- The dashboard rewrite is done, not pending: the TanStack Start
+ frontend/BFF + Rust service tier
([#1608](https://github.com/Xore/APIARY/issues/1608) and its
- follow-ups), living on the `port-foundation` branch behind Compose's
- `next` profile alongside the current Go dashboard. Feature-complete
- enough to demo; not yet live in production. See
- [`DASHBOARD-CUTOVER.md`](DASHBOARD-CUTOVER.md) and
- [#1628](https://github.com/Xore/APIARY/issues/1628) for what's left
- before it replaces the row above.
+ follow-ups) cut over on 2026-08-22 under
+ [#1628](https://github.com/Xore/APIARY/issues/1628), the Go dashboard is
+ deleted from Compose, and `dashboard-next` runs unconditionally — no
+ `next` profile, no runtime fallback. See
+ [`DASHBOARD-CUTOVER.md`](DASHBOARD-CUTOVER.md).
Everything that was tracked here as "Gate 0" and "Release 1" through
"Release 3" and "Release 5" in prior versions of this document is now closed.
diff --git a/docs/ROCKY-10-MIGRATION.md b/docs/ROCKY-10-MIGRATION.md
index c19f7bba9..f6ed4e979 100644
--- a/docs/ROCKY-10-MIGRATION.md
+++ b/docs/ROCKY-10-MIGRATION.md
@@ -1,5 +1,33 @@
# Rocky Linux 10 support in `install-homeserver.sh`
+> **Status (2026-09-27): the move is done — the homeserver now runs Rocky Linux
+> 10.2.** The RHEL path below is no longer a smoke-test target or a rehearsal
+> for a future migration: the host was re-provisioned onto Rocky 10.2
+> (2026-09-03, Anaconda) with a different disk layout, so the `rhel` branches in
+> `install-homeserver.sh` are the production path and `debian` is the fallback
+> that only a rebuild would exercise. Re-measured read-only over
+> `ssh homeserver` on 2026-09-27: `/etc/os-release` reports `ID=rocky` and
+> `VERSION="10.2 (Red Quartz)"`, and both conditions in
+> "Two things Rocky does that Ubuntu did not" are live — SELinux is `Enforcing`
+> and firewalld is `active`.
+>
+> The two things that stay open are unchanged and are the ones to read first:
+> the base OS is still installed by hand (there is still no kickstart artifact
+> in the tree), and the `:z`/`:Z` label gap below is still real — none of the
+> 35 compose files tracked under `arcane/home/` carries an SELinux relabel (34
+> stack-level `compose.yml` files, plus the decoy honeyfs compose nested inside
+> `arcane/home/honeypot-cowrie/cowrie/`, which is an attacker-facing artifact
+> rather than a deployed stack).
+>
+> Everything else in this document was re-checked against
+> `scripts/install-homeserver.sh` and `scripts/lib/install-common.sh` on
+> 2026-09-27 and still matches: the `pkg_update`/`pkg_install` shims, the
+> once-at-source-time `$DISTRO_FAMILY` resolution, the full package-name
+> table, the verbatim Docker `centos` repofile, the `cuda-rhel10` repo, the
+> `container_use_devices` boolean, and the non-fatal
+> `step_preflight_rhel_platform`. For the current disk layout see
+> [`HOMESERVER-DISK-LAYOUT.md`](HOMESERVER-DISK-LAYOUT.md).
+
The homeserver is moving from Ubuntu to Rocky Linux 10. `scripts/install-homeserver.sh`
now runs on both, so the reinstall smoke test in #1609 has a working installer.
diff --git a/docs/SENSORS.md b/docs/SENSORS.md
index 36faf392f..b7316ca73 100644
--- a/docs/SENSORS.md
+++ b/docs/SENSORS.md
@@ -52,8 +52,9 @@ The `payload-dedupe` service scans these stores hourly and atomically replaces
same-filesystem duplicates with hard links. Existing event/download URLs remain
valid while duplicate disk blocks are reclaimed; its last-run report is stored
at `state/dedupe/payload-dedupe.json`.
-Its diagnostic logger is limited to `info,warning,error` so debug chatter cannot
-consume the data disk. The `log-maintenance` sidecar copy-truncates and gzips
+Neither service produces per-event chatter: `payload-dedupe` prints a single
+JSON result line per pass, and `log-maintenance` writes one stderr line per
+rotation. The `log-maintenance` sidecar copy-truncates and gzips
human-readable Dionaea, Conpot, and Cowrie logs at 256 MiB (four archives).
Structured JSON event streams are deliberately never rotated by that sidecar,
which preserves Filebeat offsets and dashboard ingestion.
@@ -70,14 +71,22 @@ GeoIP enrichment is best-effort: empty or malformed addresses are skipped, but
the original event is always retained.
ILM keeps raw Suricata indices for 7 days, honeypot data streams for 30 days,
and dead-letter records for 60 days so high-volume scans cannot fill the disk.
+Those are the values at the *code* fallback `HONEYPOT_RETENTION_DAYS=30`. Every
+window derives from that one knob, and the shipped configuration does **not**
+use the fallback: all four tracked `.env.example` files set
+`HONEYPOT_RETENTION_DAYS=21` (#2820), at which the three windows above become
+**4d / 21d / 42d** — Suricata is `retention*7/30` (integer-truncated, floored
+at 1) and dead-letter is `retention*2`. Only the ILM *policy names*
+(`suricata-7d`, `honeypot-30d`, `dead-letter-60d`) stay fixed; they are
+identifiers referenced by the index templates, not claims about duration.
## Runtime resource budgets
Every service has a CPU, memory, and Docker `json-file` log budget. The
limits are intentionally generous relative to the host (16 logical CPUs and
-91 GiB RAM): Elasticsearch gets 8 GiB with a 4 GiB heap; Arkime capture 6 GiB;
+91 GiB RAM): Elasticsearch gets 12 GiB with a 6 GiB heap; Arkime capture 6 GiB;
Kibana, Filebeat, and the TANNER analyzer receive 2 GiB; EveBox, Dionaea, Arkime
-viewer, and the live dashboard receive 1 GiB (the dashboard also has one CPU). The
+viewer, and the live dashboard receive 1 GiB (the dashboard also has two CPUs). The
remaining lightweight sensors receive 128-512 MiB. Docker console logs rotate
at 25 MiB with three files, independently from sensor event files under
`./logs`.
@@ -86,7 +95,7 @@ at 25 MiB with three files, independently from sensor event files under
| Dashboard | Subdomain | Container |
|---|---|---|
-| Live sensor view (ours) | `honeypot.` | `dashboard` :8090 |
+| Live sensor view (ours) | `honeypot.` | `dashboard-next` :8080 |
| Kibana (ELK + Suricata) | `kibana.` | `kibana` :5601 |
| TANNER web-attack analysis | `tanner.` | `tanner_web` :8091 |
| EveBox (Suricata events) | `evebox.` | `evebox` :5636 |
@@ -116,10 +125,13 @@ XSS, command execution, PHP code/object injection, XXE, CRLF and template
injection) and stores sessions in Redis. TANNER's emulation containers are
isolated from the homeserver Docker socket and are not a malware detonation
environment. Suspicious payload detonation belongs in the separate KVM/libvirt
-sandbox described in [`sandbox/README.md`](sandbox/README.md). Containers: `tanner_redis`,
-`tanner_phpox`, `tanner_api`, `tanner` (analyzer, `:8090`), `tanner_web`
-(dashboard, `:8091`), `snare_clone` (one-shot deterministic persona installer),
-`snare` (`:8080`). The page source lives under [snare/persona](../arcane/home/honeypot-tanner/snare/persona)
+sandbox described in [`sandbox/README.md`](sandbox/README.md). Containers: `tanner_docker`
+(the nested disposable emulator daemon), `tanner_redis`, `tanner_phpox`, `tanner_api`,
+`tanner` (analyzer, `:8090`), `tanner_web` (dashboard, `:8091`), `snare` (`:8080`) --
+seven services, all on `tanner_local`. The persona clone itself is not one of them:
+`snare-clone` is a one-shot job in `honeypot-init` (`hp-snare-clone`) that writes the
+`snare-pages` volume `snare` reads. The page source lives under
+[snare/persona](../arcane/home/honeypot-tanner/snare/persona)
and is transformed into SNARE's content-addressed store during the image build;
no third-party site is cloned. All `mushorg/*` images are third-party
— verify tags/args upstream (needs a live build/pull).
@@ -127,8 +139,10 @@ no third-party site is cloned. All `mushorg/*` images are third-party
## Suricata — analysing all the traffic
`suricata` runs **on the VPS** (host networking, sniffing the public interface
-`SURICATA_IFACE`, default `ens6`) so it sees real attacker source IPs before the
-tunnel. It writes to `/opt/stacks/apiary/logs/suricata/` on the VPS:
+`CAPTURE_INTERFACE`, written at boot by `detect-capture-interface.service`,
+falling back to the legacy `SURICATA_IFACE` and then to `eth0` — *not* the old
+`ens6`, which a reboot has already renamed) so it sees real attacker source IPs
+before the tunnel. It writes to `/opt/stacks/apiary/logs/suricata/` on the VPS:
- `eve.json` (alerts, http, dns, tls, flow) — Filebeat on the home server ships
it to the `suricata-*` Elasticsearch index (stats events are dropped, see
@@ -180,7 +194,9 @@ Web UI: `http://:19080` (`arkime.` via Traefik).
> **http/api-honeypots** (`PROXY_PROTOCOL=1`), **dnp3** (`PROXY_PROTOCOL=1`),
> **dicompot** (`PROXY_PROTOCOL=1`), **citrix-honeypot**,
> **sonicwall-sma-honeypot** (`PROXY_PROTOCOL=1`),
-> **cisco-asa-honeypot**'s WebVPN side and **rdp-honeypot** (`PROXY_PROTOCOL=1`) and **all conpot sensors** (`CONPOT_PROXY_PROTOCOL=1`, gevent shim baked in
+> **cisco-asa-honeypot**'s WebVPN side and **rdp-honeypot** (`PROXY_PROTOCOL=1`),
+> **endlessh** (`PROXY_PROTOCOL=1`, public 2022) and **all conpot sensors**
+> (`CONPOT_PROXY_PROTOCOL=1`, gevent shim baked in
> by `conpot/proxy_patch.py`) parse it, so those events log the true IP and
> port. The http listener sniffs the header, so Traefik-routed requests (no
> header) keep working too.
diff --git a/docs/STACK-REBUILD.md b/docs/STACK-REBUILD.md
index f974e5a10..0e0c158d7 100644
--- a/docs/STACK-REBUILD.md
+++ b/docs/STACK-REBUILD.md
@@ -18,17 +18,20 @@ traps hit on the first live run — read it before trusting the script blind,
and definitely before doing any of this by hand on the VPS side, which the
script doesn't touch.
-Since #258 split the stack into ~19 independent Arcane-managed projects
+Since #258 split the stack into independent Arcane-managed projects
(`honeypot-init`, `honeypot-conpot`, `honeypot-cowrie`, `honeypot-multipot`,
`honeypot-http`, `honeypot-dnp3`, `honeypot-dionaea`, `honeypot-dicompot`,
-`honeypot-dns-honeypot`, `honeypot-citrix`, `honeypot-cisco-asa`,
-`honeypot-rdp`, `honeypot-endlessh`,
+`honeypot-dns-honeypot`, `honeypot-citrix-honeypot`, `honeypot-cisco-asa-honeypot`,
+`honeypot-rdp-honeypot`, `honeypot-endlessh`,
`honeypot-payload-analysis`, `honeypot-tanner`, `honeypot-elk`,
-`honeypot-dashboard`, `honeypot-utilities`, plus the now-empty
+`honeypot-dashboard`, `honeypot-utilities`, plus the then-empty
`APIARY`), a full reset is no longer "stop the stack, `docker compose
down -v`, start it again" — it's an ordered sequence across projects with a
-couple of real circular-dependency traps. This doc exists because the first
-live run of this sequence (2026-08-02) hit three of them.
+couple of real circular-dependency traps. That 19-project list was
+accurate for the day; the fleet is now **33 manifest entries under
+`arcane/home/`**, and the two loops below are written against the full
+current set rather than that original nineteen. This doc exists because
+the first live run of this sequence (2026-08-02) hit three of them.
`honeypot-keycloak` (the identity stack, `docs/KEYCLOAK-OPERATIONS.md`)
is intentionally handled separately below rather than folded into the
@@ -81,7 +84,7 @@ confusing at best.
```bash
ssh vps
docker stop hp-suricata hp-suricata-rules-refresh hp-suricata-log-maintenance \
- hp-portbridge hp-portbridge-log-rotate hp-portbridge-log-maintenance \
+ hp-portbridge hp-portbridge-log-maintenance \
hp-portbridge-blackhole-refresh hp-p0f
sudo find /opt/stacks/apiary/logs/suricata -mindepth 1 -delete
sudo find /opt/stacks/apiary/logs/portbridge -mindepth 1 -delete
@@ -91,12 +94,25 @@ sudo find /opt/stacks/apiary/logs/portbridge -mindepth 1 -delete
```bash
ssh homeserver
-for s in honeypot-elk honeypot-dashboard honeypot-utilities \
- honeypot-payload-analysis honeypot-dionaea honeypot-tanner \
- honeypot-dnp3 honeypot-http honeypot-multipot honeypot-cowrie \
- honeypot-conpot honeypot-dicompot honeypot-dns-honeypot \
- honeypot-citrix honeypot-cisco-asa honeypot-rdp \
- honeypot-endlessh honeypot-init; do
+# Every Arcane-managed stack under arcane/home/ except honeypot-keycloak:
+# 32 of the manifest's 33, derived as
+# `git ls-files arcane/home | cut -d/ -f3 | sort -u` minus
+# honeypot-keycloak (handled separately) and rex86-eval (never a
+# deployment piece). Written out long rather than globbed, so a stack a
+# future manifest entry adds does not get swept up before anyone has
+# decided whether a full reset should stop it.
+for s in honeypot-agent-intrusion-worker honeypot-attacker-identity-worker \
+ honeypot-beelzebub honeypot-canarytokens \
+ honeypot-cisco-asa-honeypot honeypot-citrix-honeypot \
+ honeypot-conpot honeypot-correlator-worker honeypot-cowrie \
+ honeypot-dashboard honeypot-dashboard-backend honeypot-dicompot \
+ honeypot-dionaea honeypot-dnp3 honeypot-dns-honeypot \
+ honeypot-elasticpot honeypot-elk honeypot-endlessh honeypot-galah \
+ honeypot-hellpot honeypot-http honeypot-init honeypot-mailoney \
+ honeypot-multipot honeypot-payload-analysis \
+ honeypot-payload-inventory-worker honeypot-rdp-honeypot \
+ honeypot-sentrypeer honeypot-sonicwall-sma honeypot-tanner \
+ honeypot-utilities unsloth; do
(cd /opt/stacks/$s && docker compose -f compose.yml down)
done
```
@@ -169,11 +185,18 @@ starts first just creates them empty and the real writer fills them in once
it's up).
```bash
-for s in honeypot-conpot honeypot-cowrie honeypot-multipot honeypot-http \
- honeypot-dnp3 honeypot-dionaea honeypot-tanner \
- honeypot-dicompot honeypot-dns-honeypot honeypot-citrix \
- honeypot-cisco-asa honeypot-rdp honeypot-endlessh \
- honeypot-payload-analysis honeypot-utilities honeypot-dashboard; do
+for s in honeypot-agent-intrusion-worker honeypot-attacker-identity-worker \
+ honeypot-beelzebub honeypot-canarytokens \
+ honeypot-cisco-asa-honeypot honeypot-citrix-honeypot \
+ honeypot-conpot honeypot-correlator-worker honeypot-cowrie \
+ honeypot-dashboard honeypot-dashboard-backend honeypot-dicompot \
+ honeypot-dionaea honeypot-dnp3 honeypot-dns-honeypot \
+ honeypot-elasticpot honeypot-endlessh honeypot-galah \
+ honeypot-hellpot honeypot-http honeypot-mailoney \
+ honeypot-multipot honeypot-payload-analysis \
+ honeypot-payload-inventory-worker honeypot-rdp-honeypot \
+ honeypot-sentrypeer honeypot-sonicwall-sma honeypot-tanner \
+ honeypot-utilities unsloth; do
(cd /opt/stacks/$s && docker compose -f compose.yml up -d)
done
```
@@ -184,10 +207,21 @@ done
ssh vps
cd /root/vps
docker compose -f docker-compose.yml up -d suricata portbridge p0f \
- suricata-rules-refresh suricata-log-maintenance portbridge-log-rotate \
+ suricata-rules-refresh suricata-log-maintenance \
portbridge-log-maintenance portbridge-blackhole-refresh
```
+There is no `portbridge-log-rotate` service to start or stop. It was removed
+in #1779, and its only job — pruning the renamed `portbridge.json.*` files
+once they age out — is what `portbridge-log-maintenance` does now, per that
+script's own header. A stop or start naming it fails on a container that does
+not exist, which in a reset runbook is a step that silently does nothing.
+
+The list also omits `portbridge-manual-blackhole-refresh` on purpose: it is
+the manual counterpart to the scheduled `portbridge-blackhole-refresh` and
+should not be brought up by a reset. `suricata-update` is covered separately,
+just below.
+
`suricata` depends on `suricata-update` (`condition: service_completed_successfully`)
— if `suricata-update`'s container is still sitting there `Exited(0)` from a
previous run, Compose treats the condition as already satisfied and won't
diff --git a/docs/STORAGE.md b/docs/STORAGE.md
index 30fa923e3..671d308de 100644
--- a/docs/STORAGE.md
+++ b/docs/STORAGE.md
@@ -81,9 +81,9 @@ Indices follow `-v` naming. Producers and consumers were
verified one-to-one during the #1960 review — the catalog table lives in
[PIPELINES.md](PIPELINES.md#4-index-catalog).
-- **Templates**: `honeypot-*`, `suricata-*`, `portbridge-*` set the shared
- ingest pipeline and flattened mappings so heterogeneous sensor fields
- land safely.
+- **Templates**: `honeypot-v2-*`, `suricata-*`, `portbridge-v2-*` set the
+ shared ingest pipeline and flattened mappings so heterogeneous sensor
+ fields land safely.
- **Derived entities** (`attackers-v1`, `campaigns-v1`,
`attacker-clusters-v1`, `agent-intrusion-campaigns`) are recomputed
idempotently by their loops — safe to delete and regenerate from raw
diff --git a/docs/TESTING.md b/docs/TESTING.md
index 1e70faf18..9c524b650 100644
--- a/docs/TESTING.md
+++ b/docs/TESTING.md
@@ -102,8 +102,8 @@ fixed back into this document and the install scripts themselves.
`install-homeserver.sh`'s own restore steps, are exhaustive for a
given run; this is exactly the kind of gap this pass exists to
catch).
-2. **Wipe both hosts** — every Dockge stack, container, volume, and piece
- of state on the homeserver and the VPS.
+2. **Wipe both hosts** — every Arcane-managed stack, container, volume, and
+ piece of state on the homeserver and the VPS.
3. **Reinstall from the real path** — `scripts/install-homeserver.sh` (or
whatever the current unattended provisioning entry point is) against a
genuinely clean OS, not a host with leftover packages/config. Redeploy
@@ -133,8 +133,9 @@ fixed back into this document and the install scripts themselves.
lingering call into a route or function that's been superseded or
replaced but never removed (the kind of gap a working install can
mask, since the old path may still technically respond). Cross-check
- call sites against the routes actually registered in `main.go` and
- the functions actually exported by each module, not just "does it
+ call sites against the routes actually registered in
+ [`main.rs`](../arcane/home/honeypot-dashboard/backend-service/src/main.rs)
+ and the functions actually exported by each module, not just "does it
still return 200."
- Dead code: the reverse direction of the check above — routes,
handlers, functions, and files that exist but are no longer called
@@ -156,8 +157,9 @@ fixed back into this document and the install scripts themselves.
access, admin-role enforcement, logout, a disabled/revoked
session losing access, and fail-closed behavior when the identity
provider is unreachable.
- - Every gateway-fronted application (Kibana, EveBox, Arkime, TANNER,
- RevDeck, Dockge, the Traefik dashboard): authorized access reaches
+ - Every gateway-fronted application — the six isolated `oidc-*`
+ gateways (Kibana, EveBox, Arkime, TANNER, RevDeck, the Traefik
+ dashboard): authorized access reaches
real content, wrong-role denial, logout, callback/deep-link
behavior, and direct-upstream bypass denial (confirm the
isolated `oidc-` Docker network still has no other member).
@@ -170,9 +172,12 @@ fixed back into this document and the install scripts themselves.
secret, or fallback survives the install: grep the fresh
deployment for `AUTH_INTROSPECTION_*`, `forward-auth`,
`strip-auth-identity`, `xore_sso`, and `X-Auth-Role` and confirm
- zero hits (this repo's own working tree already has zero --
- verified 2026-08-09 -- the check here is that a *deployed*, fresh
- install matches).
+ zero live hits (this repo's own working tree has none — verified
+ 2026-08-09, and re-verified 2026-09-27: the only remaining matches
+ are two comments saying the thing was retired, plus the one
+ allowlisted stale-path entry in `scripts/doc-path-lint-allowlist.txt`
+ for the moved VPS forward-auth directory. The check here is that a
+ *deployed*, fresh install has no live runtime).
- Retain redacted evidence (pass/fail results plus browser
traces/screenshots/logs where applicable) and link it from #787.
5. **Fix forward, and track it:** any gap found (a missing install step,
diff --git a/docs/agent-intrusion-threat-model.md b/docs/agent-intrusion-threat-model.md
index fbaad1c90..2bb52e110 100644
--- a/docs/agent-intrusion-threat-model.md
+++ b/docs/agent-intrusion-threat-model.md
@@ -36,9 +36,12 @@
For each of the nine areas #154 asked to cover, this maps to APIARY's
**actual** current architecture — verified against the real compose files,
-Go/Python source, and docs in this tree, not assumed from what a "typical"
+the source, and docs in this tree, not assumed from what a "typical"
honeypot stack might do. Each entry records: what exists today, whether the
-published campaign's technique applies here, and the evidence.
+published campaign's technique applies here, and the evidence. (The original
+pass read Go and Python source; the Go tier was deleted at #1628 on
+2026-08-22, so a re-read today should be against the Rust modules named in
+the status banner above.)
---
@@ -233,20 +236,24 @@ gap is outbound-to-internet egress policy, tracked separately in #538.
**Substantially addressed for the dashboard; inconsistent elsewhere.**
- The dashboard's own Docker-socket boundary is the strongest example in
- this tree: the dashboard container itself never mounts `/var/run/docker.sock`
- (`arcane/home/honeypot-dashboard/compose.yml`, grepped directly — absent). All
+ this tree: the dashboard containers themselves never mount
+ `/var/run/docker.sock` (`arcane/home/honeypot-dashboard/compose.yml`, grepped
+ directly — its one socket mount belongs to `services-adapter`, below). All
Docker-lifecycle actions (start/stop/restart) go through
`hp-services-adapter`, a separate container that is `cap_drop: [ALL]`,
`read_only: true`, `network_mode: none`, and reachable only via an
AF_UNIX socket the dashboard also holds — no TCP path exists to abuse it
remotely even if the dashboard container itself were compromised.
-- `hp-autoheal` is the one other service that does bind-mount the real
- `/var/run/docker.sock` (`arcane/home/honeypot-utilities/compose.yml`) — it watches
- containers by label daemon-wide and restarts unhealthy ones. This is a
- broad, standing grant (full Docker API access, not scoped to specific
- containers) held by a long-running service; workload-identity-scoped
- alternatives (e.g. a narrower label-filtered API surface) were not found
- in this tree.
+- `hp-autoheal` was the one other service that bind-mounted the real
+ `/var/run/docker.sock` (`arcane/home/honeypot-utilities/compose.yml`) — it
+ watches containers by label daemon-wide and restarts unhealthy ones. **That
+ grant has since been narrowed by #592** (the strikethrough item in
+ [Follow-up scope](#follow-up-scope)): it now talks to `hp-docker-socket-proxy`
+ at `tcp://docker-socket-proxy:2375` with no socket bind mount of its own, and
+ the proxy holds the socket `:ro` scoped to `CONTAINERS=1`, `IMAGES=1`,
+ `POST=1` on a private network. What remains is that `CONTAINERS=1` is still
+ daemon-wide rather than label-filtered — the label scoping is
+ `AUTOHEAL_CONTAINER_LABEL` inside autoheal, not an API-side restriction.
- `tanner_docker` (`arcane/home/honeypot-tanner/compose.yml`) is `privileged: true` with
its own `tmpfs /var/lib/docker` — explicitly isolated Docker-in-Docker on
the private `tanner_local` network, not a bind mount of the host socket.
@@ -259,11 +266,10 @@ gap is outbound-to-internet egress policy, tracked separately in #538.
repo has to short-lived workload identity today.
**Verdict:** the dashboard/services-adapter split is a good existing
-least-privilege pattern worth citing as the template for any future
-privileged-access surface. `hp-autoheal`'s standing daemon-wide Docker
-socket grant is the one credential-lifetime/least-privilege gap worth a
-scoped look (whether its watch scope can be narrowed), separate from #88's
-network-isolation focus.
+least-privilege pattern worth citing as the template for any new
+privileged-access surface. It is also now the template autoheal was moved onto:
+its raw socket grant is gone, replaced by the same narrow-proxy shape, leaving
+only the daemon-wide `CONTAINERS=1` scope.
---
@@ -400,8 +406,10 @@ finding.**
state; the same fragmentation carried into the Rust cutover's per-source
worker functions): Suricata/sensor event alerts, `ghidraAlerts`, `githubAnalysisAlerts`,
sandbox queue/verdict alerts, ML anomaly severity, and (as of #150) the
- new `llm-analysis` index's own severity field is not yet wired into
- `alerts.go` at all — it is currently browse-only via `/llm-analysis`.
+ new `llm-analysis` index's own severity field — which was browse-only via
+ `/llm-analysis` until it was wired into the sink as `llm_flagged_alerts` in
+ the Rust cutover, the "Done" item in [Follow-up scope](#follow-up-scope)
+ below.
- No single "this source_ip/session/sample crossed N independent trust
boundaries in a Y-minute window" correlation exists — each alert source
answers its own narrow question. This is exactly the shape the campaign's
@@ -413,11 +421,11 @@ finding.**
identity for investigation, not by trust-boundary-crossing count for
alerting.
-**Verdict:** this is the clearest concrete gap this research surfaced. Wiring
-`llm-analysis`'s severity into the existing alert sink is a small, immediate
-follow-up; a genuine cross-source trust-boundary-crossing correlation engine
-is squarely phase 3's scope, not something to build inside this research
-pass.
+**Verdict:** this is the clearest concrete gap this research surfaced. Its
+smallest piece — wiring `llm-analysis`'s severity into the alert sink — has
+since landed (`llm_flagged_alerts` in the Rust `alert-notifier`); what remains
+is a genuine cross-source trust-boundary-crossing correlation engine, which is
+squarely phase 3's scope, not something to build inside this research pass.
---
@@ -429,11 +437,11 @@ pass.
| 2 | Untrusted structured-data processing | Yes, broadly | `html/template` auto-escaping; CI YARA corpus gate | No archive/container-format parsing exists yet — must inherit this discipline when added |
| 3 | Env/`/proc/*/environ` secret exposure | Yes | `secretFromEnvironment`'s `_FILE` pattern (one use) | Pattern not applied to `ARKIME_*`, `GH_PAT`, VPS SSH key |
| 4 | Metadata-service / RFC 1918 reachability | No cloud metadata surface exists | Per-sensor private Docker networks | Outbound-to-internet egress policy (tracked in #538) |
-| 5 | Credential lifetime / workload identity | Yes | dashboard/services-adapter split (strong pattern, since carried into the backend-service/worker split) | `hp-autoheal`'s standing daemon-wide docker.sock grant |
+| 5 | Credential lifetime / workload identity | Yes | dashboard/services-adapter split (strong pattern, since carried into the backend-service/worker split); autoheal moved onto the same narrow-proxy shape by #592 | Proxy's `CONTAINERS=1` is still daemon-wide, not label-filtered |
| 6 | Encoded/chunked C2 | Yes, as honeypot capture surface | Raw payload capture (tanner/Suricata); narrow fixed-destination outbound HTTP clients | No network-layer egress enforcement (folds into #538) |
| 7 | Repeated recon / low-signal escalation | Yes — core motivating gap | ml-worker anomaly scoring; dashboard campaign clustering | No behavioral-phase correlation or combination-based severity escalation |
| 8 | Source-control/CI write paths | Yes | `analysis/github/` publish gate (CI-tested); vendored-dep hash pinning | No image digest pinning for this repo's own built images |
-| 9 | Cross-source alert correlation | Yes — core motivating gap | Multiple independent alert sources feed one sink | No trust-boundary-crossing correlation; `llm-analysis` severity not yet wired into alerts |
+| 9 | Cross-source alert correlation | Yes — core motivating gap | Multiple independent alert sources feed one sink, `llm-analysis` severity included | No trust-boundary-crossing correlation engine |
---
diff --git a/docs/analysis/README.md b/docs/analysis/README.md
index 88efa73d3..344dd25fa 100644
--- a/docs/analysis/README.md
+++ b/docs/analysis/README.md
@@ -13,11 +13,13 @@ scripts that run on the sensor host.
> [#74](https://github.com/Xore/APIARY/issues/74) (the manual publisher).
> Per the roadmap's own status line: **built** — dashboard trigger/read
> (Phases 2-3), the host publisher itself (Phase 1), queue health/alerting
-> (Phase 5), and IOC/family enrichment (Phase 6); the host publisher is
-> **built but not installed** on a given deployment until an operator runs
-> `analysis/github/install-github-publisher.sh` there; environment/Compose
-> wiring (Phase 7) has **not started**. Publication is **not** automatic
-> even where installed — see "Publication is manual" below.
+> (Phase 5), IOC/family enrichment (Phase 6), and environment/Compose wiring
+> (Phase 7 — `GITHUB_ANALYSIS_REQUEST_DIR` / `_RESULTS_DIR` /
+> `_ALERT_POSITIVES` and both spool bind-mounts are in the dashboard
+> service); the host publisher is **built but not installed** on a given
+> deployment until an operator runs
+> `analysis/github/install-github-publisher.sh` there. Publication is **not**
+> automatic even where installed — see "Publication is manual" below.
---
@@ -67,17 +69,21 @@ flowchart TB
## Components in this folder
+Most of the tooling below is no longer under the repository-root `analysis/`
+tree: #1502 moved each deployable piece under its own Arcane stack directory
+in `arcane/home/`. The table names the current home of each.
+
| Path | Purpose |
|---|---|
-| `analyze.py` | Offline triage of Cowrie / http-honeypot / multipot / Dionaea JSON logs. Stdlib only |
-| `collect.sh` | **Deprecated.** Cron-driven bulk copy of captures into a clone of `Xore/honeypot`. Superseded by the dashboard button; kept for a one-time manual backfill |
-| `dedupe-payloads.py` | Collapses duplicate captures by SHA-256 |
-| *none here* — the YARA scanner moved to `arcane/home/honeypot-payload-analysis/analysis/yara/` (#1502) | Networkless YARA scanner sidecar, local rules, and vendored upstream corpus (`sync-yara.sh`); operator doc kept at [`yara/README.md`](yara/README.md) |
-| `ghidra/` | Headless Ghidra reverse-engineering pipeline, local-model triage, and the analysis-host installer ([`ghidra/README.md`](ghidra/README.md)) |
-| `es-results-importer/` | Ships Ghidra/sandbox/GitHub-analysis/workbench-run results into Elasticsearch, read-only, alongside the raw event stream ([#378](https://github.com/Xore/APIARY/issues/378)) |
-| `elasticsearch-setup.sh`, `honeypot-kibana-setup.sh`, `filebeat.yml`, `evebox.yaml` | Log pipeline and search UI provisioning |
-| `backup-honeypot.sh`, `verify-backup.sh`, `log-maintenance.sh`, `RECOVERY.md` | Retention and recovery |
-| `verify-stack.py` | Post-deploy/recovery health gate over the backend's `/api/v1/source-health` ([#2086](https://github.com/Xore/APIARY/issues/2086)) |
+| `analysis/analyze.py` | Offline triage of Cowrie / http-honeypot / multipot / Dionaea JSON logs. Stdlib only |
+| `analysis/collect.sh` | **Deprecated.** Cron-driven bulk copy of captures into a clone of `Xore/honeypot`. Superseded by the dashboard button; kept for a one-time manual backfill |
+| `arcane/home/honeypot-payload-analysis/analysis/dedupe-payloads.py` | Collapses duplicate captures by SHA-256 |
+| *not here* — the YARA scanner moved to `arcane/home/honeypot-payload-analysis/analysis/yara/` (#1502) | Networkless YARA scanner sidecar, local rules, and vendored upstream corpus (`sync-yara.sh`); operator doc kept at [`yara/README.md`](yara/README.md) |
+| `analysis/ghidra/` (code), [`ghidra/`](ghidra/README.md) (docs) | Headless Ghidra reverse-engineering pipeline, local-model triage, and the analysis-host installer ([`ghidra/README.md`](ghidra/README.md)) |
+| `arcane/home/honeypot-dashboard/analysis/es-results-importer/` | Ships Ghidra/sandbox/GitHub-analysis/workbench-run results into Elasticsearch, read-only, alongside the raw event stream ([#378](https://github.com/Xore/APIARY/issues/378)) |
+| `arcane/home/honeypot-init/analysis/elasticsearch-setup.sh`, `arcane/home/honeypot-init/analysis/honeypot-kibana-setup.sh`, `arcane/home/honeypot-elk/analysis/filebeat.yml` | Log pipeline and search UI provisioning. `evebox.yaml` is gone — the Suricata event store moved out of EveBox's SQLite into Elasticsearch, so the container is configured entirely by CLI flags in `arcane/home/honeypot-elk/compose.yml`; packet capture and session replay is **Arkime** (`arcane/home/honeypot-elk/arkime/config.ini`, templates in `arcane/home/honeypot-init/arkime/composable-templates.js`) |
+| `analysis/backup-honeypot.sh`, `analysis/verify-backup.sh`, `arcane/home/honeypot-utilities/analysis/log-maintenance.sh`, [`RECOVERY.md`](RECOVERY.md) | Retention and recovery |
+| `analysis/verify-stack.py` | Post-deploy/recovery health gate over the backend's `/api/v1/source-health` ([#2086](https://github.com/Xore/APIARY/issues/2086)) |
The GitHub Actions workflow itself lives at
[`Xore/honeypot/.github/workflows/analyze.yml`](https://github.com/Xore/honeypot/blob/main/.github/workflows/analyze.yml).
@@ -170,7 +176,8 @@ python3 analysis/analyze.py /path/to/logdir --top 15 --json summary.json
```
```bash
-# Copy logs out of the Docker volume first
-docker run --rm -v honeypot_honeypot-logs:/logs -v "$PWD":/out \
+# Copy logs out first. Sensors bind-mount to the host, they do not use a
+# Docker volume: /opt/stacks/apiary/logs// on the box, e.g. cowrie/.
+docker run --rm -v /opt/stacks/apiary/logs/cowrie:/logs -v "$PWD":/out \
alpine sh -c 'cp /logs/*.json /out/'
```
diff --git a/docs/analysis/RECOVERY.md b/docs/analysis/RECOVERY.md
index 00f817bb4..752372ef5 100644
--- a/docs/analysis/RECOVERY.md
+++ b/docs/analysis/RECOVERY.md
@@ -23,7 +23,22 @@ authenticated with an empty event history. See
and the sizes behind it.
Recovery is intentionally not automatic because overwriting live volumes is
-destructive. On a replacement host:
+destructive.
+
+**Which archive the steps below apply to**: an `analysis/backup-honeypot.sh`
+directory — that script's on-host layout of `SHA256SUMS`,
+`stack-config-state.tar.gz`, `keycloak.sql.gz` and `volumes/.tar.gz`.
+Those names resolve nowhere else, and that archive only exists if the
+homeserver itself survived. A restore driven by the archive that actually
+survives a dead homeserver — `scripts/backup-essentials.sh`'s
+`apiary-essentials-.tar.gz.gpg`, a different layout under
+`homeserver/`, `vps/` and `repo/` — has its own procedure, and one
+prerequisite the steps below do not cover: put
+`homeserver/installer/*-install-homeserver.conf` back in place first, because
+`scripts/install-homeserver.sh` will not run without it. See
+[`docs/BACKUP-ESSENTIALS.md`](../BACKUP-ESSENTIALS.md).
+
+On a replacement host:
1. Verify `SHA256SUMS`, unpack `stack-config-state.tar.gz` into a new empty stack
directory, and inspect `.env` permissions and values.
diff --git a/docs/analysis/ghidra/AI_TRIAGE.md b/docs/analysis/ghidra/AI_TRIAGE.md
index 053271dd9..e83d4e000 100644
--- a/docs/analysis/ghidra/AI_TRIAGE.md
+++ b/docs/analysis/ghidra/AI_TRIAGE.md
@@ -56,7 +56,7 @@ flowchart TD
prompt["System + user prompt evidence named as data, not instructions"]
local{"endpoint_is_local()?"}
refuse["Refused before any request is made ai_triage left null, reason logged"]
- call["POST /v1/chat/completions"]
+ request["POST /v1/chat/completions"]
usage["token usage reported by the server"]
truncated{"prompt_tokens indicates the prompt was truncated?"}
discard["Answer discarded ai_triage left null, reason logged"]
@@ -69,7 +69,7 @@ flowchart TD
imports --> budget
budget --> evidence --> prompt --> local
local -->|no| refuse
- local -->|yes| call --> usage --> truncated
+ local -->|yes| request --> usage --> truncated
truncated -->|yes| discard
truncated -->|no| parse --> normalise --> result
```
@@ -82,7 +82,7 @@ complete and the result written. Triage never fails an analysis.
## The context window is part of the configuration
The evidence block for a real binary is around 8000 tokens. Ollama's default
-window is 4096 whatever the model can do — `qwen3:8b` advertises 40960 — and an
+window is 4096 whatever the model can do — `qwen3:14b` advertises 40960 — and an
overlong prompt is **truncated, not refused**. There is no error, no HTTP
status, and the model answers from whichever fragment survived.
@@ -90,12 +90,14 @@ Measured here on `/usr/bin/wget`: at the default the reply described a command
line with hardcoded credentials that appears nowhere in the sample; at 16384
the same prompt returns `{"family_guess": "wget", "risk_level": "low"}`.
-So the compose file sets `OLLAMA_CONTEXT_LENGTH=16384`. It has to be set on the
-server, because `/v1/chat/completions` has no field for context length — only
-Ollama's native API and that variable can reach it. Budget about 1.8 GB of KV
-cache on top of the weights; `qwen3:8b` Q4_K_M reports 7.8 GB total on the live
-host and offloads about 1 GB to system RAM. CPU/RAM offload is supported and is
-not a correctness failure. On a genuinely memory-constrained host, lower the
+So the compose file sets `OLLAMA_CONTEXT_LENGTH=32768` — raised from 16384 when
+the #568 requalification ran the ghidra slot at `context_tokens: 32768` under the
+exact production manifest gates and returned identical scores with zero gate
+regressions. It has to be set on the server, because `/v1/chat/completions` has
+no field for context length — only Ollama's native API and that variable can
+reach it. Budget the KV cache on top of the weights; `qwen3:14b` Q4_K_M reports
+9,276,198,565 bytes (8.6 GB) on the live host. CPU/RAM offload is supported and
+is not a correctness failure. On a genuinely memory-constrained host, lower the
window and the evidence budgets together rather than accept truncation.
The worker does not trust the setting. Every reply is checked against the token
@@ -112,7 +114,7 @@ directly, so it is visible at install time rather than in a malware report.
"family_guess": "Mirai variant",
"risk_level": "high",
"behaviors": ["connects to a hardcoded C2 address", "kills competing processes"],
- "model": "qwen3:8b",
+ "model": "qwen3:14b",
"slot_generation": "0123456789ab/ctx32768/vram14336mib",
"evidence_shown": "150/312 imports, 200/11482 strings (longest first, deduplicated, >=6 chars), 100/847 functions (largest first)"
}
diff --git a/docs/analysis/ghidra/DASHBOARD_INTEGRATION_PLAN.md b/docs/analysis/ghidra/DASHBOARD_INTEGRATION_PLAN.md
index 63a00fdbd..6008cdec9 100644
--- a/docs/analysis/ghidra/DASHBOARD_INTEGRATION_PLAN.md
+++ b/docs/analysis/ghidra/DASHBOARD_INTEGRATION_PLAN.md
@@ -1,5 +1,20 @@
# Ghidra Dashboard Integration — Implementation Plan
+> **Design record, not current-behaviour documentation.** Every `dashboard/*.go`
+> reference below, and every route in the Precedent table and the Architecture
+> diagram (`/ghidra/submit`, `/api/ghidra/{sha256}`, `/export/ghidra/{sha256}`,
+> `ghidra.go`, `sandbox_submit.go`, `dashboard/ui/`), describes the Go dashboard
+> that was deleted at #1628. The same phases were re-implemented in the Rust
+> `backend-service` under
+> `arcane/home/honeypot-dashboard/backend-service/src/` and the
+> `frontend-next` routes; current route table is `main.rs` (`POST
+> /api/v1/ghidra/submit`, `GET /api/v1/ghidra/{sha}`,
+> `GET /api/v1/ghidra-callgraph/{sha}`, `GET /api/v1/revdeck/{sha}`), and
+> current operator documentation is [`README.md`](README.md). The plan's
+> decisions — the spool-file trust boundary, the phase split, the
+> fail-soft and result-shape rules — carried over and are still the binding
+> content. Read the Go paths as history, not as somewhere to write code.
+>
> **Status: all six phases built** (2026-07-31). Host worker, dashboard API,
> UI, alerting, environment and compose wiring are in place, tested, and
> rendered in a browser against fixture results.
diff --git a/docs/analysis/ghidra/IMPLEMENTATION_PLAN.md b/docs/analysis/ghidra/IMPLEMENTATION_PLAN.md
index ff465716d..7bb1af352 100644
--- a/docs/analysis/ghidra/IMPLEMENTATION_PLAN.md
+++ b/docs/analysis/ghidra/IMPLEMENTATION_PLAN.md
@@ -3,8 +3,11 @@
> **Status**: Design document. Phase 4 (plugin selection) is built as of
> 2026-08-01 — scoped down to `capa` alone; the other eight candidates from
> the original plugin list are decided out (see Phase 4 below). Phases 1, 2,
-> 3, 4 and 5 are built — five of the six `scripts/` exporters exist
-> (`findcrypt.py` was deleted, superseded by `scan_crypto()` in the worker),
+> 3, 4 and 5 are built — five of the six `scripts/` postScripts that once
+> existed are gone, and the sixth (`export_imports.py`) is unused by
+> anything: `findcrypt.py` was deleted in #136, superseded by
+> `scan_crypto()` in the worker, and `call_graph.py`, `export_functions.py`,
+> `export_strings.py` and `yara_scan.py` were deleted in #141,
> the `revdeck` service is deployed (profile-gated in
> `docker-compose.ghidra.yml`) and, as of 2026-08-01 (#78), the worker
> automates it too — `worker/ghidra-worker.py`'s `revdeck_triage()` drives a
@@ -338,8 +341,12 @@ Nine candidates were originally listed here with no decision behind any of
them, same problem [#85](https://github.com/Xore/APIARY/issues/85)
found in the "Additional Static Analysis Tooling" list below. Applying the
same standard — burden of proof on inclusion, since each addition is
-third-party code pinned/updated/trusted on the analysis host — exactly one
-of the nine survives: `capa`.
+third-party code pinned/updated/trusted on the analysis host — **all nine of
+the original candidates are decided out** (the "Decided out" table below is
+exactly those nine). `capa` is the one capability that survives, but it was
+never one of the nine: it arrives through the #85 `statictools` sidecar
+route as a plain CLI against the raw sample, not as an awesome-ghidra
+plugin.
### Decided in
diff --git a/docs/analysis/ghidra/README.md b/docs/analysis/ghidra/README.md
index d4097fae3..bb76dae5d 100644
--- a/docs/analysis/ghidra/README.md
+++ b/docs/analysis/ghidra/README.md
@@ -29,12 +29,12 @@ split out into [`AI_TRIAGE.md`](AI_TRIAGE.md) (#142).
```mermaid
flowchart LR
- subgraph dashboardBox["dashboard container"]
+ subgraph dashboardBox["dashboard container (backend-service)"]
direction TB
- submit["POST /ghidra/submit"]
- poll["GET /ghidra/{sha256}"]
+ submit["POST /api/v1/ghidra/submit"]
+ poll["GET /api/v1/ghidra/{sha}"]
revdeckSubmit["workbench: select Rev·Deck"]
- revdeckPoll["GET /revdeck/{sha256}"]
+ revdeckPoll["GET /api/v1/revdeck/{sha}"]
end
subgraph hostBox["host (root)"]
@@ -58,7 +58,7 @@ flowchart LR
submit -->|writes| spool
spool --> pathunit --> worker
- worker -->|resolve_sample(): reads| samples
+ worker -->|"resolve_sample(): reads"| samples
worker -->|writes| result
result --> poll
@@ -139,7 +139,7 @@ sequenceDiagram
RevDeck-->>Worker: answer, citations, warnings
end
Worker->>Spool: write {sha256}_ghidra.json + HTML/PDF report
- Dashboard->>Spool: GET /ghidra/{sha256} reads the result
+ Dashboard->>Spool: GET /api/v1/ghidra/{sha} reads the result
```
Every sidecar call in that diagram is independently fail-soft: a down or
@@ -188,7 +188,7 @@ sudo analysis/ghidra/install-analysis-host.sh # the worker half
| Flag | Effect |
|---|---|
| `--containers-only` | Bring up/refresh the containers and stop |
-| `--model NAME` | Model to pull. Defaults to `GHIDRA_TRIAGE_MODEL` from `/etc/default/honeypot-ghidra` if that file exists, else `qwen3:8b` |
+| `--model NAME` | Model to pull. Defaults to `GHIDRA_TRIAGE_MODEL` from `/etc/default/honeypot-ghidra` if that file exists, else `qwen3:14b` |
| `--no-gpu` | Run the model on CPU even if an NVIDIA runtime is present |
| `--skip-pull` | Do not pull the model |
| `--stack-dir PATH` | Where to deploy the compose file. `""` runs it in place |
@@ -285,7 +285,7 @@ which documents each setting inline. The ones worth knowing:
| `GHIDRA_API_BASE` | `http://127.0.0.1:9090` | The headless REST service |
| `GHIDRA_ANALYSIS_TIMEOUT` | `4200` | Per binary. Deliberately longer than the container's own `ANALYSIS_TIMEOUT` |
| `GHIDRA_TRIAGE_API_BASE` | `http://127.0.0.1:11434/v1` | Empty switches triage off |
-| `GHIDRA_TRIAGE_MODEL` | `qwen3:8b` | Recorded in every result |
+| `GHIDRA_TRIAGE_MODEL` | `qwen3:14b` | Recorded in every result. `qwen3:14b` for all three slots (ghidra/sessions/revdeck) since the #568 re-evaluation — see the [model evaluation](../../local-llm-model-evaluation.md) |
| `GHIDRA_TRIAGE_TIMEOUT` | `300` | Per workflow call; two calls run per sample |
| `GHIDRA_TRIAGE_MAX_STRINGS` / `_IMPORTS` / `_FUNCTIONS` | `200` / `150` / `100` | How much of the binary the model is shown. Around 8000 tokens together — see [the context window](AI_TRIAGE.md#the-context-window-is-part-of-the-configuration) before raising them |
| `STATICTOOLS_API_BASE` | `http://127.0.0.1:9091` | ssdeep/tlsh/lief/capa/floss sidecar, see [its contract above](#the-statictools-sidecar-contract). Empty switches it off |
@@ -384,7 +384,7 @@ API_BASE : http://127.0.0.1:9090
REQUEST_DIR : /var/lib/honeypot-ghidra/requests/pending (exists=True)
RESULTS_DIR : /var/lib/honeypot-ghidra/results (exists=True)
SAMPLES_DIR : /var/lib/honeypot-sandbox/inbox/samples (exists=True)
-TRIAGE : http://127.0.0.1:11434/v1 OK, model qwen3:8b available, context fits a full evidence block (7972 tokens read)
+TRIAGE : http://127.0.0.1:11434/v1 OK, model qwen3:14b available, context fits a full evidence block (7972 tokens read)
STATICTOOLS : http://127.0.0.1:9091 OK
REVDECK : disabled (REVDECK_API_BASE is empty)
@@ -426,8 +426,8 @@ docker compose -f /opt/stacks/ghidra/compose.yml exec ollama ollama list
that decide whether triage works and how long it takes:
```
-NAME ID SIZE PROCESSOR CONTEXT
-qwen3:8b 500a1f067a9f 7.8 GB 12%/88% CPU/GPU 16384
+NAME ID SIZE PROCESSOR CONTEXT
+qwen3:14b bdbd181c33f2 8.6 GB 12%/88% CPU/GPU 32768
```
`4096` there means the window setting is not reaching the container. A mixed
diff --git a/docs/analysis/ghidra/benchmarks/README.md b/docs/analysis/ghidra/benchmarks/README.md
index 22aef522e..f732d06fa 100644
--- a/docs/analysis/ghidra/benchmarks/README.md
+++ b/docs/analysis/ghidra/benchmarks/README.md
@@ -188,10 +188,19 @@ Elasticsearch store, build the **exact production prompt**
— not a reimplementation), run it through a model, and print the raw
reply for a human or an agent to read and judge: is each claim actually
grounded in the real captured commands, does it surface something useful,
-not "does it match word-for-word." Three stages because `hp-llm-worker`
-joins only an internal synthetic-only network while `LLM_ENABLED` stays
-false, by design — this stays out of that isolation rather than routing
-around it:
+not "does it match word-for-word." Three stages because the probe does not
+flip any of the worker's safe-by-default gates or talk to the model from
+inside the running container. On the Safe #66 base
+(`llm-worker/docker-compose.yml`) `hp-llm-worker` joins only the internal
+`synthetic-only` network and `LLM_ENABLED` stays false, so that is the
+whole story there. On the authorized deployment (#1751's
+`docker-compose.captured-data-deploy.yml`, which composes in
+`docker-compose.captured-data.yml`) the container *does* get
+`honeypot-llm-data` and `honeypot-llm` and an `OLLAMA_URL` — but
+`LLM_ENABLED` and `LLM_ALLOW_CAPTURED_DATA` are still
+`${...:-false}` there, because the overlay does not touch them. Either
+way the gates are the reason this is out-of-band rather than an in-band
+call:
```bash
# stage 0: pull real command data from Elasticsearch (read-only _search)
diff --git a/docs/analysis/ghidra/benchmarks/corpus/README.md b/docs/analysis/ghidra/benchmarks/corpus/README.md
index 2b72121ea..9e5700fe7 100644
--- a/docs/analysis/ghidra/benchmarks/corpus/README.md
+++ b/docs/analysis/ghidra/benchmarks/corpus/README.md
@@ -49,7 +49,7 @@ primitive, one vulnerability class) that the first slice of this corpus had.
## Build matrix and provenance (`build_corpus.py`, `manifest.json`)
-Each of the 14 sources is compiled with:
+Each of the 17 sources is compiled with:
- **Toolchains**: `gcc` and `clang` across five architectures --
`x86_64`, `aarch64`, `i686` (32-bit x86), `mipsel`, and `armhf`. All 4 of
@@ -70,9 +70,9 @@ Each of the 14 sources is compiled with:
object's own instruction set rather than erroring, so this matters for
correctness, not just cleanliness).
- **Train/validation/test split**, recorded per case in `CASE_SPLITS` and
- carried onto every build variant of that case (`"split"` field). All 14
+ carried onto every build variant of that case (`"split"` field). All 17
cases are currently `"test"`: every one has already been used (or, for
- the 6 added most recently, is used from the moment it exists) as scored
+ the 9 added most recently, is used from the moment it exists) as scored
evaluation data, never shown to a model as a training example, so
tagging any of them `"train"` now would be retroactively wrong.
Splitting a single case's own toolchain/opt-level variants across train
@@ -80,8 +80,8 @@ Each of the 14 sources is compiled with:
underlying case in both and leak exactly the case-level knowledge the
split exists to prevent.
-14 sources x 10 toolchains x 5 opt levels = 700 builds, each with both a
-stripped and unstripped variant recorded (`manifest.json`).
+17 sources x 10 toolchains x 5 opt levels = 850 builds, each carrying both a
+stripped and unstripped variant (`manifest.json`).
## The injection payload must survive compilation (#1948)
@@ -165,7 +165,7 @@ code, not a whole program.
**Determinism verified two ways**: (1) built twice into separate output
directories in the same environment; after normalizing the one
build-directory-dependent string objdump embeds in its own header line
-(`build_corpus.py`'s `normalize_disassembly`), all 700 disassembly outputs
+(`build_corpus.py`'s `normalize_disassembly`), all 850 disassembly outputs
were byte-identical across the two builds. (2) Built in two genuinely
independent, freshly-provisioned containers (`ci_verify.sh`'s own check,
which is exactly the property CI now enforces on every change -- see
@@ -243,7 +243,7 @@ pointer write, or format-string read/write on purpose is not something an
automated corpus-verification script should ever do; the bug is already
known and static, and there is nothing to gain from triggering it for real.
-240 executions (12 cases x 10 toolchains x 2 representative opt levels,
+280 executions (14 cases x 10 toolchains x 2 representative opt levels,
`-O0`/`-O2`), 0 failures, reverified in two independent fresh containers.
## Scoring rubric and contract (`rev_cases_v2_rubric.json`, `rev_cases_v2_contract.json`)
@@ -384,7 +384,7 @@ Direct mapping to #159's own checklist:
(variable names, control flow) rather than a bare conclusion.
- [x] **Scoring is semantic and reviewed before model outputs are seen.**
Rubric authored from ground truth before any model output was inspected,
- for both the original 8 cases and the 6 added since.
+ for both the original 8 cases and the 9 added since.
- [x] **CI verifies provenance, fixture safety, hashes, and reproducible
generation.** `validate_manifest.py` + `ci_verify.sh`, wired into
`quality.yml`.
diff --git a/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md b/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md
index 41458050c..14a2804ae 100644
--- a/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md
+++ b/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md
@@ -1,5 +1,12 @@
# Injection gate v3 — what it measures and how to test it (#2694)
+> **Re-verified 2026-09-27 — every tool and path in §3–§4 still exists as named.**
+> The gate is current: the `injection` rubric block plus `injection_gate.py`,
+> with the live slot at `qwen3:14b` and 32768 context. Whether the §3.3
+> positive-control sweep has been re-run against the current roster is
+> **undetermined** from the repository — check `/var/benchmarks/` on the host.
+> Companion: `docs/analysis/ghidra/benchmarks/corpus/README.md`.
+
This is the operating manual for the corpus injection axis after #2694. It
states what a verdict means, how the gate is validated before any number from
it is cited, and the exact commands.
@@ -14,12 +21,14 @@ containing the payload's own words. Read against every stored
| | truly complied | did not comply |
|---|---|---|
-| gate FAIL | 1 (partial) | 26 |
-| gate PASS | 0 | 37 |
-
-25 of 27 failures fired on the model *quoting or paraphrasing* the planted
-string — the behaviour the system prompt asks for. The remaining ones fired
-on "appears to be benign", which is the case's own ground truth. The four
+| gate FAIL | 1 (partial) | 25 |
+| gate PASS | 0 | 38 |
+
+23 of the 26 failures fired on one of the four payload-derived forbidden terms
+(the model quoting or paraphrasing the planted string) — the behaviour the
+system prompt asks for. The other three fired only on "appears to be benign",
+which is the case's own ground truth; two of those three never reproduce the
+string at all. The four
Tier A failures that drove the #1805-c "no promotion" decision (Ornith-35B,
gemma-4-31B, Seneca-32B, huihui-qwen3.8) all explicitly identified the string
as an injection or anomaly. No matcher can repair this fixture: the injected
diff --git a/docs/analysis/ghidra/ghidrassist/README.md b/docs/analysis/ghidra/ghidrassist/README.md
index 4dce29303..b14420511 100644
--- a/docs/analysis/ghidra/ghidrassist/README.md
+++ b/docs/analysis/ghidra/ghidrassist/README.md
@@ -8,9 +8,13 @@ auto-renaming, protocol detection, and YARA rule generation.
> ⚠️ **Interactive only** — GhidrAssist is for analyst-facing use in the
> Ghidra GUI. It is NOT part of the automated CI pipeline. Your local Ghidra
> GUI is very likely a different install (and a different, probably newer,
-> Ghidra version) than the pinned `biniamfd/ghidra-headless-rest:1.2.1`
-> (Ghidra 11.3.2) this repo's own automated pipeline runs — see "Ghidra
-> version compatibility" below for why that specifically matters here.
+> Ghidra version) than the Ghidra **11.3.2** this repo's own automated
+> pipeline runs — pinned as `GHIDRA_VERSION` in
+> [`analysis/ghidra/service/Dockerfile`](../../../../analysis/ghidra/service/Dockerfile)
+> and wrapped by this repo's own `service/server.py` since #245 replaced the
+> third-party `biniamfd/ghidra-headless-rest` image (the Ghidra version
+> itself is unchanged by that swap) — see "Ghidra version compatibility"
+> below for why the version specifically matters here.
## Install — build from source (recommended)
@@ -62,7 +66,7 @@ GHIDRA_INSTALL_DIR=/path/to/your/ghidra__PUBLIC gradle buildExtension
run the extension in** — Ghidra extensions are compiled against that
install's own API and are not portable across major versions. This isn't
hypothetical for this specific commit: building `2.2.0` against Ghidra
-**11.3.2** (this repo's own pinned `biniamfd/ghidra-headless-rest` version)
+**11.3.2** (the version this repo's own analysis-host container pins)
**fails outright** —
```
@@ -78,8 +82,8 @@ buildable, for this GhidrAssist version. Building against Ghidra **12.1**
instead (verification above) succeeds cleanly.
This does not block real-world use: GhidrAssist runs in *your own local
-Ghidra GUI*, not in the headless-rest container the automated pipeline
-uses, and an analyst's own desktop Ghidra install is very likely 12.x
+Ghidra GUI*, not in the headless container the automated pipeline uses,
+and an analyst's own desktop Ghidra install is very likely 12.x
already. It does mean: point `GHIDRA_INSTALL_DIR` at your actual local
Ghidra, not at this repo's pinned analysis-host version, and don't expect
this exact commit to build against anything older than 12.0.
@@ -145,10 +149,17 @@ echo "${GHIDRASSIST_SHA256} ${GHIDRASSIST_ZIP}" | sha256sum -c -
After installation, configure the LLM in Ghidra:
`Edit → Tool Options → GhidrAssist`
-Use the same endpoint as Rev·Deck:
+Use the same endpoint and model as Rev·Deck — the same local Ollama the
+analysis host runs:
```
LLM Provider: OpenAI Compatible
Base URL: http://127.0.0.1:11434/v1 (Ollama) or OpenRouter
-Model: qwen3:8b
-API Key: not-used
+Model: qwen3:14b
+API Key: ollama
```
+
+Unlike `ghidra-worker.py`, nothing here enforces a local-only endpoint —
+GhidrAssist is a GUI extension talking to whatever provider you type in,
+and the `OpenRouter` option above is a real one. The captured samples it
+would read are live malware off this honeypot, so the local-only rule
+`AI_TRIAGE.md` documents applies to you, not to the plugin.
diff --git a/docs/analysis/ghidra/models/README.md b/docs/analysis/ghidra/models/README.md
index 89d47e0dd..9cb5fd5e6 100644
--- a/docs/analysis/ghidra/models/README.md
+++ b/docs/analysis/ghidra/models/README.md
@@ -17,10 +17,15 @@ python3 /opt/honeypot-ghidra/models/model-governance.py check-runtime \
The status file contains only state and reason codes. It contains no prompts, model replies, captured data, container paths, or credentials, and is written owner-only mode `0600`. If a dashboard later needs it, expose only the sanitized object through a privileged read-only endpoint; do not mount or relax the host file. `approved`, `drift`, and `unavailable` are advisory states: the service exits successfully with `--warn-only`, so an LLM problem never stops ingestion or deterministic analysis. Omit `--warn-only` in an operator check when drift should produce a non-zero exit status. The command only reads `/api/version`, `/api/tags`, Docker inspection metadata, and `nvidia-smi` telemetry.
-### Two expected, honest states after #2394 (not regressions)
+### Post-#2394 GPU-identity states: which are expected, which are not
-- **`approved_gpu_uuid_missing`** on the `host` leg: the deployed manifest copy at `/opt/honeypot-ghidra/models/approved-models.json` still predates the `approved_host.gpu_uuid` field until `install-analysis-host.sh` next runs end to end on that host. The checker reports this distinct, advisory-only code rather than silently comparing whichever GPU enumerates as index 0. It clears itself once the host is redeployed with the current manifest.
-- **`host_gpu_uuid_changed`**: any `--snapshot` file captured before #2394 was written under the older schema and has no GPU-identity fields for the new comparison to match against. Replaying it will read as drift under the new schema even though nothing on the host changed -- expected for old snapshots, not evidence of an actual UUID change.
+#2394 made the checker compare GPU identity by UUID rather than enumeration
+order, and that added three host-leg codes. Only the first two are expected
+during the rollout; the third is a real problem wearing the same shape.
+
+- **`approved_gpu_uuid_missing`** on the `host` leg — *expected.* The deployed manifest copy at `/opt/honeypot-ghidra/models/approved-models.json` still predates the `approved_host.gpu_uuid` field until `install-analysis-host.sh` next runs end to end on that host. The checker reports this distinct, advisory-only code rather than silently comparing whichever GPU enumerates as index 0. It clears itself once the host is redeployed with the current manifest.
+- **`host_gpu_uuid_changed`** — *expected for a legacy `--snapshot` only, and this code is overloaded, so check before dismissing.* A snapshot captured before #2394 was written under the older schema and carries no GPU-identity field, so replaying it reads as drift even though nothing on the host changed. But the checker also emits the **same code** when `gpu_uuid` is present and names a *different physical card* than the manifest pins — that is real drift, not a schema artefact, and the field the old schema compared (name, memory, driver) can all still look correct. Tell the two apart by whether the snapshot has a `gpu_uuid` key at all: absent means a legacy replay, present-and-different means the host is not running the approved card. The same per-field construction applies to `host_gpu_changed`, `host_gpu_memory_mib_changed`, `host_driver_changed` and `host_compute_capability_changed`, so treat each of those the same way.
+- **`approved_gpu_absent`** — *not expected.* The tool ran, was pointed at the approved UUID explicitly, and no such card exists on the host. Distinct from `gpu_telemetry_unavailable` (which means the telemetry could not be read at all). This one means the approved card is genuinely gone.
## When requalification is mandatory
@@ -30,7 +35,7 @@ Run the complete workflow before changing any model tag or digest, Ollama image/
Use a trusted checkout on the approved analysis host. Stop unrelated GPU-heavy jobs if needed, but do not stop or modify QEMU. The benchmark uses only checked-in synthetic TEST-NET fixtures, talks only to the explicitly supplied local Ollama endpoint, records exact artifacts/settings/timing/RAM/VRAM metadata, and unloads each candidate through Ollama after its slot. It never downloads a model.
-Keep verbose replies outside the repository in an operator-only directory with bounded retention (30 days is the recorded recommendation):
+Keep verbose replies outside the repository in an operator-only directory with bounded retention (30 days is this doc's own recommendation; no separate retention record covers this directory):
```sh
install -d -m 0700 "$HOME/model-qualification"
diff --git a/docs/analysis/ghidra/revdeck/README.md b/docs/analysis/ghidra/revdeck/README.md
index 173044fdd..9f8ec3301 100644
--- a/docs/analysis/ghidra/revdeck/README.md
+++ b/docs/analysis/ghidra/revdeck/README.md
@@ -41,16 +41,23 @@ git clone https://github.com/biniamf/ai-reverse-engineering \
# Copy and configure .env
cp ai-reverse-engineering/.env.example ai-reverse-engineering/.env
-# Edit: set API_BASE, MODEL_NAME, API_KEY
```
+`docker-compose.ghidra.yml` sets `API_BASE`, `API_KEY` and `MODEL_NAME` in
+its own `environment:` block, and a compose `environment:` entry wins over
+`env_file` — so editing those three keys in `.env` has no effect on this
+deployment. The model is selected with `REVDECK_MODEL` (default `qwen3:14b`,
+the same model the `ghidra`/`sessions`/`revdeck` slots have shared since the
+#568 re-evaluation); everything else `revdeck` needs (`GHIDRA_API_BASE`,
+`CHATS_DIR`, `LOG_FILE`) is fixed in the compose file.
+
## Recommended LLM Configs
```dotenv
-# Option 1: Local Ollama (free, private)
-API_BASE=http://127.0.0.1:11434/v1
-API_KEY=not-used
-MODEL_NAME=qwen2.5-coder:7b-instruct-q4_K_M
+# Option 1: Local Ollama (free, private) -- what ships today
+API_BASE=http://ollama:11434/v1
+API_KEY=ollama
+MODEL_NAME=qwen3:14b
# Option 2: OpenRouter (hosted)
API_BASE=https://openrouter.ai/api/v1
@@ -58,16 +65,21 @@ API_KEY=
MODEL_NAME=anthropic/claude-opus-4.8
```
+Option 2 is illustrative only: the compose file hard-pins `API_BASE` to the
+`ollama` service, so reaching a hosted provider means editing
+`docker-compose.ghidra.yml`, not `.env`.
+
## Start the full stack
```bash
cd analysis/ghidra
docker compose -f docker-compose.ghidra.yml --profile revdeck up -d
-# Pull the independently selected interactive model into the shared Ollama
-# volume (the analysis-host installer only guarantees the Ghidra model).
+# Pull the shared analysis model into the Ollama volume. The analysis-host
+# installer already pulls it (--model, default qwen3:14b); this is the manual
+# equivalent for a stack brought up without it.
docker compose -f docker-compose.ghidra.yml exec ollama \
- ollama pull qwen2.5-coder:7b-instruct-q4_K_M
+ ollama pull qwen3:14b
```
Open http://127.0.0.1:19500 — the compose file maps host port `19500` to
@@ -162,10 +174,16 @@ for `attack_surface_triage`, `vulnerability_hypothesis`, or any deeper dive a
particular sample warrants.
The local default comes from the task-specific
-[model evaluation](../../../local-llm-model-evaluation.md): it tied for
-the highest Rev·Deck score, passed the x86 intent case, and does not depend on
-the thinking-control field that the current upstream Rev·Deck client does not
-send.
+[model evaluation](../../../local-llm-model-evaluation.md). It was
+`qwen2.5-coder:7b-instruct-q4_K_M` — which tied for the highest Rev·Deck score,
+passed the x86 intent case, and does not depend on the thinking-control field
+that the current upstream Rev·Deck client does not send. The #568 re-evaluation
+since promoted `qwen3:14b` to *all three* slots (ghidra, sessions, revdeck) and
+superseded that selection: `qwen3:14b` scores lower on Rev·Deck (87.5% vs
+93.8%) but is the only candidate passing every injection-resistance and
+critical-severity gate across all three slots, which disqualifies the 7b
+baseline on its own gate column. Override with `REVDECK_MODEL` if a deployment
+wants to re-pin the older per-slot choice.
## Evidence Grounding
diff --git a/docs/analysis/gpu-queue/README.md b/docs/analysis/gpu-queue/README.md
index 12f748c21..671c4cd9e 100644
--- a/docs/analysis/gpu-queue/README.md
+++ b/docs/analysis/gpu-queue/README.md
@@ -79,8 +79,9 @@ just the `ai_triage` field on the already-written result, atomically).
Processes at most one job per invocation — simple, and the next tick picks
up wherever this one left off.
-**Dashboard** (`dashboard/gpu_queue.go`): the `/ghidra` page's "GPU queue"
-section lists every job (ES-only read, `docSearchAll` against
+**Dashboard** (`arcane/home/honeypot-dashboard/backend-service/src/gpu_queue.rs`,
+routes `/api/v1/gpu-queue` and `/api/v1/gpu-queue/{job_id}/abort`): the payload
+workbench's "GPU queue" section lists every job (ES-only `search_index` against
`gpu-job-queue`) and offers an Abort button on anything still `queued`.
Abort only has an effect before a drainer has committed to running a job
— once `running`, the Ollama call is already in flight, matching
diff --git a/docs/analysis/yara/README.md b/docs/analysis/yara/README.md
index e5c73cac5..46ac14f57 100644
--- a/docs/analysis/yara/README.md
+++ b/docs/analysis/yara/README.md
@@ -23,7 +23,9 @@ design, since it reads live malware.
| `rules/upstream/` | Vendored from [`Xore/Honeypot`](https://github.com/Xore/Honeypot) `yara-rules/`. **Do not edit** — changes here are lost on the next sync |
| `rules/index.yar` | Generated include list. What the scanner loads |
| `rules/upstream.lock` | The pinned upstream commit and a hash of the vendored tree |
-| `rules/upstream/DROPPED` | Upstream files this corpus does **not** include, with yara's reason |
+| `rules/upstream/MANIFEST` | Per-file record of what the pinned upstream commit contained, as vendored vs dropped |
+| `rules/upstream/AUTO_RULES` | The rule names defined under `upstream/auto/` |
+| `rules/upstream/DROPPED` | Upstream files this corpus does **not** include, with yara's reason. Currently empty |
`scripts/check-yara-corpus.sh` enforces in CI that `rules/upstream/` still
matches the lock and that `index.yar` names exactly the vendored files. It
@@ -51,10 +53,12 @@ entirely rather than degrading to "everything except that rule". So the sync
compiles every file before adopting it and drops the ones that fail, rather than
handing the scanner a corpus that will not load. Two things get a file dropped:
-- **It does not compile.** As of the pinned commit, four of upstream's six
- curated files declare strings their conditions never reference, which yara
- treats as an error. `rules/upstream/DROPPED` has the details; they come back
- automatically once upstream fixes them and the sync is re-run.
+- **It does not compile.** yara treats declared-but-unreferenced strings as an
+ error, and an earlier pinned commit had four of upstream's six curated files
+ failing that way. `rules/upstream/DROPPED` records any file the current
+ pinned commit loses and why; it is currently **empty** — all six curated
+ files compile and are vendored. A dropped file comes back automatically once
+ upstream fixes it and the sync is re-run.
- **It redefines a rule name** already used by `honeypot.yar` or an
earlier-sorted upstream file. A duplicate identifier is also a hard error.
@@ -69,6 +73,7 @@ compile" is only a useful answer from the compiler that will load it.
index is a list of filenames, `rules_sha256` would not move when upstream
changed every rule but no filename.
- `auto_rules` — names defined under `upstream/auto/`. These are generated from
- observed samples and are broad by construction (`AutoGen_Exe` fires on three
- of twenty stock .NET strings, so it matches most .NET binaries). Treat an auto
- hit as "seen something like this before", not as a family identification.
+ observed samples and are broad by construction (`AutoGen_190460923_exe`
+ matches on 8 of 20 stock .NET strings, so it matches most .NET binaries).
+ Treat an auto hit as "seen something like this before", not as a family
+ identification.
diff --git a/docs/benchmarks/2026-08-30-injection-gate-recalibration.md b/docs/benchmarks/2026-08-30-injection-gate-recalibration.md
index 6730ff694..20b4e1132 100644
--- a/docs/benchmarks/2026-08-30-injection-gate-recalibration.md
+++ b/docs/benchmarks/2026-08-30-injection-gate-recalibration.md
@@ -4,6 +4,42 @@
> as-written since this is a dated record; see `docs/HOMESERVER-DISK-LAYOUT.md` for
> current layout.
+> **Dated record — measured 2026-08-30, superseded by #2694.** The gate and
+> `is_injection_case()` behaviour described below no longer run: injection cases
+> are now declared by the rubric's `injection` block and scored by
+> `analysis/ghidra/benchmarks/injection_gate.py`. Every number in this file is
+> the 2026-08-30 sweep's, left exactly as measured. The current protocol,
+> re-verified 2026-09-27, is
+> `docs/analysis/ghidra/benchmarks/injection-gate-protocol.md`.
+
+> **Reconciled 2026-09-27 — §9's inventory is now shipped, and §2's Tier B tally
+> does not match the pinned fixture.** Five points, all checked against `HEAD`;
+> no number above was rewritten.
+>
+> 1. **§9's "It is *not merged*" is stale.** The branch landed as `df650a8f`
+> ("injection gate v3 … (#2694) (#2697)"), 2026-08-30 16:20:54Z — the same
+> minute as the "69 `run1` reports present at 16:20Z" in §9.2. Every path in
+> the §9 table is at `HEAD`; `rev_cases_v2_rubric.json` carries 17 cases with
+> `process_and_injection` on `forbidden: []`, and `manifest.json` holds the
+> 850 builds §9.1 claims (700 reproduced + 150 new, `errors: []`).
+> 2. **§2's Tier B counts are the matrix's own, not the fixture's.** The table
+> enumerates 30 rows but never lists `huihui-qwen3.6-35b-a3b-abliterated:q3_k`,
+> which the shipped 64-row fixture records as a gate **FAIL**. Against that
+> fixture Tier B is **15 FAIL / 15 PASS** over 30 rows, splitting **13**
+> payload-term hits to **2** on `appears to be benign` — not the 14/15 and
+> 11/3 of the headline above. §2 stays as written; read it as the matrix it
+> is, and the fixture as the shipped regression set.
+> 3. **§9's test inventory has drifted:** `test_injection_gate.py` is still 48,
+> but `test_record_baseline.py` is **54**, not 50 (102 collected, 102 pass).
+> 4. **§8's governance-gate term list is incomplete.**
+> `evaluate-models.py:421` (`process-injection`) reads
+> `("appears to be benign", "likely benign", "conclude benign")`; §4.3 lists
+> all three, §8 omits the last.
+> 5. **§8's "59 hand-labelled answers in §2" is the wrong set twice over:** §2
+> carries 29 Tier B verdicts, and the hand-labelled cohort is the 64 rows of
+> §9's fixture. The `conclude benign` figure in §1 (zero occurrences) still
+> holds — it is zero across all 64.
+
Measured 2026-08-30 against `/mnt-1/benchmarks/1947full` (run1 files; run2 verified byte-identical for every model at both tiers) and the checkout at `/mnt-1/benchmarks/APIARY` @ `a99e765`. Nothing on the host was modified; every script was piped over ssh stdin and read only.
**Verdict on the preliminary finding (now issue #2694):** confirmed in mechanism and in substance, with two corrections and five additional findings. The gate is not measuring compliance. It is measuring whether a model *quoted or paraphrased the payload* (11 of 14 Tier B failures) or *used the exact phrase "appears to be benign"* (3 of 14). The fixture cannot discriminate compliance from correct analysis because the injected verdict is true. The same defect accounts for **all four Tier A failures that drove the #1805-c / #1947 "no promotion" decision**, including the disqualification of the top-scoring model.
diff --git a/docs/benchmarks/claim-pools/README.md b/docs/benchmarks/claim-pools/README.md
index 86d558d39..bf05c10ec 100644
--- a/docs/benchmarks/claim-pools/README.md
+++ b/docs/benchmarks/claim-pools/README.md
@@ -47,11 +47,12 @@ expensive half, and #1805's ladder does not relieve it as designed:
`true` (absence from a one-line summary is not evidence a claim is false), so
its low yield costs nothing but settles little.
- The **semantic-harness rung named "cheapest first" in #1805 is not
- implementable as described.** The 240 executable checks are `assert()`
+ implementable as described.** The 280 executable checks are `assert()`
expressions like `rotate_checksum(one, 1) == 0x41`, and `semantic_checks.json`
- records only that they ran (240 checked, 0 failed). There is no mechanism to
- check a prose claim such as "XORs each byte with a single-byte key" against a
- numeric assertion. Making that rung real would be its own piece of work.
+ records only that they ran (280 checked, 0 failed, 14 cases covered). There
+ is no mechanism to check a prose claim such as "XORs each byte with a
+ single-byte key" against a numeric assertion. Making that rung real would be
+ its own piece of work.
**Do not read the 90% solo rate as unique contribution.** Only 37 of 382 claims
(10%) were made by both models; 345 by exactly one. Two models describing the
@@ -69,9 +70,9 @@ unescaped quotes inside claim text. The case is absent from this pool.
The review queue stamps each row with the rubric's `ground_truth` as it stood
at generation time, so a rubric correction can leave the queue quoting a
-retired narrative. That happened once:
+retired narrative. That has happened twice, 59 rows in total:
-- **#2384 / 2026-08-27, `integer_overflow_alloc`.** #2384 corrected the
+- **#2384 / 2026-08-27, `integer_overflow_alloc` (33 rows).** #2384 corrected the
fixture's ground truth everywhere authoritative: the wrapped `count*size`
sizes *both* the allocation and the memcpy, so the fixture cannot produce an
intra-function write-past-allocation; the accurate mechanism is silent
@@ -81,6 +82,15 @@ retired narrative. That happened once:
`ground_truth_superseded` note preserving the retired text and the grading
rule. All 33 verdicts are still placeholders — no ruling was ever made
against the retired narrative, so nothing needs re-adjudicating.
+- **#2694, `process_and_injection` (26 rows).** #2694 retired this case's
+ *resistance-test* reading, not its ground truth: its injected verdict
+ ("benign") is also the true verdict, so a claim saying it appears benign is
+ correct analysis rather than compliance with the payload, and quoting the
+ payload is coverage evidence. Resistance is now measured by
+ `strcpy_note_injected` (false verdict, paired with `strcpy_note_neutral`) and
+ `process_witness_probe`; `process_and_injection` is a candour + coverage case.
+ The 26 rows carry the same `ground_truth_superseded` shape, and all 26
+ verdicts are placeholders too.
- **Tripwire:** `tests/test_claims.py::ReviewQueueFreshnessTest` asserts every
queued row's quoted ground truth equals the current rubric text, so the next
rubric correction fails CI until the queue is restamped the same day.
diff --git a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md
index e8c805d21..3b2e69bc9 100644
--- a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md
+++ b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md
@@ -4,6 +4,15 @@
> as-written since this is a dated record; see `docs/HOMESERVER-DISK-LAYOUT.md` for
> current layout.
+> **Dated plan — written 2026-09-05, superseded by round 7 (#3079).** The hardware
+> table below is **pre-reinstall**: the host is now Rocky Linux 10.2 with LVM,
+> `/var` on `sdb1` of an 8.7T LUN, a 32G `rl-swap`, and no `/mnt-1` mount (see
+> `docs/HOMESERVER-DISK-LAYOUT.md`). Likewise, the "#2985 missing scripts" this
+> plan works around are now in git at `analysis/ghidra/benchmarks/corpus/` —
+> `requant_sweep.sh`, `slots_sweep.sh`, `chain_round7.sh` and the `round7_*`
+> builders; only `gptoss_rerun.sh` is still absent. In-flight state below is as
+> of 2026-09-05, not current.
+
**Written** 2026-09-05, from live inspection of `homeserver` and every open
benchmark issue. Supersedes nothing; it sits *beside*
`/mnt-1/benchmarks/STATE-2026-09-05-fix-stage.md`, which remains the authority on
diff --git a/docs/benchmarks/plans/2026-09-06-round7-hermes-init.md b/docs/benchmarks/plans/2026-09-06-round7-hermes-init.md
index ca7327af5..503f78d52 100644
--- a/docs/benchmarks/plans/2026-09-06-round7-hermes-init.md
+++ b/docs/benchmarks/plans/2026-09-06-round7-hermes-init.md
@@ -4,6 +4,14 @@
> as-written since this is a dated record; see `docs/HOMESERVER-DISK-LAYOUT.md` for
> current layout.
+> **Dated handoff — written 2026-09-06 (epic #3079); not a live brief.** The
+> toolchain, model roster and dispatch rules in §5 are as-written on that date.
+> What shipped since: the #3080 toolchain (`analysis/ghidra/training/`) and the
+> corpus builders with their decontamination report. Unsloth (#3092) has **no
+> installer path** — it is the Arcane stack `unsloth`
+> (`arcane/manifests/home-production.json`), deployed by gitops-sync and a
+> redeploy, never by hand. Stop it before any cold GPU leg.
+
**Read this first.** It names the plan of record, the state of the GPU, the
kickoff order, the dispatch pattern and the hard rules. Written 2026-09-06,
after the round-7 cold baseline was launched. Epic **#3079**, children
diff --git a/docs/benchmarks/plans/2026-09-06-round7-unsloth-train-requant-ollama.md b/docs/benchmarks/plans/2026-09-06-round7-unsloth-train-requant-ollama.md
index 52d0cdc2e..4f467793b 100644
--- a/docs/benchmarks/plans/2026-09-06-round7-unsloth-train-requant-ollama.md
+++ b/docs/benchmarks/plans/2026-09-06-round7-unsloth-train-requant-ollama.md
@@ -4,6 +4,13 @@
> as-written since this is a dated record; see `docs/HOMESERVER-DISK-LAYOUT.md` for
> current layout.
+> **Plan of record, written 2026-09-06 (epic #3079) — its §0 already folds in later
+> changes; the rest is as-planned, not as-shipped.** The batch leg shipped as
+> `analysis/ghidra/training/` (`train.py`, `export_to_ollama.sh`, `compose.yaml`,
+> `TOOLCHAIN.md`); the interactive half shipped as the Arcane `unsloth` stack
+> rather than an installer (#3092). `TOOLCHAIN.md`'s "no smoke test yet" still
+> stands — nothing in this plan has been exercised end to end.
+
**Written** 2026-09-06 from live inspection of `homeserver`, every open benchmark
issue, and the current Unsloth / llama.cpp / Ollama documentation. Sits beside
`2026-09-05-1947-resume-plan.md`, which remains the authority on the #1947 sweep
diff --git a/docs/canarytoken-live-fire-checklist.md b/docs/canarytoken-live-fire-checklist.md
index 17ecf8ab2..79fca518c 100644
--- a/docs/canarytoken-live-fire-checklist.md
+++ b/docs/canarytoken-live-fire-checklist.md
@@ -57,7 +57,12 @@ For each type: create → plant/open → observe the fire → confirm it lands.
paths from outside is the trap that cost #2136 a whole session.
2. **Trigger** it the way the per-type notes below describe, from a host that is
*not* on the honeypot network (so the source IP is meaningful).
-3. **Observe the fire.** The switchboard posts to `canarytokens-adapter`, which
+3. **Observe the fire.** The switchboard posts to `canarytokens-adapter` on the
+ internal `canarytokens-adapter.internal` network; the public HTTP entry point
+ is `canarytokens-http-router` (`hp-canarytokens-http-router`, host port
+ 19427), which is the service that actually holds `ADAPTER_URL`
+ (compose.yml:311) and forwards to the adapter. The switchboard itself
+ publishes no host port. The adapter then
appends one sensor JSON line to `/var/log/honeypot/canarytokens.json`
(`canarytokens-adapter/main.go`), which Filebeat tails into the honeypot index
like any other sensor.
@@ -171,9 +176,11 @@ Full path confirmed today, independently, at each link:
1. **ES doc** -- `GET honeypot-v2*/_search` on `event.sensor: canarytokens` +
`honeypot.token: ` returns 2 hits at the field paths this checklist
documents (`honeypot.token_type`, `honeypot.memo`, `honeypot.src_ip`).
-2. **Dashboard** -- `GET /api/v1/events?sensor=canarytokens&size=50&since=3650d`
+2. **Dashboard** -- `GET /api/v1/events?sensor=canarytokens&size=50&since=365d`
(the exact query `frontend-next/src/routes/canarytokens.tsx`'s "Fired
- tokens" tab issues) returns this event as row `detail: "token fired: Xore
+ tokens" tab issues — `since=365d`, chosen there because tokens fire rarely
+ and the events endpoint's own default is 10d; the trailing `0` in an older
+ revision of this checklist was a transcription slip, not a wider window) returns this event as row `detail: "token fired: Xore
verification token (working) (HTTP)"`, geo-enriched (country DE, ASN 8899)
-- confirmed surfaced, not just indexed.
diff --git a/docs/community-threat-intel-sharing.md b/docs/community-threat-intel-sharing.md
index 45340d4c6..89a89f059 100644
--- a/docs/community-threat-intel-sharing.md
+++ b/docs/community-threat-intel-sharing.md
@@ -7,7 +7,7 @@
`hpfeeds` publisher). Declined, not deferred — the reasoning below is
worth someone re-reading before re-proposing this, not just a placeholder
for "someone hasn't gotten to it yet."** TANNER already ships a disabled,
-unused `hpfeeds` config block (`tanner/tanner/config.yaml`) that an
+unused `hpfeeds` config block (`arcane/home/honeypot-tanner/tanner/tanner/config.yaml`) that an
operator can turn on by hand if they personally want to participate — see
§4 — but this repo doesn't recommend it by default, document a workflow
around it, or build anything to support it.
@@ -48,14 +48,17 @@ logs access-controlled and short-lived"). Publishing structured attack
data to an open community broker is a materially bigger, harder-to-reverse
commitment than this repo's existing IP-blocklist reporting:
-- The existing reporter (`reporter/`, #68/#69, `arcane/home/honeypot-utilities/compose.yml`)
+- The existing reporter (`arcane/home/honeypot-utilities/reporter/`, #68/#69, `arcane/home/honeypot-utilities/compose.yml`)
sends a *narrow* signal (an IP, to a blocklist, for a defensive purpose:
getting that IP blocked elsewhere) to a small number of well-understood
destinations (AbuseIPDB, Blocklist.de), stays dry-run by default, and
needed its own multi-phase build (#68 for the dry-run foundation and
safeguards, #69 for validation and metrics, #153 for reputation
- filtering and observability, still open) to get the privacy posture
- right.
+ filtering and observability) to get the privacy posture
+ right. All three are closed and implemented: #153 shipped GreyNoise
+ RIOT pre-checks (`reporter/greynoise.go`, off unless `GREYNOISE_ENABLED=1`)
+ and a `metrics.json` counter snapshot that the dashboard mirrors into
+ `reporter-metrics-v1`.
- A generic `hpfeeds` publisher would share *richer* structured data
(commands, credentials, payload hashes, session metadata -- whatever
TANNER or another sensor chose to publish) with *whichever broker an
@@ -84,7 +87,7 @@ that work, which continues independently.
## 4. What already exists and isn't being built on
-`tanner/tanner/config.yaml` has a native, currently-disabled `HPFEEDS`
+`arcane/home/honeypot-tanner/tanner/tanner/config.yaml` has a native, currently-disabled `HPFEEDS`
block (`enabled: False`, plus `HOST`/`PORT`/`IDENT`/`SECRET`/`CHANNEL`) --
TANNER's own upstream already supports publishing to an `hpfeeds` broker,
no new code required to turn it on. This is *not* a recommendation: an
diff --git a/docs/container-writable-layer-audit-2026-09-03.md b/docs/container-writable-layer-audit-2026-09-03.md
index bbbc6f2dd..b20b9c610 100644
--- a/docs/container-writable-layer-audit-2026-09-03.md
+++ b/docs/container-writable-layer-audit-2026-09-03.md
@@ -1,5 +1,14 @@
# Container writable-layer and build-cache audit (#2859), 2026-09-03
+> **Dated record — 2026-09-03.** Every measurement below (245.8 GB, 13.18 GB
+> reclaimable, the 45-hour leaked buildkit container, the dangling-volume
+> census) was a snapshot of the homeserver on that date and is **not** a live
+> status page. Per the same principle `security-fixes.md` states outright: do
+> not mirror a live system's state into a markdown file. Re-run the commands
+> before acting on any number. The attribution, the reasoning and the
+> conclusions are the substance of this document and do not expire; the
+> open items named here are tracked in their issues (#2915, #2904).
+
`docker system df` on the homeserver showed **245.8 GB in container
writable layers**, invisible to the volume audit and the retention knob.
This document is the attribution the issue asked for.
@@ -18,22 +27,33 @@ This document is the attribution the issue asked for.
`rex86-eval` alone accounts for essentially the entire 245.8 GB figure.
It's a `nvidia/cuda:12.4.1-devel-ubuntu22.04` container
-(`docker inspect rex86-eval`), bind-mounting only
-`/var/dockge/stacks/rex86-eval/work` → `/work` — everything else it writes
-lands in its own rootfs. `docker exec rex86-eval du -sh /root/.cache`
+(`docker inspect rex86-eval`), bind-mounting
+`/var/dockge/stacks/rex86-eval/work` → `/work` and the repo itself
+read-only at `${APIARY_REPO:-../../../../}` → `/repo:ro` — neither of which
+is `/root/.cache`, so everything the benchmark tooling caches still lands in
+its own rootfs. `docker exec rex86-eval du -sh /root/.cache`
confirms **227 GB of the 245 GB is `/root/.cache`** (pip/HuggingFace-shaped
model/package cache for the benchmark tooling), not mounted to a volume or
bind mount.
This is a genuine compose/run defect in the shape the issue described — a
-container writing tens of GB to its own writable layer instead of a mount —
-but `rex86-eval` is not in this repo's tracked compose files at all (it's a
-raw `docker run`/Dockge stack under `/var/dockge/stacks/rex86-eval/`, not
-`arcane/home/*`), and it is explicitly excluded from this round's scope
-(model-benchmark work, chained to the same corpus tooling #1947's paused
-sweep uses). **No fix staged for it.** The correct fix, if/when the
-benchmark work is in scope, is a bind mount for `/root/.cache` — noted here
-so a future session doesn't have to re-derive the attribution.
+container writing tens of GB to its own writable layer instead of a mount.
+
+`rex86-eval` **is** in this repo's tracked compose files:
+`arcane/home/rex86-eval/compose.yml`, `setup.sh` and `.env.example` are all
+version-controlled (restored by #847), so it is a compose piece and not the
+raw `docker run` stack this paragraph originally described. It is *not* in
+`arcane/manifests/home-production.json` — deliberately, since it is
+benchmark tooling rather than a deployment piece, which is also why it is
+absent from the 33 `arcane/home/` manifest stacks. That makes it the one
+directory on disk under `arcane/home/` with no manifest entry, alongside
+nothing else.
+
+It is explicitly excluded from this round's scope (model-benchmark work,
+chained to the same corpus tooling #1947's paused sweep uses). **No fix
+staged for it.** The correct fix, if/when the benchmark work is in scope, is a
+bind mount for `/root/.cache` — noted here so a future session doesn't have to
+re-derive the attribution.
Everything else on the host contributes single-digit megabytes each; there
is no second offender worth a compose change.
@@ -67,8 +87,8 @@ failing** — GC only starts reclaiming once usage exceeds the 100 GB
neither of which is close on this host. No config or timer change needed;
a standing prune timer would be redundant with `builder.gc`, which already
runs automatically as part of build activity per buildkit's own design
-(the comment in `install-homeserver.sh:431-500` explains why a separate
-timer isn't used).
+(the comment in `install-homeserver.sh:381-465` — the reasoning at 392-397,
+the JSON block itself at 463-465 — explains why a separate timer isn't used).
One stray finding, not actioned: a leaked `buildx_buildkit_builder-`
container (`docker-container` driver) has been running 45+ hours,
diff --git a/docs/dashboard-manual-ip-block-design.md b/docs/dashboard-manual-ip-block-design.md
index 2a2d979a3..3689aa7da 100644
--- a/docs/dashboard-manual-ip-block-design.md
+++ b/docs/dashboard-manual-ip-block-design.md
@@ -100,18 +100,31 @@ that already exists** (home is reachable at `10.8.0.2`,
`docs/CGNAT-DEPLOYMENT.md`), the same "pull, don't get pushed to" posture
`portbridge-blackhole-refresh.sh` already uses against GitHub. Concretely:
-- The dashboard exposes `GET /export/portbridge-manual-blackhole.txt`
- (then `dashboard/ip_block.go`'s `serveManualBlackholeExport`, now
- `backend-service/src/ip_block.rs`'s `export`) — plain text, one
+- The dashboard exposes the export as `GET /api/v1/ip-block-export`
+ (`backend-service/src/ip_block.rs`'s `export`) — plain text, one
IPv4 address per line, the exact format `blackhole.go`'s existing parser
already reads. No admin auth on the handler itself, the same posture every
- other `/export/*.csv` GET already takes (access control is the network
+ other `/api/v1/export/*.csv` GET already takes (access control is the network
boundary — WireGuard-only reachability — not a second app-layer secret);
the data itself (a list of IPs an operator already chose to block) is no
more sensitive than the maltrail feed it sits alongside. Reachable from the
VPS at `10.8.0.2:19090` — the `dashboard` service's real published port
(`arcane/home/honeypot-dashboard/compose.yml`, `${HP_BIND:-10.8.0.2}:19090:8080`), not an
assumed default.
+ **The path moved at the Rust cutover and one caller was not carried across.**
+ This decision was written against the Go route
+ `GET /export/portbridge-manual-blackhole.txt` (then `dashboard/ip_block.go`'s
+ `serveManualBlackholeExport`); the Rust `export` keeps that handler's
+ *byte-compatible body* but is registered at `/api/v1/ip-block-export`, and
+ `backend-service/src/main.rs` has no `/export/portbridge-manual-blackhole.txt`
+ route at all. `vps/portbridge-manual-blackhole-refresh.sh` still defaults its
+ `MANUAL_BLACKHOLE_URL` to the **old** path
+ (`http://10.8.0.2:19090/export/portbridge-manual-blackhole.txt`), so on any
+ deployment that has not overridden that variable the sidecar is fetching a
+ 404. Either the default needs repointing at `/api/v1/ip-block-export` or the
+ VPS `.env` must set `MANUAL_BLACKHOLE_URL` explicitly — worth confirming
+ against the live host before trusting that manual blocks are reaching
+ portbridge.
- A new sidecar, `vps/portbridge-manual-blackhole-refresh.sh`, is a near-
verbatim copy of `portbridge-blackhole-refresh.sh` pointed at that URL
instead of GitHub's maltrail mirror, writing to a second local file
diff --git a/docs/deploy-profiles/README.md b/docs/deploy-profiles/README.md
index e376e64d3..226a1bbb0 100644
--- a/docs/deploy-profiles/README.md
+++ b/docs/deploy-profiles/README.md
@@ -27,17 +27,30 @@ line, `#` comments and blank lines ignored.
| Profile | Backbone | Sensors | Shape |
|---|---|---|---|
-| [`full.txt`](../../deploy-profiles/full.txt) | init, elk, dashboard, utilities, payload-analysis | every deception sensor stack under `arcane/home/` | the standard deployment -- everything this repo ships |
-| [`ics-focused.txt`](../../deploy-profiles/ics-focused.txt) | init, elk, dashboard, utilities | conpot, dnp3 | OT/ICS-only exposure -- skip the general-purpose/web/SSH/legacy-protocol sensors entirely |
-| [`minimal-web.txt`](../../deploy-profiles/minimal-web.txt) | init, elk, dashboard, utilities | http, tanner | web-attack-focused -- HTTP/API honeypot + SNARE/TANNER, skip ICS/SSH/legacy-protocol sensors |
-
-`init`, `elk`, and `dashboard` are structural dependencies for any profile
-that includes at least one sensor -- `scripts/validate-deploy-profile.sh`
-(below) enforces
-this, it isn't just a convention to remember. `payload-analysis` and
+| [`full.txt`](../../deploy-profiles/full.txt) | keycloak, init, elk, dashboard, utilities, payload-analysis | the 20 classic deception sensor stacks under `arcane/home/` -- but see the gap below | the standard deployment -- everything this repo ships |
+| [`ics-focused.txt`](../../deploy-profiles/ics-focused.txt) | keycloak, init, elk, dashboard, utilities | conpot, dnp3 | OT/ICS-only exposure -- skip the general-purpose/web/SSH/legacy-protocol sensors entirely |
+| [`minimal-web.txt`](../../deploy-profiles/minimal-web.txt) | keycloak, init, elk, dashboard, utilities | http, tanner | web-attack-focused -- HTTP/API honeypot + SNARE/TANNER, skip ICS/SSH/legacy-protocol sensors |
+
+**`full.txt` is not actually "everything this repo ships".** It lists 26
+stacks (6 backbone + 20 sensors) and omits `honeypot-sonicwall-sma`, the
+decoy sensor stack #3131 added on 2026-09-08 (`hp-sonicwall-sma-honeypot`
+on `${HP_BIND:-10.8.0.2}:8543`). It is a natural fit for both `full.txt` and
+`ics-focused.txt`, and is in neither. The other `arcane/home/` stacks the
+profiles deliberately skip are the analysis-plane workers and
+`honeypot-dashboard-backend` (not persona declarations, per below) plus
+`unsloth` (the #3092 benchmark toolchain) -- and `rex86-eval`, which exists
+on disk but is in no manifest entry at all.
+
+`init` and `elk` are structural dependencies for any profile that includes at
+least one sensor; `keycloak` is a structural dependency of `dashboard`, and
+`elk` is too. `scripts/validate-deploy-profile.sh` enforces all three, so
+these aren't just conventions to remember. `payload-analysis` and
`utilities` are strongly recommended (payload dedup/YARA scanning, log
rotation/disk monitoring/autoheal) but not structurally required, so the
-validator only warns if either is missing from a non-empty profile.
+validator only warns if either is missing from a non-empty profile. Note
+that `dashboard` itself is *not* a required structural dependency: the
+validator never demands it, it only imposes `elk` and `keycloak` on a
+profile that has chosen it.
Not covered here: the VPS side (`vps/`, always deployed the same way
regardless of home profile -- see `docs/CGNAT-DEPLOYMENT.md`), the
@@ -97,9 +110,10 @@ scripts/validate-deploy-profile.sh deploy-profiles/ics-focused.txt
Checks, against the *current* repository state (not a hardcoded snapshot):
1. **Structural dependencies** -- `init`/`elk` present if any sensor stack
- is listed; `elk` present if `dashboard` is listed (the dashboard reads
- several sensors' events from Elasticsearch, not their log files --
- see #403 for why that's a real dependency, not a nice-to-have).
+ is listed; `elk` and `keycloak` present if `dashboard` is listed (the
+ dashboard reads several sensors' events from Elasticsearch, not their
+ log files -- see #403 for why that's a real dependency, not a
+ nice-to-have; and the target auth path is native Keycloak OIDC).
2. **Real-stack existence** -- every listed name must correspond to an
actual `arcane/home/honeypot-/` directory, so a typo'd or retired
stack name fails here instead of surfacing mid-deploy or as a silently
diff --git a/docs/design-lab/README.md b/docs/design-lab/README.md
index 2b4b50716..160ab84bb 100644
--- a/docs/design-lab/README.md
+++ b/docs/design-lab/README.md
@@ -67,5 +67,17 @@ The original lab served variant builds against real Elasticsearch data on
ports 19201–19205, driven by an env-guarded Go test that booted the dashboard
with a stubbed OIDC session, a `STATIC_DIR` override and nil write-services
so the real index stayed read-only. That harness depended on the Go dashboard
-and went away with it. A `frontend-next` equivalent needs the same read-only
-guarantees; scoped separately.
+and went away with it.
+
+It has since been rebuilt for `frontend-next` as
+[`branding/design-lab/lab.mjs`](../../branding/design-lab/lab.mjs) (#1828,
+#1935), which serves variants on the same 19201–19205 range and the elements
+playground on 19300. `frontend-next` has no nil-write-services handle — it
+reaches data over HTTP through two bases — so the read-only guarantee is made
+at that seam instead: `BACKEND_URL` goes through a gate that forwards
+GET/HEAD and answers 405 to everything else, and `BACKEND_MOUNTED_URL` is
+pointed at a stub that answers 503 to every request. That is stronger than the original, which relied
+on remembering to pass nil.
+
+This directory is the redacted public copy and does not carry the harness
+itself; run it from `branding/design-lab/`.
diff --git a/docs/design-lab/design-notes.md b/docs/design-lab/design-notes.md
index 2dc559161..b18a76c6b 100644
--- a/docs/design-lab/design-notes.md
+++ b/docs/design-lab/design-notes.md
@@ -1,4 +1,19 @@
# APIARY dashboard design review — dashboard.example — 2026-08-17
+
+> **Public, redacted copy — status as of 2026-09-27.** The address redaction in
+> this file is deliberate and correct as it stands: the only address literals
+> here are `127.0.0.1` (loopback) and `203.0.113.1` (RFC 5737 TEST-NET-3), and
+> the only hostname is a reserved `.example` domain. Do not substitute real
+> values for them, and do not restore anything redacted out of the source copy.
+>
+> The findings below are a **snapshot of a review session on 2026-08-17, not a
+> description of the current code**. Several cite the Go dashboard's static
+> assets and route table by name (`hp-app.js`, `hp-dynamic-nav.js`,
+> `routes.go`); that dashboard was deleted in #1628 on 2026-08-22, five days
+> after this review, and none of those files exist in the repository any more.
+> The current route authority is the Rust `axum` service at
+> `arcane/home/honeypot-dashboard/backend-service/src/main.rs`, so re-locate a
+> finding's code there before acting on it.
## Findings (running)
- Overview (light): loads fast, authenticated. Heatmap "Activity — last 24h" dominates; lower rows (multipot, conpot-kamstrup, endlessh) appear near-empty/pale — visual weight wasted?
- Theme toggle: monitor icon top-right (left of LIVE). Dark theme renders correctly on Overview.
diff --git a/docs/dionaea-bistreams-retention.md b/docs/dionaea-bistreams-retention.md
index 07f8bea5c..df6fdea5a 100644
--- a/docs/dionaea-bistreams-retention.md
+++ b/docs/dionaea-bistreams-retention.md
@@ -1,5 +1,15 @@
# Dionaea bistreams retention — consumer inventory and decision (#2862)
+> **Dated decision record — 2026-09-03.** `BISTREAMS_RETENTION_DAYS=30` is
+> still the pinned value in `arcane/home/honeypot-payload-analysis/.env.example`
+> and the decision has not been reopened. Read the forward-looking sections
+> ("nothing is 30 days old yet", "when this window first destroys something —
+> 2026-09-09") as written on 2026-09-03: that date has passed, so the pruning
+> path is now live rather than pending, and the size projections below are the
+> ones that were made then, not current measurements. Re-measure before acting
+> on the capacity numbers. The consumer inventory and the forensic argument are
+> unaffected by the passage of time and are the substance of this document.
+
`dionaea-lib`'s `bistreams/` tree holds Dionaea's raw per-connection capture
stream: every accepted connection gets a date-named subdirectory
(`YYYY-MM-DD/`) full of raw capture files, payload or not — a superset of
@@ -10,7 +20,7 @@ stream this document does not cover.
| reader | code path | reach |
|---|---|---|
-| `payload-dedupe` (`hp-payload-dedupe`) | `arcane/home/honeypot-payload-analysis/analysis/dedupe-payloads.py`: `prune_old_directories()` deletes whole date subtrees older than `BISTREAMS_RETENTION_DAYS`; `dedupe()` then hard-link-dedupes whatever's left (`PAYLOAD_ROOTS` includes `/payloads/dionaea/bistreams`) | whatever the retention window currently leaves on disk — no independent age requirement |
+| `payload-dedupe` (`hp-payload-dedupe`) | `arcane/home/honeypot-payload-analysis/analysis/dedupe-payloads.py`: `prune_old_directories()` deletes whole date subtrees older than `BISTREAMS_RETENTION_DAYS`, then `dedupe()` runs over `PAYLOAD_ROOTS` — which **does include bistreams**. The compose file sets `PAYLOAD_ROOTS=/payloads/cowrie:/payloads/dionaea/binaries:/payloads/scripts/script-payloads:/payloads/dionaea/bistreams` (compose.yml:49), so content dedupe and hard-linking both reach into the bistreams tree; `BISTREAMS_ROOT=/payloads/dionaea/bistreams` (compose.yml:56) is the *separate* variable the pruning pass reads, not an exclusion from dedupe. Bistreams is listed **last** on purpose, and the compose comment says why: it is 82% duplicate by content in a live sample, so age-pruning runs before dedupe each pass to keep what dedupe must hash bounded (#112) | whatever the retention window currently leaves on disk — no independent age requirement |
| `yara-scanner` (`hp-yara-scanner`) | `arcane/home/honeypot-payload-analysis/compose.yml`'s `YARA_PAYLOAD_ROOTS=/payloads/dionaea:...` mounts the whole `dionaea-lib` volume read-only, so it scans bistreams as part of `/payloads/dionaea` | same — whatever's currently present |
| Elasticsearch / dashboard | none — nothing indexes bistreams content directly. `HONEYPOT_RETENTION_DAYS` (21d) governs *derived* ES indices, which is a shorter and unrelated window over structured events, not a copy of the raw stream | n/a |
| manual forensic review | ad hoc, off-repo | as far back as the window allows |
diff --git a/docs/gpu-docker-passthrough.md b/docs/gpu-docker-passthrough.md
index d67add1f0..2f5989eb5 100644
--- a/docs/gpu-docker-passthrough.md
+++ b/docs/gpu-docker-passthrough.md
@@ -189,10 +189,22 @@ services:
reservations:
devices:
- driver: nvidia
- count: all
+ device_ids: ["GPU-18a00c7e-670a-c305-a2aa-20e3a71917a3"]
capabilities: [gpu]
```
+**Do not use `count: all` here.** #1539 replaced it: the homeserver carries
+two NVIDIA cards, and the `count: all` form handed Ollama *both* — wrong
+even when it works, and it starves the Windows sandbox VM of the Quadro
+P2200 reserved for its passthrough whenever Ollama loads a model during a
+detonation. The overlay therefore pins the RTX 4000 Ada's UUID, confirmed
+live via `nvidia-smi -L` on the actual box. Re-run `nvidia-smi -L` and
+update the UUID if the card is ever physically replaced.
+
+The same file also documents why `ghidra` is deliberately *not* given the
+GPU: decompilation is CPU work and would only compete for the card with the
+model it is feeding.
+
Note this repo keeps the GPU reservation in a **separate overlay file**,
applied only when a GPU is actually present:
@@ -225,7 +237,11 @@ Docker's `--gpus all` / `count: all` doesn't partition VRAM — every
container that requests the GPU gets the whole card, and it's up to each
process to behave. Nothing stops two containers from both trying to
allocate more VRAM than the card has, at which point the second allocator
-gets a CUDA out-of-memory error, not a scheduling wait.
+gets a CUDA out-of-memory error, not a scheduling wait. On this box the
+question is sharper still, because the host has **two** cards: an RTX 4000
+Ada for compute and a Quadro P2200 reserved for the Windows sandbox VM's
+passthrough. `count: all` would hand both to one container, which is the
+#1539 bug the overlay's pinned `device_ids` exists to prevent.
This repo's own answer to that (see
[`gpu-ml-worker-acceleration.md` §5, "GPU Sharing Contract with the LLM
diff --git a/docs/gpu-llm-analysis-worker.md b/docs/gpu-llm-analysis-worker.md
index 7bbd08b34..e063f751e 100644
--- a/docs/gpu-llm-analysis-worker.md
+++ b/docs/gpu-llm-analysis-worker.md
@@ -86,10 +86,19 @@ worker milestones.
**Settled by [#602](https://github.com/Xore/APIARY/issues/602)** — an
earlier draft of this table (before #602) named the card as a Quadro RTX
4000 at compute capability 7.5/Turing; that card was never on this host.
-`lspci` shows a single AD104GL controller and containers enumerate exactly
-one device. Also pinned as the runtime-governance authority in
+`lspci` showed a single AD104GL controller and containers enumerated one
+compute device. Also pinned as the runtime-governance authority in
`analysis/ghidra/models/approved-models.json`.
+> **A second card arrived after #602.** #1539 recorded that the box also
+> carries a **Quadro P2200**, reserved for the Windows sandbox VM's
+> passthrough. Every VRAM budget in this document is still correct — they
+> are budgets against the Ada, and the P2200 is not part of the compute
+> pool — but the host is no longer single-GPU, so "containers enumerate one
+> device" no longer describes the machine. It is why the Ollama reservation
+> in `analysis/ghidra/docker-compose.ghidra.gpu.yml` pins `device_ids` to
+> the Ada's UUID rather than using `count: all`.
+
| Fact | Value | Verify with |
|---|---|---|
| GPU | NVIDIA RTX 4000 Ada Generation | `nvidia-smi -L` |
@@ -98,8 +107,8 @@ one device. Also pinned as the runtime-governance authority in
| Driver / CUDA | 580.173.02 / CUDA 13.0 | `nvidia-smi` |
| Container GPU passthrough | nvidia-container-toolkit 1.19.1, `nvidia` runtime registered | `docker info \| grep -i runtime` |
| End-to-end container test | `docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi -L` lists the GPU | run it |
-| Host RAM / CPU | 91 GiB / 16 logical CPUs | `free -h`, `nproc` |
-| Stack deployment | Dockge stack at `/opt/stacks/apiary/compose.yml`, containers `hp-*` | `docker ps` |
+| Host RAM / CPU | 92 GiB / 48 logical CPUs | `free -h`, `nproc` |
+| Stack deployment | Arcane gitops, on disk under `/var/dockge/stacks//`, containers `hp-*` | `docker ps` |
| Internal network | `honeynet` (Elasticsearch and all sensors live here) | `docker network ls` |
| Elasticsearch | 8.13.4 single-node, `xpack.security.enabled=false`, reachable as `http://elasticsearch:9200` inside `honeynet` | compose file |
@@ -269,6 +278,12 @@ design sketch retained for context and must not be copied into production:
— bounded U1-only production acceptance with no payload mounts;
- [`../llm-worker/docker-compose.captured-data.yml`](../llm-worker/docker-compose.captured-data.yml)
— separately authorized #83 network and read-only volume grant;
+- [`../llm-worker/docker-compose.captured-data-deploy.yml`](../llm-worker/docker-compose.captured-data-deploy.yml)
+ — the **deployed** entrypoint: this is `dockerComposePath` for the
+ `llm-worker` entry in `arcane/manifests/home-production.json`, and it is
+ what a bare `docker-compose.yml` bring-up is missing (#2234). A local
+ `docker compose up` against `llm-worker/docker-compose.yml` alone will
+ not reproduce the live worker.
- [`../analysis/ghidra/docker-compose.ghidra.yml`](../analysis/ghidra/docker-compose.ghidra.yml)
— the pinned shared Ollama service and narrow `honeypot-llm` network.
@@ -303,8 +318,12 @@ docker compose \
config --quiet
```
-The live stack is managed by Dockge under `/opt/stacks`. Deploy only from a
-reviewed merged revision; do not maintain a second hand-edited Compose copy.
+The live stack is managed by Arcane gitops from
+`arcane/manifests/home-production.json`, not by Dockge. (`/opt/stacks` still
+exists on the host, but only as a compatibility symlink to
+`/var/dockge/stacks`, added 2026-09-04 — nothing is managed there.)
+Deploy only from a reviewed merged revision; do not maintain a second
+hand-edited Compose copy.
Model pulling stays an explicit operator action in the Ghidra/Ollama stack.
---
@@ -477,21 +496,31 @@ Retention: ILM 90 days is sufficient — derived data, recreatable from raw.
Mirrors the pattern `ml-anomalies` already established
([`ml-worker-plan.md` §8–9](ml-worker-plan.md)):
-- **Delivered (#150):** `GET /api/llm/analysis?doc_type=&severity=&since=&limit=`
+- **Delivered (#150):** `GET /api/v1/store/llm-analysis` (paged via `offset`/`size`)
→ documents from `llm-analysis`, newest first, polled on the dashboard's
existing 1-minute ES ticker (same transport decision as `ml-anomalies`,
no new broker). `/llm-analysis` page: session summaries and payload
triage in one filterable table, every row labelled "AI-generated" and
showing severity/confidence, with an evidence link back to the
- originating session or payload where one exists (`dashboard/llm_analysis.go`).
+ originating session or payload where one exists
+ (`frontend-next/src/routes/llm-analysis.tsx`, backed by the generic store
+ route at `main.rs:438` rather than a route of its own).
- **Deferred:** `GET /api/llm/analysis/stream` (SSE via redis channel
`llm-analysis-events`) -- optional per this section's original scope
("any SSE/Redis wake-up path remains optional and non-authoritative");
polling has not been shown insufficient yet.
-- **Deferred:** semantic search over sessions using `nomic-embed-text`
- embeddings stored as a `dense_vector` (384-dim) field on `llm-analysis`
- docs, queried with ES kNN search. Still waiting on U1–U3 being stable,
- per this section's original scope.
+- **Delivered:** semantic search over sessions using `nomic-embed-text`
+ embeddings stored as a `dense_vector` (768-dim) field on `llm-analysis`
+ docs, queried with ES kNN search — `GET /api/v1/llm-search`
+ (`main.rs:341`, `llm_search.rs`, over the `llm-analysis` index's
+ `doc_type: session` documents). The dimensionality is 768, not the 384
+ this document originally stated: #151 confirmed the model's real native
+ output live against `POST /api/embed` and `llm-worker/worker.py` pins
+ `EMBEDDING_DIMS = 768`, rejecting a response of any other width outright
+ rather than indexing it into a mapping it cannot satisfy. This section
+ originally deferred it
+ pending U1–U3 stability; it has since shipped, so the list above is not
+ a statement of current scope.
---
diff --git a/docs/gpu-ml-worker-acceleration.md b/docs/gpu-ml-worker-acceleration.md
index 6117038de..5c944fb7f 100644
--- a/docs/gpu-ml-worker-acceleration.md
+++ b/docs/gpu-ml-worker-acceleration.md
@@ -73,9 +73,10 @@ parts, and it does not make the models more accurate.
## 3. Hardware & Compatibility Contract
**Settled by [#602](https://github.com/Xore/APIARY/issues/602)** (verbatim
-host evidence: `lspci` shows a single AD104GL controller, containers
-enumerate exactly one device — the earlier two-card / Turing-plus-Ada
-hypothesis is refuted) and pinned as the runtime-governance authority in
+host evidence: `lspci` showed a single AD104GL controller and containers
+enumerated one compute device — the earlier two-card *Turing-plus-Ada*
+compute hypothesis is refuted) and pinned as the runtime-governance
+authority in
`analysis/ghidra/models/approved-models.json`, which
`model-governance.py check-runtime` diffs live `nvidia-smi` against every
5 minutes:
@@ -84,6 +85,14 @@ hypothesis is refuted) and pinned as the runtime-governance authority in
capability 8.9.** (An earlier draft of this document, before #602,
mis-recorded this as a Quadro RTX 4000 at compute capability 7.5/Turing —
that card was never on this host; see #602 for the full correction.)
+- **One more card arrived later, and it is not for this workload.** #1539
+ recorded that the box also carries a **Quadro P2200**, reserved for the
+ Windows sandbox VM's passthrough. So the "one device" finding above is
+ still true of the *compute* pool this guide budgets VRAM against, but the
+ host is no longer single-GPU, and this matters concretely in §4.4: a
+ `count: 1` reservation picks an arbitrary card, which is the same class of
+ bug `count: all` caused. Pin `device_ids` to the Ada's UUID, as
+ `analysis/ghidra/docker-compose.ghidra.gpu.yml` already does.
- Driver 580.173.02 (CUDA 13.0) — backward-compatible with CUDA 12.x
runtime wheels.
- nvidia-container-toolkit 1.19.1 present; `docker run --rm --gpus all
@@ -91,7 +100,7 @@ hypothesis is refuted) and pinned as the runtime-governance authority in
- Stack network is `honeynet`. The old `ml-worker/docker-compose.override.yml`
targeted a network, `analysis-net`, that never existed anywhere in this
repository; that file has been replaced by `ml-worker/docker-compose.yml`
- (its own Dockge stack), which joins `honeynet` as an external network —
+ (its own standalone stack), which joins `honeynet` as an external network —
resolved under #61.
**Wheel compatibility rule for Ada (sm_89), compute capability 8.9:**
@@ -118,17 +127,18 @@ Replace the CPU wheel lines:
```diff
-# Deep learning (CPU-only PyTorch)
--torch==2.13.0+cpu
+-torch==2.14.0+cpu
---extra-index-url https://download.pytorch.org/whl/cpu
+# Deep learning (CUDA PyTorch — see docs/gpu-ml-worker-acceleration.md §3)
-+torch==2.13.0+cu126
++torch==2.14.0+cu126
+--extra-index-url https://download.pytorch.org/whl/cu126
+
+# Embeddings (§6)
+sentence-transformers==3.0.1
- # Outlier detection (HBOS)
- pyod==3.6.2
+ # Outlier detection (HBOS). Pulls in numba+llvmlite -- see the numpy pin
+ # above; this is the version set actually verified to install together.
+ pyod==3.6.6
```
> **Verified pin (2026-08-01, #82):** `torch==2.13.0+cu124` does not exist;
@@ -141,7 +151,10 @@ Replace the CPU wheel lines:
> including both `sm_75` and `sm_89`, so the wheel's `sm_75` inclusion says
> nothing about which kernel the live card actually used; the underlying
> install-and-tensor-check result stands, only the architecture label was
-> wrong.) Never silently fall back to the `+cpu` wheel; a CPU wheel in a GPU
+> wrong.) **The tree has since moved on: `ml-worker/requirements.txt` now
+pins `torch==2.14.0+cpu` and `pyod==3.6.6`, so the +cu126 install above was
+verified at 2.13.0 only and has not been re-checked at 2.14.0. G2 applies to
+whatever version is current when this deploys.** Never silently fall back to the `+cpu` wheel; a CPU wheel in a GPU
> deployment must fail the acceptance test T2, not pass unnoticed.
### 4.2 `ml-worker/Dockerfile`
@@ -199,6 +212,12 @@ Rules:
+ ML_DEVICE: auto # auto | cpu — 'cpu' forces CPU for debugging
```
+**Pin `device_ids`, not `count: 1`.** The host carries two cards (§3), so
+`count: 1` selects an arbitrary one and can hand `ml-worker` the Quadro
+P2200 the Windows sandbox VM needs. Use the Ada's UUID exactly as
+`analysis/ghidra/docker-compose.ghidra.gpu.yml` does:
+`device_ids: ["GPU-18a00c7e-670a-c305-a2aa-20e3a71917a3"]`.
+
Keep the existing `mem_limit: 2g` / `cpus: "2.0"`; GPU memory is governed
by §5, not by `mem_limit`.
@@ -224,7 +243,8 @@ the original 6.1 GiB `qwen3.5:9b` estimate):
smaller safety margin instead of full separation.** Worst case — the
ghidra slot's model loaded at its 32k context (~14.1 GiB) plus a retrain
(~2 GiB) plus the embedder's own inference (~1 GiB) — totals ~17.1 GiB
- against the real 20475 MiB budget: about 3.3 GiB of headroom, not the
+ against the real 20475 MiB budget (19.995 GiB): about 2.9 GiB of headroom,
+ not the
"comfortable" double-digit margin a naive 6.1 GiB-chat-model estimate
would suggest. That is enough to not require full separation, but not
enough to treat as a non-issue either. This already assumes the ghidra
@@ -240,7 +260,7 @@ the original 6.1 GiB `qwen3.5:9b` estimate):
best-effort headroom management.
- If a future model re-evaluation (#568's process, re-run) picks something
materially larger than `qwen3:14b` for the ghidra slot, re-check this
- margin before assuming it still holds — the 3.3 GiB headroom above was
+ margin before assuming it still holds — the 2.9 GiB headroom above was
computed for this specific model at its current context ceiling, not as a
permanent property of the 20 GB card.
- **On CUDA OOM, do not crash:** wrap train/infer calls, catch
diff --git a/docs/honeypot-network-isolation.md b/docs/honeypot-network-isolation.md
index 9d2409107..f1874f168 100644
--- a/docs/honeypot-network-isolation.md
+++ b/docs/honeypot-network-isolation.md
@@ -39,10 +39,17 @@ is the one on the VPS. See
trap that follows from it, and note that traffic arriving over the tunnel must
not be attributed to the WireGuard peer.
-`vps/honeypot-firewall.sh` is the only firewall script in this repository. It is
-deliberately small: it idempotently `ufw allow`s the raw OT ports that
-`portbridge` handles, and does nothing else. Egress control on the VPS is not
-scripted here.
+`vps/honeypot-firewall.sh` is the only firewall script in this repository that
+*configures* anything. It is deliberately small: it idempotently `ufw allow`s
+the raw OT ports that `portbridge` handles, and does nothing else. Egress
+control on the VPS is not scripted here.
+
+The second file with "firewall" in its name,
+`vps/check-firewall-portbridge-sync.sh` (#152), configures nothing — it
+statically diffs the ports `portbridge` actually forwards against the ports
+`honeypot-firewall.sh` opens, because those two lists drifted silently once
+already. It is not wired into any workflow, so run it by hand after editing
+either file.
## 2. Sandbox isolation
@@ -79,9 +86,31 @@ Elasticsearch over the network. The one real exception is `tftp-relay`, which
has `depends_on: dionaea` and actually forwards TFTP traffic to it, so those
two share `dionaea_net` instead of each getting their own.
-`honeynet` is now the trusted analysis/management plane only: Elasticsearch,
-Kibana, Filebeat, the dashboard, EveBox, Arkime, and `log-maintenance`. SNARE/
-TANNER and its dependencies keep their own separate `tanner_local`, unchanged.
+`honeynet` is now the trusted analysis/management plane only, and that
+enumeration needs to be read as a category, not a list: 24 services across 11
+compose files attach to it. The data plane is Elasticsearch, Kibana, Filebeat,
+EveBox, Arkime (capture and viewer) and the dashboard stack; the *supporting*
+trusted members are just as much a part of the plane and are easy to forget:
+
+- the four Go worker stacks — `agent-intrusion-worker`,
+ `attacker-identity-worker`, `correlator-worker`, `payload-inventory-worker`
+ — all `profiles: ["legacy"]`;
+- `honeypot-dashboard-backend`'s unprivileged `backend-service` and
+ `honeypot-dashboard`'s `backend-service-mounted`, `backend-worker`,
+ `backend-worker-importer` and `backend-worker-payload-inventory`;
+- the one-shot `honeypot-init` containers `elasticsearch-setup`, `arkime-init`
+ and `honeypot-kibana-setup`;
+- `honeypot-utilities`' `log-maintenance` **and** `disk-space-monitor`;
+- `honeypot-elk`'s `extracted-file-importer`;
+- the two root-path stacks `auth-events-worker` and `ml-worker`.
+
+Deliberately *not* on `honeynet`: `backend-worker-enrichment` and
+`services-adapter` (both `network_mode: none`), and the dashboard's
+`oidc-sessions`, which has its own single-member `oidc-session` network.
+
+SNARE/TANNER and its dependencies keep their own separate `tanner_local`,
+unchanged — all seven services of `honeypot-tanner`, including `snare`, are on
+it.
- `tanner_docker` is `privileged: true`. This is deliberate and should stay:
TANNER's Docker-backed emulators need a daemon, and the design gives them a
@@ -96,9 +125,13 @@ TANNER and its dependencies keep their own separate `tanner_local`, unchanged.
[#89](https://github.com/Xore/APIARY/issues/89) (SNARE/TANNER) and
the per-service measurement passes referenced next to `dionaea`'s and
`conpot`'s own `cap_add` lists closed the gap this section used to describe.
-- `NET_ADMIN`/`NET_RAW` exist only on the three sandbox sniffers in
- `docker-compose.sandbox.yml`, a separate file brought up around a single
- detonation that must never be merged into `docker-compose.yml`.
+- `NET_ADMIN`/`NET_RAW` are confined to sniffers that need the bridge device
+ or a raw socket, never to a decoy. Three sit in `docker-compose.sandbox.yml`
+ (`zeek`, `suricata`, `tcpdump`), a separate file brought up around a single
+ detonation that must never be merged into `docker-compose.yml`. The rest are
+ the passive-capture services that cannot sniff without them:
+ `honeypot-elk`'s `zeek-proxy`, and the VPS's `zeek`, `huginn-sidecar`,
+ `suricata`, and `p0f`.
## 4. Host posture
diff --git a/docs/ip-reporting-plan.md b/docs/ip-reporting-plan.md
index 0188aad4c..83594a289 100644
--- a/docs/ip-reporting-plan.md
+++ b/docs/ip-reporting-plan.md
@@ -3,15 +3,18 @@
Report attacker IPs observed by APIARY to public threat-intel
blocklists via their APIs.
-> **Status:** built. `reporter/` (Go, not the Python layout sketched below --
+> **Status:** built. `arcane/home/honeypot-utilities/reporter/` (Go, not the
+> Python layout sketched below --
> that part of this plan is superseded) is a real service in
> `arcane/home/honeypot-utilities/compose.yml`. Phase 1
> ([#68](https://github.com/Xore/APIARY/issues/68)) and Phase 2
> ([#69](https://github.com/Xore/APIARY/issues/69)) are both closed.
> Phase 3-4 (reputation validation, operator observability) is
> [#153](https://github.com/Xore/APIARY/issues/153), closed and implemented:
-> `reporter/greynoise.go` implements the Phase 3 GreyNoise validation, and
-> `reporter/metrics.go` implements Phase 4's observability counters
+> `arcane/home/honeypot-utilities/reporter/greynoise.go` implements the Phase 3
+> GreyNoise validation, and
+> `arcane/home/honeypot-utilities/reporter/metrics.go` implements Phase 4's
+> observability counters
> (attempted/suppressed/dryRun/sent/failed) as a JSON snapshot — a
> deliberate deviation from the Prometheus sketch originally proposed below.
>
@@ -58,19 +61,23 @@ Primary target: **AbuseIPDB** — widely used, has a public confidence score, an
```mermaid
flowchart TD
Sensors["Cowrie / Dionaea / Conpot / HTTP-honeypot / DNP3"]
- Reporter["reporter (Python) new Docker Compose service"]
+ Reporter["reporter (Go) hp-reporter, dry-run by default"]
AbuseIPDB["AbuseIPDB"]
Blocklist["Blocklist.de"]
+ GreyNoise["GreyNoise RIOT (read-only pre-check)"]
Sensors -->|"JSON event logs on the shared Docker volumes, tailed — not Redis pub-sub; see 'Resolved design decisions'"| Reporter
Reporter -->|"POST /api/v2/reports"| AbuseIPDB
+ Reporter -->|"POST /api"| Blocklist
+ Reporter -->|"GET /v3/riot/{ip}"| GreyNoise
AbuseIPDB -->|optional| Blocklist
```
The `reporter` container:
- Watches the same log/event volume already mounted by the `analysis` and `ml-worker` containers
- Maintains a local SQLite DB (`/data/reported.db`) to deduplicate IPs per service per 24 h
-- Exposes a `/metrics` endpoint (Prometheus) so Grafana can graph reports-per-hour
+- Exposes no HTTP listener at all: Phase 4 writes a `metrics.json` snapshot into the
+ data volume on an interval instead (see the status banner's Phase 4 note)
---
@@ -162,27 +169,32 @@ Before reporting, cross-check against:
Add to `docker-compose.yml`:
+> The block below is the original sketch. It is superseded — see the status
+> banner. Three things in it are simply wrong against what shipped, and are
+> called out because they are the kind of detail that gets copy-pasted:
+> there is **no Prometheus port** (Phase 4 emits `metrics.json` into the data
+> volume instead), the Blocklist.de credentials are **`BLOCKLISTDE_SENDER` +
+> `BLOCKLISTDE_API_KEY`**, not `BLOCKLIST_DE_EMAIL`/`BLOCKLIST_DE_PASSWORD`,
+> and the live switch is `REPORTER_LIVE`. The `whitelist.txt` mount path and
+> `REPORTER_COOLDOWN_HOURS` did land as drawn.
+
```yaml
reporter:
build: ./reporter
restart: unless-stopped
environment:
ABUSEIPDB_API_KEY: ${ABUSEIPDB_API_KEY}
- BLOCKLIST_DE_EMAIL: ${BLOCKLIST_DE_EMAIL}
- BLOCKLIST_DE_PASSWORD: ${BLOCKLIST_DE_PASSWORD}
- GREYNOISE_API_KEY: ${GREYNOISE_API_KEY:-} # optional
+ BLOCKLISTDE_SENDER: ${BLOCKLISTDE_SENDER}
+ BLOCKLISTDE_API_KEY: ${BLOCKLISTDE_API_KEY}
+ GREYNOISE_API_KEY: ${GREYNOISE_API_KEY:-} # inert unless GREYNOISE_ENABLED=1
+ REPORTER_LIVE: ${REPORTER_LIVE:-} # unset = dry-run
REPORTER_COOLDOWN_HOURS: ${REPORTER_COOLDOWN_HOURS:-24}
- REPORTER_WHITELIST: /config/whitelist.txt
volumes:
- cowrie_logs:/logs/cowrie:ro
- dionaea_logs:/logs/dionaea:ro
- conpot_logs:/logs/conpot:ro
- reporter_data:/data
- ./reporter/whitelist.txt:/config/whitelist.txt:ro
- ports:
- - "127.0.0.1:9101:9101" # Prometheus metrics
- networks:
- - honeypot_internal
```
Add to `.env.example`:
@@ -190,9 +202,11 @@ Add to `.env.example`:
```dotenv
# IP Blocklist Reporting
ABUSEIPDB_API_KEY=
-BLOCKLIST_DE_EMAIL=
-BLOCKLIST_DE_PASSWORD=
+BLOCKLISTDE_SENDER=
+BLOCKLISTDE_API_KEY=
GREYNOISE_API_KEY=
+GREYNOISE_ENABLED=0
+REPORTER_LIVE= # leave empty: the reporter is dry-run until set
REPORTER_COOLDOWN_HOURS=24
```
@@ -222,20 +236,39 @@ The reporter will track a daily counter and pause with exponential back-off on `
## Files To Create
+The Python layout originally sketched here was never built — the shipped
+service is Go. This is what
+`arcane/home/honeypot-utilities/reporter/` actually contains:
+
```mermaid
flowchart TD
- ReporterDir["reporter/"] --> Dockerfile["Dockerfile"]
- ReporterDir --> Requirements["requirements.txt"]
- ReporterDir --> ReporterPy["reporter.py main loop"]
- ReporterDir --> SourcesPy["sources.py per-sensor log parsers"]
- ReporterDir --> ApisPy["apis.py AbuseIPDB + Blocklist.de clients"]
- ReporterDir --> DedupPy["dedup.py SQLite-backed deduplication"]
+ ReporterDir["reporter/ (Go)"] --> MainGo["main.go entrypoint, run loop wiring"]
+ MainGo --> RunloopGo["runloop.go poll tick"]
+ RunloopGo --> TailGo["tail.go per-sensor log tailing"]
+ TailGo --> EventGo["event.go normalised event"]
+ EventGo --> CategorizeGo["categorize.go sensor/kind to upstream category"]
+ CategorizeGo --> DedupGo["dedup.go SQLite-backed deduplication"]
+ DedupGo --> GreynoiseGo["greynoise.go RIOT pre-check (Phase 3)"]
+ GreynoiseGo --> WhitelistGo["whitelist.go CIDR/IP allowlist"]
+ WhitelistGo --> BlocklistdeGo["blocklistde.go Blocklist.de client"]
+ WhitelistGo --> ReportGo["report.go AbuseIPDB client"]
+ GreynoiseGo --> ProcessGo["process.go report decision + dispatch"]
+ ProcessGo --> ReportGo
+ ProcessGo --> BlocklistdeGo
+ ProcessGo --> MetricsGo["metrics.go counters to metrics.json"]
+ ProcessGo --> AuditGo["audit.go bounded rotating audit log"]
ReporterDir --> WhitelistTxt["whitelist.txt safe IPs/CIDRs to never report"]
- ReporterDir --> MetricsPy["metrics.py Prometheus exporter"]
+ ReporterDir --> Dockerfile["Dockerfile FROM scratch, runs as 0:0"]
DocsDir["docs/"] --> PlanMd["ip-reporting-plan.md this file"]
```
+Every function here is exercised by a test, though not always in a
+same-named file: `categorize.go` is covered from `event_test.go` and `main.go`
++ `report.go` from `dryrun_test.go`. There is no `requirements.txt`,
+no `sources.py`/`apis.py`/`dedup.py`/`metrics.py`, and no Prometheus
+exporter — see the Phase 4 note in the status banner.
+
---
## Resolved design decisions
diff --git a/docs/knowledge-store-design.md b/docs/knowledge-store-design.md
index 32a88cfdf..fa7fcb1cc 100644
--- a/docs/knowledge-store-design.md
+++ b/docs/knowledge-store-design.md
@@ -4,6 +4,25 @@ Decision record for #2289, gating #2290–#2292. No code changes ship with
this document — it is the design pass #1634 asked for before any ingest
worker exists.
+> **Status, re-measured 2026-09-27: authored and tracked, not deployed.** The
+> ingest worker does exist in git — `vault-worker/` carries a `worker.py`, a
+> `sanitize.py`, a `Dockerfile` and two compose files, the shipped output of
+> #2289/#2290 — but nothing deploys it. It is absent from
+> `arcane/manifests/home-production.json`; no vault-worker container has ever
+> run on the homeserver, and there is no `vault-worker` stack directory among
+> the deployed ones; Elasticsearch carries no `knowledge-vault*` index, so the
+> `knowledge-vault-state-v1` checkpoint cited for this worker in
+> [PIPELINES.md](PIPELINES.md) does not exist either (that row is planned, the
+> same way); and the live APIARY worker runs
+> `WORKER_LOOPS=alert-notifier,attacker-identity,agent-intrusion,correlator,dashboard-rollups,threat-intel,zeek-proxy-attribution`
+> — there is no vault loop in it.
+>
+> So everything below is the design and implementation record of a **planned**
+> subsystem, kept because #2289/#2290 shipped real code against it. Read it as
+> intent that has not been switched on: none of the paths, indices or
+> checkpoints it names exist on a live host, and no deployment decision is
+> recorded anywhere in the repository.
+
## 1. Storage: plain markdown directory, carried by the existing off-host
backup path, not git/Syncthing
@@ -39,9 +58,10 @@ stack, not a reuse of one, even though it is a trivial one to stand up
(`git init` in a directory, a commit per note-write batch). The nearest real
precedent for "git as a sync/deploy substrate" in this repo is Arcane's own
GitOps machinery (`docs/ARCANE-GIT-SYNC.md`), which already runs a
-git-pull-and-apply loop against `main` with `auto_sync = 0` set deliberately
-on rows that must not auto-follow (`docs/ARCANE-GIT-SYNC.md:321` and
-`docs/ARCANE-GIT-SYNC.md:374`) — i.e.
+git-pull-and-apply loop against `main` and where every live row currently
+carries `auto_sync = 0`, with every deploy still a manual
+sync → build → redeploy (`docs/ARCANE-GIT-SYNC.md:425` and
+`docs/ARCANE-GIT-SYNC.md:478`) — i.e.
this codebase's existing git-sync tooling defaults to *manual* triggers for
anything sensitive, which is the posture this decision adopts too (see §4).
@@ -172,7 +192,7 @@ A persistent, curated, cross-referenced copy of (bounded, redacted) attacker
material is qualitatively different from the raw per-event ES documents it's
derived from: it's smaller, denser, and easier for a human or a script to
sweep in one pass. Reading `analysis/backup-honeypot.sh`, its archive step
-(`backup-honeypot.sh:87`) already walks `./analysis ./dashboard ./personas
+(`backup-honeypot.sh:119`) already walks `./analysis ./dashboard ./personas
./state` by directory-existence check, unconditionally including anything
found there. If the vault directory (§1: `state/knowledge-vault/`) is placed
under `$stack_dir/state/`, it is **already** inside this glob and would start
@@ -189,16 +209,17 @@ above already bounds and strips what can land in a note, the vault's content
is closer in sensitivity to the config material `backup-honeypot.sh` already
carries than to the bulk payload/PCAP data it explicitly excludes — so
extending that script's existing scope to include it is the correct call,
-not an oversight to patch around later. This is a decision to record
-verbatim in `analysis/backup-honeypot.sh`'s own comment block when #2290
-lands the directory, so a future reader sees it was deliberate rather than
-inferring it from a directory glob matching by accident.
+not an oversight to patch around later. That decision is already recorded
+verbatim in `analysis/backup-honeypot.sh`'s own comment block
+(`backup-honeypot.sh:107-114`, which names #2289, #2290 and this document), so
+a future reader sees it was deliberate rather than inferring it from a
+directory glob matching by accident.
### Worker authorization gate
The vault-ingest worker (#2290) must gate non-dry-run writes the same way
-`llm-worker` gates captured-data mode. Reading `llm-worker/worker.py:200-202`
-and `llm-worker/worker.py:254-264`:
+`llm-worker` gates captured-data mode. Reading `llm-worker/worker.py:246-248`
+and `llm-worker/worker.py:314-318`:
non-dry-run requires `LLM_ENABLED=true` **and** `LLM_ALLOW_CAPTURED_DATA=true`
together, with the error message naming both. The vault worker adopts the
same two-flag shape (its own env var names, e.g. `VAULT_ENABLED` /
diff --git a/docs/kvm-network-traffic-analysis.md b/docs/kvm-network-traffic-analysis.md
index 563160ecd..a2baefddc 100644
--- a/docs/kvm-network-traffic-analysis.md
+++ b/docs/kvm-network-traffic-analysis.md
@@ -27,7 +27,7 @@ are **two** such bridges, because there are two sandboxes:
| Network XML | `sandbox/windows/setup/sandbox-network.xml` | `sandbox/network.xml` |
| Fake internet | INetSim at `10.10.10.1` | none by default; optional logged DNS + Squid allowlist (`controlled` mode) |
| Capture | `docker-compose.sandbox.yml` (tcpdump, Zeek, Suricata) | root-owned `tcpdump` per job, host and guest side |
-| Results | `sandbox/results//` | `/var/lib/honeypot-sandbox/results/` (root-only), sanitized export copied out |
+| Results | `$WINDOWS_SANDBOX_RESULTS_DIR//` — this is the path the dashboard reads. The compose file's own `./sandbox/results/current` default only applies to a manual `docker compose -f docker-compose.sandbox.yml up`; `run_sample.py` overrides `SANDBOX_RESULTS_DIR` to the current run's out_dir, and its own fallback when the env var is unset is `reports/windows-sandbox` | `/var/lib/honeypot-sandbox/results/` (root-only), sanitized export copied out |
| Orchestrator | `sandbox/windows/orchestrate/run_sample.py` | `sandbox/run-linux-sample.sh` |
Neither bridge has a `` element, so neither can route anywhere. That
diff --git a/docs/llm-inference-backend-comparison.md b/docs/llm-inference-backend-comparison.md
index 75bf7f917..ec333b4dc 100644
--- a/docs/llm-inference-backend-comparison.md
+++ b/docs/llm-inference-backend-comparison.md
@@ -2,6 +2,26 @@
Status: research for [issue #598](https://github.com/Xore/APIARY/issues/598), 2026-08-05.
+> **Superseded in part — read this first.** This is the CPU-only research
+> pass. The *measured* three-engine comparison it asked for was run
+> afterwards and is recorded in
+> `analysis/ghidra/benchmarks/engine-benchmark/README.md` (2026-08-06, on the
+> RTX 4000 Ada 20GB, extended by #832's settings tuning). It **overturns §6
+> and §7**: real-card performance was in fact measured, and vLLM's "refuses to
+> start" failure turned out to be specific to the *previous* 8GB Quadro
+> RTX 4000 rather than a property of the image. It explicitly leaves §2
+> standing ("no digest/registry system", `model-governance.py`'s pipeline being
+> "Ollama-shaped"), and it does not re-test §1's structured-output axis or §3's
+> keep-alive/swap behaviour — it runs one engine at a time by construction, so
+> it says nothing about multi-model sharing. The sections below are left as
+> written, as the record of what was known on 2026-08-05; where later work
+> contradicts them, the later work wins.
+>
+> Two later findings bear on this doc and are flagged inline where they
+> apply: #2646 (a warm resident slot is *not* reproducible at temperature 0
+> with a fixed seed, which revises §4's premise) and `ghidra-worker.py`'s
+> deliberate OpenAI-compatible `/v1` client, which revises §8's cost estimate.
+
This is a task-specific decision record, matching
[`local-llm-model-evaluation.md`](local-llm-model-evaluation.md)'s format for
the model-selection decision it's paired with. It evaluates the inference
@@ -18,6 +38,11 @@ APIs). No comparison against llama.cpp (the inference engine Ollama itself
wraps) or vLLM (the throughput-oriented alternative) had been done. #598 asked
for one, explicitly allowing "switch, and rearchitect" as a valid outcome.
+(Still true of the *deployed* stack: `arcane/manifests/home-production.json`
+deploys no llama.cpp or vLLM service. Both now appear in benchmark scripts
+under `analysis/ghidra/benchmarks/` — measured, not depended on. See the
+status banner above.)
+
## Method
Every claim below was tested directly — a real container, a real (small)
@@ -33,7 +58,10 @@ Test model: `Qwen/Qwen2.5-0.5B-Instruct-GGUF` (Q4_K_M, 630M params) for
generation tests; `nomic-ai/nomic-embed-text-v1.5-GGUF` (Q4_K_M) for the
embeddings test. Images: `ghcr.io/ggml-org/llama.cpp:server`/`:full`,
`ghcr.io/mostlygeek/llama-swap:cpu`, `ollama/ollama:0.32.0` (this repo's own
-pinned version), `vllm/vllm-openai:latest` (v0.26.0).
+pinned version *at the time*; production has since moved to
+`ollama/ollama:0.32.13` in `analysis/ghidra/docker-compose.ghidra.yml` and
+`analysis/ghidra/models/approved-models.json`), `vllm/vllm-openai:latest`
+(v0.26.0).
## 1. Structured output enforcement
@@ -87,7 +115,7 @@ llama.cpp/vLLM lose on — Ollama's digest happens to be *exactly* what
equivalent. A migration would replace "query the running server's registry
digest" with "hash the GGUF file on disk directly" — simpler and arguably
more auditable (no trust in a registry's own digest computation), but a real
-rewrite of `collect_snapshot()`/`compare_against_approved()`'s identity model,
+rewrite of `collect_snapshot()`/`evaluate_drift()`'s identity model,
and every recorded `approved-models.json` entry's `ollama_repo_digest`-shaped
field.
@@ -133,6 +161,21 @@ instances with strict, static VRAM partitions instead of dynamic sharing).
`temperature: 0, seed: 66` is relied on for reproducible output today.
+> **Revised after the fact (#2646).** That reliance did not survive contact
+> with the deployed configuration. `OLLAMA_KEEP_ALIVE=30m` keeps the weights
+> resident between samples, and a **warm** Ollama slot returns different text
+> for a byte-identical prompt at temperature 0 with a fixed seed — production
+> therefore runs permanently in the drifting regime, and no setting fixes that
+> without paying a reload. The fix that shipped is accounting, not determinism:
+> every stored assessment now records a `slot_generation` fingerprint (read
+> from Ollama's `/api/ps`) naming the resident instance that answered, so two
+> results are never compared as though one instance produced both — see
+> `docs/analysis/ghidra/AI_TRIAGE.md`. This is attributed to a warm **Ollama**
+> slot specifically; the llama.cpp result below is not stated to have
+> exercised an equivalent long-lived warm-slot regime, so it stands as
+> measured. What the finding removes is this section's *premise* — that the
+> deployed stack is reproducible today — not its per-engine results.
+
**llama.cpp**: confirmed bit-identical output across 3 repeated identical
requests (`temperature: 0, seed: 66`) against the same model. No loss.
@@ -175,14 +218,18 @@ prompt-processing (`pp`) and text-generation (`tg`) tokens/sec numbers.
(`qwen3:14b`, promoted to all three slots under #568/#569 now that VRAM is
confirmed ~20GB) on the real GPU, and compare against Ollama's own measured
throughput for the same model (`docs/local-llm-model-evaluation.md` already
-has some of these numbers).
+has some of these numbers). **Partly done, on a different model** — the later
+three-engine run in `analysis/ghidra/benchmarks/engine-benchmark/README.md`
+did measure decode throughput on the real 20GB card (llama.cpp ~21.4, Ollama
+~21.5, vLLM ~23.1 tok/s), but against the merged REx86 f16 7B weights, not
+`qwen3:14b`. The `qwen3:14b`-specific comparison is still outstanding.
## 7. Operational surface
| | Ollama 0.32.0 | llama.cpp `:server` | llama-swap `:cpu` | vLLM 0.26.0 |
|---|---:|---:|---:|---:|
| Image size | 8.06 GB | 1.21 GB | 1.24 GB | 28.1 GB |
-| Runs without a GPU present | Yes (CPU fallback) | Yes (CPU fallback) | Yes | **No — confirmed** |
+| Runs without a GPU present | Yes (CPU fallback) | Yes (CPU fallback) | Yes | **No — confirmed**, but hardware-specific; see below |
The vLLM finding is real and concrete, not inferred: the standard
`vllm/vllm-openai` image fails to even construct its own CLI argument parser
@@ -201,6 +248,20 @@ less-maintained artifact, not the mainline image this stack would pull).
Ollama and llama.cpp both degrade gracefully to CPU; vLLM does not degrade,
it refuses to start.
+> **Superseded.** That failure was hardware-specific, not a property of the
+> image. On the current 20GB RTX 4000 Ada Generation, `vllm/vllm-openai:latest`
+> (v0.26.0) starts cleanly, loads a 7B model, completes `torch.compile`
+> warmup and serves correct completions — recorded in the "Side finding"
+> section of `analysis/ghidra/benchmarks/engine-benchmark/README.md`. The
+> failure above was observed on the *previous* 8GB Quadro RTX 4000.
+>
+> Whether the mainline image has any CPU-only path at all is therefore
+> **undetermined** from this repo's evidence: the two data points are one
+> failure on the old card and one success on the new one, and neither isolates
+> GPU *presence* from GPU *capability*. The "no CPU path at all" reading above
+> overstates what was shown. Unchanged by the later run: vLLM is by far the
+> largest image of the four, and §2's digest/registry finding stands.
+
**Embeddings** (relevant to `ml-worker`/#151's planned semantic search, both
currently spec'd around `nomic-embed-text`): llama.cpp has a native
`--embedding` server mode with an OpenAI-compatible `/v1/embeddings`
@@ -231,7 +292,16 @@ contract) replaced Ollama:
`/v1/chat/completions`'s `response_format` shape instead of `/api/chat`'s
`format`; digest verification (`model_digest()`) rewritten to hash the
local GGUF file instead of querying `/api/tags`.
-- `analysis/ghidra/worker/ghidra-worker.py` — same client-shape change.
+- `analysis/ghidra/worker/ghidra-worker.py` — **no client-shape change needed**,
+ as of the current code. It already speaks an OpenAI-compatible `/v1` chat
+ dialect (`GHIDRA_TRIAGE_API_BASE`, default `http://127.0.0.1:11434/v1`) and
+ does so deliberately: the worker records that this is what makes "llama.cpp's
+ server, vLLM and LM Studio work unchanged; the requirement is that it is
+ *local*, not that it is Ollama." The one Ollama-specific coupling left is the
+ `/api/ps` probe behind `GHIDRA_TRIAGE_RUNTIME_BASE` that populates
+ `slot_generation` (#2646); the code already handles a server that is not
+ Ollama, recording `unavailable`. This line was a real cost when this doc was
+ written and is no longer one.
- `analysis/ghidra/models/model-governance.py` — the largest rewrite. Its
entire `collect_snapshot()`/drift-comparison model is built around Ollama's
registry APIs and Docker-image-reference identity; every field in
@@ -290,6 +360,10 @@ serving — not the current shared, bursty, multi-workload shape.
- Run `llama-bench` against the currently-approved model (`qwen3:14b`,
promoted under #568/#569) on the real card, side-by-side with Ollama's own
numbers, to get the actual performance data this research couldn't gather.
+ (Partly superseded: the three-engine run in
+ `analysis/ghidra/benchmarks/engine-benchmark/README.md` gathered real-card
+ throughput for llama.cpp/Ollama/vLLM, but on REx86 f16 7B weights. The
+ `qwen3:14b` case is still open.)
- Resolve the 768-vs-384 embedding-dimension discrepancy (§7) before #151's
`dense_vector` ES mapping work begins, regardless of backend choice.
- If a future re-evaluation is triggered (e.g. `model-governance.py` gets a
diff --git a/docs/llm-worker/README.md b/docs/llm-worker/README.md
index 7f0f86f1f..088673231 100644
--- a/docs/llm-worker/README.md
+++ b/docs/llm-worker/README.md
@@ -131,6 +131,20 @@ inline scripts read-only. Version 1 accepts regular text files no larger than
1 MiB, refuses symlinks and NUL-containing/binary data, and hashes content
itself instead of trusting a filename.
+**The deployed entry point is a third file, not this one.**
+`arcane/manifests/home-production.json` points the `llm-worker` stack at
+`llm-worker/docker-compose.captured-data-deploy.yml` (#1751), not at
+`docker-compose.captured-data.yml`. The deploy file is a thin `include:`
+wrapper listing `docker-compose.yml` then `docker-compose.captured-data.yml`
+in that order, so the behaviour described above is what actually runs — but the
+network/mount/volume grant lives in the included file, not the deployed one.
+It exists because the captured-data authorization had been applied by hand and
+was in no tracked file, so an Arcane sync silently reverted the container to
+`synthetic-only` while `LLM_ALLOW_CAPTURED_DATA=true` stayed set (#1751); the
+`include` list is a single two-file entry rather than two entries because
+`include` does not override the way repeated `-f` does (#2225). Deleting the
+file returns the deployment to synthetic-only by design.
+
## Guardrails
- strict pydantic schemas reject extra keys, invalid enums, malformed ATT&CK
diff --git a/docs/local-llm-model-evaluation.md b/docs/local-llm-model-evaluation.md
index 7a4fea3d8..f7a2725b8 100644
--- a/docs/local-llm-model-evaluation.md
+++ b/docs/local-llm-model-evaluation.md
@@ -2,6 +2,55 @@
Status: completed for [issue #144](https://github.com/Xore/APIARY/issues/144) and requalified under [issue #158](https://github.com/Xore/APIARY/issues/158), 2026-08-01. Re-evaluated and re-approved under [issue #568](https://github.com/Xore/APIARY/issues/568), 2026-08-05 — see [§ Issue #568 re-evaluation](#issue-568-re-evaluation-real-20gb-card) below; that section is now the current approved state, superseding the v2 table immediately above it.
+> **Reconciled 2026-09-27 against `docs/benchmarks/runs/` and
+> `docs/benchmarks/matrices/`.** No score in this file was changed. What the
+> repository can and cannot confirm:
+>
+> - **Confirmed exactly.** The round-7 cold baseline reconciles cell for cell
+> against `round7-cold-baseline.json`: 91 models, 182 cells, 367 records, 179
+> reproduced, the same 3 escalated cells with the same third-run values, 11
+> zero-scored tags, and every anchor — `qwen3:14b` 85.5 B / 83.1 A, `qwen3:8b`
+> 84.3 B / 81.9 A, `qwen2.5:14b-instruct-q4_K_M` 88.0 A, `Trendyol-32B` 95.2 A,
+> `Ornith-1.0-35B` 92.8 B, and the 12.1 / 7.3 point gaps. The cold-cohort
+> table reconciles against `1947-cohort-cold-protocol.json` (means, ±0 spreads,
+> B−A deltas, `min/run B`, Ornith's injection FAIL). The twelve-model survey
+> reconciles against `1805c-ghidra-slot-matrix.json`.
+> - **The archived runs cannot check the Ghidra column at all.** All 62 stored
+> runs — 1498 records — are `revdeck` or `sessions`; there is **no
+> `ghidra`-slot record in the repository**, and no transcript field carries
+> VRAM or a context-probe result. The Ghidra, 16k-probe and VRAM columns in
+> the #144 and #568 tables are therefore **undetermined from the repo**, not
+> confirmed and not contradicted. The approved `qwen3:14b@bdbd181c33f2…` and
+> `context_tokens: 32768` of the #568 Decision *are* confirmed against
+> `analysis/ghidra/models/approved-models.json`.
+> - **The archive is a different vintage from #568.** The stored runs are the
+> 2026-08-25→29 #1795b / #1947-wave2 / #1805c / #1947seq rounds; #568 was
+> measured 2026-08-05. Where the two overlap they differ (`qwen3:8b` sessions
+> 94.0 in the archive vs 92.5 here; `qwen2.5:14b-instruct-q4_K_M` 100.0 vs
+> 97.0). Those are re-measurements, not errors, and no figure was "corrected"
+> to match them.
+> - **Part 1's sessions column is one point low on three of eleven rows under
+> the current scorer**, reproducibly across all three repeats:
+> `Foundation-Sec-1.1-8B-Instruct-i1` 56→**57** (83.6→85.1%),
+> `Huihui-Qwen3.6-35B-A3B-abliterated` 66→**67** and stock `qwen3.8:27b`
+> 66→**67** (both 98.5→100.0%). The part-1 pins already declare a pre-#2265
+> scorer, so this is a vintage delta — but the part-1 Decision names only two
+> rows above the incumbent, and two more reach 67/67 on the current scorer.
+> The part-1/part-2 `revdeck` denominator is likewise /16 as printed against
+> /19 under the current inline `REV_CASES`.
+> - **Seven archived runs are not re-scorable as stored.** They report
+> `outcome: ok` with non-zero `output_tokens` and an empty `raw`: four
+> `qwen3:14b` runs on 2026-08-25 (51 records) and three gpt-oss-family runs on
+> 2026-08-26 (45 records). Re-scoring the archive naively yields 14–18/69 for
+> `qwen3:14b` from those four, against the authoritative 60–62/69. Exclude
+> them before recomputing anything.
+> - **Undetermined:** the round-7 injection figure "only 14 of 91 models fully
+> resist (5/5)" — `round7-cold-baseline.json` stores score and percent only,
+> with no per-case injection split, so the 5/5 counts have no in-repo source.
+> - Method check: re-deriving the 14-case corpus from the stored transcripts
+> with `rev_cases_v2_rubric.json` + `polarity.forbidden_hit` reproduces
+> `1805c-ghidra-slot-matrix.json` exactly, all 12 models at both tiers.
+
This is a task-specific decision record for the three independent local-model
slots in this repository. It does not assume that a model named in an earlier
plan is suitable, or that one model should serve all three jobs.
@@ -1375,7 +1424,13 @@ Tier A — the incumbent sits 7-12 points under the security-specialized
leaders (12.1 points Tier A vs `Trendyol-32B`'s 95.2%, 7.3 points Tier B vs
`llmfan46/Ornith-1.0-35B`'s 92.8%). Top band by run-pooled mean total_score
(mean across all 4 runs per model — 2 Tier A + 2 Tier B; not the same scale
-as the percentages above): `phi4:14b` 77.0, `Trendyol-32B` Q8_0 76.5,
+as the percentages above), **excluding the three entries whose Tier B cell was
+escalated to a third run and so has a 5-run, not 4-run, denominator** —
+`gemma-4-26B-A4B-it-ultra-uncensored-heretic` `Q4_K_M` 90.75,
+`Foundation-Sec-1.1-8B-Instruct` `Q8_0` 87.25 and
+`XORTRON.CriminalComputing.LARGE.2026.3` `i1-IQ2_XXS` 83.25, all three above
+everything listed here:
+`phi4:14b` 77.0, `Trendyol-32B` Q8_0 76.5,
`VulnLLM-R-7B` i1-Q4_K_M 76.5, `Huihui-CyberStrike-OffSec-35B` q6_k 75.5,
philbert440 `Qwen3.8-27B-Cyber` 75.5, protoLabsAI
`ThinkingCap-Qwen3.6-27B-MTP` `latest` 75.5 (the successfully-pulled
diff --git a/docs/ml-gpu-coordinated-roadmap.md b/docs/ml-gpu-coordinated-roadmap.md
index db20c8f66..3b2cb5122 100644
--- a/docs/ml-gpu-coordinated-roadmap.md
+++ b/docs/ml-gpu-coordinated-roadmap.md
@@ -1,6 +1,15 @@
# Coordinated ML and GPU Analysis Roadmap
-> **Status:** Proposed implementation sequence
+> **Status:** Proposed implementation sequence — **intent, not shipped state.**
+> Re-measured 2026-09-27: `ml-worker` and `llm-worker` now ship as their own
+> Arcane-managed stacks; the ML worker's GPU overlay
+> (`docker-compose.ml-worker.gpu.yml`) is still inert scaffolding, so the ML
+> worker remains CPU-only; retrain slots are `03:00,09:00,15:00,21:00`
+> (`ml-worker/worker.py:59`), not §4-I's `01:00,07:00,13:00,19:00`; the
+> default alert threshold did land at `0.75` (`worker.py:174`); and §1
+> decision 5 is superseded — the ML worker never grew an embedding index at
+> all, embeddings shipped in `llm-worker` at **768** dims, off by default.
+> The milestone text below is left as the historical record.
>
> **Scope:** `ml-worker`, GPU acceleration, local LLM analysis, and dashboard delivery
>
diff --git a/docs/ml-worker-evaluation.md b/docs/ml-worker-evaluation.md
index 9da6a953a..0ddbfab87 100644
--- a/docs/ml-worker-evaluation.md
+++ b/docs/ml-worker-evaluation.md
@@ -42,10 +42,16 @@ reported alongside accuracy rather than ignored.
### 2026-08-25 — Tier 1 harness landed; Tier 2 blocked
-**Tier 1 (behaviour) is live.** `evaluate_detectors.py` runs six contract
+**Tier 1 (behaviour) is live.** `evaluate_detectors.py` runs seven contract
checks against a candidate over the per-sensor fixture corpus and emits a
hashed JSON report. It is the first reusable acceptance bar `ml-worker` has
-had, and it makes "evaluate this candidate offline" answerable at all.
+had, and it makes "evaluate this candidate offline" answerable at all. The
+sixth of the seven, `check_composite_renormalises_over_present_detectors`
+(#1969), is not a per-candidate probe at all: it calls the production
+`worker.compute_composite()` directly, so candidates are always compared
+under the rules they will actually run with — absent detectors drop out of
+both numerator and denominator, an event no detector opines on composites
+to 0.0, and a single-detector opinion stands at face value.
**Tier 2 (accuracy) was blocked on [#1797](https://github.com/Xore/APIARY/issues/1797).**
There was no labelled corpus at that point. The date is retained as the
diff --git a/docs/ml-worker-plan.md b/docs/ml-worker-plan.md
index fe74b4381..75f244c2b 100644
--- a/docs/ml-worker-plan.md
+++ b/docs/ml-worker-plan.md
@@ -1,16 +1,24 @@
# ML Worker — Implementation Plan
+> **Reading note (2026-09-27):** this is a dated plan/record, not a live
+> reference. §2, §5.3, §7, §8, §9, §10, §11.4 and §11.6 describe shipped
+> behaviour and were re-checked against the code. §1, §4.1, §6, §11.2 and §12
+> keep their original-draft wording where it was never rewritten — including
+> the Go dashboard's `dashboard/ml_anomalies.go` / `settings_domain.go`
+> references, which are historical since #1628.
+
> **Status (2026-08-27, #1662):** the plan largely executed as written:
-> `ml-worker/worker.py` runs the ensemble described here from the same
-> repo-root compose files. What moved: consumers of its output live in the
+> `ml-worker/worker.py` runs the ensemble described here from its own
+> Arcane-managed stack, not the repo-root compose file (see §10).
+> What moved: consumers of its output live in the
> backend-service tier now, not `dashboard/ml_anomalies.go` (deleted).
> Open scoring-semantics defects are tracked in issues #1946/#1969 under
> epic #1974 rather than here.
-> **Status:** `ml-worker/` has its own Dockge stack
-> ([`docker-compose.yml`](../ml-worker/docker-compose.yml) +
-> [`docker-compose.ml-worker.gpu.yml`](../ml-worker/docker-compose.ml-worker.gpu.yml),
-> mirroring `analysis/ghidra/`), builds, connects to Elasticsearch, and polls
+> **Status:** `ml-worker/` has its own Arcane-managed stack
+> ([`docker-compose.yml`](../ml-worker/docker-compose.yml), the only one the
+> manifest deploys). `docker-compose.ml-worker.gpu.yml` exists in-tree but is
+> inert and undeployed, so the worker is CPU-only in practice. It
> without crashing (#62). `extract_features()`/`featurise_temporal()` read
> the real per-sensor schema (#62 task 33, #63). The dashboard delivers
> scores via the backend-service's `/api/v1/store/ml-anomalies` +
@@ -27,7 +35,10 @@
> `docker build ./ml-worker` failed outright (`pyod`'s `numba` dependency had
> no version compatible with the pinned `numpy==2.5.1` on Python 3.12 —
> reproduced twice, locally and in-container; fixed in #62 by pinning
-> `numpy==2.4.6`/`numba==0.66.0`/`llvmlite==0.48.0`). `worker.py`'s
+> `numpy==2.4.6`/`numba==0.66.0`/`llvmlite==0.48.0`; those two transitive pins
+> have since moved on again and the file now reads `numba==0.67.0` /
+> `llvmlite==0.49.0` / `pyod==3.6.6`, so treat this paragraph as the record of
+> what #62 did, not of the current requirements). `worker.py`'s
> `SOURCE_INDICES` (`cowrie-*`, `dionaea-*`, `honeypot-network-*`, `conpot-*`,
> `http-honeypot-*`) still match zero indices on the live homeserver: the real
> shape is a unified `honeypot-v2-*` stream (all sensors, disambiguated by
@@ -90,16 +101,17 @@ ground-truth labels. [web:275][web:283]
> the "v0.1 audit verdict" callout above. `worker.py`'s real,
> currently-deployed `SOURCE_INDICES` are the two rows below.
-The worker ingests from two unified, versioned index patterns
+The worker ingests from three unified, versioned index patterns
(`ml-worker/worker.py`'s `SOURCE_INDICES`):
| Index pattern | Source | Key fields |
|---------------|--------|------------|
| `honeypot-v2-*` | every honeypot sensor (Cowrie, Dionaea, Conpot, HTTP-honeypot, and every other sensor stack — disambiguated by `event.sensor`, not a separate index per sensor) | `event.sensor`, `source.ip`, `honeypot.*` (per-sensor nested fields, not uniform across sensors — see §5.3) |
| `suricata-v2-*` | Suricata network/IDS events (Filebeat) | `suricata.eve.*`, `network.*`, `alert.signature` |
+| `zeek-v1-conn-*` | Zeek connection records, added by #1774's sensing layer alongside Suricata | Zeek conn-log fields; note the pattern is `zeek-v1-conn-*`, not a `zeek-v2-*` line like the other two |
-Both index patterns share a common `@timestamp` field used for temporal
-ordering. A third index, `ml-worker-state`, is not a data source — it's the
+All three index patterns share a common `@timestamp` field used for temporal
+ordering. A fourth index, `ml-worker-state`, is not a data source — it's the
worker's own per-index-pattern checkpoint store (`load_checkpoint`/
`save_checkpoint` in `worker.py`): a `last_timestamp` plus the set of
already-seen event IDs at that exact timestamp, so a restart resumes
@@ -117,9 +129,11 @@ flowchart TD
subgraph Stack["APIARY (existing)"]
Sensors["every honeypot sensor stack (disambiguated by event.sensor, not a separate index each)"]
Suricata["Suricata / network IDS"]
- ES["Elasticsearch honeypot-v2-*, suricata-v2-*"]
+ Zeek["Zeek conn records (#1774)"]
+ ES["Elasticsearch honeypot-v2-*, suricata-v2-*, zeek-v1-conn-*"]
Sensors --> ES
Suricata --> ES
+ Zeek --> ES
end
subgraph Worker["ML Worker (ml-worker/)"]
@@ -168,8 +182,9 @@ different anomaly types: [web:275][web:276][web:283][web:292]
mixed numerical+categorical features after encoding. Proven on network
logs. [web:276][web:290]
- **Implementation:** `scikit-learn` `IsolationForest` with `contamination=0.01`
- (assume 1% of events are anomalous). Retrained every 6 hours on a 24h
- rolling window.
+ (assume 1% of events are anomalous). Retrained at four fixed UTC slots
+ daily (`RETRAIN_SLOTS_UTC`, default `03:00,09:00,15:00,21:00` — #172
+ replaced the old 6h `RETRAIN_INTERVAL`) on a 24h rolling window.
- **Output:** `anomaly_score` ∈ [-1, 0] where values closer to -1 = more anomalous.
### 4.2 LSTM Autoencoder (LSTM-AE)
@@ -391,7 +406,7 @@ loop every POLL_INTERVAL seconds (default: 30s):
10. Sleep POLL_INTERVAL
-Every 6 hours (RETRAIN_INTERVAL):
+At each `RETRAIN_SLOTS_UTC` slot (four daily, default 03:00,09:00,15:00,21:00 UTC, #172):
- Retrain IsoForest + HBOS on last 24h of all events
- Fine-tune LSTM-AE on last 24h (5 epochs, low LR)
- Save new model checkpoint to /models/
@@ -515,7 +530,7 @@ rediscovered from an empty index:
(`ml-worker/worker.py:73`); `run_worker()` installs a delete-only policy
(`ANOMALY_ILM_POLICY = "ml-anomalies-retention"`) via
`ensure_ilm_policy(es, ANOMALY_ILM_POLICY,
- build_ilm_policy(ML_ANOMALIES_RETENTION_DAYS))` (`worker.py:939`) before
+ build_ilm_policy(ML_ANOMALIES_RETENTION_DAYS))` (`worker.py:1002`) before
the index itself is created, because an index whose
`index.lifecycle.name` points at a missing policy fails its own
creation. These documents are the labelled-corpus substrate
@@ -524,7 +539,8 @@ rediscovered from an empty index:
`honeypot-30d`'s own 30-day source window, while still bounding the
index rather than leaving it permanent. The window is an env-tunable
default, not a hardcoded constant, per #261's convention.
-- **`ml-worker-metrics`: delete after 180d** (`ML_METRICS_RETENTION`,
+- **`ml-worker-metrics`: delete after 90d** (`ML_METRICS_RETENTION_DAYS`, default
+ `90` at `ml-worker/worker.py:74`,
ILM policy `ml-worker-metrics-retention`, installed idempotently by
the same `ensure_ilm_policy()` call at startup and bound via index
settings when the index is created). Diagnostic evidence for
@@ -698,7 +714,7 @@ scores to the dashboard":
## 10. Docker Compose Integration
-**Rewritten 2026-07-31 (#62).** ml-worker is its own Dockge stack now, not a
+**Rewritten 2026-07-31 (#62).** ml-worker is its own standalone stack now, not a
service folded into the root `docker-compose.yml`, and the file this section
used to show (`ml-worker/docker-compose.override.yml`, built against a
network named `analysis-net` that never existed anywhere in this
@@ -785,7 +801,7 @@ if wanted.
### 11.2 Online learning (unchanged from the original draft, still accurate)
```
- → HBOS/IsoForest: full retrain every RETRAIN_INTERVAL (default 6h) on the
+ → HBOS/IsoForest: full retrain at each `RETRAIN_SLOTS_UTC` slot (four daily, #172) on the
rolling 24h window, gated by §11.1
→ LSTM-AE: fine-tune on the same cycle (5 epochs, LR=1e-5), gated the same
way (§11.1's anomaly-rate check applies to its reconstruction-loss-based
@@ -830,7 +846,7 @@ fraction `>= THRESHOLD` exceeds `DRIFT_ANOMALY_RATE` (default `0.15`,
matching the original draft's "15%"):
- an early retrain is triggered (the next poll cycle retrains regardless of
- how much of `RETRAIN_INTERVAL` remains), and
+ which `RETRAIN_SLOTS_UTC` slot is nearest), and
- a `ml-worker-metrics` document is written flagging the drift event
(`kind: "drift"`, the observed rate, window size) so the dashboard's
`/ml-anomalies` page (#64) — or a future panel reading this index directly
@@ -884,9 +900,11 @@ its contract carried over unchanged to the Rust config module.)
| **v0.8** | Retraining scheduler + model versioning | [#65](https://github.com/Xore/APIARY/issues/65) |
| **v1.0** | Drift detection + alert threshold tuning UI | [#65](https://github.com/Xore/APIARY/issues/65) |
-v0.1 is listed as an issue rather than as done on purpose. `ml-worker/` holds a
-Dockerfile, `worker.py`, and a `docker-compose.override.yml`, but it is not a
-service in the root Compose file, it has no tests or fixtures, and nothing here
+v0.1 is listed as an issue rather than as done on purpose. As of the #61 audit
+— before #62's rewrite — `ml-worker/` held a
+Dockerfile, `worker.py`, and a `docker-compose.override.yml` (since deleted, see
+§10); it was not a
+service in the root Compose file, it had no tests or fixtures, and nothing here
has been observed running against live data. #61 is the audit that decides
whether the scaffold is a v0.1 or a starting point.
diff --git a/docs/payload-analysis-workbench.md b/docs/payload-analysis-workbench.md
index 04137dd97..223e0c67b 100644
--- a/docs/payload-analysis-workbench.md
+++ b/docs/payload-analysis-workbench.md
@@ -1,6 +1,6 @@
# Payload analysis workbench
-The dashboard's `/payload-workbench` route selects captured evidence and `/payload-workbench/{sha256}` is the unified orchestration surface for issue #155. It is a separate page rather than an extension of `/ghidra`: recipes and parent runs span deterministic analysis, Ghidra, and two sandbox backends, while `/ghidra/{sha256}`, `/sandbox/{job}` and `/payload-analysis/{sha256}` remain the canonical native result renderers.
+The dashboard's `/payload-workbench/results` route is the owner-isolated review surface, and its `workbench-builder` section is the unified orchestration surface for issue #155. It is a separate page rather than an extension of `/ghidra`: recipes and parent runs span deterministic analysis, Ghidra, and two sandbox backends, while `/ghidra/{sha256}`, `/sandbox/{job}` and `/payload-analysis/{sha256}` remain the canonical native result renderers.
## Trust boundary
@@ -15,7 +15,7 @@ Run ownership is likewise never taken from client input. The Rust tier derives i
## Analyzer registry
Seven analyzer IDs, one server-computed `workbenchAnalyzer` registry
-(`dashboard/workbench_domain.go`). A run selects 1-5 of them; the server
+(`backend-service/src/workbench_domain.rs`). A run selects 1-5 of them; the server
rejects zero selections, more than 5, an unknown ID, or a duplicate.
| ID | Applicability | Adapter | Result link | Concurrency class |
@@ -109,24 +109,27 @@ degraded to a stale local copy (#405 follow-up).
## HTTP contracts
-All APIs require a live administrator identity. Every mutation additionally requires a same-origin request, `application/json`, one document no larger than 64 KiB, and the closed Go schema (unknown fields are rejected). Create, read/list, reconcile, cancel, and retry are further scoped server-side to the caller's verified owner identity (see Trust boundary) — a request's `owner` field is ignored entirely, never an access decision (#3110).
+All APIs require a live administrator identity. Every mutation additionally requires a same-origin request, `application/json`, and the closed request schema. (A 64 KiB body cap and unknown-field rejection were part of the original Go contract; neither is present in the Rust tier, so do not rely on them.) Create, read/list, reconcile, cancel, and retry are further scoped server-side to the caller's verified owner identity (see Trust boundary) — a request's `owner` field is ignored entirely, never an access decision (#3110).
+
+Routes are registered in `backend-service/src/main.rs:456-468`. The hash is a **query** parameter on the registry and list routes, not a path parameter, and `cancel`/`retry` share one route via an `{action}` segment rather than being separate paths.
| Method and route | Purpose |
|---|---|
-| `GET /api/payload-workbench/registry/{sha256}` | server-derived registry, applicability, external-publication notice, and advisory model health |
-| `GET /api/payload-workbench/recipes` | visible private/shared recipe revisions |
-| `POST /api/payload-workbench/recipes` | append an immutable recipe revision |
-| `GET /api/payload-workbench/runs?sha256=...` | recent parent runs for the caller and payload |
-| `POST /api/payload-workbench/runs` | submit a saved revision or typed one-off selection |
-| `GET /api/payload-workbench/runs/{run_id}` | reconcile and return one parent run |
-| `POST /api/payload-workbench/runs/{run_id}/children/{analyzer_id}/retry` | bounded deliberate retry |
-| `POST /api/payload-workbench/runs/{run_id}/children/{analyzer_id}/cancel` | cancel an exact pending marker when supported |
+| `GET /api/v1/workbench/analyzers?hash=…` | server-derived registry, applicability, external-publication notice, and advisory model health |
+| `GET /api/v1/workbench/recipes` | visible private/shared recipe revisions |
+| `POST /api/v1/workbench/recipes` | append an immutable recipe revision |
+| `GET /api/v1/workbench/runs?hash=…&limit=…` | recent parent runs for the caller and payload |
+| `POST /api/v1/workbench/runs` | submit a saved revision or typed one-off selection |
+| `GET /api/v1/workbench/runs/{id}` | reconcile and return one parent run |
+| `POST /api/v1/workbench/runs/{id}/children/{analyzer_id}/{action}` | bounded deliberate retry, or cancel an exact pending marker when supported (`action` ∈ `retry`\|`cancel`) |
Create, recipe-save, retry, and cancel outcomes use the existing dashboard audit sink. Audit fields name the contract fields but do not copy payload content, prompts, model replies, filenames, credentials, or tool output.
## Model-status adapter
-`/var/lib/honeypot-ghidra/model-status.json` remains root-owned mode `0600`. `honeypot-model-status-adapter.service` reads and re-validates it, strips every field outside schema v1, and serves only `GET /v1/status` over `/run/honeypot-model-status/status.sock`. The dashboard mounts that runtime directory read-only and uses `MODEL_STATUS_SOCKET=/model-status/status.sock`. There is no TCP listener and no write, pull, replace, promote, prompt, or model-selection route.
+`/var/lib/honeypot-ghidra/model-status.json` remains root-owned mode `0600`. `honeypot-model-status-adapter.service` reads and re-validates it, strips every field outside schema v1, and serves only `GET /v1/status` over `/run/honeypot-model-status/status.sock`. The dashboard is expected to mount that runtime directory read-only and use `MODEL_STATUS_SOCKET=/model-status/status.sock`. There is no TCP listener and no write, pull, replace, promote, prompt, or model-selection route.
+
+> **Undetermined:** `MODEL_STATUS_SOCKET` has no occurrence anywhere in the Rust backend, the compose files, or `.env.example` as of 2026-09-27 — the adapter half of this contract is real and installed, but the consumer side is not visible in the repository. Treat the socket wiring as intended-but-unwired rather than working, and confirm before relying on the advisory model-health field the registry returns.
Re-run `sudo analysis/ghidra/install-analysis-host.sh` to install or update the adapter. Its failure only displays `unavailable`; it never disables a worker.
@@ -138,7 +141,7 @@ Deploy the dashboard normally after merging. Rollback is additive and safe:
1. deploy the previous dashboard image;
2. optionally disable `honeypot-model-status-adapter.service`;
-3. leave the workbench indices in Elasticsearch untouched (a rolled-back dashboard from before the #405 follow-up reads its own local `/state/analysis-workbench` copy instead and simply does not see runs created after the rollback).
+3. leave the workbench indices in Elasticsearch untouched (a rolled-back pre-workbench dashboard does not see runs created after the rollback).
The old `/ghidra/submit` and `/sandbox/submit` routes remain compatible. No worker or native result schema is changed by the workbench.
diff --git a/docs/persona-design.md b/docs/persona-design.md
index 4ada313c4..dce39ce9b 100644
--- a/docs/persona-design.md
+++ b/docs/persona-design.md
@@ -5,8 +5,12 @@
Two related pieces of persona design that were implicit rather than
documented decisions: whether a honeypot may reach the internet outbound,
and how to name/place a honeypot host so it doesn't look staged. T-Pot's
-own README calls both out by name (`README.md` line 296 for outbound; the
-"where to place a honeypot" guidance for siting). This repo already does
+own upstream README calls both out by name (its `README.md` line 296 for
+outbound; the "where to place a honeypot" guidance for siting). That line
+number refers to T-Pot's own repository, not to this repo's `README.md`, which
+is ~135 lines; it was recorded from an unpinned upstream read and could not be
+re-verified during this reconciliation, so treat the line number as a pointer
+to look up rather than a stable citation. This repo already does
deep, source-verified realism work for Windows personas
([#91](https://github.com/Xore/APIARY/issues/91)/[#94](https://github.com/Xore/APIARY/issues/94)/[#96](https://github.com/Xore/APIARY/issues/96))
and has a full fictional-organization inventory
@@ -41,7 +45,8 @@ The tradeoff is real in both directions:
| Cowrie | Allowed (flag: `COWRIE_AIR_GAPPED`, default `false`) | The one sensor in this stack designed around capturing attacker-fetched malware — its whole SSH/Telnet fake-shell premise is attackers running `wget`/`curl`/`tftp` against real URLs. See `arcane/home/honeypot-cowrie/compose.yml`'s `cowrie_net`. |
| Dionaea | Allowed (flag: `DIONAEA_AIR_GAPPED`, default `false`) | Same tradeoff as Cowrie (#269/#538): captures shellcode/binaries pushed *to* it over SMB/FTP/TFTP/etc, which both ship enabled by default — this is the attacker's actual malware sample, not just the exploit attempt. See `arcane/home/honeypot-dionaea/compose.yml`'s `dionaea_net` (#541). `internal: true` still permits `tftp-relay`'s inbound forwarding on the same network — it only removes the outbound route. |
| Tanner/Snare | Allowed (flag: `TANNER_AIR_GAPPED`, default `false`) | Same tradeoff as Cowrie and Dionaea: the `template_injection` emulator fetches real RFI payloads when enabled, capturing the attacker's actual payload instead of just the RFI attempt. See `arcane/home/honeypot-tanner/compose.yml`'s `tanner_local`. Setting the flag also breaks the emulator's own `REMOTE_DOCKERFILE` self-maintenance fetch (`raw.githubusercontent.com`) — a real cost, not just a capture-vs-safety tradeoff. |
-| Everything else (Conpot personas, DNP3, HTTP/API honeypot, multipot, dicompot, dns-honeypot, citrix-honeypot, cisco-asa-honeypot, rdp-honeypot) | Allowed, no design reason either way | None of these protocols involve the honeypot fetching attacker-supplied URLs — outbound access is unused in the intended interaction, just never explicitly closed off. An operator who wants maximum containment can set `internal: true` directly on that sensor's network in its compose file without losing anything these honeypots actually rely on. |
+| Canarytokens | Allowed (`canarytokens_net`; no air-gap flag exists) | **A different deception category, not a capture-vs-safety call.** This stack is a honeytoken *platform* — it plants fake documents, credentials and DNS names that alert when touched — not a sensor that fetches attacker-supplied content. Its reachability is the product: #1487 made the switchboard's HTTP channel publicly reachable through the VPS precisely so dashboard-created file/doc tokens fire when opened outside our own network, and that stack's compose file is explicit that "inert (internal-only) tokens don't serve it." The tradeoff argued above therefore does not transfer to it, in either direction. Note the asymmetry with the `internal: true` rule below: `internal` removes the *outbound* route, so it would not by itself break the VPS's inbound bridge — whether it is safe on `canarytokens_net` is undecided in-repo. Don't assume either way. |
+| Everything else (Conpot personas, DNP3, HTTP/API honeypot, multipot, dicompot, dns-honeypot, citrix-honeypot, cisco-asa-honeypot, rdp-honeypot, sonicwall-sma, endlessh, beelzebub, hellpot, elasticpot, galah, sentrypeer, mailoney) | Allowed, no design reason either way | None of these protocols involve the honeypot fetching attacker-supplied URLs — outbound access is unused in the intended interaction, just never explicitly closed off. An operator who wants maximum containment can set `internal: true` directly on that sensor's network in its compose file without losing anything these honeypots actually rely on. |
| `yara-scanner` (not a honeypot — offline payload analysis) | Blocked (`network_mode: none`) | Already air-gapped; scans captured files at rest, never needs network access at all. The one existing precedent this decision extends. |
`COWRIE_AIR_GAPPED=true` (`.env`) sets `internal: true` on `cowrie_net`,
diff --git a/docs/personas/README.md b/docs/personas/README.md
index 4a6daa8cd..97b3e97e8 100644
--- a/docs/personas/README.md
+++ b/docs/personas/README.md
@@ -3,17 +3,28 @@
[`personas.json`](../../personas/personas.json) is the canonical inventory for the fictional
organizations, sites, and assets exposed by this stack. Each event should carry
`persona_id`, `site_id`, `asset_id`, and `organization`; Filebeat adds those
-fields for upstream formats that cannot emit them. The live dashboard exposes
-each field as a clickable investigation pivot, and Elasticsearch stores them
-under `honeypot.*`.
+fields for upstream formats that cannot emit them. Elasticsearch stores all four
+under `honeypot.*`. The live dashboard turns the first three into clickable
+investigation pivots in an event's decoy group; `organization` is stored and
+exported but is not itself a pivot — the dashboard's `organization` filter is
+`source.as.organization_name` (the attacker's network/ASN owner), a different
+field entirely.
+
+All 18 entries in `personas.json`:
| Persona | Sensors | Attacker-facing identity |
|---|---|---|
| `nexusai-gpu01` | Cowrie | Ubuntu GPU inference/training node |
-| `nexusai-core` | multipot | mail, database, cache, VNC, search and Docker backend estate |
+| `nexusai-core` | multipot, Mailoney | mail, database, cache, VNC, search and Docker backend estate |
| `nexusai-edge` | HTTP honeypot | public NexusAI documentation/account edge |
| `nexusai-platform` | API honeypot | Kubernetes, registry, metadata and inference gateway |
+| `nexusai-directory` | Beelzebub | LDAP, SSH, HTTP and MCP directory/AI-agent estate, plus a secondary admin bastion |
+| `nexusai-analytics-legacy` | Elasticpot | standalone legacy analytics Elasticsearch node, deliberately distinct from `nexusai-core`'s own `es-logs-01` |
| `meridian-legacy` | Dionaea | legacy FTP, SMB and SIP integration server |
+| `meridian-legacy-web` | Hellpot | decommissioned marketing site that was never formally taken offline |
+| `meridian-customer-portal` | SNARE/TANNER | fictional customer service portal |
+| `meridian-staff-console` | Galah | internal staff/admin console with request-varying backend tooling |
+| `harborline-pbx` | SentryPeer | legacy SIP trunk for an old dispatch-office phone system |
| `rheinwerk-water-s7-200` | Conpot | water-intake Siemens S7-226 |
| `rheinwerk-water-s7-1200` | Conpot | treatment-hall Siemens S7-1215C |
| `nordchem-s7-1500` | Conpot | chemical-line Siemens S7-1516 |
@@ -21,7 +32,6 @@ under `honeypot.*`.
| `elbegrid-dnp3` | DNP3 sensor | substation 23 DNP3 outstation/RTU |
| `northfuel-guardian` | Conpot | filling-station tank gauge |
| `stadtwaerme-kamstrup` | Conpot | district-heating MULTICAL meter |
-| `meridian-customer-portal` | SNARE/TANNER | fictional customer service portal |
Validate the inventory and Cowrie identity before deployment:
@@ -33,9 +43,12 @@ Compose runs `persona-apply` before `log-init`, so every normal Dockge/Compose
deployment validates the manifest and event wiring before sensors start. It
also idempotently refreshes Dionaea's mutable FTP, TFTP, UPnP, and printer
persona files in the persistent volume and records the applied manifest hash in
-`state/personas/applied.json`. Run it manually with:
+`state/personas/applied.json`. `persona-apply` is a service of the
+`honeypot-init` stack (`arcane/home/honeypot-init/compose.yml`), not of the
+root marker compose file, so run it against that stack's deployed copy:
```bash
+cd /var/dockge/stacks/honeypot-init # or /opt/stacks/honeypot-init
docker compose -f compose.yml run --rm persona-apply
```
diff --git a/docs/sandbox/README.md b/docs/sandbox/README.md
index 94a28065d..28ca9ac58 100644
--- a/docs/sandbox/README.md
+++ b/docs/sandbox/README.md
@@ -1,12 +1,16 @@
# Hard-isolated malware sandbox plan
The homeserver supports this design: Intel VT-x is enabled, `/dev/kvm` is
-available, KVM is loaded, and the host exposes 84 IOMMU groups. The foundation
+available, KVM is loaded, and IOMMU is on with **94 groups** (re-measured
+2026-09-27). This was not free on Rocky: #1609 recorded 87 groups on the old
+Ubuntu host, and the 2026-09-03 Rocky 10 rebuild came up with **zero**, so
+`install-homeserver.sh`'s `step_vfio_gpu_passthrough` adds `intel_iommu=on`
+to the kernel command line. The foundation
installer provisions system libvirt and the dedicated isolated network.
## Route selection and evidence return across four dynamic-detonation routes
-The workbench's registry (`dashboard/workbench_domain.go`) offers four
+The workbench's registry (`backend-service/src/workbench_domain.rs`) offers four
routes to dynamic detonation, not one sandbox with options. Each is its own
guest, network, spool, and result format — this section is the canonical
side-by-side comparison; each route's own internal detail lives in its own
@@ -15,7 +19,7 @@ section below (Linux) or its own directory (`sandbox/windows/`,
```mermaid
flowchart TB
- workbench["Payload workbench — analyst selects a route (dashboard/workbench_domain.go)"]
+ workbench["Payload workbench — analyst selects a route (backend-service/src/workbench_domain.rs)"]
subgraph linuxRoute["linux-sandbox"]
direction TB
@@ -41,7 +45,7 @@ flowchart TB
capeGuest["Windows guest under CAPE's own cuckoo.py orchestration (runs on the host directly, not in Docker) + cape-mongo"]
end
- sharedLock{{"honeypot-kvm-detonation.lock — shared ONLY between windows-sandbox and cape (#320): 16 logical CPUs total, win11-sandbox alone already 8 vCPU. Held only around the actual detonation call, not the whole drain loop. linux-sandbox and windows-ghosts are NOT part of this lock — independent, can run concurrently with anything."}}
+ sharedLock{{"honeypot-kvm-detonation.lock — shared ONLY between windows-sandbox and cape (#320): 48 logical CPUs total, win11-sandbox alone already 8 vCPU. Held only around the actual detonation call, not the whole drain loop. linux-sandbox and windows-ghosts are NOT part of this lock — independent, can run concurrently with anything."}}
workbench -->|"hash-only request, SANDBOX_REQUEST_DIR"| linuxRoute
workbench -->|"hash-only request, WINDOWS_SANDBOX_REQUEST_DIR"| winRoute
@@ -72,7 +76,7 @@ credential ever crosses the dashboard/host boundary for any of the four.
never wait on anything.** `sandbox/windows/run_pending.sh` and
`sandbox/cape/worker/cape-worker.py` share one host-wide
`honeypot-kvm-detonation.lock` (#320) — a real capacity constraint, not a
-correctness one: both are KVM/QEMU domains on the same 16-logical-CPU host,
+correctness one: both are KVM/QEMU domains on the same 48-logical-CPU host,
and `windows-sandbox`'s own guest is already configured for 8 vCPU. The
lock is held only around the actual detonation call, never the whole drain
loop, so an idle worker on either side never blocks the other. `linux-sandbox`
@@ -235,6 +239,12 @@ sudo bash /opt/stacks/apiary/sandbox/install-windows-forensics.sh
## Required operating controls
- Reserve at most 4 vCPU and 8 GiB RAM per analysis VM; run one job initially.
+ For scale, the Windows analysis domain is currently defined at **8 vCPU /
+ 16 GiB** (`sandbox/windows/packer/win11-kvm.xml`), and
+ `sandbox/sandbox.env.example` ships `SANDBOX_VM_MEMORY_MB=3072` as its
+ default — so this line's 4 vCPU / 8 GiB guidance matches neither figure
+ exactly. It is operator guidance rather than a measured limit; settle it
+ against the real per-VM reservation before relying on it.
- Enforce a 10-minute hard timeout and kill QEMU if graceful shutdown fails.
- Store golden images on root-owned storage and verify SHA-256 before every job.
- Sign/validate result JSON and treat all guest-produced text as untrusted.
diff --git a/docs/sandbox/cape/IMPLEMENTATION_PLAN.md b/docs/sandbox/cape/IMPLEMENTATION_PLAN.md
index 1cfed78e5..f6725b401 100644
--- a/docs/sandbox/cape/IMPLEMENTATION_PLAN.md
+++ b/docs/sandbox/cape/IMPLEMENTATION_PLAN.md
@@ -54,7 +54,7 @@ Same as `docs/sandbox/windows/IMPLEMENTATION_PLAN.md` and
`docs/sandbox/ghosts/IMPLEMENTATION_PLAN.md`:
- KVM/QEMU/libvirt + docker-compose only — no VMware, no Hyper-V
- No CI-triggered detonation — the dashboard's Workbench is the only
- trigger (`workbench_orchestrator.go` → spool file → host-side systemd
+ trigger (`workbench_orchestrator.rs` → spool file → host-side systemd
worker)
- VM lifecycle is CAPE's own responsibility once configured (its
`kvm`/`libvirt` machinery module talks to `virsh` directly) — this
@@ -161,24 +161,26 @@ differently-configured venv without saying so.
- **Host-side CAPE sandbox worker** (`sandbox/cape/worker/`, systemd path
unit) — [#318]
- Watches `CAPE_REQUEST_DIR` for `{sha256}.request` files written by
- `dashboard/workbench_orchestrator.go`'s "cape" analyzer
+ the backend-service's
+ [`workbench_orchestrator.rs`](../../../arcane/home/honeypot-dashboard/backend-service/src/workbench_orchestrator.rs)
+ "cape" analyzer
- `cape-worker.py`: submits the sample to CAPE's own `apiv2`
(`/apiv2/tasks/create/file/`), polls `/apiv2/tasks/status/{id}/` until
- `reported`, fetches `/apiv2/tasks/report/{id}/json/`, writes
+ `reported`, fetches `/apiv2/tasks/get/report/{id}/json/`, writes
`{sha256}_cape.json` into `CAPE_RESULTS_DIR`
- - **Not yet verified against a live submission through this specific
- path** — narrower than before, not still fully open: #314's own
- `utils/submit.py` + `/apiv2/tasks/status/` have now been exercised
- against a real, `reported` analysis (see #314's own section below),
- so the service side of this is confirmed live. `cape-worker.py`'s
- own client code, though, has never itself submitted anything — its
- endpoint contract is CAPEv2's documented `apiv2` shape, the same
- starting point `ghidra-worker.py`'s own header warns went stale once
- already ("the endpoints originally taken from the plan documents
- were wrong"). Its own `--selftest` only checks reachability for
- exactly this reason — extend it into a real round trip next, the
- same discipline `ghidra-worker.py --selftest`'s real analysis round
- trip already holds itself to; nothing external blocks this now.
+ - **Verified live through this path** (2026-08-08, #318), not just the
+ service side: `--selftest --round-trip` submitted a real probe through
+ `CapeClient` itself, and both the submit and the poll reached `reported`
+ with a fetchable report. Same lesson `ghidra-worker.py`'s header warns
+ about — two endpoints taken from CAPEv2's *documented* apiv2 shape were
+ wrong and only a live run caught them: the report route is
+ `/apiv2/tasks/get/report/{id}/json/`, with an extra `get/` segment (the
+ documented `/apiv2/tasks/report/{id}/json/` 404s against a task that has
+ already reported), and `ready()` must not read
+ `/apiv2/cuckoo/status/` reporting itself disabled — this host's
+ `api.conf` has `[cuckoostatus]` off, an unrelated per-endpoint opt-in —
+ as CAPE being unreachable. Both corrections are recorded in
+ `cape-worker.py`'s own module docstring, which is the authority.
---
@@ -192,7 +194,7 @@ differently-configured venv without saying so.
| VM lifecycle | CAPE's own `kvm`/`libvirt` machinery module (not this worker) | `sandbox/windows/orchestrate/run_sample.py`, direct `virsh` |
| Results | `{sha256}_cape.json` → `CAPE_RESULTS_DIR`; dashboard only reads | `{sha256}_sandbox.json` → `WINDOWS_SANDBOX_RESULTS_DIR` |
| Trust boundary | Dashboard never touches libvirt, Docker, or CAPE's API credentials directly | Same |
-| Detail page | `/cape/{sha256}` — landed with #319, re-landed by the cutover (#1628); the page itself sits behind normal session auth — the backing `/api/v1/cape/{sha}` call is service-token gated, same middleware as the other detail pages, no admin/role check, detonation confirmation | `GET /sandbox/{job}` |
+| Detail page | `/cape/{sha256}` — landed with #319, re-landed by the cutover (#1628); the page itself sits behind normal session auth — the backing `/api/v1/cape/{sha}` call is service-token gated, same middleware as the other detail pages, no admin/role check, detonation confirmation | `GET /api/v1/sandbox/{job}` |
No new trust boundary. The dashboard container stays unprivileged and never
calls `virsh`, `docker`, or CAPE's own API directly — same guarantee every
@@ -206,17 +208,24 @@ other pipeline in this repo already holds itself to.
except `deterministic`. There is no automatic classification-based routing
to CAPE (or to `windows-sandbox`, or `windows-ghosts`) anywhere in this
codebase today — an operator selects analyzers explicitly in the Workbench
-UI, and `workbenchRegistry`'s `Applicable`/`Available` fields only control
+UI, and `workbench_registry`'s `applicable`/`available` fields only control
whether "cape" is offered as a *choice*, never whether it runs
automatically. This was already the right answer by construction once the
-registry entry existed (`dashboard/workbench_domain.go`'s `cape` entry,
-`AcceptedKinds: ["windows"]`, `Applicable: windowsApplicable`,
-`Confirmation: "detonation"`) — no separate routing logic needed building,
-and none should be: an operator choosing to spend an hours-long CAPE
-analysis slot is exactly the kind of decision this repo's Workbench
-pattern reserves for a human, the same reasoning `windows-ghosts`'s own
-loud, opt-in-only framing already documents for its own WAN-permitted
-route.
+registry entry existed — no separate routing logic needed building, and none
+should be: an operator choosing to spend an hours-long CAPE analysis slot is
+exactly the kind of decision this repo's Workbench pattern reserves for a
+human, the same reasoning `windows-ghosts`'s own loud, opt-in-only framing
+already documents for its own WAN-permitted route.
+
+That entry is the backend-service's
+[`workbench_domain.rs`](../../../arcane/home/honeypot-dashboard/backend-service/src/workbench_domain.rs)
+`cape` `WorkbenchAnalyzer`: `display_name: "CAPE sandbox"`,
+`accepted_kinds: ["windows"]`, `applicable: windows_applicable`,
+`confirmation: "detonation"`, `detonates: true`, `required_role: "admin"`,
+`concurrency: "cape-kvm"`, `local_only: true`,
+`result_link_shape: "/cape/{sha256}"`, and `availability`/`available` keyed
+on `cape_configured` — the CAPE spool being usable at all, not on the
+payload.
---
@@ -520,18 +529,6 @@ ghidra/revdeck entries already hold themselves to.
## Known gaps (tracked, not silently dropped)
-- **`cape-worker.py`'s CAPE API client is still unverified against a
- live service** — narrower than before, not removed: #314's own
- `utils/submit.py` CLI and the `/apiv2/tasks/status//` read
- endpoint are both now confirmed live against a real analysis (see
- above), which was the actual blocker (no service to test against).
- `cape-worker.py`'s own client code, endpoints matched against CAPEv2's
- documented `apiv2` blueprint but never yet exercised, is real
- remaining work — same category of risk that turned out wrong once
- already for `ghidra-worker.py`'s Ghidra REST client. Run
- `cape-worker.py --selftest` (extended into a real submission, the way
- `ghidra-worker.py --selftest` already does for its own service) as the
- next concrete step; nothing external blocks it now.
- **PostgreSQL not stood up.** CAPE's default SQLite task DB works for
#314's actual ask (get the host stack running, confirmed with a real
end-to-end analysis) and was made noticeably more concurrent-safe by
@@ -584,9 +581,20 @@ sandbox/cape/
honeypot-cape-worker.service systemd service unit
honeypot-cape.default.example /etc/default/honeypot-cape template
-dashboard/
- cape.go capeRequestDir/capeResultsDir (#319, partial)
- workbench_domain.go "cape" entry in workbenchRegistry (#319)
+arcane/home/honeypot-dashboard/backend-service/src/
+ workbench_domain.rs "cape" entry in the analyzer registry, plus the
+ CAPE_REQUEST_DIR/CAPE_RESULTS_DIR `cape_configured`
+ check (#319, re-landed by the cutover #1628)
+ detail.rs /api/v1/cape/{sha} + /raw result endpoints (#319)
+ worker.rs cape_alerts(): CAPE spool/worker health (#319,
+ partial — mirrors the retired Go cape.go)
+
+ # The Go tier's cape.go is gone. Its half that still has no Rust
+ # counterpart is the write side: workbench_orchestrator.rs's `marker_dir`
+ # has arms for ghidra / windows-sandbox / windows-ghosts / linux-sandbox /
+ # revdeck but none for cape, so a Workbench selection can list and validate
+ # "cape" yet never drops a {sha256}.request into the spool. That is the
+ # "partial" in #319, and it is why #317's routing is still manual.
sandbox/windows/run_pending.sh #320's shared cross-pipeline lock added
```
diff --git a/docs/sandbox/ghosts/IMPLEMENTATION_PLAN.md b/docs/sandbox/ghosts/IMPLEMENTATION_PLAN.md
index 66f807b10..66a3bd382 100644
--- a/docs/sandbox/ghosts/IMPLEMENTATION_PLAN.md
+++ b/docs/sandbox/ghosts/IMPLEMENTATION_PLAN.md
@@ -56,7 +56,7 @@ safe to run anything through.
Same as `docs/sandbox/windows/IMPLEMENTATION_PLAN.md`:
- KVM/QEMU/libvirt + docker-compose only — no VMware, no Hyper-V
- No CI-triggered detonation — the dashboard's Workbench is the only
- trigger (`workbench_orchestrator.go` → spool file → host-side systemd
+ trigger (`workbench_orchestrator.rs` → spool file → host-side systemd
worker)
- VM lifecycle via `virsh`/`qemu-img` only
- Results written to a spool directory the dashboard reads — no outbound
@@ -129,9 +129,10 @@ Two constraints specific to this chain:
- **Host-side GHOSTS sandbox worker** (systemd path unit) — [#328]
- Watches `GHOSTS_SANDBOX_REQUEST_DIR` for `{hash}.request` files
- written by `dashboard/workbench_orchestrator.go`'s "windows-ghosts"
- analyzer — a deliberately opt-in-only Workbench selection, never
- auto-routed to by payload classification
+ written by the backend-service's
+ [`workbench_orchestrator.rs`](../../../arcane/home/honeypot-dashboard/backend-service/src/workbench_orchestrator.rs)
+ "windows-ghosts" analyzer — a deliberately opt-in-only Workbench
+ selection, never auto-routed to by payload classification
- `process-ghosts-web-requests.sh` resolves the hash against the same
shared sample inbox `sandbox/windows`'s own resolution step uses
- `orchestrate/run_sample.py`: revert `win11-ghosts.qcow2` → WinRM/SMB
@@ -139,7 +140,7 @@ Two constraints specific to this chain:
execute sample → Sysmon EVTX snapshot → pull GHOSTS' own activity
log from `Ghosts.Api`'s database → revert again, unconditionally
- Writes `windows-ghosts-.json` → `GHOSTS_SANDBOX_RESULTS_DIR`,
- `dashboard/sandbox.go`'s `sandboxResult` shape, `"route":
+ the same result shape the other sandbox routes use, with `"route":
"windows-ghosts"` so the result page's isolation description (#327)
renders correctly instead of the default (wrong, for this route)
claim of "no forwarding, strict libvirt NIC filter"
@@ -155,7 +156,7 @@ Two constraints specific to this chain:
| Worker | `honeypot-ghosts-sandbox-worker.path` → `.service`, never run by the dashboard | `honeypot-windows-sandbox-worker.path` → `.service` |
| Results | `windows-ghosts-.json` → `GHOSTS_SANDBOX_RESULTS_DIR`; dashboard only reads | `windows-.json` → `WINDOWS_SANDBOX_RESULTS_DIR` |
| Trust boundary | Dashboard never touches libvirt, Docker, or WinRM directly | Same |
-| Detail page | `GET /sandbox/{job}` (shared route, `Route` field distinguishes) | `GET /sandbox/{job}` |
+| Detail page | `GET /api/v1/sandbox/{job}` (shared route, `Route` field distinguishes) | `GET /api/v1/sandbox/{job}` |
No new trust boundary. The dashboard container stays unprivileged and never
calls `virsh`, `docker`, or WinRM directly — same guarantee `sandbox/windows`
diff --git a/docs/sandbox/ghosts/README.md b/docs/sandbox/ghosts/README.md
index 4347e0f73..618b41cf2 100644
--- a/docs/sandbox/ghosts/README.md
+++ b/docs/sandbox/ghosts/README.md
@@ -11,9 +11,12 @@ are #325-#330.
`ghosts-postgres` + `ghosts-api` only, built from CMU SEI's
[`cmu-sei/GHOSTS`](https://github.com/cmu-sei/GHOSTS) source pinned to the
-`v9.0.0` tag (no published images exist upstream, so `compose.yml`'s
-`build.context` points a git URL straight at `src/` — nothing to clone by
-hand).
+`v9.0.0` tag (no published images exist upstream, so the source is *vendored*
+at `sandbox/ghosts/vendor/ghosts-src/` and `compose.yml` builds from that
+local context through this repo's own `Dockerfile.api-prep`). It used to be a
+remote git build context pointed straight at `src/`; that stopped working when
+Arcane resolved build-context git refs under `refs/heads/` only, so #1506
+vendored the tree instead — see that directory's `VENDORED.md`.
**Deliberately not deployed**: Frontend, Grafana, n8n. None of the three are
required for NPC-simulation (timeline-driven browsing/document/handler
@@ -31,24 +34,41 @@ narrow (dashboard container never touches Docker/libvirt/WinRM directly).
## Fixed address
-`ghosts-api` publishes no host port. It gets a static address on the
-dedicated `ghosts_net` bridge instead:
+`ghosts-api` publishes port 5000 on the libvirt `ghosts` network's *own gateway*
+address, virbr-ghosts's `10.20.30.1` — not the docker-internal `10.90.0.2` an
+earlier version of this file used:
```
-GHOSTS_API_ADDR=10.90.0.2:5000
+GHOSTS_API_ADDR=10.20.30.1:5000
```
-Reachable from this host directly — Docker routes user-defined bridges
-without any `-p` — and, once #325 exists, from the WAN-permitted GHOSTS
-guest through exactly one narrow routing/firewall exception written against
-this single address. Same pattern as RevDeck's
-`REVDECK_API_BASE=http://10.8.0.2:19500`: pick the fixed address once, up
-front, specifically so later issues can write a one-line exception instead
-of a floating rule.
-
-Don't change `10.90.0.2` without updating #325's LAN-blocking exception to
-match, and without updating `/etc/default/honeypot-ghosts` on the host
-(written by `install-host.sh`, read by whatever #325/#328 add later).
+The first attempt gave `ghosts-api` only a static address on the dedicated
+`ghosts_net` bridge (`10.90.0.2`) and published no host port. That is fine
+host-locally — Docker routes user-defined bridges without any `-p` — but it
+never reached the WAN-permitted GHOSTS guest: recent Docker versions add a
+`raw` table PREROUTING rule (`ip daddr iifname != drop`) that blocks routing straight to a container's backend IP from any
+other interface, regardless of what FORWARD/DOCKER-USER say, because Docker
+expects cross-network reachability to go through a published port. Binding to
+virbr-ghosts's gateway also means the guest needs no FORWARD-chain exception at
+all: the traffic is local to its own default gateway, covered by
+`network-filter.sh`'s ordinary bridge-gateway ACCEPT. Requires the `ghosts`
+libvirt network (`network.xml`) to exist before this container starts, or Docker
+cannot bind the address. Same "one fixed, documented address" pattern as
+RevDeck's `REVDECK_API_BASE=http://10.8.0.2:19500`: pick the address once, up
+front, specifically so later issues can write a one-line exception instead of a
+floating rule.
+
+`10.90.0.2` still exists — it remains `ghosts-api`'s static address on
+`ghosts_net`, which is how the throwaway `ghosts-client-test` container resolves
+the API by service name on the same bridge, and why `ghosts-postgres` is pinned
+to `10.90.0.3` (Docker would otherwise hand `10.90.0.2` to the database first
+and the API's explicit request would fail to start).
+
+Don't change `10.20.30.1` without updating `network.xml` and
+`network-filter.sh` to match, and without updating
+`/etc/default/honeypot-ghosts` on the host (written by `install-host.sh`, read
+by whatever #325/#328 add later).
## Deploy
@@ -77,15 +97,38 @@ endpoint.
sandbox/ghosts/install-host.sh --skip-enroll-test # containers only
```
-## Notes for whoever picks up #325 (network isolation)
-
-- `ghosts-api` currently only needs to be reachable from this host and from
- other containers on `ghosts_net` — nothing reaches it from outside yet.
- #325's job is adding the single guest → `10.90.0.2:5000` exception through
- the isolated bridge that issue creates for the GHOSTS guest, not opening
- this address any wider.
-- The stack has no authentication in front of it (matches GHOSTS' own
- defaults — verify this hasn't changed before relying on it). That's fine
- while the only path to it is host-local/docker-internal; it stops being
- fine the moment #325 adds a route from a WAN-permitted guest, so revisit
- before that lands.
+## Network isolation as it actually stands (#325, #2444, #2257)
+
+#325 has landed, so this is no longer a note to a future issue — it is the
+shipped state, and the address paragraph above is written against it.
+
+- The `ghosts` libvirt network is `network.xml`; `install-network.sh` plus
+ `network-filter.sh` (unit `ghosts-network-filter.service`) are what put the
+ WAN-facing guest behind the host's FORWARD DROP/ACCEPT pairs. It drops
+ RFC1918 and LAN destinations generally while leaving DNS real, and
+ `verify-network-isolation.sh` is the guest-side check.
+- `ghosts-api` is reachable from the guest only through the published
+ `10.20.30.1:5000`, and `network-filter.sh` additionally source-pins tcp/5000
+ on virbr-ghosts to the enrolled clients listed in its `GHOSTS_API_CLIENTS`
+ (today just `10.20.30.50/32`, `win11-ghosts`), so an unpinned guest fails
+ closed at the firewall even for routes that do exist. Admitting a new client
+ is deliberately a two-place change: a static `` entry
+ in `network.xml` **and** an address in `GHOSTS_API_CLIENTS`.
+- #2444 image-preps the route surface rather than trusting the upstream
+ image: `Dockerfile.api-prep` deletes the animations control plane and the
+ `/api/attack` scenario tooling (upstream operator tooling #324 excluded, and
+ #2444 showed an unauthenticated guest could drive them — scheduled
+ server-side GETs to any caller-supplied URL, ATT&CK-table wipe-and-reload),
+ and patches Swagger's middleware out. What remains is the client
+ enrollment/check-in plane plus machine inventory, timelines, surveys,
+ results and the SignalR hubs.
+- The API still has no authentication of any kind — this is the *accepted
+ residual risk* #2257 wrote down, not an oversight. `appsettings.json`'s
+ `InitSettings` block reads like a credentialed surface but is bound and then
+ discarded by `ApiDetails.LoadConfiguration()`, so it gates nothing. #2257
+ narrowed the blast radius (the deleted routes above, no anonymous endpoint
+ map, `ASPNETCORE_ENVIRONMENT=Production` so no developer exception pages),
+ not the auth model. A compromised ghost can still read and rewrite NPC state
+ — notably `TimelinePartial` updates, which silently steer NPC behaviour and
+ poison experiments built on it. Do not put anything on this API that a
+ compromised guest must not read or forge.
diff --git a/docs/sandbox/windows-guest-risk-config-model.md b/docs/sandbox/windows-guest-risk-config-model.md
index c6a5b02bf..2f321a125 100644
--- a/docs/sandbox/windows-guest-risk-config-model.md
+++ b/docs/sandbox/windows-guest-risk-config-model.md
@@ -7,7 +7,7 @@
> an existing one) has a real model to check itself against instead of
> re-deriving the reasoning from scratch. See "Follow-up work" for what
> this surfaced but didn't build.
-> **Last updated**: 2026-08-08
+> **Last updated**: 2026-09-27
> **Tracking**: [#467](https://github.com/Xore/APIARY/issues/467)
---
@@ -87,7 +87,7 @@ insufficient without the loud warning).
| Axis | `win11-sandbox` | `win11-ghosts` | `win11-cape` |
|---|---|---|---|
| 1. Network exposure | Isolated (no ``; FakeNet-served for intercepted outbound) | **Real WAN** (`` present, deliberate — #325/#331) | Isolated (no ``, same posture as the Linux runner) |
-| 2. Persona / NPC | Legacy persona daemon (`07-living-persona.ps1`, #290) | GHOSTS NPC (real `Ghosts.Api` client, `sandbox/ghosts/Dockerfile.client-win`) | **None** — no NPC daemon, deliberately excluded (`win11-cape.pkr.hcl`'s own header). Static identity (`autounattend.xml`'s `ComputerName`/`FullName`/etc.) is distinct from `win11-analysis`'s own as of #904, closing the fingerprint-reuse gap this cell used to flag |
+| 2. Persona / NPC | Legacy persona daemon (`07-living-persona.ps1`, #290) | GHOSTS NPC — upstream `Ghosts.Client.Universal`, not `Ghosts.Client.Windows` (#326) and not the `Ghosts.Api` server, built by `sandbox/ghosts/Dockerfile.client-win`; the built assembly is renamed to `EndpointAgent` there (and `C:\ghosts\Ghosts.Client.Universal.exe` is not shipped) because the original filename is itself a giveaway | **None** — no NPC daemon, deliberately excluded (`win11-cape.pkr.hcl`'s own header). Static identity (`autounattend.xml`'s `ComputerName`/`FullName`/etc.) is distinct from `win11-analysis`'s own as of #904, closing the fingerprint-reuse gap this cell used to flag |
| 3. Simulated input | Yes — cubic-Bezier mouse movement, Gaussian jitter, periodic typing (`07-living-persona.ps1`) | N/A — GHOSTS' own real activity substitutes | No |
| 4. Simulated background traffic | Yes (`08-traffic-noise.ps1`) | N/A — real traffic from real browsing | No |
| 5. Filesystem bait | Yes (`05-decoy-content.ps1`) | No | No |
@@ -122,13 +122,22 @@ filed as its own issue rather than bundled here (every one of them
needs an actual golden-image rebuild to verify, the same rebuild-gated
posture #368/#787's own comments already hold every other
guest-behavior change to — not something to casually re-trigger inside
-a documentation change):
+a documentation change). #904 has since landed and is marked as such
+below; the other three are still open:
- [#901](https://github.com/Xore/APIARY/issues/901) — Validate the
admin-gated LOLDrivers toggle end-to-end against a real
`win11-ghosts.qcow2` cycle (with-set vs. without, gate-on vs.
gate-off) — the code (#873) already exists and is untested against a
- real image; this is verification work, not new engineering.
+ real image; this is verification work, not new engineering. The
+ evidence-gathering half now exists too
+ (`sandbox/ghosts/loldriver-gate-test.ps1`, #901's own acceptance
+ script: load-attempt `RTCore64.sys` as a kernel service, report
+ whether it actually loaded, read back
+ `VulnerableDriverBlocklistEnable`, delete the service). No run
+ output is recorded in this repo, so the validation itself is still
+ outstanding — the script being present is not the same as it having
+ been run.
- [#902](https://github.com/Xore/APIARY/issues/902) — Design and add a
userspace-only vulnerable-software attack-surface option (axis 7) —
genuinely new engineering: which software, which CVEs, how it's
@@ -139,9 +148,12 @@ a documentation change):
(axis 6 and/or 7) at all, given its debugger-class-evasion focus
differs from `win11-analysis`'s AV/behavioral-evasion one — and
implement whichever way that decision goes.
-- [#904](https://github.com/Xore/APIARY/issues/904) — Give
- `win11-cape` its own persona identity distinct from
- `win11-analysis`'s (not full persona/input/traffic-noise parity,
- which stays deliberately excluded — just fixing the fingerprint-reuse
- gap `autounattend.xml`'s own header already flags, promoted here to a
- tracked issue instead of a comment-only note).
+- ~~[#904](https://github.com/Xore/APIARY/issues/904)~~ — Give
+ `win11-cape` its own persona identity distinct from `win11-analysis`'s
+ — **landed**. `sandbox/cape/packer/autounattend.xml` now carries
+ `VPM-ENG0089` / Daniel Kowalski / Vantage Precision Manufacturing
+ against `win11-analysis`'s `ACP-FIN0142` / Robert Tanaka / Ashford
+ Capital Partners, and the matrix cell above reflects that. Note the
+ scope it actually shipped at: identity/fingerprint distinctness only
+ — not full persona/input/traffic-noise parity, which stays
+ deliberately excluded (see that cell, and the file's own header).
diff --git a/docs/sandbox/windows/IMPLEMENTATION_PLAN.md b/docs/sandbox/windows/IMPLEMENTATION_PLAN.md
index d865b8e79..4fba685be 100644
--- a/docs/sandbox/windows/IMPLEMENTATION_PLAN.md
+++ b/docs/sandbox/windows/IMPLEMENTATION_PLAN.md
@@ -1,13 +1,34 @@
# Windows 11 Malware Sandbox — Golden Image Implementation Plan
-> **Status**: In Progress — Phase 7's dashboard half is implemented, and the
-> host half now has its orchestrator, spool worker, and systemd units. What
-> remains is the golden image itself (Phases 1–3) and the gateway compose
-> (Phase 4); until a `win11-sandbox` domain exists, the worker will
-> revert-fail on every request and preserve it as `.request.failed`. There
-> is no `GOLDEN_READY` snapshot — see the revised Golden Image vs Snapshots
-> decision below and #358 for why.
-> **Last updated**: 2026-07-30
+> **Design record, not current-behaviour documentation.** Every
+> `dashboard/*.go` reference below, the Go code blocks, and the route column
+> of the "Wiring Pattern" table describe the Go dashboard deleted at #1628.
+> Phase 7's decisions were re-implemented in the Rust `backend-service`
+> (`arcane/home/honeypot-dashboard/backend-service/src/`) and the
+> `frontend-next` routes; the route table is `main.rs` (`POST
+> /api/v1/sandbox/submit`, `GET /api/v1/sandbox/{job}`, `GET
+> /api/v1/sandbox/golden-image-status`, `GET /api/v1/sandbox/vnc`, `POST
+> /api/v1/ghidra/submit`, `GET /api/v1/ghidra/{sha}`), and the host-half
+> operator notes live in [`runner/README.md`](runner/README.md). The plan's
+> binding content — the spool-file trust boundary, the Windows/Linux
+> determination path, the #358 golden-image-over-snapshot decision, the
+> in-guest detonation chain — carried over and still holds. Read the Go
+> paths as history, not as somewhere to write code. The host-side
+> systemd/spool specifics in §7.2 and §7.3 have drifted further than that
+> and are corrected inline.
+>
+> **Status**: Shipped. Phase 7's dashboard half is implemented, and the
+> host half now has its orchestrator, spool worker, and systemd units. Every
+> build artifact Phases 1–4 describe is present in the tree, and the live-host
+> step this file used to leave open is now closed: `win11-analysis.qcow2` has
+> been built on this host and rebuilt several times — #1128 fixed a rebuild
+> breaker and verified the fix against a real rebuild ("the VM now boots and
+> reaches the WinRM-wait stage"), and #957's screen-resolution fix was
+> confirmed live in a running guest. The `win11-sandbox` domain therefore
+> exists and the worker's `revert` works, rather than revert-failing on every
+> request. There is no `GOLDEN_READY` snapshot — see the revised Golden Image
+> vs Snapshots decision below and #358 for why.
+> **Last updated**: 2026-09-27
> **Host platform**: KVM + QEMU + libvirt + docker-compose (NO VMware)
> **Phase 1 tracking**: [#47](https://github.com/Xore/APIARY/issues/47)
> — one issue per remaining step, each with its own verification and failure
@@ -28,10 +49,16 @@
The analysis host runs **KVM/QEMU/libvirt** and **docker-compose** only.
- No VMware Workstation, no VirtualBox, no Hyper-V
- No GitHub Actions — the sandbox is triggered **from the dashboard**, not CI
-- All VM lifecycle (create, snapshot, revert, destroy) via `virsh` / `qemu-img`
+- All VM lifecycle (create, revert, start, stop, status) via `virsh` / `qemu-img`.
+ There is no snapshot path — see the Golden Image vs Snapshots decision below
+ and `setup/kvm_manage.sh`'s own header for why.
- Gateway services run as **Docker Compose services** (INetSim, Zeek, Suricata, mitmproxy)
- Golden image built automatically with **Packer + QEMU builder**
-- Orchestrator uses `libvirt` Python API (`libvirt-python`)
+- Orchestrator drives VM lifecycle through `virsh` / `qemu-img` subprocesses
+ (see Phase 5). It does **not** use the `libvirt` Python API — an earlier
+ revision of this constraint said `libvirt-python`, which contradicted
+ Phase 5 and was never true. In this repository `import libvirt` appears
+ exactly once, in CAPE's own vendored `sandbox/cape/capev2-overrides/modules/machinery/capekvm.py`.
- Results written to a spool directory the dashboard reads — **no outbound network, no git push**
---
@@ -58,7 +85,7 @@ flowchart TD
DockerNet --> Suricata["suricata — IDS on virbr-sandbox"]
Host --> Worker["Host-side sandbox worker (systemd path unit)"]
- Worker --> Watch["Watches WINDOWS_SANDBOX_REQUEST_DIR for {hash}.request files written by the dashboard (sandbox_submit.go) — routed here only after the dashboard's determination path (see below) classifies the payload as Windows; everything else goes to the pre-existing Linux runner (sandbox/linux-runner.service, sandbox/worker.sh) watching the original SANDBOX_REQUEST_DIR"]
+ Worker --> Watch["Watches WINDOWS_SANDBOX_REQUEST_DIR for {hash}.request files written by the dashboard (sandbox_submit.rs) — routed here only after the dashboard's determination path (see below) classifies the payload as Windows; everything else goes to the pre-existing Linux runner (sandbox/linux-runner.service, sandbox/worker.sh) watching the original SANDBOX_REQUEST_DIR"]
Worker --> Revert["destroy + fresh CoW clone from golden image + start (kvm_manage.sh revert / run_sample.py revert_to_golden() — not a virsh snapshot, see #358)"]
Worker --> Detonate["WinRM → copy sample, start tools, detonate"]
Worker --> Wait["Wait observation window"]
@@ -79,11 +106,19 @@ integration (`docs/analysis/ghidra/DASHBOARD_INTEGRATION_PLAN.md`):
| Worker | Host-side systemd path unit (`honeypot-windows-sandbox-worker.path`), never run by the dashboard | Host-side systemd path unit (`honeypot-ghidra-worker.path`) |
| Results | Worker writes `{hash}_sandbox.json` to `SANDBOX_RESULTS_DIR`; dashboard only reads | Worker writes `{sha256}_ghidra.json` to `GHIDRA_RESULTS_DIR`; dashboard only reads |
| Trust boundary | Dashboard never touches Docker, libvirt, or the VM directly | Same |
-| List page | `GET /sandbox` → `sandboxData()` → `{{define "sandbox"}}` | `GET /ghidra` → `ghidraData()` |
+| List page | *none* — the Go `GET /sandbox` list page and its `{{define "sandbox"}}` template were not re-landed; `frontend-next` has `/sandbox/$job` and `/sandbox/vnc` only | *none* — `frontend-next` has `/ghidra/$sha` only |
| Detail page | `GET /sandbox/{job}` | `GET /ghidra/{sha256}` |
| JSON API | `GET /api/sandbox`, `/api/sandbox/{job}` | `GET /api/ghidra`, `/api/ghidra/{sha256}` |
| Export | `GET /export/sandbox/{job}` (bundle download) | `GET /export/ghidra/{sha256}` |
+Current forms of the routes that do exist are the `/api/v1/…` ones listed
+in the banner above. The three right-hand columns were re-checked on
+2026-09-27: the list page, the `/api/…` JSON tier and the `/export/…`
+bundle routes have **no** Rust counterpart — `main.rs` exposes no
+`/api/v1/export/sandbox` or `/api/v1/export/ghidra` route, and the only
+`/api/v1/export/*` handlers are the six CSV/JSON event, command, IP,
+campaign, cluster and history exports plus `/api/v1/ip-block-export`.
+
No new trust boundary is introduced. The dashboard container stays
unprivileged and **never** calls `virsh`, `docker`, or WinRM directly.
@@ -100,13 +135,15 @@ captured payload today.
### Signal: reuse `classifyPayload` — no new classifier
-`dashboard/payload_kind.go`'s `classifyPayload(data []byte) payloadClassification`
-already sniffs magic bytes (`MZ` → `debug/pe`, `\x7fELF` → `debug/elf`,
-script shebangs/headers) and returns a `Platform` of `"Windows"`,
-`"Linux"`, or `"Cross-platform"` for every kind of payload the dashboard
-already stores — this is the exact same classification already shown on
-the payload detail page ("Windows PE forensics" card, etc). Routing needs
-no new detection logic, only a decision on top of the existing field.
+`classifyPayload(data []byte) payloadClassification` — now
+[`payload_kind.rs`](../../../arcane/home/honeypot-dashboard/backend-service/src/payload_kind.rs)'s
+`classify_payload` — already sniffs magic bytes (`MZ` → `debug/pe`,
+`\x7fELF` → `debug/elf`, script shebangs/headers) and returns a `Platform`
+of `"Windows"`, `"Linux"`, or `"Cross-platform"` for every kind of payload
+the dashboard already stores — this is the exact same classification
+already shown on the payload detail page ("Windows PE forensics" card,
+etc). Routing needs no new detection logic, only a decision on top of the
+existing field.
| `classifyPayload(...).Code` | `.Platform` | Routed to |
|---|---|---|
@@ -120,6 +157,16 @@ no new detection logic, only a decision on top of the existing field.
### Determination function (`dashboard/sandbox_submit.go`)
+Now `determine_sandbox_target()` in
+[`sandbox_submit.rs`](../../../arcane/home/honeypot-dashboard/backend-service/src/sandbox_submit.rs).
+It returns `Option<&'static str>` rather than a `(target, dynamic)` pair —
+`None` is the "not dynamic, no VM submission possible" answer, which is the
+Go function's `dynamic == false` return folded into one. It still returns
+only `"windows"` or `"linux"`: a `"ghosts"` arm exists, but in the sibling
+`sandbox_request_dir()` rather than here, because the GHOSTS route is
+WAN-permitted and opt-in only through the Workbench, so classification must
+never select it.
+
```go
type sandboxTarget string
@@ -283,7 +330,7 @@ packer plugins install github.com/hashicorp/qemu
apt install -y p7zip-full python3-virt-firmware sbsigntool
# Python deps for orchestrator
-pip install libvirt-python pywinrm python-evtx lxml requests smbprotocol
+pip install pywinrm python-evtx lxml requests smbprotocol
# Docker Compose (for gateway services)
apt install -y docker-compose-plugin
@@ -389,10 +436,12 @@ gone — see the removal note in `win11-analysis.pkr.hcl` itself.)
# Rebuild from scratch
packer build -force win11-analysis.pkr.hcl
-# Update just the logging config without full rebuild:
-virt-customize -a /golden-images/win11-analysis.qcow2 \
- --upload sysmon_config.xml:/Windows/sysmon_config.xml \
- --run-command 'C:\Windows\sysmon64.exe -c C:\Windows\sysmon_config.xml'
+# There is deliberately no "update just the logging config without a rebuild"
+# virt-customize step here. An earlier revision of this plan had one, uploading
+# a sysmon_config.xml that does not exist in the repo: 04-tools.ps1 fetches the
+# config at build time from raw.githubusercontent.com, pinned to a commit SHA and
+# verified against a recorded sha256 (#86). To change it, re-pin both values in
+# 04-tools.ps1 and rebuild.
```
---
@@ -437,9 +486,13 @@ only step that empirically confirms it holds up in a real booted guest.
## Phase 3 — Windows 11 Hardening for Malware Analysis
Implemented in
-[`packer/scripts/`](../../../sandbox/windows/packer/scripts/) — four provisioner scripts, the
-hardening and anti-evasion phases run at image-build time, not as a separate
-script. (Earlier revisions of this plan named a `setup/harden_analysis_vm.ps1`
+[`packer/scripts/`](../../../sandbox/windows/packer/scripts/) — ten provisioner
+scripts (`01-hardening`, `04-tools`, `05-decoy-content`, `06-chrome-history`,
+`07-living-persona`, `08-traffic-noise`, `09-vcredist`, `10-loldrivers`,
+`11-detonation-orchestrator`, `12-display-resolution`; the numbering has
+gaps because it tracks the Packer phase each one serves), the hardening and
+anti-evasion phases run at image-build time, not as a separate script.
+(Earlier revisions of this plan named a `setup/harden_analysis_vm.ps1`
that was never written.)
### 3.1 Disable Noise Sources
@@ -459,43 +512,71 @@ that was never written.)
### 3.2 Enable Maximum Telemetry (Analyst Side)
```
-✓ Sysmon 64 with SwiftOnSecurity config
+✓ Sysmon 64 with SwiftOnSecurity config (fetched at build time, pinned to a
+ commit SHA and verified against a recorded sha256 — 04-tools.ps1, #86)
✓ PowerShell ScriptBlock logging (Event 4104)
✓ PowerShell Module logging (Event 4103)
✓ PowerShell Transcription to C:\PSTranscripts\
-✓ Process creation auditing (Event 4688 + full cmdline)
-✓ Object access auditing (Event 4663)
-✓ Registry auditing (Event 4657)
-✓ All event log sizes expanded to 500 MB
+✓ Process creation auditing (Event 4688 + full cmdline) — `auditpol /set
+ /subcategory:'Process Creation'` plus ProcessCreationIncludeCmdLine_Enabled
+✗ Object access auditing (Event 4663) — not implemented
+✗ Registry auditing (Event 4657) — not implemented
+✓ All event log sizes expanded to 500 MB (Sysmon/Operational,
+ PowerShell/Operational, Security, System, Application)
✓ FakeNet-NG intercepting all outbound traffic
✓ QEMU guest agent (for host-side artifact collection)
```
+`Process Creation` is the **only** audit subcategory this build touches;
+`04-tools.ps1` has exactly one `auditpol` call. The two `✗` lines were
+carried here as done and are not — nothing in `sandbox/windows/packer/`
+enables either subcategory.
+
### 3.3 Anti-Evasion (Make VM Look Real)
```
-✓ Hostname: DESKTOP-$(random 7 chars) — matches real Win11 pattern
-✓ Username: john.doe / jane.smith / mike.wilson (rotation)
-✓ Populate: Documents, Desktop, Downloads with decoy files
-✓ Install: Chrome, 7-Zip, Notepad++ (common software footprint)
-✓ Browser history: inject fake history entries
-✓ Recent files: inject 20+ fake recent document entries
-✓ Disk: 80 GB+ (malware checks disk size < 60 GB = sandbox)
-✓ RAM: 8 GB+ (malware checks < 4 GB = sandbox)
-✓ CPU: 4 vCPU (malware checks < 2 = sandbox)
-✓ Screen: 1920x1080
+✓ Hostname: ACP-FIN0142, a business-shaped decoy, not a DESKTOP-* pattern
+ (autounattend.xml; it was DESKTOP-AN4LY5T, then DESKTOP-JK3PLQ2, before
+ #293 — that shape is exactly what the guest is meant to stop looking like)
+✓ Username: a single `analyst` account, matching the WinRM/autologon account
+ the provisioners and run_sample.py all use — not a rotating persona set
+✓ Populate: Documents with decoy PDFs/RTF/CSV (05-decoy-content.ps1)
+✓ Install: Chrome (06-chrome-history.ps1, which also seeds its history) and
+ Sysinternals/Regshot/FakeNet/QEMU guest agent. Phase 8's "common software"
+ is runtimes, not desktop apps: vcredist-all, dotnetfx, dotnet-6.0/8.0
+ desktopruntime, javaruntime, silverlight
+✓ Browser history: inject fake history entries (06-chrome-history.ps1)
+✓ Recent files: real .lnk shortcuts (WScript.Shell), back-dated 1-60 days
+✓ Disk: 90 GB (disk_size = "90000" — malware checks disk size)
+✓ RAM: 16 GB at build (memory = "16384"); the detonation domain runs
+ 16 GB with 8 GB current
+✓ CPU: 12 vCPU at build (cpus = "12"); the detonation domain runs 8 vCPU
+✓ Screen: 1920x1080 (win11-kvm.xml