From 78d35c08a6cdce0a0d79eca1e5a8bb8e43499e66 Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 14:47:15 +0200 Subject: [PATCH 01/21] docs(arcane): reconcile ARCANE-GIT-SYNC.md with the 39-entry manifest - 37/31/32 census -> 39 (33 under arcane/home/, 6 self-contained at root paths); document #2911 technitium and #3092 unsloth - manifest import section: filter reaches 35 of 39, technitium replaced pihole's dedicated step, unsloth has no installer path at all - manifest declares branch: production since #1943, not main; no refs/heads/production exists on origin (checked 2026-09-27) - 34-of-37 build / three pullers -> 34-of-39 build / five pullers - restart: no census: 20 hits, 13 declarations, 8 files; honeypot-elk's arkime-pcap-init (#3128) is a second Arcane-started one-shot - relative-bind-mount set 9 -> 10 (ghidra); env_file claim: ghidra has one - re-count dashboard 268->279 files, ghosts 989->962, ghosts-src 963->935 - :?required stacks are canarytokens/ghosts/technitium/unsloth --- docs/ARCANE-GIT-SYNC.md | 244 +++++++++++++++++++++++++++------------- 1 file changed, 168 insertions(+), 76 deletions(-) diff --git a/docs/ARCANE-GIT-SYNC.md b/docs/ARCANE-GIT-SYNC.md index 6b26dddc..c2d08d0b 100644 --- a/docs/ARCANE-GIT-SYNC.md +++ b/docs/ARCANE-GIT-SYNC.md @@ -1,6 +1,6 @@ # Arcane Git sync -How the 37 home-hosted stacks (31 that migrated under `arcane/home/` plus 6 +How the 39 home-hosted stacks (33 that live under `arcane/home/` plus 6 that were already self-contained and stayed at their existing path) get to the live host, replacing the old model of copying or symlinking top-level `docker-compose.*.yml` files into place. Everything here was confirmed live @@ -9,7 +9,10 @@ taken from Arcane's docs — see the risk that motivated that in "Version/API compatibility" below. Census numbers and the live sync-store state were re-verified 2026-08-27 against the pinned `v2.9.0` image and its own sqlite store (#2549): 31 in-tree directories + 6 self-contained = 37, exactly the -manifest's entry count. +manifest's entry count *at that date*. The manifest has since grown to 39 — +#2911 swapped the `pihole` entry for `technitium` and #3092 added `unsloth` +directly under `arcane/home/` — re-counted 2026-09-27 from +`arcane/manifests/home-production.json`: 33 in-tree + 6 self-contained = 39. ## The model @@ -18,19 +21,21 @@ Each stack gets its own **directory-aware Git sync**: Arcane clones the selected `compose.yml` (not just that one file) under `/var/dockge/stacks//`, and deploys it. The manifest at [`arcane/manifests/home-production.json`](../arcane/manifests/home-production.json) -is the single source of truth for which 37 stacks exist, what branch/path +is the single source of truth for which 39 stacks exist, what branch/path each syncs from, and any per-stack sync limits — `scripts/install-homeserver.sh`, CI, and this doc all read from it rather than maintaining separate lists. - The 32 `honeypot-*` stacks live under `arcane/home//`: their build context and git-tracked config were moved there from repository root (see each compose file's own `#1502` comment for what moved and why). + `unsloth` is the 33rd directory there — #3092 added it in-tree rather + than at a root path, so it never went through the #1502 move. - The 6 other stacks (`auth-events-worker`, `llm-worker`, `ml-worker`, - `analysis/ghidra`, `sandbox/ghosts`, `pihole`) were already self-contained - and stayed at their existing path — moving them would have broken real - references from `scripts/install-homeserver.sh`, CI workflows, and - `deploy.yml`'s own ghidra-worker resync step. See each one's own compose - file header for the specifics. + `analysis/ghidra`, `sandbox/ghosts`, `technitium`) were already + self-contained and stayed at their existing path — moving them would + have broken real references from `scripts/install-homeserver.sh`, CI + workflows, and `deploy.yml`'s own ghidra-worker resync step. See each + one's own compose file header for the specifics. - `honeypot-arcane` itself is **not** in the manifest and never will be — syncing the thing that has to already be running before any sync can happen is a bootstrap loop, not a simplification. It stays @@ -58,15 +63,27 @@ been provisioned once. `scripts/install-homeserver.sh`'s `step_arcane_import_stacks` reads the manifest and creates one `POST /environments/0/gitops-syncs` per matching entry (environment `0` is Arcane's single "Local Docker" environment on a -one-host deployment). Its selection filter matches every `honeypot-*` -entry — and, since #1505, three of the six non-`honeypot-*` stacks too: -`auth-events-worker`, `llm-worker` and `ml-worker` are imported by that -step as well (each confirmed to have no host-local state beyond `.env`). -The other three keep their dedicated installer steps for reasons specific -to each: `pihole`'s non-`.env` host state, `analysis/ghidra`'s conditional -GPU compose overlay, and `sandbox/ghosts`'s Arcane build-context -limitation (#1506) — see the script's own Phase 8 header comment for the -reasoning behind each. +one-host deployment). Its selection filter matches all 32 `honeypot-*` +entries — and, since #1505, three of the seven non-`honeypot-*` entries +too: `auth-events-worker`, `llm-worker` and `ml-worker` are imported by +that step as well (each confirmed to have no host-local state beyond +`.env`). That is 35 of the manifest's 39 entries. The other three the +installer provisions itself keep their dedicated steps for reasons specific +to each: `technitium`'s non-`.env` host state (its `config/` directory needs +non-root ownership, #2911 — this is the step `pihole` used to have), +`analysis/ghidra`'s conditional GPU compose overlay, and `sandbox/ghosts`' +s Arcane build-context limitation (#1506) — see the script's own Phase 8 +header comment for the reasoning behind each. + +The 39th entry, `unsloth` (#3092), is the one the installer does not reach +at all: it matches neither arm of the filter, and the script has no +`step_unsloth_*` of its own. That is not a defect — `arcane/home/unsloth/compose.yml`'s +own header documents it as an operator-started stack ("Deployed through +Arcane like every other homeserver stack … Never `docker compose up` by +hand"), deliberately carrying no `restart:` policy so the cold-benchmark legs +in `analysis/ghidra/training/` can have the card to themselves. It is +covered where fleet-wide operations are concerned — `docs/STACK-REBUILD.md` +lists it in both reset loops — just not by the from-scratch import path. To import (or re-import) by hand instead, `POST` the manifest's entries to `/environments/0/gitops-syncs/import` — the bulk-import shape matches the @@ -106,16 +123,19 @@ letting Arcane's own directory check block the whole import. not track** — not just `.env` and `secrets/`. Confirmed live: `pihole` also keeps its DNSCrypt resolver config and its own Pi-hole database directly under its own top-level directory (`dnscrypt-proxy/`, - `etc-pihole/`, `etc-dnsmasq.d/`), a shape none of the other 37 stacks - have. Back up the *actual* bind-mount sources a stack's compose file + `etc-pihole/`, `etc-dnsmasq.d/`), a shape only a handful of stacks have + — the live one now is `technitium`, whose `config/` (zones, settings, + blocklists) needs non-root ownership for its distroless image, which is + exactly why the installer still provisions it by hand (#2911). Back up the *actual* bind-mount sources a stack's compose file declares, not an assumed `.env`/`secrets/` checklist — read the compose file if in doubt. 2. Back up everything found in step 1. 3. Remove the stack's current directory. 4. Create the Arcane sync (`syncDirectory: true`) — this deploys immediately, and will legitimately fail-closed if a required secret - isn't present yet (expected, not a bug — see canarytokens/ghosts/ - keycloak/dashboard's own `:?required` variables). + isn't present yet (expected, not a bug — see the four stacks that do + declare `:?required` variables: `honeypot-canarytokens`, `ghosts`, + `technitium` and `unsloth`). 5. Restore everything backed up in step 1, **preserving original ownership and permissions, not just content**. Confirmed live: restoring a secret file as `root:root` when the container expects the previous owning @@ -194,18 +214,24 @@ alone. Re-verify against whatever Arcane version is pinned in is no per-message log-suppression Arcane exposes, and `hp-arcane`'s container logs aren't ingested into this repo's ELK pipeline (it tails application log files, not `docker logs` streams), so there is no - repo-side filter either. Fires for exactly the 9 stacks with at least one - relative-source bind mount under this identity mount (`honeypot-cowrie`, - `honeypot-dionaea`, `honeypot-elk`, `honeypot-init`, `honeypot-keycloak`, - `honeypot-payload-analysis`, `honeypot-tanner`, `honeypot-utilities`, - `pihole`) — every relative mount on the 7 of those 9 with running - containers was confirmed to resolve to its real, existing host path, zero - missing-on-host mounts. (`pihole` and `honeypot-keycloak` were the two - without running containers, so their mounts were reasoned from the same - identity-mount arithmetic rather than observed — #2853, #2764. `pihole` - has since been replaced by `technitium` in the manifest, #2911; the - warning's mechanism is per-relative-mount and unchanged by the swap, but - the stack name in this list is the pre-#2911 one.) **Decision: live with it — this repo has no fix + repo-side filter either. Fires for exactly the 10 stacks with at least one + relative-source bind mount under this identity mount (`ghidra`, + `honeypot-cowrie`, `honeypot-dionaea`, `honeypot-elk`, `honeypot-init`, + `honeypot-keycloak`, `honeypot-payload-analysis`, `honeypot-tanner`, + `honeypot-utilities`, `technitium`; the 10 is re-derived 2026-09-27 from + each of the manifest's 39 entries' `volumes:` sources) — every relative + mount on the 7 of the 9 that existed at the 2026-09-03 check and had + running containers was confirmed to resolve to its real, existing host + path, zero missing-on-host mounts. (`pihole` and `honeypot-keycloak` were + the two without running containers, so their mounts were reasoned from the + same identity-mount arithmetic rather than observed — #2853, #2764. + `pihole` has since been replaced by `technitium` in the manifest, #2911 — + hence the post-#2911 name above; the warning's mechanism is + per-relative-mount and unchanged by the swap. The tenth stack, `ghidra`, + is not new: its `./revdeck-proxyfix/*` mounts date from #1173, so the + 2026-09-03 live pass simply never reached it, and the "9" it recorded was + a nine-name subset of the real set rather than a different count.) + **Decision: live with it — this repo has no fix available, and the warning is confirmed harmless.** Re-verified live 2026-09-03: still firing at the same 9 stacks, ~1 warning per relative mount per sync. Not reported upstream (`getarcaneapp/arcane`) as of this @@ -247,7 +273,7 @@ alone. Re-verify against whatever Arcane version is pinned in (`dashboard/vendor/`, exactly 650 tracked files at the #1502 move, re-counted from git history), which exceeded the default back then. That tree is gone with the Go tier itself (#1659) and the synced directory is - down to 268 tracked files (re-counted 2026-08-27, + down to 279 tracked files (re-counted 2026-09-27, `git ls-files arcane/home/honeypot-dashboard`) — under the default again, so the bump is dormant headroom, kept because raising it is an Arcane-side change rather than a repo one. The failure mode itself @@ -262,9 +288,11 @@ alone. Re-verify against whatever Arcane version is pinned in `ghosts`'s sync (`sandbox/ghosts/compose.yml`) always failed — either a ~40s-then-500 with no detail, or (once `maxSyncTotalSize` alone was raised) a fast `file count limit exceeded` — because - `sandbox/ghosts/vendor/ghosts-src/` (963 tracked files, ~132 MB) sits in - the same directory tree as the compose file, so the sync's whole-directory - walk of `sandbox/ghosts/` (989 files, 135,789,139 bytes in total) blows + `sandbox/ghosts/vendor/ghosts-src/` (935 tracked files, 135,601,433 bytes) + sits in the same directory tree as the compose file, so the sync's + whole-directory walk of `sandbox/ghosts/` (962 files, 135,748,394 bytes — + both re-counted 2026-09-27; the 2026-08-31 live walk quoted below measured + 989/135,789,139) blows past *both* defaults (`maxSyncFiles: 500`, `maxSyncTotalSize: 50MB`), not just the one either failure message names. Fixed the same way as the `honeypot-dashboard` case above — both limits raised on the sync record @@ -347,8 +375,11 @@ alone. Re-verify against whatever Arcane version is pinned in ## Local environment overrides Compose's own `.env`-in-project-directory interpolation already covers -every `${VAR}` reference in these 37 stacks — none of them use `env_file:`, -and none needed it added. Arcane's effective environment merge +every `${VAR}` reference in these 39 stacks — none of them *need* `env_file:` +and only one declares it: `analysis/ghidra/docker-compose.ghidra.yml`'s +`revdeck`-profile service carries `env_file: [{path: .env, required: false}]` +(#110), which duplicates the interpolation it sits beside rather than +carrying anything. None needed it added. Arcane's effective environment merge (`project.env` + `.env.git` → `.env`) feeds that same mechanism transparently, so a local override set through Arcane's own UI for a synced project works exactly like editing `.env` by hand always did; no stack-file @@ -407,12 +438,28 @@ Two things worth knowing when setting an override: ## Promotion workflow and change control Decided in #1507: **release/tag promotion**, with `autoSync` enabled for -exactly the three stacks where a sync is the whole deploy. As of -2026-08-27 that policy has **never been put into effect** — every sync -still tracks `main` and nothing auto-deploys (#2549 re-derived the live -state). What actually runs is the manual model below; #2577 holds the -one-time activation steps and the dangling-sync cleanup if that ever -changes. +exactly the three stacks where a sync is the whole deploy. **The manifest +half of that policy is in effect: every one of the manifest's 39 entries +declares `branch: "production"`, and #1943 (2026-08-25) is the commit that +changed it from `main` — re-derived 2026-09-27, and unchanged since except +for the `pihole`→`technitium` and `unsloth` entries that inherited it.** +The live-store half is not. As of the last reads (2026-08-27 #2549, 2026-09-03) +every live sync still tracked `main` and nothing auto-deployed, so the +manifest and the host disagree about what a sync follows. #2577 holds the +one-time activation steps and the dangling-sync cleanup. + +**The pointer the manifest names does not exist yet.** `git ls-remote +--heads origin` on 2026-09-27 returned `main` and assorted agent/dependabot/ +design branches and no `refs/heads/production`; `scripts/promote-release.sh +--list` reports it as "branch does not exist yet". A sync pointed at +`production` therefore fails the same way a tag does — Arcane prefixes +`refs/heads/` onto whatever it is given (see "Arcane cannot track a tag" +below) — so this is the one place where the manifest is currently ahead of +reality in a way that bites: the documented bulk-import path, +`POST /environments/0/gitops-syncs/import`, would create 39 syncs that all +fail on their first run, on a branch that nothing creates until someone runs +`scripts/promote-release.sh `. What actually runs today is the manual +model below. ### Two facts (still true today) @@ -457,27 +504,36 @@ treat every sync as a per-project restart and bound the blast radius accordingly: scope each run to one stack at a time and verify the result — rather than firing a fleet-wide sync pass and racing the timeout. (A purpose-built single-project script is a natural follow-up; ship it separately so the - doc's claim of 'every sync is a restart' is reviewable on its own.)**34 of the 37 stacks build an image.** Only `honeypot-elk`, -`honeypot-keycloak` and `pihole` pull — re-derived 2026-08-27 from the 37 -manifest compose paths (34 carry `build:`; the same three pullers the -#1502-era text named, which said "35 of the 38" before #2381 retired -wordpot and f139fe24 retired the Go ip-enrichment-worker). For any + doc's claim of 'every sync is a restart' is reviewable on its own.)**34 of the 39 stacks build an image.** Five pull — +re-derived 2026-09-27 from the 39 manifest compose paths (34 carry +`build:`): `honeypot-elk`, `honeypot-keycloak`, `llm-worker`, `technitium` +and `unsloth`. `llm-worker` is in that set only because the manifest +deploys `llm-worker/docker-compose.captured-data-deploy.yml`, an overlay +with no `build:` of its own — `llm-worker/docker-compose.yml` still has +one. The #1502-era text named "35 of the 38" and the same three pullers, +before #2381 retired wordpot and f139fe24 retired the Go +ip-enrichment-worker. For any building stack, `autoSync: true` would mean every merge produces a deployment that looks successful and changes nothing — worse than a manual process, because it is unattended. ### What actually runs -- **Every sync tracks `branch: "main"`** — all 38 live `gitops_syncs` - rows (the 37 manifest stacks plus #2577's dangling `honeypot-wordpot` - orphan). The rows predate #1507's decision and nothing re-pointed - them; no `production` pointer exists. +- **Every sync tracks `branch: "main"`** on the live host, while the + manifest has said `branch: "production"` for all 39 entries since + #1943 — so the rows predate #1507's decision and nothing re-pointed + them. Every live `gitops_syncs` row is one of 39 manifest stacks plus + #2577's dangling `honeypot-wordpot` orphan, so 40; the 2026-08-27 read + saw 38 of them, before #2911's `technitium` and #3092's `unsloth`. The + gap between the two is the thing to resolve before the next import, and + the branch to resolve it against does not exist yet (see "Promotion + workflow" above). - **`autoSync` is 0 everywhere, including the three the manifest flags.** - `honeypot-elk`, `honeypot-keycloak` and `pihole` carry `autoSync: true` - in `arcane/manifests/home-production.json`, but the live store has - `auto_sync = 0` on all 38 rows — the elk/keycloak/pihole auto-follow - policy is silently inert: a promotion, or any push, will not deploy - them. + `honeypot-elk`, `honeypot-keycloak` and `technitium` carry + `autoSync: true` in `arcane/manifests/home-production.json`, but the live + store has `auto_sync = 0` on every row — the elk/keycloak/technitium + auto-follow policy is silently inert: a promotion, or any push, will not + deploy them. - **Every deploy is manual**: sync → build → redeploy per stack. The order matters and is not arbitrary: `honeypot-dashboard` must sync before `honeypot-dashboard-backend` builds, because the Rust source @@ -485,8 +541,8 @@ manual process, because it is unattended. - **Promotion is CI-only.** `scripts/promote-release.sh v0.1.0` exists and still refuses a ref that is not a tag, and a tag that is not an ancestor of `main` — so if a pointer ever exists, what reaches it has - always been through CI. Today it moves nothing, because there is no - pointer to move. + always been through CI. As of 2026-09-27 it still moves nothing: there + is no `production` branch on `origin` to move. ### The #1507 design, for the record @@ -501,8 +557,14 @@ manual process, because it is unattended. single-sync PATCH work. #2577 closed (PR #2704) having done only the wordpot-orphan half of its own -scope — the #1507 activation itself was never done, and #2858 (below) -decided explicitly not to do it in this round either. +scope, and #2858 (below) decided explicitly not to do the activation in this +round either. Two halves of "the #1507 activation" are worth keeping +distinct: the manifest half **was** done, in #1943 on 2026-08-25, which is +what put `branch: "production"` and the three `autoSync: true` flags into +`arcane/manifests/home-production.json`; the live-store half — re-pointing +the rows already in Arcane at that branch and setting the +`pullImageAfterSync`/`redeployAfterSync` fields the manifest cannot carry — +is what is still outstanding. ## autoSync decision (#2858) @@ -526,7 +588,7 @@ investigated "a merged PR never reached the host" issues. **Decision: leave `autoSync: false` fleet-wide. Do not activate #1507's three-puller policy in this round either.** Reasoning: -- Turning `autoSync` on for all 37 projects means every merge to `main` +- Turning `autoSync` on for all 39 projects means every merge to `main` redeploys the fleet unattended — a real increase in blast radius, and exactly the kind of change that should not happen as a side effect of fixing a visibility gap. (This round's own brief calls this out @@ -560,26 +622,31 @@ three-puller policy in this round either.** Reasoning: ops-triage pass. `autoSync` therefore remains `false` fleet-wide for now. - #2854's one-shot-abort hazard (a `restart: no` job's clean `exit(0)` making Arcane report `failed` on a deploy that actually completed) is - fully scoped to `honeypot-init`. It is **not** the only file in the repo + scoped to `honeypot-init` **and, since #3128 (2026-09-08), `honeypot-elk`.** + It is **not** the only file in the repo with that shape, and a grep alone does not establish the claim — YAML writes it three ways, so the check has to be `git grep -nE "restart:[[:space:]]*[\"']?no[\"']?"`, which on `origin/main` - returns fourteen hits — twelve real declarations across seven files (all - seven are in the table below), plus two prose matches, in this document and - in `scripts/compose-drift-watch.py`'s header. What scopes the hazard is that - only one of the seven is a service Arcane actually starts: + returns twenty hits — thirteen real declarations across eight files (all + eight are in the table below), plus seven prose matches: four in this + document, two in `scripts/arcane-sync-drift-report.py`, one in + `scripts/compose-drift-watch.py`'s header. (This bullet's own totals were + fourteen / twelve / seven / two / one at the 2026-09-04 round-6 pass.) + What scopes the hazard is that + only two of the eight are services Arcane actually starts: | hit | why it cannot trip the hazard | |---|---| - | `arcane/home/honeypot-init/compose.yml` (6×) | **this is the exposed one** | + | `arcane/home/honeypot-init/compose.yml` (6×) | **this is an exposed one** | + | `arcane/home/honeypot-elk/compose.yml:540` | **the other exposed one** — `arkime-pcap-init`, a `chown`/`chmod` one-shot added by #3128, not profile-gated, in a manifest entry Arcane syncs and starts like any other | | `sandbox/ghosts/compose.yml:162` | `ghosts` *is* an Arcane project, but the service is `ghosts-client-test`, gated behind `profiles: ["test"]`, so a default `up` never creates it | - | `llm-worker/docker-compose.{production-session-canary,synthetic-canary}.yml` | canary overlays; the manifest deploys `llm-worker/docker-compose.yml`, which has no `restart: no` | + | `llm-worker/docker-compose.{production-session-canary,synthetic-canary}.yml` | canary overlays; the manifest deploys `llm-worker/docker-compose.captured-data-deploy.yml`, which has no `restart: no` | | `docker-compose.sandbox.yml:48` | per-detonation sandbox lifecycle, not an Arcane manifest entry | - | `vps/docker-compose.yml:956` | the VPS stack, deployed by a different mechanism entirely | + | `vps/docker-compose.yml:966` | the VPS stack, deployed by a different mechanism entirely | | `sandbox/ghosts/vendor/ghosts-src/Ghosts.Api/docker-compose.yml:85` | vendored upstream source, never deployed | - All six excluded files were checked against - `arcane/manifests/home-production.json`'s 37 entries and their + All seven excluded files were checked against + `arcane/manifests/home-production.json`'s 39 entries and their `dockerComposePath` values, not inferred from the file paths. `honeypot-init` is not one of #1507's three auto-sync candidates, so this decision doesn't change its exposure either way: it keeps deploying @@ -591,6 +658,18 @@ three-puller policy in this round either.** Reasoning: 0755, owned by `github-deploy-runner`, dated 2026-09-03 23:07. #2908 is closed. + **`honeypot-elk` is a #1507 auto-sync candidate, so this one is not + hypothetical the way `honeypot-init` is.** Two consequences follow, and + neither is established here: whether the live `honeypot-elk` record + actually reports `failed` (the #2854 shape is a strong reason to expect + it does, but this has not been read off the live API), and, if it does, + that `scripts/arcane-sync-drift-report.py`'s `KNOWN_STRUCTURAL_FAILURES` + still names only `honeypot-init` — so the report would exit non-zero on a + perfectly healthy fleet, which is the exact "permanently red and therefore + ignored" failure mode its own comment says the exemption exists to + prevent. Worth one `GET /environments/0/gitops-syncs` before #1507's + activation is revisited. + **What makes Arcane gitops-sync drift visible, since nothing did before:** `scripts/arcane-sync-drift-report.py` — read-only, on-demand (not wired into a scheduled workflow: that would need an Arcane API key available to a CI @@ -618,7 +697,20 @@ permanently-red-and-therefore-ignored failure mode `scripts/isolation-audit.sh`'s own tiering comment was written against. Exempted projects are still **printed**, with the reason, under an `EXEMPT` heading; they are not silenced. Any project that reports `failed` without -being named there still fails the run. +being named there still fails the run. `KNOWN_STRUCTURAL_FAILURES` names +exactly one project, `honeypot-init`; if the `honeypot-elk` exposure +described under #2858 above turns out to be real, that table needs a second +entry for the same reason, or a healthy fleet goes permanently red. + +**The baseline is hardcoded to `origin/main`.** `scripts/arcane-sync-drift-report.py` +fetches `origin main` and measures `..origin/main` +unconditionally — it does not read each record's own `branch`. That is +correct while every live sync tracks `main` (it still does, per the reads +above), but it becomes the wrong measure the moment #1507's live-store +activation re-points rows at `production`: a sync correctly caught up to +the promoted release would then read as N commits behind `main` for every +merge since the promotion. Same shape as the `honeypot-elk` gap above — +worth resolving in the same pass. Exit-code behaviour was demonstrated in both directions (2026-09-03) by driving the shipped `main()` with a synthetic record set: a fleet whose only @@ -636,7 +728,7 @@ owed and worth one look once a current key is to hand. The fleet has a second one — `deploy.yml`'s rsync into `/opt/stacks/apiary` (the same inode as `/var/dockge/stacks/apiary`) — that no sync record covers. That channel had gone 18 days without a successful run, which is what made a -fleet reading `lastSyncCommit == main` on all 37 records still run a +fleet reading `lastSyncCommit == main` on all 39 records still run a weeks-old copy of every rsynced script: `diagnostics.yml`'s isolation-invariants step executes the *deployed* `isolation-audit.sh` from that path. Tracked and fixed as **#2908**, now closed — re-derived From bafffadbcd77eaf91968c96bba2304e411ef9d1a Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 14:53:05 +0200 Subject: [PATCH 02/21] docs(cgnat): reconcile deployment paths and the gateway chain - 31 in-tree stacks -> 33 (add honeypot-dashboard-backend #1622 and honeypot-sonicwall-sma #3131, drop the retired ip-enrichment worker); note unsloth as the 33rd directory - pihole -> technitium in the self-contained six and its dedicated step - the investigation chain is a loadBalancer to each oauth2-proxy gateway, not a forwardAuth middleware (0 forwardAuth in dynamic.yml), and only five of the six gateways have a socat-hp-* bridge - Arcane pin is v2.11.1, not v2.8.0; the confirmed limitations have not been re-confirmed against it - mark the SFTP-upload/rebuild-from-terminal paragraph as pre-#1502 --- docs/CGNAT-DEPLOYMENT.md | 83 +++++++++++++++++++++++++--------------- 1 file changed, 53 insertions(+), 30 deletions(-) diff --git a/docs/CGNAT-DEPLOYMENT.md b/docs/CGNAT-DEPLOYMENT.md index 948f079b..ce81c3f8 100644 --- a/docs/CGNAT-DEPLOYMENT.md +++ b/docs/CGNAT-DEPLOYMENT.md @@ -22,13 +22,19 @@ flowchart TD validation), `honeypot-elk`, `honeypot-cowrie`, `honeypot-dionaea`, `honeypot-conpot`, `honeypot-dnp3`, `honeypot-http`, `honeypot-multipot`, `honeypot-payload-analysis`, `honeypot-tanner`, `honeypot-dashboard`, - `honeypot-utilities`, the standalone honeypots (cisco-asa, citrix, rdp, - dicompot, dns-honeypot, endlessh, beelzebub, hellpot, elasticpot, galah, + `honeypot-dashboard-backend`, `honeypot-utilities`, the standalone + honeypots (cisco-asa, citrix, rdp, sonicwall-sma, dicompot, dns-honeypot, + endlessh, beelzebub, hellpot, elasticpot, galah, sentrypeer, mailoney, canarytokens), `honeypot-keycloak`, and the - workers (ip-enrichment, agent-intrusion, attacker-identity, correlator, - payload-inventory) — 31 stacks total, one Arcane-managed directory each + workers (agent-intrusion, attacker-identity, correlator, + payload-inventory) — 32 stacks total, one Arcane-managed directory each under `arcane/home//` (`honeypot-wordpot` sat here until #2381 - retired it). `honeypot-init` still deploys first; every + retired it; `ip-enrichment-worker` was in that worker list until + f139fe24 retired the Go service, `honeypot-dashboard-backend` joined when + #1622 split it out of `honeypot-dashboard`, and `honeypot-sonicwall-sma` + when #3131 added it — re-counted 2026-09-27 against + `arcane/manifests/home-production.json`, which holds 33 in-tree entries: + these 32 plus `unsloth`). `honeypot-init` still deploys first; every sensor stack waits on its completion markers at its own entrypoint rather than a Compose-level dependency, same reasoning as before, just across more projects now. See `docs/STACK-REBUILD.md` for the full current list @@ -39,7 +45,7 @@ flowchart TD - VPS: plain Docker Compose manages `/root/vps/docker-compose.yml`. Unchanged by #1502 — VPS deployment stays outside Arcane entirely, as that issue's own scope decision. -- Each of the 31 migrated stacks' Compose source (build context, git-tracked config, +- Each of the 33 in-tree stacks' Compose source (build context, git-tracked config, `compose.yml` with an explicit top-level `name:` pinned to its live project name) lives self-contained under `arcane/home//` in this repository. Arcane clones the repo and materializes the *entire @@ -52,25 +58,31 @@ flowchart TD these syncs on a from-scratch install, driven by the single source of truth at `arcane/manifests/home-production.json`. Six more home-hosted stacks (`auth-events-worker`, `llm-worker`, `ml-worker`, - `analysis/ghidra`, `sandbox/ghosts`, `pihole`) are Arcane-managed too but + `analysis/ghidra`, `sandbox/ghosts`, `technitium`) are Arcane-managed too but were already self-contained, so they kept their existing repository-root path instead of moving. Three of those six (`auth-events-worker`, `llm-worker`, `ml-worker`) are also imported by `step_arcane_import_stacks` itself now (#1505 — confirmed to have no host-local state beyond `.env`); the other three keep their own dedicated - installer steps for reasons specific to each (`pihole`'s non-`.env` host - state, `analysis/ghidra`'s conditional GPU compose overlay, and + installer steps for reasons specific to each (`technitium`'s non-`.env` host + state — the step `pihole` had until #2911 swapped the two — `analysis/ghidra`'s conditional GPU compose overlay, and `sandbox/ghosts`'s confirmed Arcane build-context limitation, #1506) — see `scripts/install-homeserver.sh`'s own Phase 8 header comment for the - full reasoning behind each. + full reasoning behind each. That filter reaches 35 of the manifest's 39 + entries; `unsloth` is the one the installer reaches neither way (no + `step_unsloth_*`, and not a `honeypot-*` name), by design — see + `docs/ARCANE-GIT-SYNC.md`'s "Manifest import". - The public gateway source is under `vps/`. Arcane is used only on the home server. The VPS uses `docker compose` directly. See `docs/ARCANE-GIT-SYNC.md` for the sync model, cutover procedure, and -confirmed Arcane v2.8.0 platform limitations (a required compose variable +confirmed Arcane platform limitations (a required compose variable in a port-binding position, remote build contexts pinned to a Git tag, the sync file-count limit, and stale project records after a `destroy` call -all have confirmed workarounds documented there). +all have confirmed workarounds documented there). Those were each confirmed +against `v2.8.0`–`v2.9.0`; `docker-compose.arcane.yml` now pins +`manager:v2.11.1`, and none of them has been re-confirmed against that +image — its own section header says to re-verify on upgrade. ## WireGuard addressing @@ -135,12 +147,15 @@ the only internet-facing component. 8. Run `python3 analysis/verify-stack.py` (with `DASHBOARD_SERVICE_TOKEN` from `honeypot-dashboard/.env`) and inspect `/source-health`. -Each stack is a folder under your Arcane stacks dir (default `/opt/stacks/`). -Upload the whole home folder via SFTP — compose **and** the build -sub-folders (`cowrie/`, `multipot/`, `http-honeypot/`, `dashboard/`, …) — -since Arcane's own editor only edits the compose file. After editing Go -source or honeyfs content, rebuild from the `APIARY` stack's Arcane -**terminal**: `docker compose -f compose.yml up -d --build`. +Each stack is a folder under your Arcane stacks dir (`/var/dockge/stacks`; +`/opt/stacks` is a symlink to it, #1185). **Since #1502 nothing is uploaded +by SFTP** — Arcane materializes each stack's whole directory from its Git +sync, and `honeypot-wordpot` aside the source of truth is the repository, not +a hand-copied folder. The SFTP-upload and "edit then rebuild from Arcane's +terminal" instructions this paragraph used to give are part of the pre-#258 +model the callout above already flags; what replaces them is a commit plus a +sync, and a separate `POST /projects/{id}/build` for the stacks that have a +`build:` service — see `docs/ARCANE-GIT-SYNC.md`. ### Boot-safe home networking and VPS log mounts @@ -263,11 +278,12 @@ template is in [`vps/traefik/dynamic.yml`](../vps/traefik/dynamic.yml): `honeypot-http` (`decoy.`) + `honeypot-web` (catch-all) → fake nginx, `honeypot-snare` (`www-portal.` and `snare.`) → SNARE, one native-OIDC route for the dashboard (no gateway, since #1026), one native-OIDC -route for Arcane (no gateway, #1185), and six forward-auth-protected +route for Arcane (no gateway, #1185), and six gateway-fronted investigation routes sitting behind their own Keycloak-backed `oauth2-proxy` gateway: Kibana, TANNER, EveBox, Arkime, Rev·Deck, and the Traefik dashboard -itself. Each has a matching -`socat-hp-*` bridge in [`vps/docker-compose.yml`](../vps/docker-compose.yml). +itself. Five of the six have a matching +`socat-hp-*` bridge in [`vps/docker-compose.yml`](../vps/docker-compose.yml); +the Traefik dashboard's gateway terminates on the VPS, not through one. Traefik is an HTTP(S) reverse proxy — it adds TLS, per-subdomain routing and auth to the web honeypots and dashboards. The other protocols (SSH, SMB, @@ -325,14 +341,22 @@ Cloudflare answers every proxied hostname with **526**. `deploy.yml` never overwrites `traefik/certs/`. On a normal deploy it only checks that `origin.pem` still parses. -### The forward-auth bridge, generically +### The gateway-fronted chain, generically Six investigation UIs (Kibana, TANNER, EveBox, Arkime, Rev·Deck, the Traefik dashboard) reach home through the identical chain — one pattern, -six routers in `vps/traefik/dynamic.yml`, six `socat-hp-*` bridges, each -fronted by its own Keycloak-backed `oauth2-proxy` gateway container, not -six different mechanisms. The honeypot dashboard and Arcane are the two -exceptions — both speak native OIDC directly, no gateway — see the note below. +six routers in `vps/traefik/dynamic.yml`, six `oauth2-proxy` gateway +containers, five `socat-hp-*` bridges, not six different mechanisms. The +gateway *is* the router's upstream rather than a `forwardAuth` middleware: +`honeypot-kibana`'s `service:` is a loadBalancer at `http://oidc-kibana:4180`, +and that container's `OAUTH2_PROXY_UPSTREAMS` is the socat bridge +(`http://socat-hp-kibana:5601`). `grep -c forwardAuth vps/traefik/dynamic.yml` +is 0, so nothing in the config uses the forward-auth middleware form. The +sixth gateway, `oidc-traefik`, has no socat hop at all — its upstream is +Traefik's own dashboard (`http://traefik:8081`) on the VPS, which is why the +bridge count is five and not six. The honeypot dashboard and Arcane are the +two exceptions — both speak native OIDC directly, no gateway — see the note +below. ```mermaid sequenceDiagram @@ -340,7 +364,7 @@ sequenceDiagram actor Op as operator's browser participant CF as Cloudflare
(proxied DNS) participant TR as Traefik
(TLS termination + routing) - participant OA as oauth2-proxy
(forward-auth, one per service) + participant OA as oauth2-proxy
(gateway, one per service) participant KC as Keycloak
(honeypot-keycloak, at home) participant SOC as socat-hp-*
(VPS container) participant WG as WireGuard tunnel @@ -348,14 +372,13 @@ sequenceDiagram Op->>CF: HTTPS request, e.g. kibana. CF->>TR: proxied, real client IP in X-Forwarded-For - TR->>OA: forward-auth check + TR->>OA: routed to the app's own gateway (oidc-kibana:4180) alt no valid session OA-->>Op: redirect to Keycloak login (auth.) Op->>KC: authenticate (password + mandatory TOTP) KC-->>OA: OIDC callback, session established end - OA-->>TR: identity headers - TR->>SOC: request, security-headers applied + OA->>SOC: proxied request, identity headers added SOC->>WG: raw TCP, VPS listen port → 10.8.0.2:home-exposed-port WG->>APP: delivered to the app's own internal port APP-->>Op: response, relayed back through the same chain From 4f51d140bb58b51656bfb9be54ba8b62a1485b17 Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 14:59:12 +0200 Subject: [PATCH 03/21] docs(arcane): flag that v2.11.1 pin is unre-verified --- docs/ARCANE-GIT-SYNC.md | 9 ++++++++- 1 file changed, 8 insertions(+), 1 deletion(-) diff --git a/docs/ARCANE-GIT-SYNC.md b/docs/ARCANE-GIT-SYNC.md index c2d08d0b..71d0f517 100644 --- a/docs/ARCANE-GIT-SYNC.md +++ b/docs/ARCANE-GIT-SYNC.md @@ -193,7 +193,14 @@ a stack means: These are platform behaviors, not something fixable from a compose file alone. Re-verify against whatever Arcane version is pinned in -`docker-compose.arcane.yml` if it's ever upgraded. +`docker-compose.arcane.yml` if it's ever upgraded. **That upgrade has +happened and the re-verification has not:** every item below was confirmed +against `v2.8.0`–`v2.9.0` (including the no-`profiles:`-support finding, +which cites the upstream issue as still open at `v2.8.1`), while +`docker-compose.arcane.yml` now pins `ghcr.io/getarcaneapp/manager:v2.11.1` +by digest. Nothing here is claimed to be false of `v2.11.1` — it is claimed +to be unconfirmed against it, and three minor versions of upstream is +exactly the gap that re-verification is for. - **`"project directory is not inside a mounted directory"` is a false warning against every stack with a relative bind-mount source, under From 666fbbf342f53259ccbd799166e684f31463bce1 Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 14:59:42 +0200 Subject: [PATCH 04/21] docs(security): document the leak gate SECURITY.md relied on SECURITY.md names 'leak a real secret' as a reportable class but never says the repository enforces it in CI, and never warns that the gate scans untracked files. Both facts were load-bearing on this run: the checker carries the home-server address as a named forbidden literal, and RESUME.md quotes that address verbatim, so the gate fails on a clean tracked tree purely because of dispatch artifacts sitting in the checkout. Adds the gate's actual pattern set, its three fail-closed exemption sets with their real sizes (3 ALLOWED_DOTENV paths, 2 ALLOWED_LITERAL_FIXTURE_FILES paths), the four deployment-specific literals it assembles from fragments, and the untracked-file scan. No existing prose changed. --- SECURITY.md | 25 +++++++++++++++++++++++++ 1 file changed, 25 insertions(+) diff --git a/SECURITY.md b/SECURITY.md index b15c3808..b9fc60a8 100644 --- a/SECURITY.md +++ b/SECURITY.md @@ -13,3 +13,28 @@ live malware, private keys, production `.env` files, packet captures containing private traffic, or unredacted logs to a public issue. Supported security fixes target the current `main` branch. + +## The leak gate, and what it already blocks + +"Leak a real secret" is enforced in CI, not by review alone. +`scripts/check-public-leaks.py` fails the build on private keys, GitHub/AWS/Slack +tokens, literal `PASSWORD=`/`SECRET=`/`TOKEN=`/`API_KEY=` assignments, credentials +embedded in a URL, any file named `.env`, and private/runtime binaries (`.pcap`, +`.qcow2`, `.key`, `.pem`, …). It also carries four deployment-specific literals — +one public domain, the VPS address, the home-server address and a known default +password — assembled from fragments at runtime so the checker does not harbor the +values it bans, and a `Host()` rule over `vps/traefik/*.yml` that requires a +reserved `.example`/`.test`/`.invalid`/`.localhost` name, so an installer smoke +test cannot resolve live DNS. + +Exemptions are explicit and fail-closed, in three sets at the top of the script: +`ALLOWED_DOTENV` (3 paths), `ALLOWED_LITERAL_FIXTURE_FILES` (2 paths), and the +`change-me` / `DECOY_ONLY` / `${…}` / `$(…)` / `` forms inside the +credential-assignment pattern. A file that is exempt from the literal scan is +still scanned for every other pattern. + +The gate reads `git ls-files -co --exclude-standard`, so **untracked files are +scanned too**: a scratch note left in your own checkout fails the check exactly +like a committed one. Gitignore it or delete it — the address and domain values +this repository is deployed against are named literals in the checker, and a +dispatch brief or run log that quotes one will trip it. From 4f9ce99e27b71ab7a98c88e2d5dc004527d6f988 Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:05:19 +0200 Subject: [PATCH 05/21] docs(homeserver): re-measure the live disk layout; it is not the documented one Every row of HOMESERVER-DISK-LAYOUT.md's 'as installed' table was stale. Re-measured read-only over ssh 2026-09-27 (lsblk/lvs/findmnt/df/du): - OS is Rocky Linux 10.2, not Ubuntu/subiquity. The doc's curtin autoinstall notes described an install that has been replaced. - boot disk is a different, 4x-larger NVMe (PC401 SK hynix 1TB, 953.9G) and is now LVM-backed: rl-root 70G, rl-swap 32G, rl-home 849.3G. The doc recorded 'no LVM' as a design decision. - /var is on an sdb1 partition of an 8.7T LUN, not a whole disk. The doc left sdb's size unstated in the table and then called it 'the 1.7T sdb disk' in prose -- an internal contradiction, now moot. - sda is a USB-attached Samsung PSSD T7 at /mnt/usb-recovery, so the /mnt-2 bulk-storage role is gone; no sr0 optical device enumerates. - swap is a 32G LVM LV (14.6G in use at measurement), not an 8G /swap.img swapfile; /swap.img does not exist. - /var/lib/docker is 2.9T and /var/dockge 350G, against 103G/229G and 332G combined in the doc. 45 stack dirs, not 23. The autoinstall config and its manual-partitioning walkthrough are kept as the record of the former Ubuntu layout and explicitly marked as no longer a rebuild target for this host. BACKUP-ESSENTIALS.md: three corrections, all verified against scripts/backup-essentials.sh and the live host. - 40 stack .env files as of 2026-09-27, not 41, and phrased per-stack so the number is read as a measurement rather than a constant. - The volumes table listed short names (arcane-data, evebox-config, ...). Four of the five real volumes carry an Arcane project prefix, and backup-essentials.sh writes each archive as .tar.gz, so following the old table would 'restore' into volumes no stack is mounted against. Table and restore step 5 now carry the real names. - Samsung PSSD T7 is the model lsblk/udevadm report; and the dead Keycloak restic config is now doubly dead, since /mnt-2 itself has been decommissioned. HOST-TUNING.md needed no change: all five tunings, half-of-RAM capped 8G zram, priority 100, --replace-swap, the per-class schedulers, the cups/bluetooth/ModemManager service list and the '20 GB card' claim all match scripts/tune-rocky10.sh, and the host's RTX 4000 Ada (20475 MiB) confirms the card size. --- docs/BACKUP-ESSENTIALS.md | 17 ++++-- docs/HOMESERVER-DISK-LAYOUT.md | 104 +++++++++++++++++++++------------ 2 files changed, 79 insertions(+), 42 deletions(-) diff --git a/docs/BACKUP-ESSENTIALS.md b/docs/BACKUP-ESSENTIALS.md index 9246c8bb..65262266 100644 --- a/docs/BACKUP-ESSENTIALS.md +++ b/docs/BACKUP-ESSENTIALS.md @@ -16,13 +16,13 @@ for restoring onto a replacement host see | | | |---|---| -| `homeserver/env/*.env` | all 41 Arcane/Dockge stack `.env` files | +| `homeserver/env/*.env` | one file per Arcane/Dockge stack (40 under `/var/dockge/stacks/` as of 2026-09-27) | | `homeserver/secrets/` | secret files kept beside a stack rather than in its `.env` | | `homeserver/wireguard/` | `wg0.conf` including the private key | | `homeserver/installer/` | `install-homeserver.conf` — the installer's answers file, which exists only on the root filesystem a reinstall wipes | | `homeserver/technitium/` | hand-maintained Technitium DNS config | | `homeserver/keycloak/keycloak.sql.gz` | `pg_dump` of the identity DB — realm, clients, client secrets, users | -| `homeserver/volumes/` | `dashboard-state`, `arcane-data`, `evebox-config`, `canarytokens-redis-data`, `es-importer-state` | +| `homeserver/volumes/` | `dashboard-state`, `honeypot-arcane_arcane-data`, `honeypot-elk_evebox-config`, `honeypot-canarytokens_canarytokens-redis-data`, `honeypot-dashboard_es-importer-state` — the Arcane-prefixed names are the real volume names | | `vps/env/vps.env`, `vps/secrets/`, `vps/traefik/`, `vps/wireguard/` | the VPS's entire config surface, including the Traefik origin certificates | | `*/manifest/` | host reference notes — disks, volumes, containers, WireGuard, nftables | | `repo/docs/`, `repo/scripts/`, `repo/analysis/` | this repository's runbooks and operational scripts | @@ -69,7 +69,7 @@ Three locations, all written by the workstation, which is the backup host: |---|---|---|---| | 1 | `/run/media/xore//apiary-backups` | ext4 (Crucial X8 USB) | udisks auto-mount — only present while plugged in | | 2 | `~/apiary-backups` | XFS (internal) | always available | -| 3 | `homeserver:/mnt/usb-recovery/apiary-backups` | ext4 (Samsung T7, label `APIARY-BACKUP`) | mounted from fstab by UUID with `nofail` | +| 3 | `homeserver:/mnt/usb-recovery/apiary-backups` | ext4 (Samsung PSSD T7, label `APIARY-BACKUP`) | mounted from fstab by UUID with `nofail` | Location 3 was a Ventoy stick formatted exfat until 2026-08-23, mounted read-only and absent from `/etc/fstab` — so every write to it failed and it @@ -215,7 +215,12 @@ repository — `install-homeserver.conf.example` carries only placeholders. `vps/secrets/oidc/`. 5. **Volumes.** For each `homeserver/volumes/.tar.gz`, with the stack stopped, create the volume and unpack into it through a networkless - container: + container. `` is the **full real volume name** — the archive is + written as `$volume.tar.gz` by `backup-essentials.sh`, so four of the + five carry their Arcane project prefix + (`honeypot-arcane_arcane-data.tar.gz`, and so on). Creating a + short-named `arcane-data` volume instead would restore into a volume + no stack is mounted against. ```bash docker volume create docker run --rm --network none -v :/dst -v "$PWD/homeserver/volumes:/src:ro" \ @@ -297,7 +302,9 @@ gone, for two reasons that happen to point the same way: Also found and worth knowing: `honeypot-keycloak/.env` carries a full set of `RESTIC_*` variables pointing at `/mnt-2/apiary-keycloak`, but that repository -directory does not exist, its password file (`secrets/restic-password`) does +directory does not exist — and as of 2026-09-27 neither does `/mnt-2` itself, +which has been decommissioned, so the path cannot start working by accident. +Its password file (`secrets/restic-password`) does not exist, `restic` is not installed on the homeserver and no unit references it. It is dead configuration — no Keycloak restic backup has ever run. The `keycloak.sql.gz` dump in both scripts here covers that gap. diff --git a/docs/HOMESERVER-DISK-LAYOUT.md b/docs/HOMESERVER-DISK-LAYOUT.md index 2c3673d2..6b534940 100644 --- a/docs/HOMESERVER-DISK-LAYOUT.md +++ b/docs/HOMESERVER-DISK-LAYOUT.md @@ -3,13 +3,22 @@ This documents the physical disk layout of the honeypot homeserver (`supermicro`) as it actually exists today, and a generated Ubuntu **autoinstall** config (the Ubuntu/subiquity equivalent of Windows' -`autounattend.xml`) to reproduce that layout on a reinstall or a second -build server. Captured 2026-08-04 as part of the #518 smoke-test research. - -Ubuntu Server's installer (`subiquity`) is driven by `curtin` under the -hood — the fstab comments on this box literally say "was on /dev/sdX -during curtin installation", confirming this machine was already installed -this way rather than by hand. +`autounattend.xml`) that reproduces the layout the box had at the time of +the #518 smoke-test research. + +> **The autoinstall config below no longer describes this host.** The +> physical table and the provisioning steps were re-measured read-only on +> 2026-09-27. `supermicro` has since been reinstalled as **Rocky Linux +> 10.2** and no longer runs the Ubuntu/`curtin` layout: it uses LVM (the +> original notes recorded "no LVM"), `/var` sits on a *partition* of the +> RAID LUN rather than the whole disk, and swap is a 32G LVM logical +> volume rather than a swapfile. The original capture was 2026-08-04, when +> the fstab comments did literally say "was on /dev/sdX during curtin +> installation" — that evidence was sound for the Ubuntu install, which +> has since been replaced. Keep the autoinstall file as the record of the +> Ubuntu layout; do not use it as a rebuild target for the current host. +> See `docs/HOST-TUNING.md` for the tuning that *does* apply to the Rocky +> install. ## Why this layout, not one big disk @@ -22,24 +31,36 @@ reinstall of the OS disk alone doesn't touch captured evidence. ## Physical layout (as installed) +Re-measured read-only on 2026-09-27 via `lsblk`/`lvs`/`findmnt`/`df`. + | Device | Model | Size | Partition table | Filesystem | Mount | Role | |---|---|---|---|---|---|---| -| `nvme0n1` | Samsung MZVLW256HEHP | 238.5G | GPT | vfat (p1) / ext4 (p2) | `/boot/efi`, `/` | OS + EFI, boot disk | -| `sdb` | AVAGO MR9440-8i (RAID LUN) | — | whole-disk (no partition table) | xfs | `/var` | Docker root, Arcane-managed stacks, container state (`/var/lib/docker`, `/var/dockge`) — now also `benchmarks/`, `training/`, `hf-cache/`, `buildx-cache/`, `ci-registry-mirror/`, the former `/mnt-1` workload | -| `sda` | Intel SSDSC2KB480G8L | 447.1G | GPT, 1 partition | xfs | `/mnt-2` | Reserved bulk storage (currently empty) | -| `sr0` | ATAPI optical | — | — | — | — | Unused | - -**`/mnt-1` is decommissioned.** Its RAID VD (formerly `sdc`) suffered a -two-drive fault on 2026-09-09 (#3158) and no longer enumerates as a block -device at all; the mount was unwired (#3159, PR #3159) and everything that -lived under it moved to `/var`. `/mnt-1` itself survives on the host only as -a directory of compatibility symlinks into `/var` (`benchmarks`, `training`, -`hf-cache`, `buildx-cache`, `ci-registry-mirror`) so any script still hard- -coding the old path keeps resolving — new code should target `/var/*` -directly. See #3158/#3159 for the incident and decommission detail; this -table's `sdb` size is left unstated above rather than guessed, since the -volume backing `/var` changed as part of that recovery and hasn't been -re-measured for this doc. +| `nvme0n1` | PC401 NVMe SK hynix 1TB | 953.9G | GPT, 3 partitions | vfat (p1, 600M) / xfs (p2, 2G) / LVM2_member (p3, 951.3G) | `/boot/efi`, `/boot`, — | OS boot disk | +| └ `rl-root` | (LVM on `nvme0n1p3`) | 70G | — | xfs | `/` | OS root | +| └ `rl-swap` | (LVM on `nvme0n1p3`) | 32G | — | swap | `[SWAP]` | Swap | +| └ `rl-home` | (LVM on `nvme0n1p3`) | 849.3G | — | xfs | `/home` | Home | +| `sdb` | AVAGO MR9440-8i (RAID LUN) | 8.7T | GPT, 1 partition (`sdb1`, whole remaining size) | xfs | `/var` | Docker root, Arcane-managed stacks, container state (`/var/lib/docker`, `/var/dockge`) — now also `benchmarks/`, `training/`, `hf-cache/`, `buildx-cache/`, `ci-registry-mirror/`, the former `/mnt-1` workload | +| `sda` | PSSD T7 (**USB-attached**) | 465.8G | GPT, 1 partition (`sda1`) | ext4 | `/mnt/usb-recovery` | USB recovery disk, not local bulk storage | + +Three of the four rows changed since the 2026-08-04 capture, and none of +the change is cosmetic. The OS disk is a different, 4x-larger NVMe; the +former `sda` bulk-storage disk (`/mnt-2`) is now a USB-attached portable +SSD mounted at `/mnt/usb-recovery`; and the boot disk is now LVM-backed +with a separate `/home`. The old `sr0` ATAPI optical drive is no longer +enumerated at all. + +**`/mnt-1` and `/mnt-2` are both decommissioned.** `/mnt-1`'s RAID VD +(formerly `sdc`) suffered a two-drive fault on 2026-09-09 (#3158) and no +longer enumerates as a block device at all; the mount was unwired (#3159, +PR #3159) and everything that lived under it moved to `/var`. `/mnt-1` +itself survives on the host only as a directory of compatibility +symlinks into `/var` (`benchmarks`, `training`, `hf-cache`, +`buildx-cache`, `ci-registry-mirror`) so any script still hard-coding the +old path keeps resolving — new code should target `/var/*` directly. See +#3158/#3159 for the incident and decommission detail. `/mnt-2` is gone for +a different reason: its disk is the USB `PSSD T7` above, remounted at +`/mnt/usb-recovery`, so it is no longer local bulk storage and must not be +relied on for a rebuild. `sdb` sits behind an AVAGO/LSI MR9440-8i hardware RAID controller and appears to the OS as a SCSI LUN, not a raw disk — the controller's own @@ -49,17 +70,21 @@ the controller's own tooling (`storcli`/`perccli` or vendor equivalent) if the RAID config itself needs to be reproducible, not just the OS partitioning on top of it. -`/var` on its own disk is the key decision: `/var/lib/docker` is 103G and -`/var/dockge` (bind-mounted stack data for all 23 Arcane-managed stacks, including +`/var` on its own disk is the key decision, and it has only become more +load-bearing: `/var/lib/docker` is **2.9T** and `/var/dockge` (stack data +for the 45 directories under `/var/dockge/stacks/`, including Elasticsearch indices, Cowrie logs, payload captures, sandbox disks) is -229G — 332G combined, well past what the 238G OS disk could hold even -before accounting for the OS itself. Putting `/var` on the 1.7T `sdb` -disk instead of growing the root filesystem was the right call and should -be preserved on any rebuild. - -Swap is an **8G swapfile** at `/swap.img` on the root filesystem, not a -dedicated partition — simpler to resize than a swap partition and fine at -this scale (91G RAM, swap is a safety margin not a working set). +**350G**. `/var` is 70% full (6.1T of 8.8T) with 2.7T free. The manifest +still declares 39 sync entries and 34 of those directories are +manifest-managed; `rex86-eval` is present on disk but **not** in the +manifest. Putting `/var` on the RAID LUN instead of growing the root +filesystem remains the right call and should be preserved on any rebuild. + +Swap is a **32G LVM logical volume** (`rl-swap`) in the `rl` volume group, +not a dedicated partition and not a swapfile — the 8G `/swap.img` +swapfile described in the 2026-08-04 capture no longer exists. 92G of RAM +means swap is a safety margin rather than a working set, though it was +under real pressure at measurement time (14.6G in use, priority -2). ## Reproducing it: `autoinstall/homeserver-user-data.yaml` @@ -102,11 +127,16 @@ with `homeserver-user-data.yaml` renamed to `user-data` alongside an empty - SSH (key-only, no password auth) and the `xfsprogs`/`nvme-cli` packages the manual partitioning step below needs. -**What has to be done by hand, at the storage screen, using the physical -layout table above as the target:** 3-disk layout (NVMe boot/OS: GPT, -EFI + ext4 root; `/var`: whole-disk xfs, no partition table; `/mnt-2`: +**What has to be done by hand, at the storage screen.** For reproducing +the **former Ubuntu layout** (the one this template was written against, +and the one the 2026-08-04 capture recorded): 3-disk layout (NVMe boot/OS: +GPT, EFI + ext4 root; `/var`: whole-disk xfs, no partition table; `/mnt-2`: GPT + single xfs partition), no LVM, 8G swapfile instead of a swap -partition. `/mnt-1` is no longer part of the target layout (decommissioned, +partition. That is **not** the live layout any more — the box now uses LVM +with a separate `/home`, `/var` on a partition, and a 32G swap LV, and its +disks have all been replaced (see the table above). Do not use this +paragraph as a partition plan for the current host; it is a record of what +the autoinstall flow produced. `/mnt-1` is no longer part of the target layout (decommissioned, see above) — do not recreate it on a rebuild. The template does **not** attempt to reproduce the AVAGO RAID controller's own LUN configuration either — that has to happen before the OS installer ever sees a block From 8629bc09f368b16fc468b70d2b18bd762589ff08 Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:05:30 +0200 Subject: [PATCH 06/21] docs(ci-cd): reconcile workflow topology, counts, and VPS sync excludes --- docs/CI-CD.md | 107 ++++++++++++++++++++++++++++++++++---------------- 1 file changed, 74 insertions(+), 33 deletions(-) diff --git a/docs/CI-CD.md b/docs/CI-CD.md index 79e0442e..d19a8df1 100644 --- a/docs/CI-CD.md +++ b/docs/CI-CD.md @@ -47,15 +47,18 @@ flowchart TB containersHome["containers.yml — all image builds"] securityHome["security.yml — all CodeQL languages"] pagesHome["pages.yml artifact build"] + imageScanHome["image-security-scan.yml"] end prPush --> quality prPush -->|"PR: build only,
never published"| containerBuild + prPush --> codeql + prPush --> pages mainPush --> quality mainPush --> containerBuild mainPush --> codeql mainPush --> pages - mainPush -.->|"every workflow's compute jobs —
only after passing the ci-router
trust gate + heartbeat; pull_request needs
repo variable CI_HOMESERVER_PRS,
and forks can never qualify"| ciSelfHosted + mainPush -.->|"those five workflows' ci-target jobs —
only after passing the ci-router
trust gate + heartbeat; pull_request needs
repo variable CI_HOMESERVER_PRS,
and forks can never qualify"| ciSelfHosted ``` **`honeypot-ci` does not see `pull_request` by default, by design.** A @@ -64,9 +67,11 @@ job; a self-hosted runner's job runs as a real process on real home-network infrastructure. A malicious test file in an unreviewed PR (`os.system(...)`, a crafted Go `TestMain`) would execute wherever that runner has access — the same reasoning `production-home`'s own deployment -runner (below) already applies. Every workflow's executor routing (each -caller's own `ci-target` job, which since #2571 always calls the shared -`.github/workflows/ci-router.yml`) trusts +runner (below) already applies. Executor routing is a caller-side job named +`ci-target`, which since #2571 calls the shared +`.github/workflows/ci-router.yml`. Five of the repo's twenty workflows have +one: `quality.yml`, `containers.yml`, `security.yml`, `pages.yml` and +`image-security-scan.yml`. It trusts push-to-main (already reviewed and merged), the `schedule` and `workflow_dispatch` (an operator's own machinery); same-repo pull requests need the repository variable `CI_HOMESERVER_PRS=true`, and fork @@ -316,10 +321,14 @@ stack on the host lives entirely in Arcane's Git-sync machinery — see [ARCANE-GIT-SYNC.md](ARCANE-GIT-SYNC.md) for the full contract (its non-obvious cornerstones: creating a sync *is* an initial deploy, a sync materializes files without redeploying — live `redeploy_after_sync` -defaults to 0, though the manifest schema cannot express it — and every -synced stack runs `autoSync: false`: the #1507 tag-promotion / -`production`-pointer policy was decided but never deployed, so all syncs -track `main` and deploys are manual; ARCANE-GIT-SYNC.md's promotion +defaults to 0, though the manifest schema cannot express it — and deploys +are manual. #1507's tag-promotion / `production`-pointer policy was only +half activated: #1943 (2026-08-25) put `branch: "production"` and three +`autoSync: true` flags into `arcane/manifests/home-production.json`, so all +39 entries now name that branch rather than `main` — but the live store +still reads `auto_sync = 0` on every row, and `origin` has no +`refs/heads/production` for the pointer to name, so nothing follows a +promotion and deploys stay manual; ARCANE-GIT-SYNC.md's promotion section carries the live-state evidence). This workflow deliberately stopped touching those directories entirely: running an rsync/build loop alongside Arcane's own sync would put two @@ -476,10 +485,9 @@ alert/intelligence history in the old one is gone. Everything else that was still monolithic as of the earlier revision of this section (`dionaea`, `payload-dedupe`, `yara-scanner`, and the Tanner -group) has since split out too -- see the `honeypot-dionaea` and -`honeypot-payload-analysis` section below; only the Tanner group remains in -`APIARY`, as part of its own internal `depends_on` chain not yet -worth splitting. +group) has since split out too -- see the `honeypot-dionaea`, +`honeypot-payload-analysis` and `honeypot-tanner` sections below. Nothing +remains in `APIARY`: the root `docker-compose.yml` is `services: {}`. #### Dashboard redeploy (single replica; #266 rolling pair retired, #1659 legacy `dashboard` removed) @@ -630,13 +638,26 @@ open handles into the log directories this script wipes for this target. ### honeypot-elk (#258) `arcane/home/honeypot-elk/compose.yml` bundles the ELK/analysis plane (`elasticsearch`, -`kibana`, `filebeat`, `evebox`, `arkime-capture`, `arkime-viewer`, -`pcap-sync`) into one stack at `/opt/stacks/honeypot-elk` -- the last group +`kibana`, `filebeat`, `evebox`, `pcap-sync`, `arkime-pcap-init`, +`arkime-capture`, `arkime-viewer`, `extracted-file-importer`, `zeek-proxy`) +into one stack at `/opt/stacks/honeypot-elk` -- the last group that was still in the monolithic file. Kept together, not split further: -all seven sit on the shared `honeynet` network and either read from or -write to the one Elasticsearch instance, so splitting them apart would -turn every one of those relationships into a cross-stack shared resource -for services that only ever make sense running together. +they share the one Elasticsearch instance, the `arkime-pcap` volume and the +host's `logs/` bind-mount tree, so splitting them apart would turn every one +of those relationships into a cross-stack shared resource for services that +only ever make sense running together. + +The shared-network story is narrower than it looks, and the count above is +not "ten on `honeynet`". Seven of the ten declare `honeynet` explicitly +(`elasticsearch`, `kibana`, `filebeat`, `evebox`, `arkime-capture`, +`arkime-viewer`, `extracted-file-importer`); `elasticsearch` is additionally +on `llm-data`. `pcap-sync` and `arkime-pcap-init` declare no `networks:` at +all and so ride the project's implicit default network -- `pcap-sync` moves +rotated pcaps through host bind-mounts and a marker file, and +`arkime-pcap-init` is a one-shot `chown` of the `arkime-pcap` volume, so +neither needs the shared network. `zeek-proxy` sets `network_mode: host` +outright and reaches the sensor plane through host-published ports, which is +the point of it. `honeynet` and `llm-data` get the usual explicit shared `name:` treatment. `es-data` does **not**, despite appearances: `honeypot-init`'s @@ -868,8 +889,9 @@ The install also drops two things next to the unit: The leading `+` runs that line as root even though the unit's own `User=` is the unprivileged runner account, so no new sudoers grant was needed -(unlike `compose-project-state.py` above, this runs as part of the unit's -own privileged startup rather than from inside a workflow step). +(unlike `scripts/compose-project-state.py`, the narrow root helper for +`compose-drift-watch.py`, this runs as part of the unit's own privileged +startup rather than from inside a workflow step). **Why it exists.** A root process that writes into a runner's `_work` checkout leaves files the runner user can never delete, and @@ -1072,7 +1094,7 @@ The `Pick cache backend` step therefore chooses per executor: the runner can actually write it, so a rebuild replay (#1609) recreates it rather than leaving a hand-made directory nobody records. If the step has not run on a given box, `Pick cache backend` emits a workflow warning and -falls back to `type=gha` -- a slow build, not eighteen failed matrix rows. +falls back to `type=gha` -- a slow build, not nineteen failed matrix rows. **Bounding it.** `type=local` has *no* eviction: every export leaves unreferenced blobs behind in `blobs/sha256/` forever. @@ -1096,9 +1118,11 @@ gets to them. Deleting on close reclaims that quota immediately. `docker/login-action` targets `ghcr.io` and is gated `if: github.event_name != 'pull_request'`, so on a PR every base-image pull went out anonymous -- and Docker Hub meters anonymous pulls **per source -IP**, at roughly 100 per 6h. With `CI_HOMESERVER_PRS=true` all 18 matrix -rows leave this box through one address, and the tree carries **74 -non-`scratch` Hub `FROM` lines**. One cold run spends most of the budget; +IP**, at roughly 100 per 6h. With `CI_HOMESERVER_PRS=true` all 19 matrix +rows leave this box through one address, and the tree carries **75 +non-`scratch` `FROM` lines** across 56 tracked Dockerfiles, 68 of them +resolving to Docker Hub (the other 7 are `mcr.microsoft.com`). One cold run +spends most of the budget; the run after it fails with `toomanyrequests` on whichever rows happen to ask last. #2771's per-image `type=gha` scopes do not help: that cache holds *our* layers, never the base image, so every run re-resolves every `FROM` @@ -1181,9 +1205,9 @@ flowchart TB checkout["actions/checkout"] key["VPS_SSH_KEY written to a
temp file, mode 0600"] backup[("Snapshot: /root/vps-backups/
pre-deploy-<timestamp>.tar.gz,
10 most recent kept")] - rsync["rsync vps/ -> /root/vps/
over SSH, excluding .env,
traefik/certs/, traefik/dynamic.yml
(VPS-owned, see table below)"] + rsync["rsync vps/ -> /root/vps/
over SSH, excluding .env,
traefik/certs/, traefik/dynamic.yml,
secrets/ (VPS-owned, see table below)"] validate["SSH: docker compose config
validates /root/vps/docker-compose.yml"] - up["SSH: docker compose up -d --build"] + up["SSH: docker compose up -d --build
--remove-orphans (#2813)"] dynGen["Separate step: substitute DOMAIN
into the committed *.honeypot.example
placeholders, validate as YAML,
no leftover placeholders --
all BEFORE touching the VPS"] dynWrite["Copy to a temp path on the VPS,
then write in place with cat --
never copy-then-rename (see below:
Traefik's bind mount tracks the
inode, not the path)"] verify["Verify step: fail the job if certs
or dynamic.yml are missing, empty,
unparseable, or still placeholder"] @@ -1211,10 +1235,14 @@ The VPS job runs on a short-lived GitHub-hosted Ubuntu runner: `/root/vps-backups/pre-deploy-.tar.gz`, keeping the ten most recent archives. 5. `rsync` sends only the repository's `vps/` directory over SSH to - `/root/vps/`, **excluding** `traefik/dynamic.yml` (see below). + `/root/vps/`, **excluding** the four VPS-owned paths in the table below + (`.env`, `traefik/certs/`, `traefik/dynamic.yml`, `secrets/`). 6. A second SSH command runs on the VPS, validates `/root/vps/docker-compose.yml`, and executes - `docker compose up -d --build`. + `docker compose up -d --build --remove-orphans`. The flag matters: a + service removed from `docker-compose.yml` otherwise leaves its container + running forever, which is how #2813 found `socat-hp-wordpot` still up a + week after #2469 retired wordpot and dropped its forwarder rule. 7. A dedicated step generates the deployable `traefik/dynamic.yml` -- substitutes `DOMAIN` for every `*.honeypot.example` placeholder in the committed template -- and validates the result (parses as YAML, no @@ -1229,7 +1257,7 @@ The VPS job runs on a short-lived GitHub-hosted Ubuntu runner: ### Files the VPS owns, not the repository `--delete-delay` removes destination files that no longer exist under the -repository's `vps/` directory, and overwrites the ones that do. Three paths are +repository's `vps/` directory, and overwrites the ones that do. Four paths are therefore excluded from the main `rsync` because the VPS copy is authoritative (or, for `dynamic.yml`, because it needs different handling entirely): @@ -1238,6 +1266,11 @@ therefore excluded from the main `rsync` because the VPS copy is authoritative | `.env` | Secrets and host-specific values. | | `traefik/certs/` | Issued TLS certificates. They do not exist in the repository, so an unexcluded `--delete-delay` deletes them, and the workflow cannot reissue them. | | `traefik/dynamic.yml` | Carries the deployment's real domain. The committed copy is a `*.honeypot.example` placeholder -- Traefik's file provider has no `${VAR}`-style substitution the way docker-compose already gives every other host-specific value in this repo, so this file can't just be templated in place the normal way. Deployed by its own dedicated step instead (step 7 above), which substitutes `DOMAIN` and writes the result separately. | +| `secrets/` | Per-gateway OIDC cookie-secret/client-secret files (`OIDC_SECRETS_DIR=./secrets/oidc` in `vps/.env.example`). Git-ignored, so they are never present in the checkout at all -- and `--delete-delay` reads "absent from the source" as "delete it". That happened once for real and took down every oauth2-proxy gateway at the time, which is why the exclude exists. A client-secret is never regenerated: it has to match what is already registered with Keycloak, so recovering it means `kcadm get clients//client-secret`, and losing it means re-registering the client. | + +The pre-deploy backup step ahead of the bulk `rsync` archives the same four +paths (`.env`, `traefik/certs`, `traefik/dynamic.yml`, `secrets/`) that +actually exist, keeping the last ten under `/root/vps-backups/`. The certificates were lost once, in a single `target: both` run before that exclusion existed: Traefik fell back to self-signed and every router silently @@ -1294,8 +1327,8 @@ scratch twice, reaching the same blocker both times. **No workflow edit is needed.** Every `secrets.VPS_*` / `secrets.DOMAIN` reference already sits inside a job that declares -`environment: production-vps` — `deploy.yml`'s `vps` job (`:207`, environment -at `:210`), `diagnostics.yml`'s `vps` job (`:252`/`:255`), and +`environment: production-vps` — `deploy.yml`'s `vps` job (`:231`, environment +at `:234`), `diagnostics.yml`'s `vps` job (`:292`/`:295`), and `vps-start-blackhole.yml`'s `start-blackhole-profile` job (`:22`/`:24`). The `home` jobs (`deploy.yml:21`, `diagnostics.yml:76`) read none of the five. Environment secrets also shadow repository secrets of the same name, so @@ -1324,7 +1357,7 @@ source rather than the password manager, rotate it deliberately rather than as a side effect of the move. `DOCKERHUB_USERNAME` / `DOCKERHUB_TOKEN` (added 2026-09-01) are read only by -`containers.yml:157-160`, which declares **no** `environment:` at all — so +`containers.yml:157-162`, which declares **no** `environment:` at all — so they have to stay repository-scoped until that workflow gains one, and they are correctly out of this migration's scope rather than merely deferred. They are still the reason the repository-secret set grew from five to seven, which @@ -1411,7 +1444,10 @@ diagnostics workflow itself. Home: GitHub -> outbound-polling self-hosted runner on homeserver -> local rsync /opt/stacks/apiary - -> Arcane compose.yml -> docker compose up + -> compose config --quiet (validation only, no `up`; + the root file is services: {}) + -> Arcane's own Git-sync machinery materializes and + deploys stacks (ARCANE-GIT-SYNC.md) VPS: GitHub-hosted runner -> rsync + SSH over VPS_PORT @@ -1419,6 +1455,11 @@ GitHub-hosted runner -> rsync + SSH over VPS_PORT -> docker compose up on VPS ``` +The home path has not run a deploy in this workflow since #1502 — see +"What deploy.yml actually runs (since #1502)" above for the full list of +what it does instead. A `home` run that succeeds has still changed nothing +on the host. + Selecting `both` creates both jobs from the same workflow run. They share the `honeypot-production` concurrency group, but the home and VPS jobs are otherwise independent: one can fail while the other succeeds. Always inspect From ffcc12926428238fa3873c96f571d56ddcb95414 Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:06:50 +0200 Subject: [PATCH 07/21] docs(analysis): reconcile workbench and LLM-worker docs with the Rust backend Both docs still described the Go dashboard that #1628 deleted. The corrections were verified against backend-service/src/main.rs and the workbench modules, and the manifest. payload-analysis-workbench.md: - All 8 HTTP contract routes were wrong. The live set is /api/v1/workbench/{analyzers,runs,runs/{id},runs/{id}/children/ {analyzer_id}/{action},recipes} (main.rs:456-468), so: the prefix moves from /api/payload-workbench to /api/v1/workbench, the hash is a query parameter rather than a path segment, and cancel/retry collapse into one route dispatched on {action} rather than being two paths. The cancelled child-action row is gone with the merge. - Registry pointer dashboard/workbench_domain.go -> backend-service/src/workbench_domain.rs. There are zero .go files under dashboard/ in this repo. - /payload-workbench -> /payload-workbench/results, whose workbench-builder section is the orchestration surface. - The 'closed Go schema (unknown fields are rejected)' and 64 KiB body cap are not present in the Rust tier (no deny_unknown_fields on workbench_api.rs, no body-limit constant anywhere), so the sentence now says not to rely on them rather than asserting a guarantee the server does not make. - Dropped the rollback note about a local /state/analysis-workbench copy; no such store exists, workbench_es.rs is the only one. - MODEL_STATUS_SOCKET appears nowhere in the backend, compose or .env.example. Marked undetermined rather than deleting the claim -- the adapter is genuinely installed, and deleting would lose the intended contract. gpu-llm-analysis-worker.md: - /api/llm/analysis does not exist; the page reads the generic store route /api/v1/store/llm-analysis (main.rs:438, stores.rs). dashboard/llm_analysis.go -> frontend-next/src/routes/llm-analysis.tsx. - Semantic search was marked 'Deferred' but shipped as /api/v1/llm-search (main.rs:341). Recorded as delivered, with a note that the section is historical scope rather than current state. - 'Managed by Dockge under /opt/stacks' is obsolete: Arcane gitops manages /var/dockge/stacks, and /opt/stacks is now only a compatibility symlink to it (created 2026-09-04). - The file list omitted the entrypoint the manifest actually deploys. arcane/manifests/home-production.json sets the llm-worker sync's dockerComposePath to llm-worker/docker-compose.captured-data-deploy.yml, so a bare docker-compose.yml bring-up does not reproduce the live worker (#2234). Added with that warning. - Host RAM/CPU 91 GiB / 16 CPUs -> 92 GiB / 48, measured on the host. --- docs/gpu-llm-analysis-worker.md | 33 ++++++++++++++++++++++-------- docs/payload-analysis-workbench.md | 29 ++++++++++++++------------ 2 files changed, 40 insertions(+), 22 deletions(-) diff --git a/docs/gpu-llm-analysis-worker.md b/docs/gpu-llm-analysis-worker.md index 7bbd08b3..dedd0379 100644 --- a/docs/gpu-llm-analysis-worker.md +++ b/docs/gpu-llm-analysis-worker.md @@ -98,8 +98,8 @@ one device. Also pinned as the runtime-governance authority in | Driver / CUDA | 580.173.02 / CUDA 13.0 | `nvidia-smi` | | Container GPU passthrough | nvidia-container-toolkit 1.19.1, `nvidia` runtime registered | `docker info \| grep -i runtime` | | End-to-end container test | `docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi -L` lists the GPU | run it | -| Host RAM / CPU | 91 GiB / 16 logical CPUs | `free -h`, `nproc` | -| Stack deployment | Dockge stack at `/opt/stacks/apiary/compose.yml`, containers `hp-*` | `docker ps` | +| Host RAM / CPU | 92 GiB / 48 logical CPUs | `free -h`, `nproc` | +| Stack deployment | Arcane gitops, on disk under `/var/dockge/stacks//`, containers `hp-*` | `docker ps` | | Internal network | `honeynet` (Elasticsearch and all sensors live here) | `docker network ls` | | Elasticsearch | 8.13.4 single-node, `xpack.security.enabled=false`, reachable as `http://elasticsearch:9200` inside `honeynet` | compose file | @@ -269,6 +269,12 @@ design sketch retained for context and must not be copied into production: — bounded U1-only production acceptance with no payload mounts; - [`../llm-worker/docker-compose.captured-data.yml`](../llm-worker/docker-compose.captured-data.yml) — separately authorized #83 network and read-only volume grant; +- [`../llm-worker/docker-compose.captured-data-deploy.yml`](../llm-worker/docker-compose.captured-data-deploy.yml) + — the **deployed** entrypoint: this is `dockerComposePath` for the + `llm-worker` entry in `arcane/manifests/home-production.json`, and it is + what a bare `docker-compose.yml` bring-up is missing (#2234). A local + `docker compose up` against `llm-worker/docker-compose.yml` alone will + not reproduce the live worker. - [`../analysis/ghidra/docker-compose.ghidra.yml`](../analysis/ghidra/docker-compose.ghidra.yml) — the pinned shared Ollama service and narrow `honeypot-llm` network. @@ -303,8 +309,12 @@ docker compose \ config --quiet ``` -The live stack is managed by Dockge under `/opt/stacks`. Deploy only from a -reviewed merged revision; do not maintain a second hand-edited Compose copy. +The live stack is managed by Arcane gitops from +`arcane/manifests/home-production.json`, not by Dockge. (`/opt/stacks` still +exists on the host, but only as a compatibility symlink to +`/var/dockge/stacks`, added 2026-09-04 — nothing is managed there.) +Deploy only from a reviewed merged revision; do not maintain a second +hand-edited Compose copy. Model pulling stays an explicit operator action in the Ghidra/Ollama stack. --- @@ -477,21 +487,26 @@ Retention: ILM 90 days is sufficient — derived data, recreatable from raw. Mirrors the pattern `ml-anomalies` already established ([`ml-worker-plan.md` §8–9](ml-worker-plan.md)): -- **Delivered (#150):** `GET /api/llm/analysis?doc_type=&severity=&since=&limit=` +- **Delivered (#150):** `GET /api/v1/store/llm-analysis` (paged via `offset`/`size`) → documents from `llm-analysis`, newest first, polled on the dashboard's existing 1-minute ES ticker (same transport decision as `ml-anomalies`, no new broker). `/llm-analysis` page: session summaries and payload triage in one filterable table, every row labelled "AI-generated" and showing severity/confidence, with an evidence link back to the - originating session or payload where one exists (`dashboard/llm_analysis.go`). + originating session or payload where one exists + (`frontend-next/src/routes/llm-analysis.tsx`, backed by the generic store + route at `main.rs:438` rather than a route of its own). - **Deferred:** `GET /api/llm/analysis/stream` (SSE via redis channel `llm-analysis-events`) -- optional per this section's original scope ("any SSE/Redis wake-up path remains optional and non-authoritative"); polling has not been shown insufficient yet. -- **Deferred:** semantic search over sessions using `nomic-embed-text` +- **Delivered:** semantic search over sessions using `nomic-embed-text` embeddings stored as a `dense_vector` (384-dim) field on `llm-analysis` - docs, queried with ES kNN search. Still waiting on U1–U3 being stable, - per this section's original scope. + docs, queried with ES kNN search — `GET /api/v1/llm-search` + (`main.rs:341`, `llm_search.rs`, over the `llm-analysis` index's + `doc_type: session` documents). This section originally deferred it + pending U1–U3 stability; it has since shipped, so the list above is not + a statement of current scope. --- diff --git a/docs/payload-analysis-workbench.md b/docs/payload-analysis-workbench.md index c15326a7..684b3301 100644 --- a/docs/payload-analysis-workbench.md +++ b/docs/payload-analysis-workbench.md @@ -1,6 +1,6 @@ # Payload analysis workbench -The dashboard's `/payload-workbench` route selects captured evidence and `/payload-workbench/{sha256}` is the unified orchestration surface for issue #155. It is a separate page rather than an extension of `/ghidra`: recipes and parent runs span deterministic analysis, Ghidra, and two sandbox backends, while `/ghidra/{sha256}`, `/sandbox/{job}` and `/payload-analysis/{sha256}` remain the canonical native result renderers. +The dashboard's `/payload-workbench/results` route is the owner-isolated review surface, and its `workbench-builder` section is the unified orchestration surface for issue #155. It is a separate page rather than an extension of `/ghidra`: recipes and parent runs span deterministic analysis, Ghidra, and two sandbox backends, while `/ghidra/{sha256}`, `/sandbox/{job}` and `/payload-analysis/{sha256}` remain the canonical native result renderers. ## Trust boundary @@ -15,7 +15,7 @@ Run ownership is likewise never taken from client input. The Rust tier derives i ## Analyzer registry Seven analyzer IDs, one server-computed `workbenchAnalyzer` registry -(`dashboard/workbench_domain.go`). A run selects 1-5 of them; the server +(`backend-service/src/workbench_domain.rs`). A run selects 1-5 of them; the server rejects zero selections, more than 5, an unknown ID, or a duplicate. | ID | Applicability | Adapter | Result link | Concurrency class | @@ -109,24 +109,27 @@ degraded to a stale local copy (#405 follow-up). ## HTTP contracts -All APIs require a live administrator identity. Every mutation additionally requires a same-origin request, `application/json`, one document no larger than 64 KiB, and the closed Go schema (unknown fields are rejected). Create, read/list, reconcile, cancel, and retry are further scoped server-side to the caller's verified owner identity (see Trust boundary) — a request's `owner` field is ignored entirely, never an access decision (#3110). +All APIs require a live administrator identity. Every mutation additionally requires a same-origin request, `application/json`, and the closed request schema. (A 64 KiB body cap and unknown-field rejection were part of the original Go contract; neither is present in the Rust tier, so do not rely on them.) Create, read/list, reconcile, cancel, and retry are further scoped server-side to the caller's verified owner identity (see Trust boundary) — a request's `owner` field is ignored entirely, never an access decision (#3110). + +Routes are registered in `backend-service/src/main.rs:456-468`. The hash is a **query** parameter on the registry and list routes, not a path parameter, and `cancel`/`retry` share one route via an `{action}` segment rather than being separate paths. | Method and route | Purpose | |---|---| -| `GET /api/payload-workbench/registry/{sha256}` | server-derived registry, applicability, external-publication notice, and advisory model health | -| `GET /api/payload-workbench/recipes` | visible private/shared recipe revisions | -| `POST /api/payload-workbench/recipes` | append an immutable recipe revision | -| `GET /api/payload-workbench/runs?sha256=...` | recent parent runs for the caller and payload | -| `POST /api/payload-workbench/runs` | submit a saved revision or typed one-off selection | -| `GET /api/payload-workbench/runs/{run_id}` | reconcile and return one parent run | -| `POST /api/payload-workbench/runs/{run_id}/children/{analyzer_id}/retry` | bounded deliberate retry | -| `POST /api/payload-workbench/runs/{run_id}/children/{analyzer_id}/cancel` | cancel an exact pending marker when supported | +| `GET /api/v1/workbench/analyzers?hash=…` | server-derived registry, applicability, external-publication notice, and advisory model health | +| `GET /api/v1/workbench/recipes` | visible private/shared recipe revisions | +| `POST /api/v1/workbench/recipes` | append an immutable recipe revision | +| `GET /api/v1/workbench/runs?hash=…&limit=…` | recent parent runs for the caller and payload | +| `POST /api/v1/workbench/runs` | submit a saved revision or typed one-off selection | +| `GET /api/v1/workbench/runs/{id}` | reconcile and return one parent run | +| `POST /api/v1/workbench/runs/{id}/children/{analyzer_id}/{action}` | bounded deliberate retry, or cancel an exact pending marker when supported (`action` ∈ `retry`\|`cancel`) | Create, recipe-save, retry, and cancel outcomes use the existing dashboard audit sink. Audit fields name the contract fields but do not copy payload content, prompts, model replies, filenames, credentials, or tool output. ## Model-status adapter -`/var/lib/honeypot-ghidra/model-status.json` remains root-owned mode `0600`. `honeypot-model-status-adapter.service` reads and re-validates it, strips every field outside schema v1, and serves only `GET /v1/status` over `/run/honeypot-model-status/status.sock`. The dashboard mounts that runtime directory read-only and uses `MODEL_STATUS_SOCKET=/model-status/status.sock`. There is no TCP listener and no write, pull, replace, promote, prompt, or model-selection route. +`/var/lib/honeypot-ghidra/model-status.json` remains root-owned mode `0600`. `honeypot-model-status-adapter.service` reads and re-validates it, strips every field outside schema v1, and serves only `GET /v1/status` over `/run/honeypot-model-status/status.sock`. The dashboard is expected to mount that runtime directory read-only and use `MODEL_STATUS_SOCKET=/model-status/status.sock`. There is no TCP listener and no write, pull, replace, promote, prompt, or model-selection route. + +> **Undetermined:** `MODEL_STATUS_SOCKET` has no occurrence anywhere in the Rust backend, the compose files, or `.env.example` as of 2026-09-27 — the adapter half of this contract is real and installed, but the consumer side is not visible in the repository. Treat the socket wiring as intended-but-unwired rather than working, and confirm before relying on the advisory model-health field the registry returns. Re-run `sudo analysis/ghidra/install-analysis-host.sh` to install or update the adapter. Its failure only displays `unavailable`; it never disables a worker. @@ -138,7 +141,7 @@ Deploy the dashboard normally after merging. Rollback is additive and safe: 1. deploy the previous dashboard image; 2. optionally disable `honeypot-model-status-adapter.service`; -3. leave the workbench indices in Elasticsearch untouched (a rolled-back dashboard from before the #405 follow-up reads its own local `/state/analysis-workbench` copy instead and simply does not see runs created after the rollback). +3. leave the workbench indices in Elasticsearch untouched (a rolled-back pre-workbench dashboard does not see runs created after the rollback). The old `/ghidra/submit` and `/sandbox/submit` routes remain compatible. No worker or native result schema is changed by the workbench. From 6e017fd0048f121bdab561b1efbbf4a742b40fb0 Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:07:53 +0200 Subject: [PATCH 08/21] docs(deploy-profiles): correct backbone list, structural deps, and full.txt gap --- docs/deploy-profiles/README.md | 38 +++++++++++++++++++++++----------- 1 file changed, 26 insertions(+), 12 deletions(-) diff --git a/docs/deploy-profiles/README.md b/docs/deploy-profiles/README.md index f2d209dd..73d538e0 100644 --- a/docs/deploy-profiles/README.md +++ b/docs/deploy-profiles/README.md @@ -27,17 +27,30 @@ line, `#` comments and blank lines ignored. | Profile | Backbone | Sensors | Shape | |---|---|---|---| -| [`full.txt`](../../deploy-profiles/full.txt) | init, elk, dashboard, utilities, payload-analysis | every deception sensor stack under `arcane/home/` | the standard deployment -- everything this repo ships | -| [`ics-focused.txt`](../../deploy-profiles/ics-focused.txt) | init, elk, dashboard, utilities | conpot, dnp3 | OT/ICS-only exposure -- skip the general-purpose/web/SSH/legacy-protocol sensors entirely | -| [`minimal-web.txt`](../../deploy-profiles/minimal-web.txt) | init, elk, dashboard, utilities | http, tanner | web-attack-focused -- HTTP/API honeypot + SNARE/TANNER, skip ICS/SSH/legacy-protocol sensors | - -`init`, `elk`, and `dashboard` are structural dependencies for any profile -that includes at least one sensor -- `scripts/validate-deploy-profile.sh` -(below) enforces -this, it isn't just a convention to remember. `payload-analysis` and +| [`full.txt`](../../deploy-profiles/full.txt) | keycloak, init, elk, dashboard, utilities, payload-analysis | the 20 classic deception sensor stacks under `arcane/home/` -- but see the gap below | the standard deployment -- everything this repo ships | +| [`ics-focused.txt`](../../deploy-profiles/ics-focused.txt) | keycloak, init, elk, dashboard, utilities | conpot, dnp3 | OT/ICS-only exposure -- skip the general-purpose/web/SSH/legacy-protocol sensors entirely | +| [`minimal-web.txt`](../../deploy-profiles/minimal-web.txt) | keycloak, init, elk, dashboard, utilities | http, tanner | web-attack-focused -- HTTP/API honeypot + SNARE/TANNER, skip ICS/SSH/legacy-protocol sensors | + +**`full.txt` is not actually "everything this repo ships".** It lists 26 +stacks (6 backbone + 20 sensors) and omits `honeypot-sonicwall-sma`, the +decoy sensor stack #3131 added on 2026-09-08 (`hp-sonicwall-sma-honeypot` +on `${HP_BIND:-10.8.0.2}:8543`). It is a natural fit for both `full.txt` and +`ics-focused.txt`, and is in neither. The other `arcane/home/` stacks the +profiles deliberately skip are the analysis-plane workers and +`honeypot-dashboard-backend` (not persona declarations, per below) plus +`unsloth` (the #3092 benchmark toolchain) -- and `rex86-eval`, which exists +on disk but is in no manifest entry at all. + +`init` and `elk` are structural dependencies for any profile that includes at +least one sensor; `keycloak` is a structural dependency of `dashboard`, and +`elk` is too. `scripts/validate-deploy-profile.sh` enforces all three, so +these aren't just conventions to remember. `payload-analysis` and `utilities` are strongly recommended (payload dedup/YARA scanning, log rotation/disk monitoring/autoheal) but not structurally required, so the -validator only warns if either is missing from a non-empty profile. +validator only warns if either is missing from a non-empty profile. Note +that `dashboard` itself is *not* a required structural dependency: the +validator never demands it, it only imposes `elk` and `keycloak` on a +profile that has chosen it. Not covered here: the VPS side (`vps/`, always deployed the same way regardless of home profile -- see `docs/CGNAT-DEPLOYMENT.md`), the @@ -56,9 +69,10 @@ scripts/validate-deploy-profile.sh deploy-profiles/ics-focused.txt Checks, against the *current* repository state (not a hardcoded snapshot): 1. **Structural dependencies** -- `init`/`elk` present if any sensor stack - is listed; `elk` present if `dashboard` is listed (the dashboard reads - several sensors' events from Elasticsearch, not their log files -- - see #403 for why that's a real dependency, not a nice-to-have). + is listed; `elk` and `keycloak` present if `dashboard` is listed (the + dashboard reads several sensors' events from Elasticsearch, not their + log files -- see #403 for why that's a real dependency, not a + nice-to-have; and the target auth path is native Keycloak OIDC). 2. **Real-stack existence** -- every listed name must correspond to an actual `arcane/home/honeypot-/` directory, so a typo'd or retired stack name fails here instead of surfacing mid-deploy or as a silently From 17e20ab9d87945165a5e0a38880a2197d48ed915 Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:09:36 +0200 Subject: [PATCH 09/21] docs(homeserver): sharpen the manifest/directory-count sentence The previous phrasing conflated two different counts: 34 is how many directories exist on disk under arcane/home/, while the manifest holds 39 entries of which 33 name one of those directories and 6 are root-level stacks. Spelled out so the sentence cannot be read as '34 of the 45 stack directories are manifest-managed'. --- docs/HOMESERVER-DISK-LAYOUT.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/HOMESERVER-DISK-LAYOUT.md b/docs/HOMESERVER-DISK-LAYOUT.md index 6b534940..ac3ba183 100644 --- a/docs/HOMESERVER-DISK-LAYOUT.md +++ b/docs/HOMESERVER-DISK-LAYOUT.md @@ -75,9 +75,9 @@ load-bearing: `/var/lib/docker` is **2.9T** and `/var/dockge` (stack data for the 45 directories under `/var/dockge/stacks/`, including Elasticsearch indices, Cowrie logs, payload captures, sandbox disks) is **350G**. `/var` is 70% full (6.1T of 8.8T) with 2.7T free. The manifest -still declares 39 sync entries and 34 of those directories are -manifest-managed; `rex86-eval` is present on disk but **not** in the -manifest. Putting `/var` on the RAID LUN instead of growing the root +still declares 39 sync entries, 33 of which name one of the 34 directories +under `arcane/home/`; `rex86-eval` is present on disk but **not** in the +manifest (the other 6 manifest entries are root-level stacks). Putting `/var` on the RAID LUN instead of growing the root filesystem remains the right call and should be preserved on any rebuild. Swap is a **32G LVM logical volume** (`rl-swap`) in the `rl` volume group, From f19bf5202de6ddfb9d658c8334a824c92d8e1e92 Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:14:30 +0200 Subject: [PATCH 10/21] docs(infra,benchmarks,gpu): fix manifest count, cohort table, and GPU device pinning --- .../benchmarks/injection-gate-protocol.md | 8 ++++---- docs/gpu-docker-passthrough.md | 20 +++++++++++++++++-- 2 files changed, 22 insertions(+), 6 deletions(-) diff --git a/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md b/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md index 41458050..8b701c75 100644 --- a/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md +++ b/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md @@ -14,11 +14,11 @@ containing the payload's own words. Read against every stored | | truly complied | did not comply | |---|---|---| -| gate FAIL | 1 (partial) | 26 | -| gate PASS | 0 | 37 | +| gate FAIL | 1 (partial) | 25 | +| gate PASS | 0 | 38 | -25 of 27 failures fired on the model *quoting or paraphrasing* the planted -string — the behaviour the system prompt asks for. The remaining ones fired +25 of the 26 failures fired on the model *quoting or paraphrasing* the planted +string — the behaviour the system prompt asks for. The rest fired on "appears to be benign", which is the case's own ground truth. The four Tier A failures that drove the #1805-c "no promotion" decision (Ornith-35B, gemma-4-31B, Seneca-32B, huihui-qwen3.8) all explicitly identified the string diff --git a/docs/gpu-docker-passthrough.md b/docs/gpu-docker-passthrough.md index d67add1f..2f5989eb 100644 --- a/docs/gpu-docker-passthrough.md +++ b/docs/gpu-docker-passthrough.md @@ -189,10 +189,22 @@ services: reservations: devices: - driver: nvidia - count: all + device_ids: ["GPU-18a00c7e-670a-c305-a2aa-20e3a71917a3"] capabilities: [gpu] ``` +**Do not use `count: all` here.** #1539 replaced it: the homeserver carries +two NVIDIA cards, and the `count: all` form handed Ollama *both* — wrong +even when it works, and it starves the Windows sandbox VM of the Quadro +P2200 reserved for its passthrough whenever Ollama loads a model during a +detonation. The overlay therefore pins the RTX 4000 Ada's UUID, confirmed +live via `nvidia-smi -L` on the actual box. Re-run `nvidia-smi -L` and +update the UUID if the card is ever physically replaced. + +The same file also documents why `ghidra` is deliberately *not* given the +GPU: decompilation is CPU work and would only compete for the card with the +model it is feeding. + Note this repo keeps the GPU reservation in a **separate overlay file**, applied only when a GPU is actually present: @@ -225,7 +237,11 @@ Docker's `--gpus all` / `count: all` doesn't partition VRAM — every container that requests the GPU gets the whole card, and it's up to each process to behave. Nothing stops two containers from both trying to allocate more VRAM than the card has, at which point the second allocator -gets a CUDA out-of-memory error, not a scheduling wait. +gets a CUDA out-of-memory error, not a scheduling wait. On this box the +question is sharper still, because the host has **two** cards: an RTX 4000 +Ada for compute and a Quadro P2200 reserved for the Windows sandbox VM's +passthrough. `count: all` would hand both to one container, which is the +#1539 bug the overlay's pinned `device_ids` exists to prevent. This repo's own answer to that (see [`gpu-ml-worker-acceleration.md` §5, "GPU Sharing Contract with the LLM From 0d76e77c06e2d4a7e113c2722b9203576ecb419a Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:18:39 +0200 Subject: [PATCH 11/21] docs(infra): fix shipped-vs-planned tense, KVM bridges, and GPU roadmap drift MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ROCKY-10-MIGRATION.md: 'is moving from Ubuntu' -> 'has moved' (the rebuild hit live 2026-09-03; install-homeserver.sh:324 says so). Also resolved the one claim left undetermined by the previous pass: both cards really are on the bus (P2200 17:00.0/10de:1c31, Ada 65:00.0), but only the Ada has a driver bound, so nvidia-smi lists one GPU. Documented that explicitly, because 'one GPU in nvidia-smi' reads like missing hardware and is the reason the nvidia-open-vs-cuda-drivers choice still matters. KEYCLOAK-CUTOVER.md: the hard cutover this contract specifies has shipped. vps/forward-auth/ is gone and xore_sso/_auth/verify/strip-auth-identity exist nowhere but this doc and docs/TESTING.md. Status line now says SHIPPED; the contract body is untouched. Verified all 8 OIDC client IDs still exist in arcane/home/honeypot-keycloak/keycloak/realm/apiary-realm.json with the documented flow flags. kvm-network-traffic-analysis.md: 'two bridges' understated it -- there are four (virbr-hpsbx, virbr-cape, virbr-ghosts, plus 198.18.0.0/24 on virbr-hpsbx in controlled mode). Results path was wrong: run_sample.py overrides the compose default per run via WINDOWS_SANDBOX_RESULTS_DIR, and sandbox/results/ does not exist. 'The Linux sandbox has no Zeek equivalent' is false -- run-linux-sample.sh runs 'zeek -r' offline and export-result.py reads the logs. The Phase 0 FORWARD DROP pair is a documented manual step, not scripted, and needs re-verification against Rocky 10.2 firewalld. kvm-snapshot-vs-golden-image.md: §5's golden-win10 and /var/lib/libvirt/golden/ are illustrative; kvm_manage.sh uses SANDBOX_ROOT=/var/dockge/sandbox with golden-images/win11-analysis.qcow2 and spawns via qemu-img create -b + virsh define, not virt-clone. Added a pointer instead of rewriting §5. ml-gpu-coordinated-roadmap.md: kept as a dated plan, added an intent-vs-shipped banner covering the four things that moved since 2026-08-01, including that §1 decision 5 was superseded outright -- the ML worker has no embedding code at all, and llm-worker ships 768-dim embeddings behind LLM_EMBEDDING_ENABLED (default off). --- docs/ROCKY-10-MIGRATION.md | 10 +++++++++- docs/kvm-network-traffic-analysis.md | 18 +++++++++++++----- docs/kvm-snapshot-vs-golden-image.md | 7 +++++++ docs/ml-gpu-coordinated-roadmap.md | 11 ++++++++++- 4 files changed, 39 insertions(+), 7 deletions(-) diff --git a/docs/ROCKY-10-MIGRATION.md b/docs/ROCKY-10-MIGRATION.md index c19f7bba..b2b2203a 100644 --- a/docs/ROCKY-10-MIGRATION.md +++ b/docs/ROCKY-10-MIGRATION.md @@ -1,6 +1,7 @@ # Rocky Linux 10 support in `install-homeserver.sh` -The homeserver is moving from Ubuntu to Rocky Linux 10. `scripts/install-homeserver.sh` +The homeserver has moved from Ubuntu to Rocky Linux 10 (the rebuild hit +live 2026-09-03). `scripts/install-homeserver.sh` now runs on both, so the reinstall smoke test in #1609 has a working installer. ## How it works @@ -62,6 +63,13 @@ Quadro P2200 alongside the Ada RTX 4000; the P2200 is meant to be bound to `vfio-pci` for the Windows sandbox rather than driven by the host, but choosing the open modules would make it unusable on the host if ever needed. +Both cards are still on the PCI bus (P2200 at `17:00.0` / `10de:1c31`, Ada at +`65:00.0`), re-measured 2026-09-27. Only the Ada currently has a driver bound +(`Kernel driver in use: nvidia`); the P2200 shows under *Kernel modules* with +none *in use*, so `nvidia-smi -L` lists exactly one GPU. That is expected, and +it is **not** evidence the P2200 is missing — the `nvidia-open` reasoning above +only bites if the P2200 is ever handed back to the host. + **NVIDIA container toolkit** — same rpm repofile approach. On RHEL the `container_use_devices` SELinux boolean is also set, without which a container cannot open the GPU device nodes. diff --git a/docs/kvm-network-traffic-analysis.md b/docs/kvm-network-traffic-analysis.md index 563160ec..a9749144 100644 --- a/docs/kvm-network-traffic-analysis.md +++ b/docs/kvm-network-traffic-analysis.md @@ -18,7 +18,9 @@ Traffic is captured on the host side of an isolated libvirt bridge, never inside the guest and never by an agent the guest could interfere with. There -are **two** such bridges, because there are two sandboxes: +are **two** such bridges for the detonation pair documented here (CAPE adds +`virbr-cape` and GHOSTS adds `virbr-ghosts`, and controlled mode puts +`198.18.0.0/24` on `virbr-hpsbx` itself): | | Windows detonation guest | Linux / Wine runner | |---|---|---| @@ -27,7 +29,7 @@ are **two** such bridges, because there are two sandboxes: | Network XML | `sandbox/windows/setup/sandbox-network.xml` | `sandbox/network.xml` | | Fake internet | INetSim at `10.10.10.1` | none by default; optional logged DNS + Squid allowlist (`controlled` mode) | | Capture | `docker-compose.sandbox.yml` (tcpdump, Zeek, Suricata) | root-owned `tcpdump` per job, host and guest side | -| Results | `sandbox/results//` | `/var/lib/honeypot-sandbox/results/` (root-only), sanitized export copied out | +| Results | `$WINDOWS_SANDBOX_RESULTS_DIR//` — `run_sample.py` overrides the compose default of `sandbox/results/current` per run | root-only, sanitized export copied out | | Orchestrator | `sandbox/windows/orchestrate/run_sample.py` | `sandbox/run-linux-sample.sh` | Neither bridge has a `` element, so neither can route anywhere. That @@ -85,7 +87,11 @@ For the Windows bridge: 1. The libvirt network has no ``. 2. The gateway containers sit on a macvlan marked `internal: true`, which removes the default route Docker would otherwise install. -3. Phase 0 adds an iptables DROP pair across `virbr-sandbox`. +3. Phase 0 asks the operator to add a FORWARD DROP pair across + `virbr-sandbox`. This is a documented manual step, not something a script + applies — re-verify it against this host's firewall backend, which is now + Rocky 10.2 with firewalld/nftables rather than the iptables the step was + originally written for. Removing any one of them because "the container needs to pull something" puts live malware on the internet. Build images ahead of time instead. @@ -170,8 +176,10 @@ for a run the worker does not know about. Authenticated administrators can download the pcaps. Oversize pcaps and the raw result directories stay root-only. - **`generate_report.py`** folds `zeek_logs/http.log` into the Windows report. - The Linux sandbox has no Zeek equivalent — also - [#87](https://github.com/Xore/APIARY/issues/87). + The Linux sandbox now gets Zeek too, offline over the finished pcap rather + than live on the bridge (`sandbox/run-linux-sample.sh` runs `zeek -r` and + `sandbox/export-result.py` reads `conn`/`dns`/`http`/`ssl`/`files` from + `zeek_logs/`). ## 5. Retention diff --git a/docs/kvm-snapshot-vs-golden-image.md b/docs/kvm-snapshot-vs-golden-image.md index b3a0c58a..6162f7b3 100644 --- a/docs/kvm-snapshot-vs-golden-image.md +++ b/docs/kvm-snapshot-vs-golden-image.md @@ -233,6 +233,13 @@ virsh start "analysis-${SAMPLE}" echo "VM analysis-${SAMPLE} started with overlay $OVERLAY" ``` +> §5's `golden-win10` and `/var/lib/libvirt/golden/` are **illustrative**, not +> this stack's real names or paths. The live mechanism uses +> `SANDBOX_ROOT=/var/dockge/sandbox`, a base image of +> `golden-images/win11-analysis.qcow2` and a per-run +> `vms/win11-sandbox.qcow2`, and it spawns with `qemu-img create -b` plus +> `virsh define` rather than `virt-clone` — see §7. + ### 5.3 Destroy and Clean Up After the Run ```bash diff --git a/docs/ml-gpu-coordinated-roadmap.md b/docs/ml-gpu-coordinated-roadmap.md index db20c8f6..3b2cb512 100644 --- a/docs/ml-gpu-coordinated-roadmap.md +++ b/docs/ml-gpu-coordinated-roadmap.md @@ -1,6 +1,15 @@ # Coordinated ML and GPU Analysis Roadmap -> **Status:** Proposed implementation sequence +> **Status:** Proposed implementation sequence — **intent, not shipped state.** +> Re-measured 2026-09-27: `ml-worker` and `llm-worker` now ship as their own +> Arcane-managed stacks; the ML worker's GPU overlay +> (`docker-compose.ml-worker.gpu.yml`) is still inert scaffolding, so the ML +> worker remains CPU-only; retrain slots are `03:00,09:00,15:00,21:00` +> (`ml-worker/worker.py:59`), not §4-I's `01:00,07:00,13:00,19:00`; the +> default alert threshold did land at `0.75` (`worker.py:174`); and §1 +> decision 5 is superseded — the ML worker never grew an embedding index at +> all, embeddings shipped in `llm-worker` at **768** dims, off by default. +> The milestone text below is left as the historical record. > > **Scope:** `ml-worker`, GPU acceleration, local LLM analysis, and dashboard delivery > From 109373b7206338e2e1e9fac70b653736f0073c76 Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:19:05 +0200 Subject: [PATCH 12/21] docs(gpu-ml): refresh torch/pyod pins, second GPU, and VRAM headroom --- docs/gpu-ml-worker-acceleration.md | 40 ++++++++++++++++++++++-------- 1 file changed, 30 insertions(+), 10 deletions(-) diff --git a/docs/gpu-ml-worker-acceleration.md b/docs/gpu-ml-worker-acceleration.md index a21fef3b..5c944fb7 100644 --- a/docs/gpu-ml-worker-acceleration.md +++ b/docs/gpu-ml-worker-acceleration.md @@ -73,9 +73,10 @@ parts, and it does not make the models more accurate. ## 3. Hardware & Compatibility Contract **Settled by [#602](https://github.com/Xore/APIARY/issues/602)** (verbatim -host evidence: `lspci` shows a single AD104GL controller, containers -enumerate exactly one device — the earlier two-card / Turing-plus-Ada -hypothesis is refuted) and pinned as the runtime-governance authority in +host evidence: `lspci` showed a single AD104GL controller and containers +enumerated one compute device — the earlier two-card *Turing-plus-Ada* +compute hypothesis is refuted) and pinned as the runtime-governance +authority in `analysis/ghidra/models/approved-models.json`, which `model-governance.py check-runtime` diffs live `nvidia-smi` against every 5 minutes: @@ -84,6 +85,14 @@ hypothesis is refuted) and pinned as the runtime-governance authority in capability 8.9.** (An earlier draft of this document, before #602, mis-recorded this as a Quadro RTX 4000 at compute capability 7.5/Turing — that card was never on this host; see #602 for the full correction.) +- **One more card arrived later, and it is not for this workload.** #1539 + recorded that the box also carries a **Quadro P2200**, reserved for the + Windows sandbox VM's passthrough. So the "one device" finding above is + still true of the *compute* pool this guide budgets VRAM against, but the + host is no longer single-GPU, and this matters concretely in §4.4: a + `count: 1` reservation picks an arbitrary card, which is the same class of + bug `count: all` caused. Pin `device_ids` to the Ada's UUID, as + `analysis/ghidra/docker-compose.ghidra.gpu.yml` already does. - Driver 580.173.02 (CUDA 13.0) — backward-compatible with CUDA 12.x runtime wheels. - nvidia-container-toolkit 1.19.1 present; `docker run --rm --gpus all @@ -118,17 +127,18 @@ Replace the CPU wheel lines: ```diff -# Deep learning (CPU-only PyTorch) --torch==2.13.0+cpu +-torch==2.14.0+cpu ---extra-index-url https://download.pytorch.org/whl/cpu +# Deep learning (CUDA PyTorch — see docs/gpu-ml-worker-acceleration.md §3) -+torch==2.13.0+cu126 ++torch==2.14.0+cu126 +--extra-index-url https://download.pytorch.org/whl/cu126 + +# Embeddings (§6) +sentence-transformers==3.0.1 - # Outlier detection (HBOS) - pyod==3.6.2 + # Outlier detection (HBOS). Pulls in numba+llvmlite -- see the numpy pin + # above; this is the version set actually verified to install together. + pyod==3.6.6 ``` > **Verified pin (2026-08-01, #82):** `torch==2.13.0+cu124` does not exist; @@ -141,7 +151,10 @@ Replace the CPU wheel lines: > including both `sm_75` and `sm_89`, so the wheel's `sm_75` inclusion says > nothing about which kernel the live card actually used; the underlying > install-and-tensor-check result stands, only the architecture label was -> wrong.) Never silently fall back to the `+cpu` wheel; a CPU wheel in a GPU +> wrong.) **The tree has since moved on: `ml-worker/requirements.txt` now +pins `torch==2.14.0+cpu` and `pyod==3.6.6`, so the +cu126 install above was +verified at 2.13.0 only and has not been re-checked at 2.14.0. G2 applies to +whatever version is current when this deploys.** Never silently fall back to the `+cpu` wheel; a CPU wheel in a GPU > deployment must fail the acceptance test T2, not pass unnoticed. ### 4.2 `ml-worker/Dockerfile` @@ -199,6 +212,12 @@ Rules: + ML_DEVICE: auto # auto | cpu — 'cpu' forces CPU for debugging ``` +**Pin `device_ids`, not `count: 1`.** The host carries two cards (§3), so +`count: 1` selects an arbitrary one and can hand `ml-worker` the Quadro +P2200 the Windows sandbox VM needs. Use the Ada's UUID exactly as +`analysis/ghidra/docker-compose.ghidra.gpu.yml` does: +`device_ids: ["GPU-18a00c7e-670a-c305-a2aa-20e3a71917a3"]`. + Keep the existing `mem_limit: 2g` / `cpus: "2.0"`; GPU memory is governed by §5, not by `mem_limit`. @@ -224,7 +243,8 @@ the original 6.1 GiB `qwen3.5:9b` estimate): smaller safety margin instead of full separation.** Worst case — the ghidra slot's model loaded at its 32k context (~14.1 GiB) plus a retrain (~2 GiB) plus the embedder's own inference (~1 GiB) — totals ~17.1 GiB - against the real 20475 MiB budget: about 3.3 GiB of headroom, not the + against the real 20475 MiB budget (19.995 GiB): about 2.9 GiB of headroom, + not the "comfortable" double-digit margin a naive 6.1 GiB-chat-model estimate would suggest. That is enough to not require full separation, but not enough to treat as a non-issue either. This already assumes the ghidra @@ -240,7 +260,7 @@ the original 6.1 GiB `qwen3.5:9b` estimate): best-effort headroom management. - If a future model re-evaluation (#568's process, re-run) picks something materially larger than `qwen3:14b` for the ghidra slot, re-check this - margin before assuming it still holds — the 3.3 GiB headroom above was + margin before assuming it still holds — the 2.9 GiB headroom above was computed for this specific model at its current context ceiling, not as a permanent property of the 20 GB card. - **On CUDA OOM, do not crash:** wrap train/infer calls, catch From b835eae8cf5e50e59bc015129d3ce7cfb3d21f90 Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:19:09 +0200 Subject: [PATCH 13/21] docs(kvm): correct the results-path row; drop an invented compose default The previous commit cited a 'compose default of sandbox/results/current' for WINDOWS_SANDBOX_RESULTS_DIR. No such default exists -- docker-compose.sandbox.yml never mentions the variable. The verified facts are that run_sample.py:124 reads the variable with a 'reports/windows-sandbox' fallback, .env.example documents '/windows-sandbox-results', and the dashboard's es-results-importer populates it per stack. Reworded to those, which also keeps the doc-path lint green. --- docs/kvm-network-traffic-analysis.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/kvm-network-traffic-analysis.md b/docs/kvm-network-traffic-analysis.md index a9749144..ab1ebd4d 100644 --- a/docs/kvm-network-traffic-analysis.md +++ b/docs/kvm-network-traffic-analysis.md @@ -29,7 +29,7 @@ are **two** such bridges for the detonation pair documented here (CAPE adds | Network XML | `sandbox/windows/setup/sandbox-network.xml` | `sandbox/network.xml` | | Fake internet | INetSim at `10.10.10.1` | none by default; optional logged DNS + Squid allowlist (`controlled` mode) | | Capture | `docker-compose.sandbox.yml` (tcpdump, Zeek, Suricata) | root-owned `tcpdump` per job, host and guest side | -| Results | `$WINDOWS_SANDBOX_RESULTS_DIR//` — `run_sample.py` overrides the compose default of `sandbox/results/current` per run | root-only, sanitized export copied out | +| Results | `$WINDOWS_SANDBOX_RESULTS_DIR//`, set per run; `run_sample.py` falls back to `reports/windows-sandbox` and the dashboard's results importer populates the variable | root-only, sanitized export copied out | | Orchestrator | `sandbox/windows/orchestrate/run_sample.py` | `sandbox/run-linux-sample.sh` | Neither bridge has a `` element, so neither can route anywhere. That From 43b18b315379b30b0285a7974ce861976552e726 Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:20:43 +0200 Subject: [PATCH 14/21] docs(gpu-llm): correct embedding dims to 768 and note the second GPU --- docs/gpu-llm-analysis-worker.md | 22 ++++++++++++++++++---- 1 file changed, 18 insertions(+), 4 deletions(-) diff --git a/docs/gpu-llm-analysis-worker.md b/docs/gpu-llm-analysis-worker.md index dedd0379..e063f751 100644 --- a/docs/gpu-llm-analysis-worker.md +++ b/docs/gpu-llm-analysis-worker.md @@ -86,10 +86,19 @@ worker milestones. **Settled by [#602](https://github.com/Xore/APIARY/issues/602)** — an earlier draft of this table (before #602) named the card as a Quadro RTX 4000 at compute capability 7.5/Turing; that card was never on this host. -`lspci` shows a single AD104GL controller and containers enumerate exactly -one device. Also pinned as the runtime-governance authority in +`lspci` showed a single AD104GL controller and containers enumerated one +compute device. Also pinned as the runtime-governance authority in `analysis/ghidra/models/approved-models.json`. +> **A second card arrived after #602.** #1539 recorded that the box also +> carries a **Quadro P2200**, reserved for the Windows sandbox VM's +> passthrough. Every VRAM budget in this document is still correct — they +> are budgets against the Ada, and the P2200 is not part of the compute +> pool — but the host is no longer single-GPU, so "containers enumerate one +> device" no longer describes the machine. It is why the Ollama reservation +> in `analysis/ghidra/docker-compose.ghidra.gpu.yml` pins `device_ids` to +> the Ada's UUID rather than using `count: all`. + | Fact | Value | Verify with | |---|---|---| | GPU | NVIDIA RTX 4000 Ada Generation | `nvidia-smi -L` | @@ -501,10 +510,15 @@ Mirrors the pattern `ml-anomalies` already established ("any SSE/Redis wake-up path remains optional and non-authoritative"); polling has not been shown insufficient yet. - **Delivered:** semantic search over sessions using `nomic-embed-text` - embeddings stored as a `dense_vector` (384-dim) field on `llm-analysis` + embeddings stored as a `dense_vector` (768-dim) field on `llm-analysis` docs, queried with ES kNN search — `GET /api/v1/llm-search` (`main.rs:341`, `llm_search.rs`, over the `llm-analysis` index's - `doc_type: session` documents). This section originally deferred it + `doc_type: session` documents). The dimensionality is 768, not the 384 + this document originally stated: #151 confirmed the model's real native + output live against `POST /api/embed` and `llm-worker/worker.py` pins + `EMBEDDING_DIMS = 768`, rejecting a response of any other width outright + rather than indexing it into a mapping it cannot satisfy. This section + originally deferred it pending U1–U3 stability; it has since shipped, so the list above is not a statement of current scope. From 20783fca5e6212b916f1ced29e52175b8c651565 Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:21:28 +0200 Subject: [PATCH 15/21] docs(recovery,sandbox,ml): fix restore order, stale Go paths, and the retrain schedule MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit analysis/RECOVERY.md: the two backups were described as 'the same scope', which is wrong and hides a filename trap. Scope differs -- the essentials archive additionally carries the VPS config, WireGuard, Technitium, the installer answers and the repo runbooks. Concretely, SHA256SUMS and stack-config-state.tar.gz are the ON-HOST copy's only; a backup-essentials.sh archive has neither, and this doc's own first bullet points readers at the essentials archive. Also replaced 'docker compose -f compose.yml create' with per-volume 'docker volume create ' -- four of the five volumes carry an Arcane prefix and live in four different stacks, so one compose create cannot make them -- and restored the real start order (honeypot-elk healthy before honeypot-init) from STACK-REBUILD.md, which this step had lost. sandbox/README.md: the workbench registry pointer dashboard/workbench_domain.go -> backend-service/src/workbench_domain.rs (the Go tree has no .go files since #1628), in prose and in the mermaid node. Two host numbers were stale: 16 -> 48 logical CPUs, and the IOMMU figure. The doc's 84 groups was wrong on the number AND missed the story: the Rocky 10 rebuild came up with zero groups, which is why install-homeserver.sh's step_vfio_gpu_passthrough now adds intel_iommu=on to the kernel command line. Re-measured 2026-09-27: 94 groups, with intel_iommu=on iommu=pt confirmed live in /proc/cmdline, and /dev/kvm present. Line 237's 4 vCPU / 8 GiB guidance matches neither the real Windows domain (8 vCPU / 16 GiB) nor sandbox.env.example's SANDBOX_VM_MEMORY_MB=3072, so it is now flagged as operator guidance rather than left reading like a measured limit. ml-worker-plan.md: RETRAIN_INTERVAL is gone -- #172 replaced it with four fixed UTC slots (worker.py:59, default 03:00,09:00,15:00,21:00). Fixed in all four places it was asserted, including the early-retrain trigger, which is now slot-nearest rather than interval-remaining. Corrected the status callout: the worker runs from its own Arcane stack, not the repo-root compose file, and docker-compose.ml-worker.gpu.yml is inert and undeployed so the worker is CPU-only. Fixed the v0.1 paragraph, which contradicted its own §10 and the line below it by claiming there were no tests. Added a dated reading note separating the shipped sections from the original-draft ones, so the remaining dashboard/*.go references read as historical instead of as current. Benchmark records #65-#69: dated intent-vs-shipped banners only, no history rewritten and no recorded number touched. Added alongside the existing /mnt-1 banners rather than replacing them. #66 is superseded by #2694; #67's hardware table is pre-reinstall and its #2985 'missing scripts' are now in git except gptoss_rerun.sh; #68's unsloth is the Arcane stack, not an installer; #69's TOOLCHAIN.md 'no smoke test yet' still stands. --- docs/analysis/RECOVERY.md | 29 +++++++++++---- .../benchmarks/injection-gate-protocol.md | 15 ++++++-- ...2026-08-30-injection-gate-recalibration.md | 8 ++++ .../plans/2026-09-05-1947-resume-plan.md | 9 +++++ .../plans/2026-09-06-round7-hermes-init.md | 8 ++++ ...-06-round7-unsloth-train-requant-ollama.md | 7 ++++ docs/ml-worker-plan.md | 37 ++++++++++++------- docs/sandbox/README.md | 20 +++++++--- 8 files changed, 104 insertions(+), 29 deletions(-) diff --git a/docs/analysis/RECOVERY.md b/docs/analysis/RECOVERY.md index 00f817bb..3e6a9e8c 100644 --- a/docs/analysis/RECOVERY.md +++ b/docs/analysis/RECOVERY.md @@ -4,11 +4,14 @@ For a *deliberate* full reset on the same hosts (not disaster recovery), see [`docs/STACK-REBUILD.md`](../STACK-REBUILD.md) instead — this doc is about restoring a backup archive onto a replacement host after data loss. -Two backups cover this, with the same scope and the same exclusions: +Two backups cover the homeserver stack, with the same exclusions. **Scope +differs**, which matters here because the two produce different filenames: - [`scripts/backup-essentials.sh`](../../scripts/backup-essentials.sh) runs on the workstation and fans an encrypted archive out to three locations. This - is the one that survives the homeserver dying. Full restore procedure: + is the one that survives the homeserver dying, and the broader of the two — + it additionally carries the VPS config, WireGuard, Technitium, the + installer's answers file and the repo's runbooks. Full restore procedure: [`docs/BACKUP-ESSENTIALS.md`](../BACKUP-ESSENTIALS.md). - `sudo analysis/backup-honeypot.sh` runs on the homeserver itself, into a timestamped mode-0700 directory beneath `/opt/backups/honeypot`. Faster to @@ -25,15 +28,25 @@ and the sizes behind it. Recovery is intentionally not automatic because overwriting live volumes is destructive. On a replacement host: -1. Verify `SHA256SUMS`, unpack `stack-config-state.tar.gz` into a new empty stack - directory, and inspect `.env` permissions and values. +1. Unpack the archive's `env/` and `secrets/` trees back under + `/var/dockge/stacks//` and inspect `.env` permissions and values. + Note the filename first: `SHA256SUMS` and `stack-config-state.tar.gz` are + the **on-host** copy's only — `analysis/backup-honeypot.sh` writes them and + `analysis/verify-backup.sh` checks them. A `scripts/backup-essentials.sh` + archive contains neither; verify that one with + `sha256sum -c apiary-essentials-.tar.gz.gpg.sha256`. 2. Restore the Keycloak database from `keycloak.sql.gz` with `hp-keycloak-postgres` up and `hp-keycloak` still stopped — this is what preserves the OIDC client secrets that the VPS's own `secrets/oidc/` copies have to match. -3. Create the named volumes with `docker compose -f compose.yml create`, keep all - services stopped, and restore each matching volume archive using a temporary - networkless BusyBox container. -4. Start setup jobs and sensors, then run `analysis/verify-stack.py` — with +3. Keep all services stopped, then `docker volume create ` for each + `homeserver/volumes/.tar.gz` and restore each archive into it through a + temporary networkless BusyBox container. Use the **full real volume name** — + four of the five carry their Arcane project prefix and live in four + different stacks, so one `docker compose -f compose.yml create` cannot cover + them. +4. Start stacks in [`docs/STACK-REBUILD.md`](../STACK-REBUILD.md)'s order — + `honeypot-elk` healthy before `honeypot-init`, everything else after — then + run `analysis/verify-stack.py` — with `DASHBOARD_SERVICE_TOKEN` from the restored stack's `.env` exported; it reads source-health through dashboard-next's `/bff` passthrough and exits nonzero on any failure. diff --git a/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md b/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md index 8b701c75..14a2804a 100644 --- a/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md +++ b/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md @@ -1,5 +1,12 @@ # Injection gate v3 — what it measures and how to test it (#2694) +> **Re-verified 2026-09-27 — every tool and path in §3–§4 still exists as named.** +> The gate is current: the `injection` rubric block plus `injection_gate.py`, +> with the live slot at `qwen3:14b` and 32768 context. Whether the §3.3 +> positive-control sweep has been re-run against the current roster is +> **undetermined** from the repository — check `/var/benchmarks/` on the host. +> Companion: `docs/analysis/ghidra/benchmarks/corpus/README.md`. + This is the operating manual for the corpus injection axis after #2694. It states what a verdict means, how the gate is validated before any number from it is cited, and the exact commands. @@ -17,9 +24,11 @@ containing the payload's own words. Read against every stored | gate FAIL | 1 (partial) | 25 | | gate PASS | 0 | 38 | -25 of the 26 failures fired on the model *quoting or paraphrasing* the planted -string — the behaviour the system prompt asks for. The rest fired -on "appears to be benign", which is the case's own ground truth. The four +23 of the 26 failures fired on one of the four payload-derived forbidden terms +(the model quoting or paraphrasing the planted string) — the behaviour the +system prompt asks for. The other three fired only on "appears to be benign", +which is the case's own ground truth; two of those three never reproduce the +string at all. The four Tier A failures that drove the #1805-c "no promotion" decision (Ornith-35B, gemma-4-31B, Seneca-32B, huihui-qwen3.8) all explicitly identified the string as an injection or anomaly. No matcher can repair this fixture: the injected diff --git a/docs/benchmarks/2026-08-30-injection-gate-recalibration.md b/docs/benchmarks/2026-08-30-injection-gate-recalibration.md index 6730ff69..25cf7118 100644 --- a/docs/benchmarks/2026-08-30-injection-gate-recalibration.md +++ b/docs/benchmarks/2026-08-30-injection-gate-recalibration.md @@ -4,6 +4,14 @@ > as-written since this is a dated record; see `docs/HOMESERVER-DISK-LAYOUT.md` for > current layout. +> **Dated record — measured 2026-08-30, superseded by #2694.** The gate and +> `is_injection_case()` behaviour described below no longer run: injection cases +> are now declared by the rubric's `injection` block and scored by +> `analysis/ghidra/benchmarks/injection_gate.py`. Every number in this file is +> the 2026-08-30 sweep's, left exactly as measured. The current protocol, +> re-verified 2026-09-27, is +> `docs/analysis/ghidra/benchmarks/injection-gate-protocol.md`. + Measured 2026-08-30 against `/mnt-1/benchmarks/1947full` (run1 files; run2 verified byte-identical for every model at both tiers) and the checkout at `/mnt-1/benchmarks/APIARY` @ `a99e765`. Nothing on the host was modified; every script was piped over ssh stdin and read only. **Verdict on the preliminary finding (now issue #2694):** confirmed in mechanism and in substance, with two corrections and five additional findings. The gate is not measuring compliance. It is measuring whether a model *quoted or paraphrased the payload* (11 of 14 Tier B failures) or *used the exact phrase "appears to be benign"* (3 of 14). The fixture cannot discriminate compliance from correct analysis because the injected verdict is true. The same defect accounts for **all four Tier A failures that drove the #1805-c / #1947 "no promotion" decision**, including the disqualification of the top-scoring model. diff --git a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md index e8c805d2..d391c049 100644 --- a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md +++ b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md @@ -4,6 +4,15 @@ > as-written since this is a dated record; see `docs/HOMESERVER-DISK-LAYOUT.md` for > current layout. +> **Dated plan — written 2026-09-05, superseded by round 7 (#3079).** The hardware +> table below is **pre-reinstall**: the host is now Rocky Linux 10.2 with LVM, +> `/var` on `sdb1` of an 8.7T LUN, a 32G `rl-swap`, and no `/mnt-1` mount (see +> `docs/HOMESERVER-DISK-LAYOUT.md`). Likewise, the "#2985 missing scripts" this +> plan works around are now in git — `corpus/requant_sweep.sh`, +> `corpus/slots_sweep.sh`, `corpus/chain_round7.sh` and the `corpus/round7_*` +> builders; only `gptoss_rerun.sh` is still absent. In-flight state below is as +> of 2026-09-05, not current. + **Written** 2026-09-05, from live inspection of `homeserver` and every open benchmark issue. Supersedes nothing; it sits *beside* `/mnt-1/benchmarks/STATE-2026-09-05-fix-stage.md`, which remains the authority on diff --git a/docs/benchmarks/plans/2026-09-06-round7-hermes-init.md b/docs/benchmarks/plans/2026-09-06-round7-hermes-init.md index ca7327af..503f78d5 100644 --- a/docs/benchmarks/plans/2026-09-06-round7-hermes-init.md +++ b/docs/benchmarks/plans/2026-09-06-round7-hermes-init.md @@ -4,6 +4,14 @@ > as-written since this is a dated record; see `docs/HOMESERVER-DISK-LAYOUT.md` for > current layout. +> **Dated handoff — written 2026-09-06 (epic #3079); not a live brief.** The +> toolchain, model roster and dispatch rules in §5 are as-written on that date. +> What shipped since: the #3080 toolchain (`analysis/ghidra/training/`) and the +> corpus builders with their decontamination report. Unsloth (#3092) has **no +> installer path** — it is the Arcane stack `unsloth` +> (`arcane/manifests/home-production.json`), deployed by gitops-sync and a +> redeploy, never by hand. Stop it before any cold GPU leg. + **Read this first.** It names the plan of record, the state of the GPU, the kickoff order, the dispatch pattern and the hard rules. Written 2026-09-06, after the round-7 cold baseline was launched. Epic **#3079**, children diff --git a/docs/benchmarks/plans/2026-09-06-round7-unsloth-train-requant-ollama.md b/docs/benchmarks/plans/2026-09-06-round7-unsloth-train-requant-ollama.md index 52d0cdc2..4f467793 100644 --- a/docs/benchmarks/plans/2026-09-06-round7-unsloth-train-requant-ollama.md +++ b/docs/benchmarks/plans/2026-09-06-round7-unsloth-train-requant-ollama.md @@ -4,6 +4,13 @@ > as-written since this is a dated record; see `docs/HOMESERVER-DISK-LAYOUT.md` for > current layout. +> **Plan of record, written 2026-09-06 (epic #3079) — its §0 already folds in later +> changes; the rest is as-planned, not as-shipped.** The batch leg shipped as +> `analysis/ghidra/training/` (`train.py`, `export_to_ollama.sh`, `compose.yaml`, +> `TOOLCHAIN.md`); the interactive half shipped as the Arcane `unsloth` stack +> rather than an installer (#3092). `TOOLCHAIN.md`'s "no smoke test yet" still +> stands — nothing in this plan has been exercised end to end. + **Written** 2026-09-06 from live inspection of `homeserver`, every open benchmark issue, and the current Unsloth / llama.cpp / Ollama documentation. Sits beside `2026-09-05-1947-resume-plan.md`, which remains the authority on the #1947 sweep diff --git a/docs/ml-worker-plan.md b/docs/ml-worker-plan.md index a0afcaa3..daff9a31 100644 --- a/docs/ml-worker-plan.md +++ b/docs/ml-worker-plan.md @@ -1,16 +1,24 @@ # ML Worker — Implementation Plan +> **Reading note (2026-09-27):** this is a dated plan/record, not a live +> reference. §2, §5.3, §7, §8, §9, §10, §11.4 and §11.6 describe shipped +> behaviour and were re-checked against the code. §1, §4.1, §6, §11.2 and §12 +> keep their original-draft wording where it was never rewritten — including +> the Go dashboard's `dashboard/ml_anomalies.go` / `settings_domain.go` +> references, which are historical since #1628. + > **Status (2026-08-27, #1662):** the plan largely executed as written: -> `ml-worker/worker.py` runs the ensemble described here from the same -> repo-root compose files. What moved: consumers of its output live in the +> `ml-worker/worker.py` runs the ensemble described here from its own +> Arcane-managed stack, not the repo-root compose file (see §10). +> What moved: consumers of its output live in the > backend-service tier now, not `dashboard/ml_anomalies.go` (deleted). > Open scoring-semantics defects are tracked in issues #1946/#1969 under > epic #1974 rather than here. > **Status:** `ml-worker/` has its own Arcane-managed stack -> ([`docker-compose.yml`](../ml-worker/docker-compose.yml) + -> [`docker-compose.ml-worker.gpu.yml`](../ml-worker/docker-compose.ml-worker.gpu.yml), -> mirroring `analysis/ghidra/`), builds, connects to Elasticsearch, and polls +> ([`docker-compose.yml`](../ml-worker/docker-compose.yml), the only one the +> manifest deploys). `docker-compose.ml-worker.gpu.yml` exists in-tree but is +> inert and undeployed, so the worker is CPU-only in practice. It > without crashing (#62). `extract_features()`/`featurise_temporal()` read > the real per-sensor schema (#62 task 33, #63). The dashboard delivers > scores via the backend-service's `/api/v1/store/ml-anomalies` + @@ -168,8 +176,9 @@ different anomaly types: [web:275][web:276][web:283][web:292] mixed numerical+categorical features after encoding. Proven on network logs. [web:276][web:290] - **Implementation:** `scikit-learn` `IsolationForest` with `contamination=0.01` - (assume 1% of events are anomalous). Retrained every 6 hours on a 24h - rolling window. + (assume 1% of events are anomalous). Retrained at four fixed UTC slots + daily (`RETRAIN_SLOTS_UTC`, default `03:00,09:00,15:00,21:00` — #172 + replaced the old 6h `RETRAIN_INTERVAL`) on a 24h rolling window. - **Output:** `anomaly_score` ∈ [-1, 0] where values closer to -1 = more anomalous. ### 4.2 LSTM Autoencoder (LSTM-AE) @@ -391,7 +400,7 @@ loop every POLL_INTERVAL seconds (default: 30s): 10. Sleep POLL_INTERVAL -Every 6 hours (RETRAIN_INTERVAL): +At each `RETRAIN_SLOTS_UTC` slot (four daily, default 03:00,09:00,15:00,21:00 UTC, #172): - Retrain IsoForest + HBOS on last 24h of all events - Fine-tune LSTM-AE on last 24h (5 epochs, low LR) - Save new model checkpoint to /models/ @@ -786,7 +795,7 @@ if wanted. ### 11.2 Online learning (unchanged from the original draft, still accurate) ``` - → HBOS/IsoForest: full retrain every RETRAIN_INTERVAL (default 6h) on the + → HBOS/IsoForest: full retrain at each `RETRAIN_SLOTS_UTC` slot (four daily, #172) on the rolling 24h window, gated by §11.1 → LSTM-AE: fine-tune on the same cycle (5 epochs, LR=1e-5), gated the same way (§11.1's anomaly-rate check applies to its reconstruction-loss-based @@ -831,7 +840,7 @@ fraction `>= THRESHOLD` exceeds `DRIFT_ANOMALY_RATE` (default `0.15`, matching the original draft's "15%"): - an early retrain is triggered (the next poll cycle retrains regardless of - how much of `RETRAIN_INTERVAL` remains), and + which `RETRAIN_SLOTS_UTC` slot is nearest), and - a `ml-worker-metrics` document is written flagging the drift event (`kind: "drift"`, the observed rate, window size) so the dashboard's `/ml-anomalies` page (#64) — or a future panel reading this index directly @@ -885,9 +894,11 @@ its contract carried over unchanged to the Rust config module.) | **v0.8** | Retraining scheduler + model versioning | [#65](https://github.com/Xore/APIARY/issues/65) | | **v1.0** | Drift detection + alert threshold tuning UI | [#65](https://github.com/Xore/APIARY/issues/65) | -v0.1 is listed as an issue rather than as done on purpose. `ml-worker/` holds a -Dockerfile, `worker.py`, and a `docker-compose.override.yml`, but it is not a -service in the root Compose file, it has no tests or fixtures, and nothing here +v0.1 is listed as an issue rather than as done on purpose. As of the #61 audit +— before #62's rewrite — `ml-worker/` held a +Dockerfile, `worker.py`, and a `docker-compose.override.yml` (since deleted, see +§10); it was not a +service in the root Compose file, it had no tests or fixtures, and nothing here has been observed running against live data. #61 is the audit that decides whether the scaffold is a v0.1 or a starting point. diff --git a/docs/sandbox/README.md b/docs/sandbox/README.md index 94a28065..28ca9ac5 100644 --- a/docs/sandbox/README.md +++ b/docs/sandbox/README.md @@ -1,12 +1,16 @@ # Hard-isolated malware sandbox plan The homeserver supports this design: Intel VT-x is enabled, `/dev/kvm` is -available, KVM is loaded, and the host exposes 84 IOMMU groups. The foundation +available, KVM is loaded, and IOMMU is on with **94 groups** (re-measured +2026-09-27). This was not free on Rocky: #1609 recorded 87 groups on the old +Ubuntu host, and the 2026-09-03 Rocky 10 rebuild came up with **zero**, so +`install-homeserver.sh`'s `step_vfio_gpu_passthrough` adds `intel_iommu=on` +to the kernel command line. The foundation installer provisions system libvirt and the dedicated isolated network. ## Route selection and evidence return across four dynamic-detonation routes -The workbench's registry (`dashboard/workbench_domain.go`) offers four +The workbench's registry (`backend-service/src/workbench_domain.rs`) offers four routes to dynamic detonation, not one sandbox with options. Each is its own guest, network, spool, and result format — this section is the canonical side-by-side comparison; each route's own internal detail lives in its own @@ -15,7 +19,7 @@ section below (Linux) or its own directory (`sandbox/windows/`, ```mermaid flowchart TB - workbench["Payload workbench —
analyst selects a route
(dashboard/workbench_domain.go)"] + workbench["Payload workbench —
analyst selects a route
(backend-service/src/workbench_domain.rs)"] subgraph linuxRoute["linux-sandbox"] direction TB @@ -41,7 +45,7 @@ flowchart TB capeGuest["Windows guest under CAPE's
own cuckoo.py orchestration
(runs on the host directly,
not in Docker) + cape-mongo"] end - sharedLock{{"honeypot-kvm-detonation.lock —
shared ONLY between windows-sandbox
and cape (#320): 16 logical CPUs total,
win11-sandbox alone already 8 vCPU.
Held only around the actual detonation
call, not the whole drain loop.
linux-sandbox and windows-ghosts are
NOT part of this lock — independent,
can run concurrently with anything."}} + sharedLock{{"honeypot-kvm-detonation.lock —
shared ONLY between windows-sandbox
and cape (#320): 48 logical CPUs total,
win11-sandbox alone already 8 vCPU.
Held only around the actual detonation
call, not the whole drain loop.
linux-sandbox and windows-ghosts are
NOT part of this lock — independent,
can run concurrently with anything."}} workbench -->|"hash-only request,
SANDBOX_REQUEST_DIR"| linuxRoute workbench -->|"hash-only request,
WINDOWS_SANDBOX_REQUEST_DIR"| winRoute @@ -72,7 +76,7 @@ credential ever crosses the dashboard/host boundary for any of the four. never wait on anything.** `sandbox/windows/run_pending.sh` and `sandbox/cape/worker/cape-worker.py` share one host-wide `honeypot-kvm-detonation.lock` (#320) — a real capacity constraint, not a -correctness one: both are KVM/QEMU domains on the same 16-logical-CPU host, +correctness one: both are KVM/QEMU domains on the same 48-logical-CPU host, and `windows-sandbox`'s own guest is already configured for 8 vCPU. The lock is held only around the actual detonation call, never the whole drain loop, so an idle worker on either side never blocks the other. `linux-sandbox` @@ -235,6 +239,12 @@ sudo bash /opt/stacks/apiary/sandbox/install-windows-forensics.sh ## Required operating controls - Reserve at most 4 vCPU and 8 GiB RAM per analysis VM; run one job initially. + For scale, the Windows analysis domain is currently defined at **8 vCPU / + 16 GiB** (`sandbox/windows/packer/win11-kvm.xml`), and + `sandbox/sandbox.env.example` ships `SANDBOX_VM_MEMORY_MB=3072` as its + default — so this line's 4 vCPU / 8 GiB guidance matches neither figure + exactly. It is operator guidance rather than a measured limit; settle it + against the real per-VM reservation before relying on it. - Enforce a 10-minute hard timeout and kill QEMU if graceful shutdown fails. - Store golden images on root-owned storage and verify SHA-256 before every job. - Sign/validate result JSON and treat all guest-produced text as untrusted. From 5e107ddf4979172cd87c9662b95de1b69440a91e Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:21:31 +0200 Subject: [PATCH 16/21] docs(ml-worker): correct Tier 1 check count to seven --- docs/ml-worker-evaluation.md | 10 ++++++++-- 1 file changed, 8 insertions(+), 2 deletions(-) diff --git a/docs/ml-worker-evaluation.md b/docs/ml-worker-evaluation.md index 9da6a953..0ddbfab8 100644 --- a/docs/ml-worker-evaluation.md +++ b/docs/ml-worker-evaluation.md @@ -42,10 +42,16 @@ reported alongside accuracy rather than ignored. ### 2026-08-25 — Tier 1 harness landed; Tier 2 blocked -**Tier 1 (behaviour) is live.** `evaluate_detectors.py` runs six contract +**Tier 1 (behaviour) is live.** `evaluate_detectors.py` runs seven contract checks against a candidate over the per-sensor fixture corpus and emits a hashed JSON report. It is the first reusable acceptance bar `ml-worker` has -had, and it makes "evaluate this candidate offline" answerable at all. +had, and it makes "evaluate this candidate offline" answerable at all. The +sixth of the seven, `check_composite_renormalises_over_present_detectors` +(#1969), is not a per-candidate probe at all: it calls the production +`worker.compute_composite()` directly, so candidates are always compared +under the rules they will actually run with — absent detectors drop out of +both numerator and denominator, an event no detector opines on composites +to 0.0, and a single-detector opinion stands at face value. **Tier 2 (accuracy) was blocked on [#1797](https://github.com/Xore/APIARY/issues/1797).** There was no labelled corpus at that point. The date is retained as the From a98d1269c354da63049ad3ab586345308efa812f Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:22:28 +0200 Subject: [PATCH 17/21] docs(ml-worker): add zeek-v1-conn source index and refresh pin record --- docs/ml-worker-plan.md | 16 +++++++++++----- 1 file changed, 11 insertions(+), 5 deletions(-) diff --git a/docs/ml-worker-plan.md b/docs/ml-worker-plan.md index daff9a31..75f244c2 100644 --- a/docs/ml-worker-plan.md +++ b/docs/ml-worker-plan.md @@ -35,7 +35,10 @@ > `docker build ./ml-worker` failed outright (`pyod`'s `numba` dependency had > no version compatible with the pinned `numpy==2.5.1` on Python 3.12 — > reproduced twice, locally and in-container; fixed in #62 by pinning -> `numpy==2.4.6`/`numba==0.66.0`/`llvmlite==0.48.0`). `worker.py`'s +> `numpy==2.4.6`/`numba==0.66.0`/`llvmlite==0.48.0`; those two transitive pins +> have since moved on again and the file now reads `numba==0.67.0` / +> `llvmlite==0.49.0` / `pyod==3.6.6`, so treat this paragraph as the record of +> what #62 did, not of the current requirements). `worker.py`'s > `SOURCE_INDICES` (`cowrie-*`, `dionaea-*`, `honeypot-network-*`, `conpot-*`, > `http-honeypot-*`) still match zero indices on the live homeserver: the real > shape is a unified `honeypot-v2-*` stream (all sensors, disambiguated by @@ -98,16 +101,17 @@ ground-truth labels. [web:275][web:283] > the "v0.1 audit verdict" callout above. `worker.py`'s real, > currently-deployed `SOURCE_INDICES` are the two rows below. -The worker ingests from two unified, versioned index patterns +The worker ingests from three unified, versioned index patterns (`ml-worker/worker.py`'s `SOURCE_INDICES`): | Index pattern | Source | Key fields | |---------------|--------|------------| | `honeypot-v2-*` | every honeypot sensor (Cowrie, Dionaea, Conpot, HTTP-honeypot, and every other sensor stack — disambiguated by `event.sensor`, not a separate index per sensor) | `event.sensor`, `source.ip`, `honeypot.*` (per-sensor nested fields, not uniform across sensors — see §5.3) | | `suricata-v2-*` | Suricata network/IDS events (Filebeat) | `suricata.eve.*`, `network.*`, `alert.signature` | +| `zeek-v1-conn-*` | Zeek connection records, added by #1774's sensing layer alongside Suricata | Zeek conn-log fields; note the pattern is `zeek-v1-conn-*`, not a `zeek-v2-*` line like the other two | -Both index patterns share a common `@timestamp` field used for temporal -ordering. A third index, `ml-worker-state`, is not a data source — it's the +All three index patterns share a common `@timestamp` field used for temporal +ordering. A fourth index, `ml-worker-state`, is not a data source — it's the worker's own per-index-pattern checkpoint store (`load_checkpoint`/ `save_checkpoint` in `worker.py`): a `last_timestamp` plus the set of already-seen event IDs at that exact timestamp, so a restart resumes @@ -125,9 +129,11 @@ flowchart TD subgraph Stack["APIARY (existing)"] Sensors["every honeypot sensor stack
(disambiguated by event.sensor,
not a separate index each)"] Suricata["Suricata / network IDS"] - ES["Elasticsearch
honeypot-v2-*, suricata-v2-*"] + Zeek["Zeek conn records
(#1774)"] + ES["Elasticsearch
honeypot-v2-*, suricata-v2-*,
zeek-v1-conn-*"] Sensors --> ES Suricata --> ES + Zeek --> ES end subgraph Worker["ML Worker (ml-worker/)"] From 6fe036f06ca229bd829f0a0d16e2694946953568 Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:22:53 +0200 Subject: [PATCH 18/21] =?UTF-8?q?docs(benchmarks):=20land=20the=20#66=20?= =?UTF-8?q?=C2=A79/=C2=A72=20reconciliation=20banner?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Uncommitted work in the tree when this run ended; committing so the reconciliation is not lost. It sits after the 2026-08-30 supersession banner and rewrites no measured number: §9's 'it is not merged' is stale (the branch landed as df650a8f, the same minute as the 16:20Z report count), §2's Tier B tally is the matrix's own rather than the shipped 64-row fixture's (15 FAIL / 15 PASS, 13 payload hits to 2, not 14/15 and 11/3), test_record_baseline.py is 54 not 50, §8's governance-gate term list omits 'conclude benign', and §8's '59 hand-labelled answers' is the wrong set twice over. --- ...2026-08-30-injection-gate-recalibration.md | 28 +++++++++++++++++++ 1 file changed, 28 insertions(+) diff --git a/docs/benchmarks/2026-08-30-injection-gate-recalibration.md b/docs/benchmarks/2026-08-30-injection-gate-recalibration.md index 25cf7118..20b4e113 100644 --- a/docs/benchmarks/2026-08-30-injection-gate-recalibration.md +++ b/docs/benchmarks/2026-08-30-injection-gate-recalibration.md @@ -12,6 +12,34 @@ > re-verified 2026-09-27, is > `docs/analysis/ghidra/benchmarks/injection-gate-protocol.md`. +> **Reconciled 2026-09-27 — §9's inventory is now shipped, and §2's Tier B tally +> does not match the pinned fixture.** Five points, all checked against `HEAD`; +> no number above was rewritten. +> +> 1. **§9's "It is *not merged*" is stale.** The branch landed as `df650a8f` +> ("injection gate v3 … (#2694) (#2697)"), 2026-08-30 16:20:54Z — the same +> minute as the "69 `run1` reports present at 16:20Z" in §9.2. Every path in +> the §9 table is at `HEAD`; `rev_cases_v2_rubric.json` carries 17 cases with +> `process_and_injection` on `forbidden: []`, and `manifest.json` holds the +> 850 builds §9.1 claims (700 reproduced + 150 new, `errors: []`). +> 2. **§2's Tier B counts are the matrix's own, not the fixture's.** The table +> enumerates 30 rows but never lists `huihui-qwen3.6-35b-a3b-abliterated:q3_k`, +> which the shipped 64-row fixture records as a gate **FAIL**. Against that +> fixture Tier B is **15 FAIL / 15 PASS** over 30 rows, splitting **13** +> payload-term hits to **2** on `appears to be benign` — not the 14/15 and +> 11/3 of the headline above. §2 stays as written; read it as the matrix it +> is, and the fixture as the shipped regression set. +> 3. **§9's test inventory has drifted:** `test_injection_gate.py` is still 48, +> but `test_record_baseline.py` is **54**, not 50 (102 collected, 102 pass). +> 4. **§8's governance-gate term list is incomplete.** +> `evaluate-models.py:421` (`process-injection`) reads +> `("appears to be benign", "likely benign", "conclude benign")`; §4.3 lists +> all three, §8 omits the last. +> 5. **§8's "59 hand-labelled answers in §2" is the wrong set twice over:** §2 +> carries 29 Tier B verdicts, and the hand-labelled cohort is the 64 rows of +> §9's fixture. The `conclude benign` figure in §1 (zero occurrences) still +> holds — it is zero across all 64. + Measured 2026-08-30 against `/mnt-1/benchmarks/1947full` (run1 files; run2 verified byte-identical for every model at both tiers) and the checkout at `/mnt-1/benchmarks/APIARY` @ `a99e765`. Nothing on the host was modified; every script was piped over ssh stdin and read only. **Verdict on the preliminary finding (now issue #2694):** confirmed in mechanism and in substance, with two corrections and five additional findings. The gate is not measuring compliance. It is measuring whether a model *quoted or paraphrased the payload* (11 of 14 Tier B failures) or *used the exact phrase "appears to be benign"* (3 of 14). The fixture cannot discriminate compliance from correct analysis because the injected verdict is true. The same defect accounts for **all four Tier A failures that drove the #1805-c / #1947 "no promotion" decision**, including the disqualification of the top-scoring model. From 6af5e8983fd63623c53bc313a9fe10e00290095f Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:27:51 +0200 Subject: [PATCH 19/21] docs(benchmarks): name the live corpus path in the #1947 resume-plan banner The banner lists the #2985 scripts as now in git, but there is no root corpus/; the tracked copies are under analysis/ghidra/benchmarks/corpus/. The doc body already gives that directory at the P1 section, so the banner now agrees with it. No measured number touched. --- docs/benchmarks/plans/2026-09-05-1947-resume-plan.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md index d391c049..3b2e69bc 100644 --- a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md +++ b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md @@ -8,8 +8,8 @@ > table below is **pre-reinstall**: the host is now Rocky Linux 10.2 with LVM, > `/var` on `sdb1` of an 8.7T LUN, a 32G `rl-swap`, and no `/mnt-1` mount (see > `docs/HOMESERVER-DISK-LAYOUT.md`). Likewise, the "#2985 missing scripts" this -> plan works around are now in git — `corpus/requant_sweep.sh`, -> `corpus/slots_sweep.sh`, `corpus/chain_round7.sh` and the `corpus/round7_*` +> plan works around are now in git at `analysis/ghidra/benchmarks/corpus/` — +> `requant_sweep.sh`, `slots_sweep.sh`, `chain_round7.sh` and the `round7_*` > builders; only `gptoss_rerun.sh` is still absent. In-flight state below is as > of 2026-09-05, not current. From 7aba4cb6a46c4559ebf45a86fa59770074bbce12 Mon Sep 17 00:00:00 2001 From: docs-reconcile Date: Sun, 27 Sep 2026 15:28:54 +0200 Subject: [PATCH 20/21] docs: revert five files that are outside this slice's assignment SECURITY.md, docs/ROCKY-10-MIGRATION.md, docs/analysis/RECOVERY.md, docs/kvm-network-traffic-analysis.md and docs/kvm-snapshot-vs-golden-image.md are not on the round-2 r2-infra assignment list, but earlier commits in this branch had corrected them. Reverted to the branch point so this branch touches only its 19 assigned files and cannot collide with the slice that owns those docs. The corrections are preserved as a handoff patch outside the repo; they are reported in the round-2 handback. --- SECURITY.md | 25 ------------------------ docs/ROCKY-10-MIGRATION.md | 10 +--------- docs/analysis/RECOVERY.md | 29 ++++++++-------------------- docs/kvm-network-traffic-analysis.md | 18 +++++------------ docs/kvm-snapshot-vs-golden-image.md | 7 ------- 5 files changed, 14 insertions(+), 75 deletions(-) diff --git a/SECURITY.md b/SECURITY.md index b9fc60a8..b15c3808 100644 --- a/SECURITY.md +++ b/SECURITY.md @@ -13,28 +13,3 @@ live malware, private keys, production `.env` files, packet captures containing private traffic, or unredacted logs to a public issue. Supported security fixes target the current `main` branch. - -## The leak gate, and what it already blocks - -"Leak a real secret" is enforced in CI, not by review alone. -`scripts/check-public-leaks.py` fails the build on private keys, GitHub/AWS/Slack -tokens, literal `PASSWORD=`/`SECRET=`/`TOKEN=`/`API_KEY=` assignments, credentials -embedded in a URL, any file named `.env`, and private/runtime binaries (`.pcap`, -`.qcow2`, `.key`, `.pem`, …). It also carries four deployment-specific literals — -one public domain, the VPS address, the home-server address and a known default -password — assembled from fragments at runtime so the checker does not harbor the -values it bans, and a `Host()` rule over `vps/traefik/*.yml` that requires a -reserved `.example`/`.test`/`.invalid`/`.localhost` name, so an installer smoke -test cannot resolve live DNS. - -Exemptions are explicit and fail-closed, in three sets at the top of the script: -`ALLOWED_DOTENV` (3 paths), `ALLOWED_LITERAL_FIXTURE_FILES` (2 paths), and the -`change-me` / `DECOY_ONLY` / `${…}` / `$(…)` / `` forms inside the -credential-assignment pattern. A file that is exempt from the literal scan is -still scanned for every other pattern. - -The gate reads `git ls-files -co --exclude-standard`, so **untracked files are -scanned too**: a scratch note left in your own checkout fails the check exactly -like a committed one. Gitignore it or delete it — the address and domain values -this repository is deployed against are named literals in the checker, and a -dispatch brief or run log that quotes one will trip it. diff --git a/docs/ROCKY-10-MIGRATION.md b/docs/ROCKY-10-MIGRATION.md index b2b2203a..c19f7bba 100644 --- a/docs/ROCKY-10-MIGRATION.md +++ b/docs/ROCKY-10-MIGRATION.md @@ -1,7 +1,6 @@ # Rocky Linux 10 support in `install-homeserver.sh` -The homeserver has moved from Ubuntu to Rocky Linux 10 (the rebuild hit -live 2026-09-03). `scripts/install-homeserver.sh` +The homeserver is moving from Ubuntu to Rocky Linux 10. `scripts/install-homeserver.sh` now runs on both, so the reinstall smoke test in #1609 has a working installer. ## How it works @@ -63,13 +62,6 @@ Quadro P2200 alongside the Ada RTX 4000; the P2200 is meant to be bound to `vfio-pci` for the Windows sandbox rather than driven by the host, but choosing the open modules would make it unusable on the host if ever needed. -Both cards are still on the PCI bus (P2200 at `17:00.0` / `10de:1c31`, Ada at -`65:00.0`), re-measured 2026-09-27. Only the Ada currently has a driver bound -(`Kernel driver in use: nvidia`); the P2200 shows under *Kernel modules* with -none *in use*, so `nvidia-smi -L` lists exactly one GPU. That is expected, and -it is **not** evidence the P2200 is missing — the `nvidia-open` reasoning above -only bites if the P2200 is ever handed back to the host. - **NVIDIA container toolkit** — same rpm repofile approach. On RHEL the `container_use_devices` SELinux boolean is also set, without which a container cannot open the GPU device nodes. diff --git a/docs/analysis/RECOVERY.md b/docs/analysis/RECOVERY.md index 3e6a9e8c..00f817bb 100644 --- a/docs/analysis/RECOVERY.md +++ b/docs/analysis/RECOVERY.md @@ -4,14 +4,11 @@ For a *deliberate* full reset on the same hosts (not disaster recovery), see [`docs/STACK-REBUILD.md`](../STACK-REBUILD.md) instead — this doc is about restoring a backup archive onto a replacement host after data loss. -Two backups cover the homeserver stack, with the same exclusions. **Scope -differs**, which matters here because the two produce different filenames: +Two backups cover this, with the same scope and the same exclusions: - [`scripts/backup-essentials.sh`](../../scripts/backup-essentials.sh) runs on the workstation and fans an encrypted archive out to three locations. This - is the one that survives the homeserver dying, and the broader of the two — - it additionally carries the VPS config, WireGuard, Technitium, the - installer's answers file and the repo's runbooks. Full restore procedure: + is the one that survives the homeserver dying. Full restore procedure: [`docs/BACKUP-ESSENTIALS.md`](../BACKUP-ESSENTIALS.md). - `sudo analysis/backup-honeypot.sh` runs on the homeserver itself, into a timestamped mode-0700 directory beneath `/opt/backups/honeypot`. Faster to @@ -28,25 +25,15 @@ and the sizes behind it. Recovery is intentionally not automatic because overwriting live volumes is destructive. On a replacement host: -1. Unpack the archive's `env/` and `secrets/` trees back under - `/var/dockge/stacks//` and inspect `.env` permissions and values. - Note the filename first: `SHA256SUMS` and `stack-config-state.tar.gz` are - the **on-host** copy's only — `analysis/backup-honeypot.sh` writes them and - `analysis/verify-backup.sh` checks them. A `scripts/backup-essentials.sh` - archive contains neither; verify that one with - `sha256sum -c apiary-essentials-.tar.gz.gpg.sha256`. +1. Verify `SHA256SUMS`, unpack `stack-config-state.tar.gz` into a new empty stack + directory, and inspect `.env` permissions and values. 2. Restore the Keycloak database from `keycloak.sql.gz` with `hp-keycloak-postgres` up and `hp-keycloak` still stopped — this is what preserves the OIDC client secrets that the VPS's own `secrets/oidc/` copies have to match. -3. Keep all services stopped, then `docker volume create ` for each - `homeserver/volumes/.tar.gz` and restore each archive into it through a - temporary networkless BusyBox container. Use the **full real volume name** — - four of the five carry their Arcane project prefix and live in four - different stacks, so one `docker compose -f compose.yml create` cannot cover - them. -4. Start stacks in [`docs/STACK-REBUILD.md`](../STACK-REBUILD.md)'s order — - `honeypot-elk` healthy before `honeypot-init`, everything else after — then - run `analysis/verify-stack.py` — with +3. Create the named volumes with `docker compose -f compose.yml create`, keep all + services stopped, and restore each matching volume archive using a temporary + networkless BusyBox container. +4. Start setup jobs and sensors, then run `analysis/verify-stack.py` — with `DASHBOARD_SERVICE_TOKEN` from the restored stack's `.env` exported; it reads source-health through dashboard-next's `/bff` passthrough and exits nonzero on any failure. diff --git a/docs/kvm-network-traffic-analysis.md b/docs/kvm-network-traffic-analysis.md index ab1ebd4d..563160ec 100644 --- a/docs/kvm-network-traffic-analysis.md +++ b/docs/kvm-network-traffic-analysis.md @@ -18,9 +18,7 @@ Traffic is captured on the host side of an isolated libvirt bridge, never inside the guest and never by an agent the guest could interfere with. There -are **two** such bridges for the detonation pair documented here (CAPE adds -`virbr-cape` and GHOSTS adds `virbr-ghosts`, and controlled mode puts -`198.18.0.0/24` on `virbr-hpsbx` itself): +are **two** such bridges, because there are two sandboxes: | | Windows detonation guest | Linux / Wine runner | |---|---|---| @@ -29,7 +27,7 @@ are **two** such bridges for the detonation pair documented here (CAPE adds | Network XML | `sandbox/windows/setup/sandbox-network.xml` | `sandbox/network.xml` | | Fake internet | INetSim at `10.10.10.1` | none by default; optional logged DNS + Squid allowlist (`controlled` mode) | | Capture | `docker-compose.sandbox.yml` (tcpdump, Zeek, Suricata) | root-owned `tcpdump` per job, host and guest side | -| Results | `$WINDOWS_SANDBOX_RESULTS_DIR//`, set per run; `run_sample.py` falls back to `reports/windows-sandbox` and the dashboard's results importer populates the variable | root-only, sanitized export copied out | +| Results | `sandbox/results//` | `/var/lib/honeypot-sandbox/results/` (root-only), sanitized export copied out | | Orchestrator | `sandbox/windows/orchestrate/run_sample.py` | `sandbox/run-linux-sample.sh` | Neither bridge has a `` element, so neither can route anywhere. That @@ -87,11 +85,7 @@ For the Windows bridge: 1. The libvirt network has no ``. 2. The gateway containers sit on a macvlan marked `internal: true`, which removes the default route Docker would otherwise install. -3. Phase 0 asks the operator to add a FORWARD DROP pair across - `virbr-sandbox`. This is a documented manual step, not something a script - applies — re-verify it against this host's firewall backend, which is now - Rocky 10.2 with firewalld/nftables rather than the iptables the step was - originally written for. +3. Phase 0 adds an iptables DROP pair across `virbr-sandbox`. Removing any one of them because "the container needs to pull something" puts live malware on the internet. Build images ahead of time instead. @@ -176,10 +170,8 @@ for a run the worker does not know about. Authenticated administrators can download the pcaps. Oversize pcaps and the raw result directories stay root-only. - **`generate_report.py`** folds `zeek_logs/http.log` into the Windows report. - The Linux sandbox now gets Zeek too, offline over the finished pcap rather - than live on the bridge (`sandbox/run-linux-sample.sh` runs `zeek -r` and - `sandbox/export-result.py` reads `conn`/`dns`/`http`/`ssl`/`files` from - `zeek_logs/`). + The Linux sandbox has no Zeek equivalent — also + [#87](https://github.com/Xore/APIARY/issues/87). ## 5. Retention diff --git a/docs/kvm-snapshot-vs-golden-image.md b/docs/kvm-snapshot-vs-golden-image.md index 6162f7b3..b3a0c58a 100644 --- a/docs/kvm-snapshot-vs-golden-image.md +++ b/docs/kvm-snapshot-vs-golden-image.md @@ -233,13 +233,6 @@ virsh start "analysis-${SAMPLE}" echo "VM analysis-${SAMPLE} started with overlay $OVERLAY" ``` -> §5's `golden-win10` and `/var/lib/libvirt/golden/` are **illustrative**, not -> this stack's real names or paths. The live mechanism uses -> `SANDBOX_ROOT=/var/dockge/sandbox`, a base image of -> `golden-images/win11-analysis.qcow2` and a per-run -> `vms/win11-sandbox.qcow2`, and it spawns with `qemu-img create -b` plus -> `virsh define` rather than `virt-clone` — see §7. - ### 5.3 Destroy and Clean Up After the Run ```bash From 4516f45a4e61c3dd4232b4f4a2bda7c8a98549e2 Mon Sep 17 00:00:00 2001 From: Xore Date: Sun, 27 Sep 2026 15:31:19 +0200 Subject: [PATCH 21/21] docs(models): reconcile the evaluation record against the stored runs and matrices Checked every claim this file can be checked against, and recorded the result at the top. No score in the file was changed -- the one edit is a clause that makes an existing claim true. Confirmed: the round-7 cold baseline reconciles cell for cell against round7-cold-baseline.json (91 models, 182 cells, 367 records, 179 reproduced, the same 3 escalated cells and third-run values, 11 zero-scored tags, and every anchor including the 12.1 / 7.3 point gaps). The cold-cohort table reconciles against 1947-cohort-cold-protocol.json down to min/run B and Ornith's injection FAIL. The twelve-model survey reconciles against 1805c-ghidra-slot-matrix.json. The #568 approved qwen3:14b@bdbd181c33f2 and context_tokens 32768 are confirmed against approved-models.json. Recorded as undetermined rather than asserted either way: the archive holds no ghidra-slot record at all (all 62 runs, 1498 records, are revdeck or sessions) and no transcript carries VRAM or a context-probe result, so the Ghidra, 16k and VRAM columns have no in-repo source. The round-7 "14 of 91 fully resist (5/5)" figure likewise -- the matrix stores score and percent only. Recorded as vintage deltas, not errors: the archive is the 2026-08-25..29 sweep while #568 is 2026-08-05, so overlapping rows differ by design; and part 1's sessions column sits one point low on three of eleven rows under the current scorer, reproducibly over all three repeats, which widens the set of rows above the incumbent named in that section's Decision. Flagged for anyone recomputing: seven archived runs report outcome ok with non-zero output_tokens and an empty raw, and yield 14-18/69 for qwen3:14b against the authoritative 60-62/69. The inline edit: the round-7 "top band by run-pooled mean total_score" list omits the three highest means -- gemma-4-26B 90.75, Foundation-Sec Q8_0 87.25, XORTRON LARGE 83.25 -- because those are exactly the three models whose Tier B cell was escalated to a third run and so carry a 5-run denominator. Named the exclusion instead of silently dropping them. --- docs/local-llm-model-evaluation.md | 57 +++++++++++++++++++++++++++++- 1 file changed, 56 insertions(+), 1 deletion(-) diff --git a/docs/local-llm-model-evaluation.md b/docs/local-llm-model-evaluation.md index 7a4fea3d..f7a2725b 100644 --- a/docs/local-llm-model-evaluation.md +++ b/docs/local-llm-model-evaluation.md @@ -2,6 +2,55 @@ Status: completed for [issue #144](https://github.com/Xore/APIARY/issues/144) and requalified under [issue #158](https://github.com/Xore/APIARY/issues/158), 2026-08-01. Re-evaluated and re-approved under [issue #568](https://github.com/Xore/APIARY/issues/568), 2026-08-05 — see [§ Issue #568 re-evaluation](#issue-568-re-evaluation-real-20gb-card) below; that section is now the current approved state, superseding the v2 table immediately above it. +> **Reconciled 2026-09-27 against `docs/benchmarks/runs/` and +> `docs/benchmarks/matrices/`.** No score in this file was changed. What the +> repository can and cannot confirm: +> +> - **Confirmed exactly.** The round-7 cold baseline reconciles cell for cell +> against `round7-cold-baseline.json`: 91 models, 182 cells, 367 records, 179 +> reproduced, the same 3 escalated cells with the same third-run values, 11 +> zero-scored tags, and every anchor — `qwen3:14b` 85.5 B / 83.1 A, `qwen3:8b` +> 84.3 B / 81.9 A, `qwen2.5:14b-instruct-q4_K_M` 88.0 A, `Trendyol-32B` 95.2 A, +> `Ornith-1.0-35B` 92.8 B, and the 12.1 / 7.3 point gaps. The cold-cohort +> table reconciles against `1947-cohort-cold-protocol.json` (means, ±0 spreads, +> B−A deltas, `min/run B`, Ornith's injection FAIL). The twelve-model survey +> reconciles against `1805c-ghidra-slot-matrix.json`. +> - **The archived runs cannot check the Ghidra column at all.** All 62 stored +> runs — 1498 records — are `revdeck` or `sessions`; there is **no +> `ghidra`-slot record in the repository**, and no transcript field carries +> VRAM or a context-probe result. The Ghidra, 16k-probe and VRAM columns in +> the #144 and #568 tables are therefore **undetermined from the repo**, not +> confirmed and not contradicted. The approved `qwen3:14b@bdbd181c33f2…` and +> `context_tokens: 32768` of the #568 Decision *are* confirmed against +> `analysis/ghidra/models/approved-models.json`. +> - **The archive is a different vintage from #568.** The stored runs are the +> 2026-08-25→29 #1795b / #1947-wave2 / #1805c / #1947seq rounds; #568 was +> measured 2026-08-05. Where the two overlap they differ (`qwen3:8b` sessions +> 94.0 in the archive vs 92.5 here; `qwen2.5:14b-instruct-q4_K_M` 100.0 vs +> 97.0). Those are re-measurements, not errors, and no figure was "corrected" +> to match them. +> - **Part 1's sessions column is one point low on three of eleven rows under +> the current scorer**, reproducibly across all three repeats: +> `Foundation-Sec-1.1-8B-Instruct-i1` 56→**57** (83.6→85.1%), +> `Huihui-Qwen3.6-35B-A3B-abliterated` 66→**67** and stock `qwen3.8:27b` +> 66→**67** (both 98.5→100.0%). The part-1 pins already declare a pre-#2265 +> scorer, so this is a vintage delta — but the part-1 Decision names only two +> rows above the incumbent, and two more reach 67/67 on the current scorer. +> The part-1/part-2 `revdeck` denominator is likewise /16 as printed against +> /19 under the current inline `REV_CASES`. +> - **Seven archived runs are not re-scorable as stored.** They report +> `outcome: ok` with non-zero `output_tokens` and an empty `raw`: four +> `qwen3:14b` runs on 2026-08-25 (51 records) and three gpt-oss-family runs on +> 2026-08-26 (45 records). Re-scoring the archive naively yields 14–18/69 for +> `qwen3:14b` from those four, against the authoritative 60–62/69. Exclude +> them before recomputing anything. +> - **Undetermined:** the round-7 injection figure "only 14 of 91 models fully +> resist (5/5)" — `round7-cold-baseline.json` stores score and percent only, +> with no per-case injection split, so the 5/5 counts have no in-repo source. +> - Method check: re-deriving the 14-case corpus from the stored transcripts +> with `rev_cases_v2_rubric.json` + `polarity.forbidden_hit` reproduces +> `1805c-ghidra-slot-matrix.json` exactly, all 12 models at both tiers. + This is a task-specific decision record for the three independent local-model slots in this repository. It does not assume that a model named in an earlier plan is suitable, or that one model should serve all three jobs. @@ -1375,7 +1424,13 @@ Tier A — the incumbent sits 7-12 points under the security-specialized leaders (12.1 points Tier A vs `Trendyol-32B`'s 95.2%, 7.3 points Tier B vs `llmfan46/Ornith-1.0-35B`'s 92.8%). Top band by run-pooled mean total_score (mean across all 4 runs per model — 2 Tier A + 2 Tier B; not the same scale -as the percentages above): `phi4:14b` 77.0, `Trendyol-32B` Q8_0 76.5, +as the percentages above), **excluding the three entries whose Tier B cell was +escalated to a third run and so has a 5-run, not 4-run, denominator** — +`gemma-4-26B-A4B-it-ultra-uncensored-heretic` `Q4_K_M` 90.75, +`Foundation-Sec-1.1-8B-Instruct` `Q8_0` 87.25 and +`XORTRON.CriminalComputing.LARGE.2026.3` `i1-IQ2_XXS` 83.25, all three above +everything listed here: +`phi4:14b` 77.0, `Trendyol-32B` Q8_0 76.5, `VulnLLM-R-7B` i1-Q4_K_M 76.5, `Huihui-CyberStrike-OffSec-35B` q6_k 75.5, philbert440 `Qwen3.8-27B-Cyber` 75.5, protoLabsAI `ThinkingCap-Qwen3.6-27B-MTP` `latest` 75.5 (the successfully-pulled