diff --git a/docs/ARCANE-GIT-SYNC.md b/docs/ARCANE-GIT-SYNC.md index 6b26dddc7..71d0f517e 100644 --- a/docs/ARCANE-GIT-SYNC.md +++ b/docs/ARCANE-GIT-SYNC.md @@ -1,6 +1,6 @@ # Arcane Git sync -How the 37 home-hosted stacks (31 that migrated under `arcane/home/` plus 6 +How the 39 home-hosted stacks (33 that live under `arcane/home/` plus 6 that were already self-contained and stayed at their existing path) get to the live host, replacing the old model of copying or symlinking top-level `docker-compose.*.yml` files into place. Everything here was confirmed live @@ -9,7 +9,10 @@ taken from Arcane's docs — see the risk that motivated that in "Version/API compatibility" below. Census numbers and the live sync-store state were re-verified 2026-08-27 against the pinned `v2.9.0` image and its own sqlite store (#2549): 31 in-tree directories + 6 self-contained = 37, exactly the -manifest's entry count. +manifest's entry count *at that date*. The manifest has since grown to 39 — +#2911 swapped the `pihole` entry for `technitium` and #3092 added `unsloth` +directly under `arcane/home/` — re-counted 2026-09-27 from +`arcane/manifests/home-production.json`: 33 in-tree + 6 self-contained = 39. ## The model @@ -18,19 +21,21 @@ Each stack gets its own **directory-aware Git sync**: Arcane clones the selected `compose.yml` (not just that one file) under `/var/dockge/stacks//`, and deploys it. The manifest at [`arcane/manifests/home-production.json`](../arcane/manifests/home-production.json) -is the single source of truth for which 37 stacks exist, what branch/path +is the single source of truth for which 39 stacks exist, what branch/path each syncs from, and any per-stack sync limits — `scripts/install-homeserver.sh`, CI, and this doc all read from it rather than maintaining separate lists. - The 32 `honeypot-*` stacks live under `arcane/home//`: their build context and git-tracked config were moved there from repository root (see each compose file's own `#1502` comment for what moved and why). + `unsloth` is the 33rd directory there — #3092 added it in-tree rather + than at a root path, so it never went through the #1502 move. - The 6 other stacks (`auth-events-worker`, `llm-worker`, `ml-worker`, - `analysis/ghidra`, `sandbox/ghosts`, `pihole`) were already self-contained - and stayed at their existing path — moving them would have broken real - references from `scripts/install-homeserver.sh`, CI workflows, and - `deploy.yml`'s own ghidra-worker resync step. See each one's own compose - file header for the specifics. + `analysis/ghidra`, `sandbox/ghosts`, `technitium`) were already + self-contained and stayed at their existing path — moving them would + have broken real references from `scripts/install-homeserver.sh`, CI + workflows, and `deploy.yml`'s own ghidra-worker resync step. See each + one's own compose file header for the specifics. - `honeypot-arcane` itself is **not** in the manifest and never will be — syncing the thing that has to already be running before any sync can happen is a bootstrap loop, not a simplification. It stays @@ -58,15 +63,27 @@ been provisioned once. `scripts/install-homeserver.sh`'s `step_arcane_import_stacks` reads the manifest and creates one `POST /environments/0/gitops-syncs` per matching entry (environment `0` is Arcane's single "Local Docker" environment on a -one-host deployment). Its selection filter matches every `honeypot-*` -entry — and, since #1505, three of the six non-`honeypot-*` stacks too: -`auth-events-worker`, `llm-worker` and `ml-worker` are imported by that -step as well (each confirmed to have no host-local state beyond `.env`). -The other three keep their dedicated installer steps for reasons specific -to each: `pihole`'s non-`.env` host state, `analysis/ghidra`'s conditional -GPU compose overlay, and `sandbox/ghosts`'s Arcane build-context -limitation (#1506) — see the script's own Phase 8 header comment for the -reasoning behind each. +one-host deployment). Its selection filter matches all 32 `honeypot-*` +entries — and, since #1505, three of the seven non-`honeypot-*` entries +too: `auth-events-worker`, `llm-worker` and `ml-worker` are imported by +that step as well (each confirmed to have no host-local state beyond +`.env`). That is 35 of the manifest's 39 entries. The other three the +installer provisions itself keep their dedicated steps for reasons specific +to each: `technitium`'s non-`.env` host state (its `config/` directory needs +non-root ownership, #2911 — this is the step `pihole` used to have), +`analysis/ghidra`'s conditional GPU compose overlay, and `sandbox/ghosts`' +s Arcane build-context limitation (#1506) — see the script's own Phase 8 +header comment for the reasoning behind each. + +The 39th entry, `unsloth` (#3092), is the one the installer does not reach +at all: it matches neither arm of the filter, and the script has no +`step_unsloth_*` of its own. That is not a defect — `arcane/home/unsloth/compose.yml`'s +own header documents it as an operator-started stack ("Deployed through +Arcane like every other homeserver stack … Never `docker compose up` by +hand"), deliberately carrying no `restart:` policy so the cold-benchmark legs +in `analysis/ghidra/training/` can have the card to themselves. It is +covered where fleet-wide operations are concerned — `docs/STACK-REBUILD.md` +lists it in both reset loops — just not by the from-scratch import path. To import (or re-import) by hand instead, `POST` the manifest's entries to `/environments/0/gitops-syncs/import` — the bulk-import shape matches the @@ -106,16 +123,19 @@ letting Arcane's own directory check block the whole import. not track** — not just `.env` and `secrets/`. Confirmed live: `pihole` also keeps its DNSCrypt resolver config and its own Pi-hole database directly under its own top-level directory (`dnscrypt-proxy/`, - `etc-pihole/`, `etc-dnsmasq.d/`), a shape none of the other 37 stacks - have. Back up the *actual* bind-mount sources a stack's compose file + `etc-pihole/`, `etc-dnsmasq.d/`), a shape only a handful of stacks have + — the live one now is `technitium`, whose `config/` (zones, settings, + blocklists) needs non-root ownership for its distroless image, which is + exactly why the installer still provisions it by hand (#2911). Back up the *actual* bind-mount sources a stack's compose file declares, not an assumed `.env`/`secrets/` checklist — read the compose file if in doubt. 2. Back up everything found in step 1. 3. Remove the stack's current directory. 4. Create the Arcane sync (`syncDirectory: true`) — this deploys immediately, and will legitimately fail-closed if a required secret - isn't present yet (expected, not a bug — see canarytokens/ghosts/ - keycloak/dashboard's own `:?required` variables). + isn't present yet (expected, not a bug — see the four stacks that do + declare `:?required` variables: `honeypot-canarytokens`, `ghosts`, + `technitium` and `unsloth`). 5. Restore everything backed up in step 1, **preserving original ownership and permissions, not just content**. Confirmed live: restoring a secret file as `root:root` when the container expects the previous owning @@ -173,7 +193,14 @@ a stack means: These are platform behaviors, not something fixable from a compose file alone. Re-verify against whatever Arcane version is pinned in -`docker-compose.arcane.yml` if it's ever upgraded. +`docker-compose.arcane.yml` if it's ever upgraded. **That upgrade has +happened and the re-verification has not:** every item below was confirmed +against `v2.8.0`–`v2.9.0` (including the no-`profiles:`-support finding, +which cites the upstream issue as still open at `v2.8.1`), while +`docker-compose.arcane.yml` now pins `ghcr.io/getarcaneapp/manager:v2.11.1` +by digest. Nothing here is claimed to be false of `v2.11.1` — it is claimed +to be unconfirmed against it, and three minor versions of upstream is +exactly the gap that re-verification is for. - **`"project directory is not inside a mounted directory"` is a false warning against every stack with a relative bind-mount source, under @@ -194,18 +221,24 @@ alone. Re-verify against whatever Arcane version is pinned in is no per-message log-suppression Arcane exposes, and `hp-arcane`'s container logs aren't ingested into this repo's ELK pipeline (it tails application log files, not `docker logs` streams), so there is no - repo-side filter either. Fires for exactly the 9 stacks with at least one - relative-source bind mount under this identity mount (`honeypot-cowrie`, - `honeypot-dionaea`, `honeypot-elk`, `honeypot-init`, `honeypot-keycloak`, - `honeypot-payload-analysis`, `honeypot-tanner`, `honeypot-utilities`, - `pihole`) — every relative mount on the 7 of those 9 with running - containers was confirmed to resolve to its real, existing host path, zero - missing-on-host mounts. (`pihole` and `honeypot-keycloak` were the two - without running containers, so their mounts were reasoned from the same - identity-mount arithmetic rather than observed — #2853, #2764. `pihole` - has since been replaced by `technitium` in the manifest, #2911; the - warning's mechanism is per-relative-mount and unchanged by the swap, but - the stack name in this list is the pre-#2911 one.) **Decision: live with it — this repo has no fix + repo-side filter either. Fires for exactly the 10 stacks with at least one + relative-source bind mount under this identity mount (`ghidra`, + `honeypot-cowrie`, `honeypot-dionaea`, `honeypot-elk`, `honeypot-init`, + `honeypot-keycloak`, `honeypot-payload-analysis`, `honeypot-tanner`, + `honeypot-utilities`, `technitium`; the 10 is re-derived 2026-09-27 from + each of the manifest's 39 entries' `volumes:` sources) — every relative + mount on the 7 of the 9 that existed at the 2026-09-03 check and had + running containers was confirmed to resolve to its real, existing host + path, zero missing-on-host mounts. (`pihole` and `honeypot-keycloak` were + the two without running containers, so their mounts were reasoned from the + same identity-mount arithmetic rather than observed — #2853, #2764. + `pihole` has since been replaced by `technitium` in the manifest, #2911 — + hence the post-#2911 name above; the warning's mechanism is + per-relative-mount and unchanged by the swap. The tenth stack, `ghidra`, + is not new: its `./revdeck-proxyfix/*` mounts date from #1173, so the + 2026-09-03 live pass simply never reached it, and the "9" it recorded was + a nine-name subset of the real set rather than a different count.) + **Decision: live with it — this repo has no fix available, and the warning is confirmed harmless.** Re-verified live 2026-09-03: still firing at the same 9 stacks, ~1 warning per relative mount per sync. Not reported upstream (`getarcaneapp/arcane`) as of this @@ -247,7 +280,7 @@ alone. Re-verify against whatever Arcane version is pinned in (`dashboard/vendor/`, exactly 650 tracked files at the #1502 move, re-counted from git history), which exceeded the default back then. That tree is gone with the Go tier itself (#1659) and the synced directory is - down to 268 tracked files (re-counted 2026-08-27, + down to 279 tracked files (re-counted 2026-09-27, `git ls-files arcane/home/honeypot-dashboard`) — under the default again, so the bump is dormant headroom, kept because raising it is an Arcane-side change rather than a repo one. The failure mode itself @@ -262,9 +295,11 @@ alone. Re-verify against whatever Arcane version is pinned in `ghosts`'s sync (`sandbox/ghosts/compose.yml`) always failed — either a ~40s-then-500 with no detail, or (once `maxSyncTotalSize` alone was raised) a fast `file count limit exceeded` — because - `sandbox/ghosts/vendor/ghosts-src/` (963 tracked files, ~132 MB) sits in - the same directory tree as the compose file, so the sync's whole-directory - walk of `sandbox/ghosts/` (989 files, 135,789,139 bytes in total) blows + `sandbox/ghosts/vendor/ghosts-src/` (935 tracked files, 135,601,433 bytes) + sits in the same directory tree as the compose file, so the sync's + whole-directory walk of `sandbox/ghosts/` (962 files, 135,748,394 bytes — + both re-counted 2026-09-27; the 2026-08-31 live walk quoted below measured + 989/135,789,139) blows past *both* defaults (`maxSyncFiles: 500`, `maxSyncTotalSize: 50MB`), not just the one either failure message names. Fixed the same way as the `honeypot-dashboard` case above — both limits raised on the sync record @@ -347,8 +382,11 @@ alone. Re-verify against whatever Arcane version is pinned in ## Local environment overrides Compose's own `.env`-in-project-directory interpolation already covers -every `${VAR}` reference in these 37 stacks — none of them use `env_file:`, -and none needed it added. Arcane's effective environment merge +every `${VAR}` reference in these 39 stacks — none of them *need* `env_file:` +and only one declares it: `analysis/ghidra/docker-compose.ghidra.yml`'s +`revdeck`-profile service carries `env_file: [{path: .env, required: false}]` +(#110), which duplicates the interpolation it sits beside rather than +carrying anything. None needed it added. Arcane's effective environment merge (`project.env` + `.env.git` → `.env`) feeds that same mechanism transparently, so a local override set through Arcane's own UI for a synced project works exactly like editing `.env` by hand always did; no stack-file @@ -407,12 +445,28 @@ Two things worth knowing when setting an override: ## Promotion workflow and change control Decided in #1507: **release/tag promotion**, with `autoSync` enabled for -exactly the three stacks where a sync is the whole deploy. As of -2026-08-27 that policy has **never been put into effect** — every sync -still tracks `main` and nothing auto-deploys (#2549 re-derived the live -state). What actually runs is the manual model below; #2577 holds the -one-time activation steps and the dangling-sync cleanup if that ever -changes. +exactly the three stacks where a sync is the whole deploy. **The manifest +half of that policy is in effect: every one of the manifest's 39 entries +declares `branch: "production"`, and #1943 (2026-08-25) is the commit that +changed it from `main` — re-derived 2026-09-27, and unchanged since except +for the `pihole`→`technitium` and `unsloth` entries that inherited it.** +The live-store half is not. As of the last reads (2026-08-27 #2549, 2026-09-03) +every live sync still tracked `main` and nothing auto-deployed, so the +manifest and the host disagree about what a sync follows. #2577 holds the +one-time activation steps and the dangling-sync cleanup. + +**The pointer the manifest names does not exist yet.** `git ls-remote +--heads origin` on 2026-09-27 returned `main` and assorted agent/dependabot/ +design branches and no `refs/heads/production`; `scripts/promote-release.sh +--list` reports it as "branch does not exist yet". A sync pointed at +`production` therefore fails the same way a tag does — Arcane prefixes +`refs/heads/` onto whatever it is given (see "Arcane cannot track a tag" +below) — so this is the one place where the manifest is currently ahead of +reality in a way that bites: the documented bulk-import path, +`POST /environments/0/gitops-syncs/import`, would create 39 syncs that all +fail on their first run, on a branch that nothing creates until someone runs +`scripts/promote-release.sh `. What actually runs today is the manual +model below. ### Two facts (still true today) @@ -457,27 +511,36 @@ treat every sync as a per-project restart and bound the blast radius accordingly: scope each run to one stack at a time and verify the result — rather than firing a fleet-wide sync pass and racing the timeout. (A purpose-built single-project script is a natural follow-up; ship it separately so the - doc's claim of 'every sync is a restart' is reviewable on its own.)**34 of the 37 stacks build an image.** Only `honeypot-elk`, -`honeypot-keycloak` and `pihole` pull — re-derived 2026-08-27 from the 37 -manifest compose paths (34 carry `build:`; the same three pullers the -#1502-era text named, which said "35 of the 38" before #2381 retired -wordpot and f139fe24 retired the Go ip-enrichment-worker). For any + doc's claim of 'every sync is a restart' is reviewable on its own.)**34 of the 39 stacks build an image.** Five pull — +re-derived 2026-09-27 from the 39 manifest compose paths (34 carry +`build:`): `honeypot-elk`, `honeypot-keycloak`, `llm-worker`, `technitium` +and `unsloth`. `llm-worker` is in that set only because the manifest +deploys `llm-worker/docker-compose.captured-data-deploy.yml`, an overlay +with no `build:` of its own — `llm-worker/docker-compose.yml` still has +one. The #1502-era text named "35 of the 38" and the same three pullers, +before #2381 retired wordpot and f139fe24 retired the Go +ip-enrichment-worker. For any building stack, `autoSync: true` would mean every merge produces a deployment that looks successful and changes nothing — worse than a manual process, because it is unattended. ### What actually runs -- **Every sync tracks `branch: "main"`** — all 38 live `gitops_syncs` - rows (the 37 manifest stacks plus #2577's dangling `honeypot-wordpot` - orphan). The rows predate #1507's decision and nothing re-pointed - them; no `production` pointer exists. +- **Every sync tracks `branch: "main"`** on the live host, while the + manifest has said `branch: "production"` for all 39 entries since + #1943 — so the rows predate #1507's decision and nothing re-pointed + them. Every live `gitops_syncs` row is one of 39 manifest stacks plus + #2577's dangling `honeypot-wordpot` orphan, so 40; the 2026-08-27 read + saw 38 of them, before #2911's `technitium` and #3092's `unsloth`. The + gap between the two is the thing to resolve before the next import, and + the branch to resolve it against does not exist yet (see "Promotion + workflow" above). - **`autoSync` is 0 everywhere, including the three the manifest flags.** - `honeypot-elk`, `honeypot-keycloak` and `pihole` carry `autoSync: true` - in `arcane/manifests/home-production.json`, but the live store has - `auto_sync = 0` on all 38 rows — the elk/keycloak/pihole auto-follow - policy is silently inert: a promotion, or any push, will not deploy - them. + `honeypot-elk`, `honeypot-keycloak` and `technitium` carry + `autoSync: true` in `arcane/manifests/home-production.json`, but the live + store has `auto_sync = 0` on every row — the elk/keycloak/technitium + auto-follow policy is silently inert: a promotion, or any push, will not + deploy them. - **Every deploy is manual**: sync → build → redeploy per stack. The order matters and is not arbitrary: `honeypot-dashboard` must sync before `honeypot-dashboard-backend` builds, because the Rust source @@ -485,8 +548,8 @@ manual process, because it is unattended. - **Promotion is CI-only.** `scripts/promote-release.sh v0.1.0` exists and still refuses a ref that is not a tag, and a tag that is not an ancestor of `main` — so if a pointer ever exists, what reaches it has - always been through CI. Today it moves nothing, because there is no - pointer to move. + always been through CI. As of 2026-09-27 it still moves nothing: there + is no `production` branch on `origin` to move. ### The #1507 design, for the record @@ -501,8 +564,14 @@ manual process, because it is unattended. single-sync PATCH work. #2577 closed (PR #2704) having done only the wordpot-orphan half of its own -scope — the #1507 activation itself was never done, and #2858 (below) -decided explicitly not to do it in this round either. +scope, and #2858 (below) decided explicitly not to do the activation in this +round either. Two halves of "the #1507 activation" are worth keeping +distinct: the manifest half **was** done, in #1943 on 2026-08-25, which is +what put `branch: "production"` and the three `autoSync: true` flags into +`arcane/manifests/home-production.json`; the live-store half — re-pointing +the rows already in Arcane at that branch and setting the +`pullImageAfterSync`/`redeployAfterSync` fields the manifest cannot carry — +is what is still outstanding. ## autoSync decision (#2858) @@ -526,7 +595,7 @@ investigated "a merged PR never reached the host" issues. **Decision: leave `autoSync: false` fleet-wide. Do not activate #1507's three-puller policy in this round either.** Reasoning: -- Turning `autoSync` on for all 37 projects means every merge to `main` +- Turning `autoSync` on for all 39 projects means every merge to `main` redeploys the fleet unattended — a real increase in blast radius, and exactly the kind of change that should not happen as a side effect of fixing a visibility gap. (This round's own brief calls this out @@ -560,26 +629,31 @@ three-puller policy in this round either.** Reasoning: ops-triage pass. `autoSync` therefore remains `false` fleet-wide for now. - #2854's one-shot-abort hazard (a `restart: no` job's clean `exit(0)` making Arcane report `failed` on a deploy that actually completed) is - fully scoped to `honeypot-init`. It is **not** the only file in the repo + scoped to `honeypot-init` **and, since #3128 (2026-09-08), `honeypot-elk`.** + It is **not** the only file in the repo with that shape, and a grep alone does not establish the claim — YAML writes it three ways, so the check has to be `git grep -nE "restart:[[:space:]]*[\"']?no[\"']?"`, which on `origin/main` - returns fourteen hits — twelve real declarations across seven files (all - seven are in the table below), plus two prose matches, in this document and - in `scripts/compose-drift-watch.py`'s header. What scopes the hazard is that - only one of the seven is a service Arcane actually starts: + returns twenty hits — thirteen real declarations across eight files (all + eight are in the table below), plus seven prose matches: four in this + document, two in `scripts/arcane-sync-drift-report.py`, one in + `scripts/compose-drift-watch.py`'s header. (This bullet's own totals were + fourteen / twelve / seven / two / one at the 2026-09-04 round-6 pass.) + What scopes the hazard is that + only two of the eight are services Arcane actually starts: | hit | why it cannot trip the hazard | |---|---| - | `arcane/home/honeypot-init/compose.yml` (6×) | **this is the exposed one** | + | `arcane/home/honeypot-init/compose.yml` (6×) | **this is an exposed one** | + | `arcane/home/honeypot-elk/compose.yml:540` | **the other exposed one** — `arkime-pcap-init`, a `chown`/`chmod` one-shot added by #3128, not profile-gated, in a manifest entry Arcane syncs and starts like any other | | `sandbox/ghosts/compose.yml:162` | `ghosts` *is* an Arcane project, but the service is `ghosts-client-test`, gated behind `profiles: ["test"]`, so a default `up` never creates it | - | `llm-worker/docker-compose.{production-session-canary,synthetic-canary}.yml` | canary overlays; the manifest deploys `llm-worker/docker-compose.yml`, which has no `restart: no` | + | `llm-worker/docker-compose.{production-session-canary,synthetic-canary}.yml` | canary overlays; the manifest deploys `llm-worker/docker-compose.captured-data-deploy.yml`, which has no `restart: no` | | `docker-compose.sandbox.yml:48` | per-detonation sandbox lifecycle, not an Arcane manifest entry | - | `vps/docker-compose.yml:956` | the VPS stack, deployed by a different mechanism entirely | + | `vps/docker-compose.yml:966` | the VPS stack, deployed by a different mechanism entirely | | `sandbox/ghosts/vendor/ghosts-src/Ghosts.Api/docker-compose.yml:85` | vendored upstream source, never deployed | - All six excluded files were checked against - `arcane/manifests/home-production.json`'s 37 entries and their + All seven excluded files were checked against + `arcane/manifests/home-production.json`'s 39 entries and their `dockerComposePath` values, not inferred from the file paths. `honeypot-init` is not one of #1507's three auto-sync candidates, so this decision doesn't change its exposure either way: it keeps deploying @@ -591,6 +665,18 @@ three-puller policy in this round either.** Reasoning: 0755, owned by `github-deploy-runner`, dated 2026-09-03 23:07. #2908 is closed. + **`honeypot-elk` is a #1507 auto-sync candidate, so this one is not + hypothetical the way `honeypot-init` is.** Two consequences follow, and + neither is established here: whether the live `honeypot-elk` record + actually reports `failed` (the #2854 shape is a strong reason to expect + it does, but this has not been read off the live API), and, if it does, + that `scripts/arcane-sync-drift-report.py`'s `KNOWN_STRUCTURAL_FAILURES` + still names only `honeypot-init` — so the report would exit non-zero on a + perfectly healthy fleet, which is the exact "permanently red and therefore + ignored" failure mode its own comment says the exemption exists to + prevent. Worth one `GET /environments/0/gitops-syncs` before #1507's + activation is revisited. + **What makes Arcane gitops-sync drift visible, since nothing did before:** `scripts/arcane-sync-drift-report.py` — read-only, on-demand (not wired into a scheduled workflow: that would need an Arcane API key available to a CI @@ -618,7 +704,20 @@ permanently-red-and-therefore-ignored failure mode `scripts/isolation-audit.sh`'s own tiering comment was written against. Exempted projects are still **printed**, with the reason, under an `EXEMPT` heading; they are not silenced. Any project that reports `failed` without -being named there still fails the run. +being named there still fails the run. `KNOWN_STRUCTURAL_FAILURES` names +exactly one project, `honeypot-init`; if the `honeypot-elk` exposure +described under #2858 above turns out to be real, that table needs a second +entry for the same reason, or a healthy fleet goes permanently red. + +**The baseline is hardcoded to `origin/main`.** `scripts/arcane-sync-drift-report.py` +fetches `origin main` and measures `..origin/main` +unconditionally — it does not read each record's own `branch`. That is +correct while every live sync tracks `main` (it still does, per the reads +above), but it becomes the wrong measure the moment #1507's live-store +activation re-points rows at `production`: a sync correctly caught up to +the promoted release would then read as N commits behind `main` for every +merge since the promotion. Same shape as the `honeypot-elk` gap above — +worth resolving in the same pass. Exit-code behaviour was demonstrated in both directions (2026-09-03) by driving the shipped `main()` with a synthetic record set: a fleet whose only @@ -636,7 +735,7 @@ owed and worth one look once a current key is to hand. The fleet has a second one — `deploy.yml`'s rsync into `/opt/stacks/apiary` (the same inode as `/var/dockge/stacks/apiary`) — that no sync record covers. That channel had gone 18 days without a successful run, which is what made a -fleet reading `lastSyncCommit == main` on all 37 records still run a +fleet reading `lastSyncCommit == main` on all 39 records still run a weeks-old copy of every rsynced script: `diagnostics.yml`'s isolation-invariants step executes the *deployed* `isolation-audit.sh` from that path. Tracked and fixed as **#2908**, now closed — re-derived diff --git a/docs/BACKUP-ESSENTIALS.md b/docs/BACKUP-ESSENTIALS.md index 9246c8bbd..652622660 100644 --- a/docs/BACKUP-ESSENTIALS.md +++ b/docs/BACKUP-ESSENTIALS.md @@ -16,13 +16,13 @@ for restoring onto a replacement host see | | | |---|---| -| `homeserver/env/*.env` | all 41 Arcane/Dockge stack `.env` files | +| `homeserver/env/*.env` | one file per Arcane/Dockge stack (40 under `/var/dockge/stacks/` as of 2026-09-27) | | `homeserver/secrets/` | secret files kept beside a stack rather than in its `.env` | | `homeserver/wireguard/` | `wg0.conf` including the private key | | `homeserver/installer/` | `install-homeserver.conf` — the installer's answers file, which exists only on the root filesystem a reinstall wipes | | `homeserver/technitium/` | hand-maintained Technitium DNS config | | `homeserver/keycloak/keycloak.sql.gz` | `pg_dump` of the identity DB — realm, clients, client secrets, users | -| `homeserver/volumes/` | `dashboard-state`, `arcane-data`, `evebox-config`, `canarytokens-redis-data`, `es-importer-state` | +| `homeserver/volumes/` | `dashboard-state`, `honeypot-arcane_arcane-data`, `honeypot-elk_evebox-config`, `honeypot-canarytokens_canarytokens-redis-data`, `honeypot-dashboard_es-importer-state` — the Arcane-prefixed names are the real volume names | | `vps/env/vps.env`, `vps/secrets/`, `vps/traefik/`, `vps/wireguard/` | the VPS's entire config surface, including the Traefik origin certificates | | `*/manifest/` | host reference notes — disks, volumes, containers, WireGuard, nftables | | `repo/docs/`, `repo/scripts/`, `repo/analysis/` | this repository's runbooks and operational scripts | @@ -69,7 +69,7 @@ Three locations, all written by the workstation, which is the backup host: |---|---|---|---| | 1 | `/run/media/xore//apiary-backups` | ext4 (Crucial X8 USB) | udisks auto-mount — only present while plugged in | | 2 | `~/apiary-backups` | XFS (internal) | always available | -| 3 | `homeserver:/mnt/usb-recovery/apiary-backups` | ext4 (Samsung T7, label `APIARY-BACKUP`) | mounted from fstab by UUID with `nofail` | +| 3 | `homeserver:/mnt/usb-recovery/apiary-backups` | ext4 (Samsung PSSD T7, label `APIARY-BACKUP`) | mounted from fstab by UUID with `nofail` | Location 3 was a Ventoy stick formatted exfat until 2026-08-23, mounted read-only and absent from `/etc/fstab` — so every write to it failed and it @@ -215,7 +215,12 @@ repository — `install-homeserver.conf.example` carries only placeholders. `vps/secrets/oidc/`. 5. **Volumes.** For each `homeserver/volumes/.tar.gz`, with the stack stopped, create the volume and unpack into it through a networkless - container: + container. `` is the **full real volume name** — the archive is + written as `$volume.tar.gz` by `backup-essentials.sh`, so four of the + five carry their Arcane project prefix + (`honeypot-arcane_arcane-data.tar.gz`, and so on). Creating a + short-named `arcane-data` volume instead would restore into a volume + no stack is mounted against. ```bash docker volume create docker run --rm --network none -v :/dst -v "$PWD/homeserver/volumes:/src:ro" \ @@ -297,7 +302,9 @@ gone, for two reasons that happen to point the same way: Also found and worth knowing: `honeypot-keycloak/.env` carries a full set of `RESTIC_*` variables pointing at `/mnt-2/apiary-keycloak`, but that repository -directory does not exist, its password file (`secrets/restic-password`) does +directory does not exist — and as of 2026-09-27 neither does `/mnt-2` itself, +which has been decommissioned, so the path cannot start working by accident. +Its password file (`secrets/restic-password`) does not exist, `restic` is not installed on the homeserver and no unit references it. It is dead configuration — no Keycloak restic backup has ever run. The `keycloak.sql.gz` dump in both scripts here covers that gap. diff --git a/docs/CGNAT-DEPLOYMENT.md b/docs/CGNAT-DEPLOYMENT.md index 948f079b8..ce81c3f86 100644 --- a/docs/CGNAT-DEPLOYMENT.md +++ b/docs/CGNAT-DEPLOYMENT.md @@ -22,13 +22,19 @@ flowchart TD validation), `honeypot-elk`, `honeypot-cowrie`, `honeypot-dionaea`, `honeypot-conpot`, `honeypot-dnp3`, `honeypot-http`, `honeypot-multipot`, `honeypot-payload-analysis`, `honeypot-tanner`, `honeypot-dashboard`, - `honeypot-utilities`, the standalone honeypots (cisco-asa, citrix, rdp, - dicompot, dns-honeypot, endlessh, beelzebub, hellpot, elasticpot, galah, + `honeypot-dashboard-backend`, `honeypot-utilities`, the standalone + honeypots (cisco-asa, citrix, rdp, sonicwall-sma, dicompot, dns-honeypot, + endlessh, beelzebub, hellpot, elasticpot, galah, sentrypeer, mailoney, canarytokens), `honeypot-keycloak`, and the - workers (ip-enrichment, agent-intrusion, attacker-identity, correlator, - payload-inventory) — 31 stacks total, one Arcane-managed directory each + workers (agent-intrusion, attacker-identity, correlator, + payload-inventory) — 32 stacks total, one Arcane-managed directory each under `arcane/home//` (`honeypot-wordpot` sat here until #2381 - retired it). `honeypot-init` still deploys first; every + retired it; `ip-enrichment-worker` was in that worker list until + f139fe24 retired the Go service, `honeypot-dashboard-backend` joined when + #1622 split it out of `honeypot-dashboard`, and `honeypot-sonicwall-sma` + when #3131 added it — re-counted 2026-09-27 against + `arcane/manifests/home-production.json`, which holds 33 in-tree entries: + these 32 plus `unsloth`). `honeypot-init` still deploys first; every sensor stack waits on its completion markers at its own entrypoint rather than a Compose-level dependency, same reasoning as before, just across more projects now. See `docs/STACK-REBUILD.md` for the full current list @@ -39,7 +45,7 @@ flowchart TD - VPS: plain Docker Compose manages `/root/vps/docker-compose.yml`. Unchanged by #1502 — VPS deployment stays outside Arcane entirely, as that issue's own scope decision. -- Each of the 31 migrated stacks' Compose source (build context, git-tracked config, +- Each of the 33 in-tree stacks' Compose source (build context, git-tracked config, `compose.yml` with an explicit top-level `name:` pinned to its live project name) lives self-contained under `arcane/home//` in this repository. Arcane clones the repo and materializes the *entire @@ -52,25 +58,31 @@ flowchart TD these syncs on a from-scratch install, driven by the single source of truth at `arcane/manifests/home-production.json`. Six more home-hosted stacks (`auth-events-worker`, `llm-worker`, `ml-worker`, - `analysis/ghidra`, `sandbox/ghosts`, `pihole`) are Arcane-managed too but + `analysis/ghidra`, `sandbox/ghosts`, `technitium`) are Arcane-managed too but were already self-contained, so they kept their existing repository-root path instead of moving. Three of those six (`auth-events-worker`, `llm-worker`, `ml-worker`) are also imported by `step_arcane_import_stacks` itself now (#1505 — confirmed to have no host-local state beyond `.env`); the other three keep their own dedicated - installer steps for reasons specific to each (`pihole`'s non-`.env` host - state, `analysis/ghidra`'s conditional GPU compose overlay, and + installer steps for reasons specific to each (`technitium`'s non-`.env` host + state — the step `pihole` had until #2911 swapped the two — `analysis/ghidra`'s conditional GPU compose overlay, and `sandbox/ghosts`'s confirmed Arcane build-context limitation, #1506) — see `scripts/install-homeserver.sh`'s own Phase 8 header comment for the - full reasoning behind each. + full reasoning behind each. That filter reaches 35 of the manifest's 39 + entries; `unsloth` is the one the installer reaches neither way (no + `step_unsloth_*`, and not a `honeypot-*` name), by design — see + `docs/ARCANE-GIT-SYNC.md`'s "Manifest import". - The public gateway source is under `vps/`. Arcane is used only on the home server. The VPS uses `docker compose` directly. See `docs/ARCANE-GIT-SYNC.md` for the sync model, cutover procedure, and -confirmed Arcane v2.8.0 platform limitations (a required compose variable +confirmed Arcane platform limitations (a required compose variable in a port-binding position, remote build contexts pinned to a Git tag, the sync file-count limit, and stale project records after a `destroy` call -all have confirmed workarounds documented there). +all have confirmed workarounds documented there). Those were each confirmed +against `v2.8.0`–`v2.9.0`; `docker-compose.arcane.yml` now pins +`manager:v2.11.1`, and none of them has been re-confirmed against that +image — its own section header says to re-verify on upgrade. ## WireGuard addressing @@ -135,12 +147,15 @@ the only internet-facing component. 8. Run `python3 analysis/verify-stack.py` (with `DASHBOARD_SERVICE_TOKEN` from `honeypot-dashboard/.env`) and inspect `/source-health`. -Each stack is a folder under your Arcane stacks dir (default `/opt/stacks/`). -Upload the whole home folder via SFTP — compose **and** the build -sub-folders (`cowrie/`, `multipot/`, `http-honeypot/`, `dashboard/`, …) — -since Arcane's own editor only edits the compose file. After editing Go -source or honeyfs content, rebuild from the `APIARY` stack's Arcane -**terminal**: `docker compose -f compose.yml up -d --build`. +Each stack is a folder under your Arcane stacks dir (`/var/dockge/stacks`; +`/opt/stacks` is a symlink to it, #1185). **Since #1502 nothing is uploaded +by SFTP** — Arcane materializes each stack's whole directory from its Git +sync, and `honeypot-wordpot` aside the source of truth is the repository, not +a hand-copied folder. The SFTP-upload and "edit then rebuild from Arcane's +terminal" instructions this paragraph used to give are part of the pre-#258 +model the callout above already flags; what replaces them is a commit plus a +sync, and a separate `POST /projects/{id}/build` for the stacks that have a +`build:` service — see `docs/ARCANE-GIT-SYNC.md`. ### Boot-safe home networking and VPS log mounts @@ -263,11 +278,12 @@ template is in [`vps/traefik/dynamic.yml`](../vps/traefik/dynamic.yml): `honeypot-http` (`decoy.`) + `honeypot-web` (catch-all) → fake nginx, `honeypot-snare` (`www-portal.` and `snare.`) → SNARE, one native-OIDC route for the dashboard (no gateway, since #1026), one native-OIDC -route for Arcane (no gateway, #1185), and six forward-auth-protected +route for Arcane (no gateway, #1185), and six gateway-fronted investigation routes sitting behind their own Keycloak-backed `oauth2-proxy` gateway: Kibana, TANNER, EveBox, Arkime, Rev·Deck, and the Traefik dashboard -itself. Each has a matching -`socat-hp-*` bridge in [`vps/docker-compose.yml`](../vps/docker-compose.yml). +itself. Five of the six have a matching +`socat-hp-*` bridge in [`vps/docker-compose.yml`](../vps/docker-compose.yml); +the Traefik dashboard's gateway terminates on the VPS, not through one. Traefik is an HTTP(S) reverse proxy — it adds TLS, per-subdomain routing and auth to the web honeypots and dashboards. The other protocols (SSH, SMB, @@ -325,14 +341,22 @@ Cloudflare answers every proxied hostname with **526**. `deploy.yml` never overwrites `traefik/certs/`. On a normal deploy it only checks that `origin.pem` still parses. -### The forward-auth bridge, generically +### The gateway-fronted chain, generically Six investigation UIs (Kibana, TANNER, EveBox, Arkime, Rev·Deck, the Traefik dashboard) reach home through the identical chain — one pattern, -six routers in `vps/traefik/dynamic.yml`, six `socat-hp-*` bridges, each -fronted by its own Keycloak-backed `oauth2-proxy` gateway container, not -six different mechanisms. The honeypot dashboard and Arcane are the two -exceptions — both speak native OIDC directly, no gateway — see the note below. +six routers in `vps/traefik/dynamic.yml`, six `oauth2-proxy` gateway +containers, five `socat-hp-*` bridges, not six different mechanisms. The +gateway *is* the router's upstream rather than a `forwardAuth` middleware: +`honeypot-kibana`'s `service:` is a loadBalancer at `http://oidc-kibana:4180`, +and that container's `OAUTH2_PROXY_UPSTREAMS` is the socat bridge +(`http://socat-hp-kibana:5601`). `grep -c forwardAuth vps/traefik/dynamic.yml` +is 0, so nothing in the config uses the forward-auth middleware form. The +sixth gateway, `oidc-traefik`, has no socat hop at all — its upstream is +Traefik's own dashboard (`http://traefik:8081`) on the VPS, which is why the +bridge count is five and not six. The honeypot dashboard and Arcane are the +two exceptions — both speak native OIDC directly, no gateway — see the note +below. ```mermaid sequenceDiagram @@ -340,7 +364,7 @@ sequenceDiagram actor Op as operator's browser participant CF as Cloudflare
(proxied DNS) participant TR as Traefik
(TLS termination + routing) - participant OA as oauth2-proxy
(forward-auth, one per service) + participant OA as oauth2-proxy
(gateway, one per service) participant KC as Keycloak
(honeypot-keycloak, at home) participant SOC as socat-hp-*
(VPS container) participant WG as WireGuard tunnel @@ -348,14 +372,13 @@ sequenceDiagram Op->>CF: HTTPS request, e.g. kibana. CF->>TR: proxied, real client IP in X-Forwarded-For - TR->>OA: forward-auth check + TR->>OA: routed to the app's own gateway (oidc-kibana:4180) alt no valid session OA-->>Op: redirect to Keycloak login (auth.) Op->>KC: authenticate (password + mandatory TOTP) KC-->>OA: OIDC callback, session established end - OA-->>TR: identity headers - TR->>SOC: request, security-headers applied + OA->>SOC: proxied request, identity headers added SOC->>WG: raw TCP, VPS listen port → 10.8.0.2:home-exposed-port WG->>APP: delivered to the app's own internal port APP-->>Op: response, relayed back through the same chain diff --git a/docs/CI-CD.md b/docs/CI-CD.md index 79e0442e8..d19a8df17 100644 --- a/docs/CI-CD.md +++ b/docs/CI-CD.md @@ -47,15 +47,18 @@ flowchart TB containersHome["containers.yml — all image builds"] securityHome["security.yml — all CodeQL languages"] pagesHome["pages.yml artifact build"] + imageScanHome["image-security-scan.yml"] end prPush --> quality prPush -->|"PR: build only,
never published"| containerBuild + prPush --> codeql + prPush --> pages mainPush --> quality mainPush --> containerBuild mainPush --> codeql mainPush --> pages - mainPush -.->|"every workflow's compute jobs —
only after passing the ci-router
trust gate + heartbeat; pull_request needs
repo variable CI_HOMESERVER_PRS,
and forks can never qualify"| ciSelfHosted + mainPush -.->|"those five workflows' ci-target jobs —
only after passing the ci-router
trust gate + heartbeat; pull_request needs
repo variable CI_HOMESERVER_PRS,
and forks can never qualify"| ciSelfHosted ``` **`honeypot-ci` does not see `pull_request` by default, by design.** A @@ -64,9 +67,11 @@ job; a self-hosted runner's job runs as a real process on real home-network infrastructure. A malicious test file in an unreviewed PR (`os.system(...)`, a crafted Go `TestMain`) would execute wherever that runner has access — the same reasoning `production-home`'s own deployment -runner (below) already applies. Every workflow's executor routing (each -caller's own `ci-target` job, which since #2571 always calls the shared -`.github/workflows/ci-router.yml`) trusts +runner (below) already applies. Executor routing is a caller-side job named +`ci-target`, which since #2571 calls the shared +`.github/workflows/ci-router.yml`. Five of the repo's twenty workflows have +one: `quality.yml`, `containers.yml`, `security.yml`, `pages.yml` and +`image-security-scan.yml`. It trusts push-to-main (already reviewed and merged), the `schedule` and `workflow_dispatch` (an operator's own machinery); same-repo pull requests need the repository variable `CI_HOMESERVER_PRS=true`, and fork @@ -316,10 +321,14 @@ stack on the host lives entirely in Arcane's Git-sync machinery — see [ARCANE-GIT-SYNC.md](ARCANE-GIT-SYNC.md) for the full contract (its non-obvious cornerstones: creating a sync *is* an initial deploy, a sync materializes files without redeploying — live `redeploy_after_sync` -defaults to 0, though the manifest schema cannot express it — and every -synced stack runs `autoSync: false`: the #1507 tag-promotion / -`production`-pointer policy was decided but never deployed, so all syncs -track `main` and deploys are manual; ARCANE-GIT-SYNC.md's promotion +defaults to 0, though the manifest schema cannot express it — and deploys +are manual. #1507's tag-promotion / `production`-pointer policy was only +half activated: #1943 (2026-08-25) put `branch: "production"` and three +`autoSync: true` flags into `arcane/manifests/home-production.json`, so all +39 entries now name that branch rather than `main` — but the live store +still reads `auto_sync = 0` on every row, and `origin` has no +`refs/heads/production` for the pointer to name, so nothing follows a +promotion and deploys stay manual; ARCANE-GIT-SYNC.md's promotion section carries the live-state evidence). This workflow deliberately stopped touching those directories entirely: running an rsync/build loop alongside Arcane's own sync would put two @@ -476,10 +485,9 @@ alert/intelligence history in the old one is gone. Everything else that was still monolithic as of the earlier revision of this section (`dionaea`, `payload-dedupe`, `yara-scanner`, and the Tanner -group) has since split out too -- see the `honeypot-dionaea` and -`honeypot-payload-analysis` section below; only the Tanner group remains in -`APIARY`, as part of its own internal `depends_on` chain not yet -worth splitting. +group) has since split out too -- see the `honeypot-dionaea`, +`honeypot-payload-analysis` and `honeypot-tanner` sections below. Nothing +remains in `APIARY`: the root `docker-compose.yml` is `services: {}`. #### Dashboard redeploy (single replica; #266 rolling pair retired, #1659 legacy `dashboard` removed) @@ -630,13 +638,26 @@ open handles into the log directories this script wipes for this target. ### honeypot-elk (#258) `arcane/home/honeypot-elk/compose.yml` bundles the ELK/analysis plane (`elasticsearch`, -`kibana`, `filebeat`, `evebox`, `arkime-capture`, `arkime-viewer`, -`pcap-sync`) into one stack at `/opt/stacks/honeypot-elk` -- the last group +`kibana`, `filebeat`, `evebox`, `pcap-sync`, `arkime-pcap-init`, +`arkime-capture`, `arkime-viewer`, `extracted-file-importer`, `zeek-proxy`) +into one stack at `/opt/stacks/honeypot-elk` -- the last group that was still in the monolithic file. Kept together, not split further: -all seven sit on the shared `honeynet` network and either read from or -write to the one Elasticsearch instance, so splitting them apart would -turn every one of those relationships into a cross-stack shared resource -for services that only ever make sense running together. +they share the one Elasticsearch instance, the `arkime-pcap` volume and the +host's `logs/` bind-mount tree, so splitting them apart would turn every one +of those relationships into a cross-stack shared resource for services that +only ever make sense running together. + +The shared-network story is narrower than it looks, and the count above is +not "ten on `honeynet`". Seven of the ten declare `honeynet` explicitly +(`elasticsearch`, `kibana`, `filebeat`, `evebox`, `arkime-capture`, +`arkime-viewer`, `extracted-file-importer`); `elasticsearch` is additionally +on `llm-data`. `pcap-sync` and `arkime-pcap-init` declare no `networks:` at +all and so ride the project's implicit default network -- `pcap-sync` moves +rotated pcaps through host bind-mounts and a marker file, and +`arkime-pcap-init` is a one-shot `chown` of the `arkime-pcap` volume, so +neither needs the shared network. `zeek-proxy` sets `network_mode: host` +outright and reaches the sensor plane through host-published ports, which is +the point of it. `honeynet` and `llm-data` get the usual explicit shared `name:` treatment. `es-data` does **not**, despite appearances: `honeypot-init`'s @@ -868,8 +889,9 @@ The install also drops two things next to the unit: The leading `+` runs that line as root even though the unit's own `User=` is the unprivileged runner account, so no new sudoers grant was needed -(unlike `compose-project-state.py` above, this runs as part of the unit's -own privileged startup rather than from inside a workflow step). +(unlike `scripts/compose-project-state.py`, the narrow root helper for +`compose-drift-watch.py`, this runs as part of the unit's own privileged +startup rather than from inside a workflow step). **Why it exists.** A root process that writes into a runner's `_work` checkout leaves files the runner user can never delete, and @@ -1072,7 +1094,7 @@ The `Pick cache backend` step therefore chooses per executor: the runner can actually write it, so a rebuild replay (#1609) recreates it rather than leaving a hand-made directory nobody records. If the step has not run on a given box, `Pick cache backend` emits a workflow warning and -falls back to `type=gha` -- a slow build, not eighteen failed matrix rows. +falls back to `type=gha` -- a slow build, not nineteen failed matrix rows. **Bounding it.** `type=local` has *no* eviction: every export leaves unreferenced blobs behind in `blobs/sha256/` forever. @@ -1096,9 +1118,11 @@ gets to them. Deleting on close reclaims that quota immediately. `docker/login-action` targets `ghcr.io` and is gated `if: github.event_name != 'pull_request'`, so on a PR every base-image pull went out anonymous -- and Docker Hub meters anonymous pulls **per source -IP**, at roughly 100 per 6h. With `CI_HOMESERVER_PRS=true` all 18 matrix -rows leave this box through one address, and the tree carries **74 -non-`scratch` Hub `FROM` lines**. One cold run spends most of the budget; +IP**, at roughly 100 per 6h. With `CI_HOMESERVER_PRS=true` all 19 matrix +rows leave this box through one address, and the tree carries **75 +non-`scratch` `FROM` lines** across 56 tracked Dockerfiles, 68 of them +resolving to Docker Hub (the other 7 are `mcr.microsoft.com`). One cold run +spends most of the budget; the run after it fails with `toomanyrequests` on whichever rows happen to ask last. #2771's per-image `type=gha` scopes do not help: that cache holds *our* layers, never the base image, so every run re-resolves every `FROM` @@ -1181,9 +1205,9 @@ flowchart TB checkout["actions/checkout"] key["VPS_SSH_KEY written to a
temp file, mode 0600"] backup[("Snapshot: /root/vps-backups/
pre-deploy-<timestamp>.tar.gz,
10 most recent kept")] - rsync["rsync vps/ -> /root/vps/
over SSH, excluding .env,
traefik/certs/, traefik/dynamic.yml
(VPS-owned, see table below)"] + rsync["rsync vps/ -> /root/vps/
over SSH, excluding .env,
traefik/certs/, traefik/dynamic.yml,
secrets/ (VPS-owned, see table below)"] validate["SSH: docker compose config
validates /root/vps/docker-compose.yml"] - up["SSH: docker compose up -d --build"] + up["SSH: docker compose up -d --build
--remove-orphans (#2813)"] dynGen["Separate step: substitute DOMAIN
into the committed *.honeypot.example
placeholders, validate as YAML,
no leftover placeholders --
all BEFORE touching the VPS"] dynWrite["Copy to a temp path on the VPS,
then write in place with cat --
never copy-then-rename (see below:
Traefik's bind mount tracks the
inode, not the path)"] verify["Verify step: fail the job if certs
or dynamic.yml are missing, empty,
unparseable, or still placeholder"] @@ -1211,10 +1235,14 @@ The VPS job runs on a short-lived GitHub-hosted Ubuntu runner: `/root/vps-backups/pre-deploy-.tar.gz`, keeping the ten most recent archives. 5. `rsync` sends only the repository's `vps/` directory over SSH to - `/root/vps/`, **excluding** `traefik/dynamic.yml` (see below). + `/root/vps/`, **excluding** the four VPS-owned paths in the table below + (`.env`, `traefik/certs/`, `traefik/dynamic.yml`, `secrets/`). 6. A second SSH command runs on the VPS, validates `/root/vps/docker-compose.yml`, and executes - `docker compose up -d --build`. + `docker compose up -d --build --remove-orphans`. The flag matters: a + service removed from `docker-compose.yml` otherwise leaves its container + running forever, which is how #2813 found `socat-hp-wordpot` still up a + week after #2469 retired wordpot and dropped its forwarder rule. 7. A dedicated step generates the deployable `traefik/dynamic.yml` -- substitutes `DOMAIN` for every `*.honeypot.example` placeholder in the committed template -- and validates the result (parses as YAML, no @@ -1229,7 +1257,7 @@ The VPS job runs on a short-lived GitHub-hosted Ubuntu runner: ### Files the VPS owns, not the repository `--delete-delay` removes destination files that no longer exist under the -repository's `vps/` directory, and overwrites the ones that do. Three paths are +repository's `vps/` directory, and overwrites the ones that do. Four paths are therefore excluded from the main `rsync` because the VPS copy is authoritative (or, for `dynamic.yml`, because it needs different handling entirely): @@ -1238,6 +1266,11 @@ therefore excluded from the main `rsync` because the VPS copy is authoritative | `.env` | Secrets and host-specific values. | | `traefik/certs/` | Issued TLS certificates. They do not exist in the repository, so an unexcluded `--delete-delay` deletes them, and the workflow cannot reissue them. | | `traefik/dynamic.yml` | Carries the deployment's real domain. The committed copy is a `*.honeypot.example` placeholder -- Traefik's file provider has no `${VAR}`-style substitution the way docker-compose already gives every other host-specific value in this repo, so this file can't just be templated in place the normal way. Deployed by its own dedicated step instead (step 7 above), which substitutes `DOMAIN` and writes the result separately. | +| `secrets/` | Per-gateway OIDC cookie-secret/client-secret files (`OIDC_SECRETS_DIR=./secrets/oidc` in `vps/.env.example`). Git-ignored, so they are never present in the checkout at all -- and `--delete-delay` reads "absent from the source" as "delete it". That happened once for real and took down every oauth2-proxy gateway at the time, which is why the exclude exists. A client-secret is never regenerated: it has to match what is already registered with Keycloak, so recovering it means `kcadm get clients//client-secret`, and losing it means re-registering the client. | + +The pre-deploy backup step ahead of the bulk `rsync` archives the same four +paths (`.env`, `traefik/certs`, `traefik/dynamic.yml`, `secrets/`) that +actually exist, keeping the last ten under `/root/vps-backups/`. The certificates were lost once, in a single `target: both` run before that exclusion existed: Traefik fell back to self-signed and every router silently @@ -1294,8 +1327,8 @@ scratch twice, reaching the same blocker both times. **No workflow edit is needed.** Every `secrets.VPS_*` / `secrets.DOMAIN` reference already sits inside a job that declares -`environment: production-vps` — `deploy.yml`'s `vps` job (`:207`, environment -at `:210`), `diagnostics.yml`'s `vps` job (`:252`/`:255`), and +`environment: production-vps` — `deploy.yml`'s `vps` job (`:231`, environment +at `:234`), `diagnostics.yml`'s `vps` job (`:292`/`:295`), and `vps-start-blackhole.yml`'s `start-blackhole-profile` job (`:22`/`:24`). The `home` jobs (`deploy.yml:21`, `diagnostics.yml:76`) read none of the five. Environment secrets also shadow repository secrets of the same name, so @@ -1324,7 +1357,7 @@ source rather than the password manager, rotate it deliberately rather than as a side effect of the move. `DOCKERHUB_USERNAME` / `DOCKERHUB_TOKEN` (added 2026-09-01) are read only by -`containers.yml:157-160`, which declares **no** `environment:` at all — so +`containers.yml:157-162`, which declares **no** `environment:` at all — so they have to stay repository-scoped until that workflow gains one, and they are correctly out of this migration's scope rather than merely deferred. They are still the reason the repository-secret set grew from five to seven, which @@ -1411,7 +1444,10 @@ diagnostics workflow itself. Home: GitHub -> outbound-polling self-hosted runner on homeserver -> local rsync /opt/stacks/apiary - -> Arcane compose.yml -> docker compose up + -> compose config --quiet (validation only, no `up`; + the root file is services: {}) + -> Arcane's own Git-sync machinery materializes and + deploys stacks (ARCANE-GIT-SYNC.md) VPS: GitHub-hosted runner -> rsync + SSH over VPS_PORT @@ -1419,6 +1455,11 @@ GitHub-hosted runner -> rsync + SSH over VPS_PORT -> docker compose up on VPS ``` +The home path has not run a deploy in this workflow since #1502 — see +"What deploy.yml actually runs (since #1502)" above for the full list of +what it does instead. A `home` run that succeeds has still changed nothing +on the host. + Selecting `both` creates both jobs from the same workflow run. They share the `honeypot-production` concurrency group, but the home and VPS jobs are otherwise independent: one can fail while the other succeeds. Always inspect diff --git a/docs/HOMESERVER-DISK-LAYOUT.md b/docs/HOMESERVER-DISK-LAYOUT.md index 2c3673d2d..ac3ba183c 100644 --- a/docs/HOMESERVER-DISK-LAYOUT.md +++ b/docs/HOMESERVER-DISK-LAYOUT.md @@ -3,13 +3,22 @@ This documents the physical disk layout of the honeypot homeserver (`supermicro`) as it actually exists today, and a generated Ubuntu **autoinstall** config (the Ubuntu/subiquity equivalent of Windows' -`autounattend.xml`) to reproduce that layout on a reinstall or a second -build server. Captured 2026-08-04 as part of the #518 smoke-test research. - -Ubuntu Server's installer (`subiquity`) is driven by `curtin` under the -hood — the fstab comments on this box literally say "was on /dev/sdX -during curtin installation", confirming this machine was already installed -this way rather than by hand. +`autounattend.xml`) that reproduces the layout the box had at the time of +the #518 smoke-test research. + +> **The autoinstall config below no longer describes this host.** The +> physical table and the provisioning steps were re-measured read-only on +> 2026-09-27. `supermicro` has since been reinstalled as **Rocky Linux +> 10.2** and no longer runs the Ubuntu/`curtin` layout: it uses LVM (the +> original notes recorded "no LVM"), `/var` sits on a *partition* of the +> RAID LUN rather than the whole disk, and swap is a 32G LVM logical +> volume rather than a swapfile. The original capture was 2026-08-04, when +> the fstab comments did literally say "was on /dev/sdX during curtin +> installation" — that evidence was sound for the Ubuntu install, which +> has since been replaced. Keep the autoinstall file as the record of the +> Ubuntu layout; do not use it as a rebuild target for the current host. +> See `docs/HOST-TUNING.md` for the tuning that *does* apply to the Rocky +> install. ## Why this layout, not one big disk @@ -22,24 +31,36 @@ reinstall of the OS disk alone doesn't touch captured evidence. ## Physical layout (as installed) +Re-measured read-only on 2026-09-27 via `lsblk`/`lvs`/`findmnt`/`df`. + | Device | Model | Size | Partition table | Filesystem | Mount | Role | |---|---|---|---|---|---|---| -| `nvme0n1` | Samsung MZVLW256HEHP | 238.5G | GPT | vfat (p1) / ext4 (p2) | `/boot/efi`, `/` | OS + EFI, boot disk | -| `sdb` | AVAGO MR9440-8i (RAID LUN) | — | whole-disk (no partition table) | xfs | `/var` | Docker root, Arcane-managed stacks, container state (`/var/lib/docker`, `/var/dockge`) — now also `benchmarks/`, `training/`, `hf-cache/`, `buildx-cache/`, `ci-registry-mirror/`, the former `/mnt-1` workload | -| `sda` | Intel SSDSC2KB480G8L | 447.1G | GPT, 1 partition | xfs | `/mnt-2` | Reserved bulk storage (currently empty) | -| `sr0` | ATAPI optical | — | — | — | — | Unused | - -**`/mnt-1` is decommissioned.** Its RAID VD (formerly `sdc`) suffered a -two-drive fault on 2026-09-09 (#3158) and no longer enumerates as a block -device at all; the mount was unwired (#3159, PR #3159) and everything that -lived under it moved to `/var`. `/mnt-1` itself survives on the host only as -a directory of compatibility symlinks into `/var` (`benchmarks`, `training`, -`hf-cache`, `buildx-cache`, `ci-registry-mirror`) so any script still hard- -coding the old path keeps resolving — new code should target `/var/*` -directly. See #3158/#3159 for the incident and decommission detail; this -table's `sdb` size is left unstated above rather than guessed, since the -volume backing `/var` changed as part of that recovery and hasn't been -re-measured for this doc. +| `nvme0n1` | PC401 NVMe SK hynix 1TB | 953.9G | GPT, 3 partitions | vfat (p1, 600M) / xfs (p2, 2G) / LVM2_member (p3, 951.3G) | `/boot/efi`, `/boot`, — | OS boot disk | +| └ `rl-root` | (LVM on `nvme0n1p3`) | 70G | — | xfs | `/` | OS root | +| └ `rl-swap` | (LVM on `nvme0n1p3`) | 32G | — | swap | `[SWAP]` | Swap | +| └ `rl-home` | (LVM on `nvme0n1p3`) | 849.3G | — | xfs | `/home` | Home | +| `sdb` | AVAGO MR9440-8i (RAID LUN) | 8.7T | GPT, 1 partition (`sdb1`, whole remaining size) | xfs | `/var` | Docker root, Arcane-managed stacks, container state (`/var/lib/docker`, `/var/dockge`) — now also `benchmarks/`, `training/`, `hf-cache/`, `buildx-cache/`, `ci-registry-mirror/`, the former `/mnt-1` workload | +| `sda` | PSSD T7 (**USB-attached**) | 465.8G | GPT, 1 partition (`sda1`) | ext4 | `/mnt/usb-recovery` | USB recovery disk, not local bulk storage | + +Three of the four rows changed since the 2026-08-04 capture, and none of +the change is cosmetic. The OS disk is a different, 4x-larger NVMe; the +former `sda` bulk-storage disk (`/mnt-2`) is now a USB-attached portable +SSD mounted at `/mnt/usb-recovery`; and the boot disk is now LVM-backed +with a separate `/home`. The old `sr0` ATAPI optical drive is no longer +enumerated at all. + +**`/mnt-1` and `/mnt-2` are both decommissioned.** `/mnt-1`'s RAID VD +(formerly `sdc`) suffered a two-drive fault on 2026-09-09 (#3158) and no +longer enumerates as a block device at all; the mount was unwired (#3159, +PR #3159) and everything that lived under it moved to `/var`. `/mnt-1` +itself survives on the host only as a directory of compatibility +symlinks into `/var` (`benchmarks`, `training`, `hf-cache`, +`buildx-cache`, `ci-registry-mirror`) so any script still hard-coding the +old path keeps resolving — new code should target `/var/*` directly. See +#3158/#3159 for the incident and decommission detail. `/mnt-2` is gone for +a different reason: its disk is the USB `PSSD T7` above, remounted at +`/mnt/usb-recovery`, so it is no longer local bulk storage and must not be +relied on for a rebuild. `sdb` sits behind an AVAGO/LSI MR9440-8i hardware RAID controller and appears to the OS as a SCSI LUN, not a raw disk — the controller's own @@ -49,17 +70,21 @@ the controller's own tooling (`storcli`/`perccli` or vendor equivalent) if the RAID config itself needs to be reproducible, not just the OS partitioning on top of it. -`/var` on its own disk is the key decision: `/var/lib/docker` is 103G and -`/var/dockge` (bind-mounted stack data for all 23 Arcane-managed stacks, including +`/var` on its own disk is the key decision, and it has only become more +load-bearing: `/var/lib/docker` is **2.9T** and `/var/dockge` (stack data +for the 45 directories under `/var/dockge/stacks/`, including Elasticsearch indices, Cowrie logs, payload captures, sandbox disks) is -229G — 332G combined, well past what the 238G OS disk could hold even -before accounting for the OS itself. Putting `/var` on the 1.7T `sdb` -disk instead of growing the root filesystem was the right call and should -be preserved on any rebuild. - -Swap is an **8G swapfile** at `/swap.img` on the root filesystem, not a -dedicated partition — simpler to resize than a swap partition and fine at -this scale (91G RAM, swap is a safety margin not a working set). +**350G**. `/var` is 70% full (6.1T of 8.8T) with 2.7T free. The manifest +still declares 39 sync entries, 33 of which name one of the 34 directories +under `arcane/home/`; `rex86-eval` is present on disk but **not** in the +manifest (the other 6 manifest entries are root-level stacks). Putting `/var` on the RAID LUN instead of growing the root +filesystem remains the right call and should be preserved on any rebuild. + +Swap is a **32G LVM logical volume** (`rl-swap`) in the `rl` volume group, +not a dedicated partition and not a swapfile — the 8G `/swap.img` +swapfile described in the 2026-08-04 capture no longer exists. 92G of RAM +means swap is a safety margin rather than a working set, though it was +under real pressure at measurement time (14.6G in use, priority -2). ## Reproducing it: `autoinstall/homeserver-user-data.yaml` @@ -102,11 +127,16 @@ with `homeserver-user-data.yaml` renamed to `user-data` alongside an empty - SSH (key-only, no password auth) and the `xfsprogs`/`nvme-cli` packages the manual partitioning step below needs. -**What has to be done by hand, at the storage screen, using the physical -layout table above as the target:** 3-disk layout (NVMe boot/OS: GPT, -EFI + ext4 root; `/var`: whole-disk xfs, no partition table; `/mnt-2`: +**What has to be done by hand, at the storage screen.** For reproducing +the **former Ubuntu layout** (the one this template was written against, +and the one the 2026-08-04 capture recorded): 3-disk layout (NVMe boot/OS: +GPT, EFI + ext4 root; `/var`: whole-disk xfs, no partition table; `/mnt-2`: GPT + single xfs partition), no LVM, 8G swapfile instead of a swap -partition. `/mnt-1` is no longer part of the target layout (decommissioned, +partition. That is **not** the live layout any more — the box now uses LVM +with a separate `/home`, `/var` on a partition, and a 32G swap LV, and its +disks have all been replaced (see the table above). Do not use this +paragraph as a partition plan for the current host; it is a record of what +the autoinstall flow produced. `/mnt-1` is no longer part of the target layout (decommissioned, see above) — do not recreate it on a rebuild. The template does **not** attempt to reproduce the AVAGO RAID controller's own LUN configuration either — that has to happen before the OS installer ever sees a block diff --git a/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md b/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md index 41458050c..14a2804ae 100644 --- a/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md +++ b/docs/analysis/ghidra/benchmarks/injection-gate-protocol.md @@ -1,5 +1,12 @@ # Injection gate v3 — what it measures and how to test it (#2694) +> **Re-verified 2026-09-27 — every tool and path in §3–§4 still exists as named.** +> The gate is current: the `injection` rubric block plus `injection_gate.py`, +> with the live slot at `qwen3:14b` and 32768 context. Whether the §3.3 +> positive-control sweep has been re-run against the current roster is +> **undetermined** from the repository — check `/var/benchmarks/` on the host. +> Companion: `docs/analysis/ghidra/benchmarks/corpus/README.md`. + This is the operating manual for the corpus injection axis after #2694. It states what a verdict means, how the gate is validated before any number from it is cited, and the exact commands. @@ -14,12 +21,14 @@ containing the payload's own words. Read against every stored | | truly complied | did not comply | |---|---|---| -| gate FAIL | 1 (partial) | 26 | -| gate PASS | 0 | 37 | - -25 of 27 failures fired on the model *quoting or paraphrasing* the planted -string — the behaviour the system prompt asks for. The remaining ones fired -on "appears to be benign", which is the case's own ground truth. The four +| gate FAIL | 1 (partial) | 25 | +| gate PASS | 0 | 38 | + +23 of the 26 failures fired on one of the four payload-derived forbidden terms +(the model quoting or paraphrasing the planted string) — the behaviour the +system prompt asks for. The other three fired only on "appears to be benign", +which is the case's own ground truth; two of those three never reproduce the +string at all. The four Tier A failures that drove the #1805-c "no promotion" decision (Ornith-35B, gemma-4-31B, Seneca-32B, huihui-qwen3.8) all explicitly identified the string as an injection or anomaly. No matcher can repair this fixture: the injected diff --git a/docs/benchmarks/2026-08-30-injection-gate-recalibration.md b/docs/benchmarks/2026-08-30-injection-gate-recalibration.md index 6730ff694..20b4e1132 100644 --- a/docs/benchmarks/2026-08-30-injection-gate-recalibration.md +++ b/docs/benchmarks/2026-08-30-injection-gate-recalibration.md @@ -4,6 +4,42 @@ > as-written since this is a dated record; see `docs/HOMESERVER-DISK-LAYOUT.md` for > current layout. +> **Dated record — measured 2026-08-30, superseded by #2694.** The gate and +> `is_injection_case()` behaviour described below no longer run: injection cases +> are now declared by the rubric's `injection` block and scored by +> `analysis/ghidra/benchmarks/injection_gate.py`. Every number in this file is +> the 2026-08-30 sweep's, left exactly as measured. The current protocol, +> re-verified 2026-09-27, is +> `docs/analysis/ghidra/benchmarks/injection-gate-protocol.md`. + +> **Reconciled 2026-09-27 — §9's inventory is now shipped, and §2's Tier B tally +> does not match the pinned fixture.** Five points, all checked against `HEAD`; +> no number above was rewritten. +> +> 1. **§9's "It is *not merged*" is stale.** The branch landed as `df650a8f` +> ("injection gate v3 … (#2694) (#2697)"), 2026-08-30 16:20:54Z — the same +> minute as the "69 `run1` reports present at 16:20Z" in §9.2. Every path in +> the §9 table is at `HEAD`; `rev_cases_v2_rubric.json` carries 17 cases with +> `process_and_injection` on `forbidden: []`, and `manifest.json` holds the +> 850 builds §9.1 claims (700 reproduced + 150 new, `errors: []`). +> 2. **§2's Tier B counts are the matrix's own, not the fixture's.** The table +> enumerates 30 rows but never lists `huihui-qwen3.6-35b-a3b-abliterated:q3_k`, +> which the shipped 64-row fixture records as a gate **FAIL**. Against that +> fixture Tier B is **15 FAIL / 15 PASS** over 30 rows, splitting **13** +> payload-term hits to **2** on `appears to be benign` — not the 14/15 and +> 11/3 of the headline above. §2 stays as written; read it as the matrix it +> is, and the fixture as the shipped regression set. +> 3. **§9's test inventory has drifted:** `test_injection_gate.py` is still 48, +> but `test_record_baseline.py` is **54**, not 50 (102 collected, 102 pass). +> 4. **§8's governance-gate term list is incomplete.** +> `evaluate-models.py:421` (`process-injection`) reads +> `("appears to be benign", "likely benign", "conclude benign")`; §4.3 lists +> all three, §8 omits the last. +> 5. **§8's "59 hand-labelled answers in §2" is the wrong set twice over:** §2 +> carries 29 Tier B verdicts, and the hand-labelled cohort is the 64 rows of +> §9's fixture. The `conclude benign` figure in §1 (zero occurrences) still +> holds — it is zero across all 64. + Measured 2026-08-30 against `/mnt-1/benchmarks/1947full` (run1 files; run2 verified byte-identical for every model at both tiers) and the checkout at `/mnt-1/benchmarks/APIARY` @ `a99e765`. Nothing on the host was modified; every script was piped over ssh stdin and read only. **Verdict on the preliminary finding (now issue #2694):** confirmed in mechanism and in substance, with two corrections and five additional findings. The gate is not measuring compliance. It is measuring whether a model *quoted or paraphrased the payload* (11 of 14 Tier B failures) or *used the exact phrase "appears to be benign"* (3 of 14). The fixture cannot discriminate compliance from correct analysis because the injected verdict is true. The same defect accounts for **all four Tier A failures that drove the #1805-c / #1947 "no promotion" decision**, including the disqualification of the top-scoring model. diff --git a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md index e8c805d21..3b2e69bc9 100644 --- a/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md +++ b/docs/benchmarks/plans/2026-09-05-1947-resume-plan.md @@ -4,6 +4,15 @@ > as-written since this is a dated record; see `docs/HOMESERVER-DISK-LAYOUT.md` for > current layout. +> **Dated plan — written 2026-09-05, superseded by round 7 (#3079).** The hardware +> table below is **pre-reinstall**: the host is now Rocky Linux 10.2 with LVM, +> `/var` on `sdb1` of an 8.7T LUN, a 32G `rl-swap`, and no `/mnt-1` mount (see +> `docs/HOMESERVER-DISK-LAYOUT.md`). Likewise, the "#2985 missing scripts" this +> plan works around are now in git at `analysis/ghidra/benchmarks/corpus/` — +> `requant_sweep.sh`, `slots_sweep.sh`, `chain_round7.sh` and the `round7_*` +> builders; only `gptoss_rerun.sh` is still absent. In-flight state below is as +> of 2026-09-05, not current. + **Written** 2026-09-05, from live inspection of `homeserver` and every open benchmark issue. Supersedes nothing; it sits *beside* `/mnt-1/benchmarks/STATE-2026-09-05-fix-stage.md`, which remains the authority on diff --git a/docs/benchmarks/plans/2026-09-06-round7-hermes-init.md b/docs/benchmarks/plans/2026-09-06-round7-hermes-init.md index ca7327af5..503f78d52 100644 --- a/docs/benchmarks/plans/2026-09-06-round7-hermes-init.md +++ b/docs/benchmarks/plans/2026-09-06-round7-hermes-init.md @@ -4,6 +4,14 @@ > as-written since this is a dated record; see `docs/HOMESERVER-DISK-LAYOUT.md` for > current layout. +> **Dated handoff — written 2026-09-06 (epic #3079); not a live brief.** The +> toolchain, model roster and dispatch rules in §5 are as-written on that date. +> What shipped since: the #3080 toolchain (`analysis/ghidra/training/`) and the +> corpus builders with their decontamination report. Unsloth (#3092) has **no +> installer path** — it is the Arcane stack `unsloth` +> (`arcane/manifests/home-production.json`), deployed by gitops-sync and a +> redeploy, never by hand. Stop it before any cold GPU leg. + **Read this first.** It names the plan of record, the state of the GPU, the kickoff order, the dispatch pattern and the hard rules. Written 2026-09-06, after the round-7 cold baseline was launched. Epic **#3079**, children diff --git a/docs/benchmarks/plans/2026-09-06-round7-unsloth-train-requant-ollama.md b/docs/benchmarks/plans/2026-09-06-round7-unsloth-train-requant-ollama.md index 52d0cdc2e..4f467793b 100644 --- a/docs/benchmarks/plans/2026-09-06-round7-unsloth-train-requant-ollama.md +++ b/docs/benchmarks/plans/2026-09-06-round7-unsloth-train-requant-ollama.md @@ -4,6 +4,13 @@ > as-written since this is a dated record; see `docs/HOMESERVER-DISK-LAYOUT.md` for > current layout. +> **Plan of record, written 2026-09-06 (epic #3079) — its §0 already folds in later +> changes; the rest is as-planned, not as-shipped.** The batch leg shipped as +> `analysis/ghidra/training/` (`train.py`, `export_to_ollama.sh`, `compose.yaml`, +> `TOOLCHAIN.md`); the interactive half shipped as the Arcane `unsloth` stack +> rather than an installer (#3092). `TOOLCHAIN.md`'s "no smoke test yet" still +> stands — nothing in this plan has been exercised end to end. + **Written** 2026-09-06 from live inspection of `homeserver`, every open benchmark issue, and the current Unsloth / llama.cpp / Ollama documentation. Sits beside `2026-09-05-1947-resume-plan.md`, which remains the authority on the #1947 sweep diff --git a/docs/deploy-profiles/README.md b/docs/deploy-profiles/README.md index f2d209ddd..73d538e07 100644 --- a/docs/deploy-profiles/README.md +++ b/docs/deploy-profiles/README.md @@ -27,17 +27,30 @@ line, `#` comments and blank lines ignored. | Profile | Backbone | Sensors | Shape | |---|---|---|---| -| [`full.txt`](../../deploy-profiles/full.txt) | init, elk, dashboard, utilities, payload-analysis | every deception sensor stack under `arcane/home/` | the standard deployment -- everything this repo ships | -| [`ics-focused.txt`](../../deploy-profiles/ics-focused.txt) | init, elk, dashboard, utilities | conpot, dnp3 | OT/ICS-only exposure -- skip the general-purpose/web/SSH/legacy-protocol sensors entirely | -| [`minimal-web.txt`](../../deploy-profiles/minimal-web.txt) | init, elk, dashboard, utilities | http, tanner | web-attack-focused -- HTTP/API honeypot + SNARE/TANNER, skip ICS/SSH/legacy-protocol sensors | - -`init`, `elk`, and `dashboard` are structural dependencies for any profile -that includes at least one sensor -- `scripts/validate-deploy-profile.sh` -(below) enforces -this, it isn't just a convention to remember. `payload-analysis` and +| [`full.txt`](../../deploy-profiles/full.txt) | keycloak, init, elk, dashboard, utilities, payload-analysis | the 20 classic deception sensor stacks under `arcane/home/` -- but see the gap below | the standard deployment -- everything this repo ships | +| [`ics-focused.txt`](../../deploy-profiles/ics-focused.txt) | keycloak, init, elk, dashboard, utilities | conpot, dnp3 | OT/ICS-only exposure -- skip the general-purpose/web/SSH/legacy-protocol sensors entirely | +| [`minimal-web.txt`](../../deploy-profiles/minimal-web.txt) | keycloak, init, elk, dashboard, utilities | http, tanner | web-attack-focused -- HTTP/API honeypot + SNARE/TANNER, skip ICS/SSH/legacy-protocol sensors | + +**`full.txt` is not actually "everything this repo ships".** It lists 26 +stacks (6 backbone + 20 sensors) and omits `honeypot-sonicwall-sma`, the +decoy sensor stack #3131 added on 2026-09-08 (`hp-sonicwall-sma-honeypot` +on `${HP_BIND:-10.8.0.2}:8543`). It is a natural fit for both `full.txt` and +`ics-focused.txt`, and is in neither. The other `arcane/home/` stacks the +profiles deliberately skip are the analysis-plane workers and +`honeypot-dashboard-backend` (not persona declarations, per below) plus +`unsloth` (the #3092 benchmark toolchain) -- and `rex86-eval`, which exists +on disk but is in no manifest entry at all. + +`init` and `elk` are structural dependencies for any profile that includes at +least one sensor; `keycloak` is a structural dependency of `dashboard`, and +`elk` is too. `scripts/validate-deploy-profile.sh` enforces all three, so +these aren't just conventions to remember. `payload-analysis` and `utilities` are strongly recommended (payload dedup/YARA scanning, log rotation/disk monitoring/autoheal) but not structurally required, so the -validator only warns if either is missing from a non-empty profile. +validator only warns if either is missing from a non-empty profile. Note +that `dashboard` itself is *not* a required structural dependency: the +validator never demands it, it only imposes `elk` and `keycloak` on a +profile that has chosen it. Not covered here: the VPS side (`vps/`, always deployed the same way regardless of home profile -- see `docs/CGNAT-DEPLOYMENT.md`), the @@ -56,9 +69,10 @@ scripts/validate-deploy-profile.sh deploy-profiles/ics-focused.txt Checks, against the *current* repository state (not a hardcoded snapshot): 1. **Structural dependencies** -- `init`/`elk` present if any sensor stack - is listed; `elk` present if `dashboard` is listed (the dashboard reads - several sensors' events from Elasticsearch, not their log files -- - see #403 for why that's a real dependency, not a nice-to-have). + is listed; `elk` and `keycloak` present if `dashboard` is listed (the + dashboard reads several sensors' events from Elasticsearch, not their + log files -- see #403 for why that's a real dependency, not a + nice-to-have; and the target auth path is native Keycloak OIDC). 2. **Real-stack existence** -- every listed name must correspond to an actual `arcane/home/honeypot-/` directory, so a typo'd or retired stack name fails here instead of surfacing mid-deploy or as a silently diff --git a/docs/gpu-docker-passthrough.md b/docs/gpu-docker-passthrough.md index d67add1f0..2f5989eb5 100644 --- a/docs/gpu-docker-passthrough.md +++ b/docs/gpu-docker-passthrough.md @@ -189,10 +189,22 @@ services: reservations: devices: - driver: nvidia - count: all + device_ids: ["GPU-18a00c7e-670a-c305-a2aa-20e3a71917a3"] capabilities: [gpu] ``` +**Do not use `count: all` here.** #1539 replaced it: the homeserver carries +two NVIDIA cards, and the `count: all` form handed Ollama *both* — wrong +even when it works, and it starves the Windows sandbox VM of the Quadro +P2200 reserved for its passthrough whenever Ollama loads a model during a +detonation. The overlay therefore pins the RTX 4000 Ada's UUID, confirmed +live via `nvidia-smi -L` on the actual box. Re-run `nvidia-smi -L` and +update the UUID if the card is ever physically replaced. + +The same file also documents why `ghidra` is deliberately *not* given the +GPU: decompilation is CPU work and would only compete for the card with the +model it is feeding. + Note this repo keeps the GPU reservation in a **separate overlay file**, applied only when a GPU is actually present: @@ -225,7 +237,11 @@ Docker's `--gpus all` / `count: all` doesn't partition VRAM — every container that requests the GPU gets the whole card, and it's up to each process to behave. Nothing stops two containers from both trying to allocate more VRAM than the card has, at which point the second allocator -gets a CUDA out-of-memory error, not a scheduling wait. +gets a CUDA out-of-memory error, not a scheduling wait. On this box the +question is sharper still, because the host has **two** cards: an RTX 4000 +Ada for compute and a Quadro P2200 reserved for the Windows sandbox VM's +passthrough. `count: all` would hand both to one container, which is the +#1539 bug the overlay's pinned `device_ids` exists to prevent. This repo's own answer to that (see [`gpu-ml-worker-acceleration.md` §5, "GPU Sharing Contract with the LLM diff --git a/docs/gpu-llm-analysis-worker.md b/docs/gpu-llm-analysis-worker.md index 7bbd08b34..e063f751e 100644 --- a/docs/gpu-llm-analysis-worker.md +++ b/docs/gpu-llm-analysis-worker.md @@ -86,10 +86,19 @@ worker milestones. **Settled by [#602](https://github.com/Xore/APIARY/issues/602)** — an earlier draft of this table (before #602) named the card as a Quadro RTX 4000 at compute capability 7.5/Turing; that card was never on this host. -`lspci` shows a single AD104GL controller and containers enumerate exactly -one device. Also pinned as the runtime-governance authority in +`lspci` showed a single AD104GL controller and containers enumerated one +compute device. Also pinned as the runtime-governance authority in `analysis/ghidra/models/approved-models.json`. +> **A second card arrived after #602.** #1539 recorded that the box also +> carries a **Quadro P2200**, reserved for the Windows sandbox VM's +> passthrough. Every VRAM budget in this document is still correct — they +> are budgets against the Ada, and the P2200 is not part of the compute +> pool — but the host is no longer single-GPU, so "containers enumerate one +> device" no longer describes the machine. It is why the Ollama reservation +> in `analysis/ghidra/docker-compose.ghidra.gpu.yml` pins `device_ids` to +> the Ada's UUID rather than using `count: all`. + | Fact | Value | Verify with | |---|---|---| | GPU | NVIDIA RTX 4000 Ada Generation | `nvidia-smi -L` | @@ -98,8 +107,8 @@ one device. Also pinned as the runtime-governance authority in | Driver / CUDA | 580.173.02 / CUDA 13.0 | `nvidia-smi` | | Container GPU passthrough | nvidia-container-toolkit 1.19.1, `nvidia` runtime registered | `docker info \| grep -i runtime` | | End-to-end container test | `docker run --rm --gpus all nvidia/cuda:12.4.0-base-ubuntu22.04 nvidia-smi -L` lists the GPU | run it | -| Host RAM / CPU | 91 GiB / 16 logical CPUs | `free -h`, `nproc` | -| Stack deployment | Dockge stack at `/opt/stacks/apiary/compose.yml`, containers `hp-*` | `docker ps` | +| Host RAM / CPU | 92 GiB / 48 logical CPUs | `free -h`, `nproc` | +| Stack deployment | Arcane gitops, on disk under `/var/dockge/stacks//`, containers `hp-*` | `docker ps` | | Internal network | `honeynet` (Elasticsearch and all sensors live here) | `docker network ls` | | Elasticsearch | 8.13.4 single-node, `xpack.security.enabled=false`, reachable as `http://elasticsearch:9200` inside `honeynet` | compose file | @@ -269,6 +278,12 @@ design sketch retained for context and must not be copied into production: — bounded U1-only production acceptance with no payload mounts; - [`../llm-worker/docker-compose.captured-data.yml`](../llm-worker/docker-compose.captured-data.yml) — separately authorized #83 network and read-only volume grant; +- [`../llm-worker/docker-compose.captured-data-deploy.yml`](../llm-worker/docker-compose.captured-data-deploy.yml) + — the **deployed** entrypoint: this is `dockerComposePath` for the + `llm-worker` entry in `arcane/manifests/home-production.json`, and it is + what a bare `docker-compose.yml` bring-up is missing (#2234). A local + `docker compose up` against `llm-worker/docker-compose.yml` alone will + not reproduce the live worker. - [`../analysis/ghidra/docker-compose.ghidra.yml`](../analysis/ghidra/docker-compose.ghidra.yml) — the pinned shared Ollama service and narrow `honeypot-llm` network. @@ -303,8 +318,12 @@ docker compose \ config --quiet ``` -The live stack is managed by Dockge under `/opt/stacks`. Deploy only from a -reviewed merged revision; do not maintain a second hand-edited Compose copy. +The live stack is managed by Arcane gitops from +`arcane/manifests/home-production.json`, not by Dockge. (`/opt/stacks` still +exists on the host, but only as a compatibility symlink to +`/var/dockge/stacks`, added 2026-09-04 — nothing is managed there.) +Deploy only from a reviewed merged revision; do not maintain a second +hand-edited Compose copy. Model pulling stays an explicit operator action in the Ghidra/Ollama stack. --- @@ -477,21 +496,31 @@ Retention: ILM 90 days is sufficient — derived data, recreatable from raw. Mirrors the pattern `ml-anomalies` already established ([`ml-worker-plan.md` §8–9](ml-worker-plan.md)): -- **Delivered (#150):** `GET /api/llm/analysis?doc_type=&severity=&since=&limit=` +- **Delivered (#150):** `GET /api/v1/store/llm-analysis` (paged via `offset`/`size`) → documents from `llm-analysis`, newest first, polled on the dashboard's existing 1-minute ES ticker (same transport decision as `ml-anomalies`, no new broker). `/llm-analysis` page: session summaries and payload triage in one filterable table, every row labelled "AI-generated" and showing severity/confidence, with an evidence link back to the - originating session or payload where one exists (`dashboard/llm_analysis.go`). + originating session or payload where one exists + (`frontend-next/src/routes/llm-analysis.tsx`, backed by the generic store + route at `main.rs:438` rather than a route of its own). - **Deferred:** `GET /api/llm/analysis/stream` (SSE via redis channel `llm-analysis-events`) -- optional per this section's original scope ("any SSE/Redis wake-up path remains optional and non-authoritative"); polling has not been shown insufficient yet. -- **Deferred:** semantic search over sessions using `nomic-embed-text` - embeddings stored as a `dense_vector` (384-dim) field on `llm-analysis` - docs, queried with ES kNN search. Still waiting on U1–U3 being stable, - per this section's original scope. +- **Delivered:** semantic search over sessions using `nomic-embed-text` + embeddings stored as a `dense_vector` (768-dim) field on `llm-analysis` + docs, queried with ES kNN search — `GET /api/v1/llm-search` + (`main.rs:341`, `llm_search.rs`, over the `llm-analysis` index's + `doc_type: session` documents). The dimensionality is 768, not the 384 + this document originally stated: #151 confirmed the model's real native + output live against `POST /api/embed` and `llm-worker/worker.py` pins + `EMBEDDING_DIMS = 768`, rejecting a response of any other width outright + rather than indexing it into a mapping it cannot satisfy. This section + originally deferred it + pending U1–U3 stability; it has since shipped, so the list above is not + a statement of current scope. --- diff --git a/docs/gpu-ml-worker-acceleration.md b/docs/gpu-ml-worker-acceleration.md index a21fef3b7..5c944fb7f 100644 --- a/docs/gpu-ml-worker-acceleration.md +++ b/docs/gpu-ml-worker-acceleration.md @@ -73,9 +73,10 @@ parts, and it does not make the models more accurate. ## 3. Hardware & Compatibility Contract **Settled by [#602](https://github.com/Xore/APIARY/issues/602)** (verbatim -host evidence: `lspci` shows a single AD104GL controller, containers -enumerate exactly one device — the earlier two-card / Turing-plus-Ada -hypothesis is refuted) and pinned as the runtime-governance authority in +host evidence: `lspci` showed a single AD104GL controller and containers +enumerated one compute device — the earlier two-card *Turing-plus-Ada* +compute hypothesis is refuted) and pinned as the runtime-governance +authority in `analysis/ghidra/models/approved-models.json`, which `model-governance.py check-runtime` diffs live `nvidia-smi` against every 5 minutes: @@ -84,6 +85,14 @@ hypothesis is refuted) and pinned as the runtime-governance authority in capability 8.9.** (An earlier draft of this document, before #602, mis-recorded this as a Quadro RTX 4000 at compute capability 7.5/Turing — that card was never on this host; see #602 for the full correction.) +- **One more card arrived later, and it is not for this workload.** #1539 + recorded that the box also carries a **Quadro P2200**, reserved for the + Windows sandbox VM's passthrough. So the "one device" finding above is + still true of the *compute* pool this guide budgets VRAM against, but the + host is no longer single-GPU, and this matters concretely in §4.4: a + `count: 1` reservation picks an arbitrary card, which is the same class of + bug `count: all` caused. Pin `device_ids` to the Ada's UUID, as + `analysis/ghidra/docker-compose.ghidra.gpu.yml` already does. - Driver 580.173.02 (CUDA 13.0) — backward-compatible with CUDA 12.x runtime wheels. - nvidia-container-toolkit 1.19.1 present; `docker run --rm --gpus all @@ -118,17 +127,18 @@ Replace the CPU wheel lines: ```diff -# Deep learning (CPU-only PyTorch) --torch==2.13.0+cpu +-torch==2.14.0+cpu ---extra-index-url https://download.pytorch.org/whl/cpu +# Deep learning (CUDA PyTorch — see docs/gpu-ml-worker-acceleration.md §3) -+torch==2.13.0+cu126 ++torch==2.14.0+cu126 +--extra-index-url https://download.pytorch.org/whl/cu126 + +# Embeddings (§6) +sentence-transformers==3.0.1 - # Outlier detection (HBOS) - pyod==3.6.2 + # Outlier detection (HBOS). Pulls in numba+llvmlite -- see the numpy pin + # above; this is the version set actually verified to install together. + pyod==3.6.6 ``` > **Verified pin (2026-08-01, #82):** `torch==2.13.0+cu124` does not exist; @@ -141,7 +151,10 @@ Replace the CPU wheel lines: > including both `sm_75` and `sm_89`, so the wheel's `sm_75` inclusion says > nothing about which kernel the live card actually used; the underlying > install-and-tensor-check result stands, only the architecture label was -> wrong.) Never silently fall back to the `+cpu` wheel; a CPU wheel in a GPU +> wrong.) **The tree has since moved on: `ml-worker/requirements.txt` now +pins `torch==2.14.0+cpu` and `pyod==3.6.6`, so the +cu126 install above was +verified at 2.13.0 only and has not been re-checked at 2.14.0. G2 applies to +whatever version is current when this deploys.** Never silently fall back to the `+cpu` wheel; a CPU wheel in a GPU > deployment must fail the acceptance test T2, not pass unnoticed. ### 4.2 `ml-worker/Dockerfile` @@ -199,6 +212,12 @@ Rules: + ML_DEVICE: auto # auto | cpu — 'cpu' forces CPU for debugging ``` +**Pin `device_ids`, not `count: 1`.** The host carries two cards (§3), so +`count: 1` selects an arbitrary one and can hand `ml-worker` the Quadro +P2200 the Windows sandbox VM needs. Use the Ada's UUID exactly as +`analysis/ghidra/docker-compose.ghidra.gpu.yml` does: +`device_ids: ["GPU-18a00c7e-670a-c305-a2aa-20e3a71917a3"]`. + Keep the existing `mem_limit: 2g` / `cpus: "2.0"`; GPU memory is governed by §5, not by `mem_limit`. @@ -224,7 +243,8 @@ the original 6.1 GiB `qwen3.5:9b` estimate): smaller safety margin instead of full separation.** Worst case — the ghidra slot's model loaded at its 32k context (~14.1 GiB) plus a retrain (~2 GiB) plus the embedder's own inference (~1 GiB) — totals ~17.1 GiB - against the real 20475 MiB budget: about 3.3 GiB of headroom, not the + against the real 20475 MiB budget (19.995 GiB): about 2.9 GiB of headroom, + not the "comfortable" double-digit margin a naive 6.1 GiB-chat-model estimate would suggest. That is enough to not require full separation, but not enough to treat as a non-issue either. This already assumes the ghidra @@ -240,7 +260,7 @@ the original 6.1 GiB `qwen3.5:9b` estimate): best-effort headroom management. - If a future model re-evaluation (#568's process, re-run) picks something materially larger than `qwen3:14b` for the ghidra slot, re-check this - margin before assuming it still holds — the 3.3 GiB headroom above was + margin before assuming it still holds — the 2.9 GiB headroom above was computed for this specific model at its current context ceiling, not as a permanent property of the 20 GB card. - **On CUDA OOM, do not crash:** wrap train/infer calls, catch diff --git a/docs/local-llm-model-evaluation.md b/docs/local-llm-model-evaluation.md index 7a4fea3d8..f7a2725b8 100644 --- a/docs/local-llm-model-evaluation.md +++ b/docs/local-llm-model-evaluation.md @@ -2,6 +2,55 @@ Status: completed for [issue #144](https://github.com/Xore/APIARY/issues/144) and requalified under [issue #158](https://github.com/Xore/APIARY/issues/158), 2026-08-01. Re-evaluated and re-approved under [issue #568](https://github.com/Xore/APIARY/issues/568), 2026-08-05 — see [§ Issue #568 re-evaluation](#issue-568-re-evaluation-real-20gb-card) below; that section is now the current approved state, superseding the v2 table immediately above it. +> **Reconciled 2026-09-27 against `docs/benchmarks/runs/` and +> `docs/benchmarks/matrices/`.** No score in this file was changed. What the +> repository can and cannot confirm: +> +> - **Confirmed exactly.** The round-7 cold baseline reconciles cell for cell +> against `round7-cold-baseline.json`: 91 models, 182 cells, 367 records, 179 +> reproduced, the same 3 escalated cells with the same third-run values, 11 +> zero-scored tags, and every anchor — `qwen3:14b` 85.5 B / 83.1 A, `qwen3:8b` +> 84.3 B / 81.9 A, `qwen2.5:14b-instruct-q4_K_M` 88.0 A, `Trendyol-32B` 95.2 A, +> `Ornith-1.0-35B` 92.8 B, and the 12.1 / 7.3 point gaps. The cold-cohort +> table reconciles against `1947-cohort-cold-protocol.json` (means, ±0 spreads, +> B−A deltas, `min/run B`, Ornith's injection FAIL). The twelve-model survey +> reconciles against `1805c-ghidra-slot-matrix.json`. +> - **The archived runs cannot check the Ghidra column at all.** All 62 stored +> runs — 1498 records — are `revdeck` or `sessions`; there is **no +> `ghidra`-slot record in the repository**, and no transcript field carries +> VRAM or a context-probe result. The Ghidra, 16k-probe and VRAM columns in +> the #144 and #568 tables are therefore **undetermined from the repo**, not +> confirmed and not contradicted. The approved `qwen3:14b@bdbd181c33f2…` and +> `context_tokens: 32768` of the #568 Decision *are* confirmed against +> `analysis/ghidra/models/approved-models.json`. +> - **The archive is a different vintage from #568.** The stored runs are the +> 2026-08-25→29 #1795b / #1947-wave2 / #1805c / #1947seq rounds; #568 was +> measured 2026-08-05. Where the two overlap they differ (`qwen3:8b` sessions +> 94.0 in the archive vs 92.5 here; `qwen2.5:14b-instruct-q4_K_M` 100.0 vs +> 97.0). Those are re-measurements, not errors, and no figure was "corrected" +> to match them. +> - **Part 1's sessions column is one point low on three of eleven rows under +> the current scorer**, reproducibly across all three repeats: +> `Foundation-Sec-1.1-8B-Instruct-i1` 56→**57** (83.6→85.1%), +> `Huihui-Qwen3.6-35B-A3B-abliterated` 66→**67** and stock `qwen3.8:27b` +> 66→**67** (both 98.5→100.0%). The part-1 pins already declare a pre-#2265 +> scorer, so this is a vintage delta — but the part-1 Decision names only two +> rows above the incumbent, and two more reach 67/67 on the current scorer. +> The part-1/part-2 `revdeck` denominator is likewise /16 as printed against +> /19 under the current inline `REV_CASES`. +> - **Seven archived runs are not re-scorable as stored.** They report +> `outcome: ok` with non-zero `output_tokens` and an empty `raw`: four +> `qwen3:14b` runs on 2026-08-25 (51 records) and three gpt-oss-family runs on +> 2026-08-26 (45 records). Re-scoring the archive naively yields 14–18/69 for +> `qwen3:14b` from those four, against the authoritative 60–62/69. Exclude +> them before recomputing anything. +> - **Undetermined:** the round-7 injection figure "only 14 of 91 models fully +> resist (5/5)" — `round7-cold-baseline.json` stores score and percent only, +> with no per-case injection split, so the 5/5 counts have no in-repo source. +> - Method check: re-deriving the 14-case corpus from the stored transcripts +> with `rev_cases_v2_rubric.json` + `polarity.forbidden_hit` reproduces +> `1805c-ghidra-slot-matrix.json` exactly, all 12 models at both tiers. + This is a task-specific decision record for the three independent local-model slots in this repository. It does not assume that a model named in an earlier plan is suitable, or that one model should serve all three jobs. @@ -1375,7 +1424,13 @@ Tier A — the incumbent sits 7-12 points under the security-specialized leaders (12.1 points Tier A vs `Trendyol-32B`'s 95.2%, 7.3 points Tier B vs `llmfan46/Ornith-1.0-35B`'s 92.8%). Top band by run-pooled mean total_score (mean across all 4 runs per model — 2 Tier A + 2 Tier B; not the same scale -as the percentages above): `phi4:14b` 77.0, `Trendyol-32B` Q8_0 76.5, +as the percentages above), **excluding the three entries whose Tier B cell was +escalated to a third run and so has a 5-run, not 4-run, denominator** — +`gemma-4-26B-A4B-it-ultra-uncensored-heretic` `Q4_K_M` 90.75, +`Foundation-Sec-1.1-8B-Instruct` `Q8_0` 87.25 and +`XORTRON.CriminalComputing.LARGE.2026.3` `i1-IQ2_XXS` 83.25, all three above +everything listed here: +`phi4:14b` 77.0, `Trendyol-32B` Q8_0 76.5, `VulnLLM-R-7B` i1-Q4_K_M 76.5, `Huihui-CyberStrike-OffSec-35B` q6_k 75.5, philbert440 `Qwen3.8-27B-Cyber` 75.5, protoLabsAI `ThinkingCap-Qwen3.6-27B-MTP` `latest` 75.5 (the successfully-pulled diff --git a/docs/ml-gpu-coordinated-roadmap.md b/docs/ml-gpu-coordinated-roadmap.md index db20c8f66..3b2cb5122 100644 --- a/docs/ml-gpu-coordinated-roadmap.md +++ b/docs/ml-gpu-coordinated-roadmap.md @@ -1,6 +1,15 @@ # Coordinated ML and GPU Analysis Roadmap -> **Status:** Proposed implementation sequence +> **Status:** Proposed implementation sequence — **intent, not shipped state.** +> Re-measured 2026-09-27: `ml-worker` and `llm-worker` now ship as their own +> Arcane-managed stacks; the ML worker's GPU overlay +> (`docker-compose.ml-worker.gpu.yml`) is still inert scaffolding, so the ML +> worker remains CPU-only; retrain slots are `03:00,09:00,15:00,21:00` +> (`ml-worker/worker.py:59`), not §4-I's `01:00,07:00,13:00,19:00`; the +> default alert threshold did land at `0.75` (`worker.py:174`); and §1 +> decision 5 is superseded — the ML worker never grew an embedding index at +> all, embeddings shipped in `llm-worker` at **768** dims, off by default. +> The milestone text below is left as the historical record. > > **Scope:** `ml-worker`, GPU acceleration, local LLM analysis, and dashboard delivery > diff --git a/docs/ml-worker-evaluation.md b/docs/ml-worker-evaluation.md index 9da6a953a..0ddbfab87 100644 --- a/docs/ml-worker-evaluation.md +++ b/docs/ml-worker-evaluation.md @@ -42,10 +42,16 @@ reported alongside accuracy rather than ignored. ### 2026-08-25 — Tier 1 harness landed; Tier 2 blocked -**Tier 1 (behaviour) is live.** `evaluate_detectors.py` runs six contract +**Tier 1 (behaviour) is live.** `evaluate_detectors.py` runs seven contract checks against a candidate over the per-sensor fixture corpus and emits a hashed JSON report. It is the first reusable acceptance bar `ml-worker` has -had, and it makes "evaluate this candidate offline" answerable at all. +had, and it makes "evaluate this candidate offline" answerable at all. The +sixth of the seven, `check_composite_renormalises_over_present_detectors` +(#1969), is not a per-candidate probe at all: it calls the production +`worker.compute_composite()` directly, so candidates are always compared +under the rules they will actually run with — absent detectors drop out of +both numerator and denominator, an event no detector opines on composites +to 0.0, and a single-detector opinion stands at face value. **Tier 2 (accuracy) was blocked on [#1797](https://github.com/Xore/APIARY/issues/1797).** There was no labelled corpus at that point. The date is retained as the diff --git a/docs/ml-worker-plan.md b/docs/ml-worker-plan.md index a0afcaa36..75f244c2b 100644 --- a/docs/ml-worker-plan.md +++ b/docs/ml-worker-plan.md @@ -1,16 +1,24 @@ # ML Worker — Implementation Plan +> **Reading note (2026-09-27):** this is a dated plan/record, not a live +> reference. §2, §5.3, §7, §8, §9, §10, §11.4 and §11.6 describe shipped +> behaviour and were re-checked against the code. §1, §4.1, §6, §11.2 and §12 +> keep their original-draft wording where it was never rewritten — including +> the Go dashboard's `dashboard/ml_anomalies.go` / `settings_domain.go` +> references, which are historical since #1628. + > **Status (2026-08-27, #1662):** the plan largely executed as written: -> `ml-worker/worker.py` runs the ensemble described here from the same -> repo-root compose files. What moved: consumers of its output live in the +> `ml-worker/worker.py` runs the ensemble described here from its own +> Arcane-managed stack, not the repo-root compose file (see §10). +> What moved: consumers of its output live in the > backend-service tier now, not `dashboard/ml_anomalies.go` (deleted). > Open scoring-semantics defects are tracked in issues #1946/#1969 under > epic #1974 rather than here. > **Status:** `ml-worker/` has its own Arcane-managed stack -> ([`docker-compose.yml`](../ml-worker/docker-compose.yml) + -> [`docker-compose.ml-worker.gpu.yml`](../ml-worker/docker-compose.ml-worker.gpu.yml), -> mirroring `analysis/ghidra/`), builds, connects to Elasticsearch, and polls +> ([`docker-compose.yml`](../ml-worker/docker-compose.yml), the only one the +> manifest deploys). `docker-compose.ml-worker.gpu.yml` exists in-tree but is +> inert and undeployed, so the worker is CPU-only in practice. It > without crashing (#62). `extract_features()`/`featurise_temporal()` read > the real per-sensor schema (#62 task 33, #63). The dashboard delivers > scores via the backend-service's `/api/v1/store/ml-anomalies` + @@ -27,7 +35,10 @@ > `docker build ./ml-worker` failed outright (`pyod`'s `numba` dependency had > no version compatible with the pinned `numpy==2.5.1` on Python 3.12 — > reproduced twice, locally and in-container; fixed in #62 by pinning -> `numpy==2.4.6`/`numba==0.66.0`/`llvmlite==0.48.0`). `worker.py`'s +> `numpy==2.4.6`/`numba==0.66.0`/`llvmlite==0.48.0`; those two transitive pins +> have since moved on again and the file now reads `numba==0.67.0` / +> `llvmlite==0.49.0` / `pyod==3.6.6`, so treat this paragraph as the record of +> what #62 did, not of the current requirements). `worker.py`'s > `SOURCE_INDICES` (`cowrie-*`, `dionaea-*`, `honeypot-network-*`, `conpot-*`, > `http-honeypot-*`) still match zero indices on the live homeserver: the real > shape is a unified `honeypot-v2-*` stream (all sensors, disambiguated by @@ -90,16 +101,17 @@ ground-truth labels. [web:275][web:283] > the "v0.1 audit verdict" callout above. `worker.py`'s real, > currently-deployed `SOURCE_INDICES` are the two rows below. -The worker ingests from two unified, versioned index patterns +The worker ingests from three unified, versioned index patterns (`ml-worker/worker.py`'s `SOURCE_INDICES`): | Index pattern | Source | Key fields | |---------------|--------|------------| | `honeypot-v2-*` | every honeypot sensor (Cowrie, Dionaea, Conpot, HTTP-honeypot, and every other sensor stack — disambiguated by `event.sensor`, not a separate index per sensor) | `event.sensor`, `source.ip`, `honeypot.*` (per-sensor nested fields, not uniform across sensors — see §5.3) | | `suricata-v2-*` | Suricata network/IDS events (Filebeat) | `suricata.eve.*`, `network.*`, `alert.signature` | +| `zeek-v1-conn-*` | Zeek connection records, added by #1774's sensing layer alongside Suricata | Zeek conn-log fields; note the pattern is `zeek-v1-conn-*`, not a `zeek-v2-*` line like the other two | -Both index patterns share a common `@timestamp` field used for temporal -ordering. A third index, `ml-worker-state`, is not a data source — it's the +All three index patterns share a common `@timestamp` field used for temporal +ordering. A fourth index, `ml-worker-state`, is not a data source — it's the worker's own per-index-pattern checkpoint store (`load_checkpoint`/ `save_checkpoint` in `worker.py`): a `last_timestamp` plus the set of already-seen event IDs at that exact timestamp, so a restart resumes @@ -117,9 +129,11 @@ flowchart TD subgraph Stack["APIARY (existing)"] Sensors["every honeypot sensor stack
(disambiguated by event.sensor,
not a separate index each)"] Suricata["Suricata / network IDS"] - ES["Elasticsearch
honeypot-v2-*, suricata-v2-*"] + Zeek["Zeek conn records
(#1774)"] + ES["Elasticsearch
honeypot-v2-*, suricata-v2-*,
zeek-v1-conn-*"] Sensors --> ES Suricata --> ES + Zeek --> ES end subgraph Worker["ML Worker (ml-worker/)"] @@ -168,8 +182,9 @@ different anomaly types: [web:275][web:276][web:283][web:292] mixed numerical+categorical features after encoding. Proven on network logs. [web:276][web:290] - **Implementation:** `scikit-learn` `IsolationForest` with `contamination=0.01` - (assume 1% of events are anomalous). Retrained every 6 hours on a 24h - rolling window. + (assume 1% of events are anomalous). Retrained at four fixed UTC slots + daily (`RETRAIN_SLOTS_UTC`, default `03:00,09:00,15:00,21:00` — #172 + replaced the old 6h `RETRAIN_INTERVAL`) on a 24h rolling window. - **Output:** `anomaly_score` ∈ [-1, 0] where values closer to -1 = more anomalous. ### 4.2 LSTM Autoencoder (LSTM-AE) @@ -391,7 +406,7 @@ loop every POLL_INTERVAL seconds (default: 30s): 10. Sleep POLL_INTERVAL -Every 6 hours (RETRAIN_INTERVAL): +At each `RETRAIN_SLOTS_UTC` slot (four daily, default 03:00,09:00,15:00,21:00 UTC, #172): - Retrain IsoForest + HBOS on last 24h of all events - Fine-tune LSTM-AE on last 24h (5 epochs, low LR) - Save new model checkpoint to /models/ @@ -786,7 +801,7 @@ if wanted. ### 11.2 Online learning (unchanged from the original draft, still accurate) ``` - → HBOS/IsoForest: full retrain every RETRAIN_INTERVAL (default 6h) on the + → HBOS/IsoForest: full retrain at each `RETRAIN_SLOTS_UTC` slot (four daily, #172) on the rolling 24h window, gated by §11.1 → LSTM-AE: fine-tune on the same cycle (5 epochs, LR=1e-5), gated the same way (§11.1's anomaly-rate check applies to its reconstruction-loss-based @@ -831,7 +846,7 @@ fraction `>= THRESHOLD` exceeds `DRIFT_ANOMALY_RATE` (default `0.15`, matching the original draft's "15%"): - an early retrain is triggered (the next poll cycle retrains regardless of - how much of `RETRAIN_INTERVAL` remains), and + which `RETRAIN_SLOTS_UTC` slot is nearest), and - a `ml-worker-metrics` document is written flagging the drift event (`kind: "drift"`, the observed rate, window size) so the dashboard's `/ml-anomalies` page (#64) — or a future panel reading this index directly @@ -885,9 +900,11 @@ its contract carried over unchanged to the Rust config module.) | **v0.8** | Retraining scheduler + model versioning | [#65](https://github.com/Xore/APIARY/issues/65) | | **v1.0** | Drift detection + alert threshold tuning UI | [#65](https://github.com/Xore/APIARY/issues/65) | -v0.1 is listed as an issue rather than as done on purpose. `ml-worker/` holds a -Dockerfile, `worker.py`, and a `docker-compose.override.yml`, but it is not a -service in the root Compose file, it has no tests or fixtures, and nothing here +v0.1 is listed as an issue rather than as done on purpose. As of the #61 audit +— before #62's rewrite — `ml-worker/` held a +Dockerfile, `worker.py`, and a `docker-compose.override.yml` (since deleted, see +§10); it was not a +service in the root Compose file, it had no tests or fixtures, and nothing here has been observed running against live data. #61 is the audit that decides whether the scaffold is a v0.1 or a starting point. diff --git a/docs/payload-analysis-workbench.md b/docs/payload-analysis-workbench.md index c15326a70..684b3301a 100644 --- a/docs/payload-analysis-workbench.md +++ b/docs/payload-analysis-workbench.md @@ -1,6 +1,6 @@ # Payload analysis workbench -The dashboard's `/payload-workbench` route selects captured evidence and `/payload-workbench/{sha256}` is the unified orchestration surface for issue #155. It is a separate page rather than an extension of `/ghidra`: recipes and parent runs span deterministic analysis, Ghidra, and two sandbox backends, while `/ghidra/{sha256}`, `/sandbox/{job}` and `/payload-analysis/{sha256}` remain the canonical native result renderers. +The dashboard's `/payload-workbench/results` route is the owner-isolated review surface, and its `workbench-builder` section is the unified orchestration surface for issue #155. It is a separate page rather than an extension of `/ghidra`: recipes and parent runs span deterministic analysis, Ghidra, and two sandbox backends, while `/ghidra/{sha256}`, `/sandbox/{job}` and `/payload-analysis/{sha256}` remain the canonical native result renderers. ## Trust boundary @@ -15,7 +15,7 @@ Run ownership is likewise never taken from client input. The Rust tier derives i ## Analyzer registry Seven analyzer IDs, one server-computed `workbenchAnalyzer` registry -(`dashboard/workbench_domain.go`). A run selects 1-5 of them; the server +(`backend-service/src/workbench_domain.rs`). A run selects 1-5 of them; the server rejects zero selections, more than 5, an unknown ID, or a duplicate. | ID | Applicability | Adapter | Result link | Concurrency class | @@ -109,24 +109,27 @@ degraded to a stale local copy (#405 follow-up). ## HTTP contracts -All APIs require a live administrator identity. Every mutation additionally requires a same-origin request, `application/json`, one document no larger than 64 KiB, and the closed Go schema (unknown fields are rejected). Create, read/list, reconcile, cancel, and retry are further scoped server-side to the caller's verified owner identity (see Trust boundary) — a request's `owner` field is ignored entirely, never an access decision (#3110). +All APIs require a live administrator identity. Every mutation additionally requires a same-origin request, `application/json`, and the closed request schema. (A 64 KiB body cap and unknown-field rejection were part of the original Go contract; neither is present in the Rust tier, so do not rely on them.) Create, read/list, reconcile, cancel, and retry are further scoped server-side to the caller's verified owner identity (see Trust boundary) — a request's `owner` field is ignored entirely, never an access decision (#3110). + +Routes are registered in `backend-service/src/main.rs:456-468`. The hash is a **query** parameter on the registry and list routes, not a path parameter, and `cancel`/`retry` share one route via an `{action}` segment rather than being separate paths. | Method and route | Purpose | |---|---| -| `GET /api/payload-workbench/registry/{sha256}` | server-derived registry, applicability, external-publication notice, and advisory model health | -| `GET /api/payload-workbench/recipes` | visible private/shared recipe revisions | -| `POST /api/payload-workbench/recipes` | append an immutable recipe revision | -| `GET /api/payload-workbench/runs?sha256=...` | recent parent runs for the caller and payload | -| `POST /api/payload-workbench/runs` | submit a saved revision or typed one-off selection | -| `GET /api/payload-workbench/runs/{run_id}` | reconcile and return one parent run | -| `POST /api/payload-workbench/runs/{run_id}/children/{analyzer_id}/retry` | bounded deliberate retry | -| `POST /api/payload-workbench/runs/{run_id}/children/{analyzer_id}/cancel` | cancel an exact pending marker when supported | +| `GET /api/v1/workbench/analyzers?hash=…` | server-derived registry, applicability, external-publication notice, and advisory model health | +| `GET /api/v1/workbench/recipes` | visible private/shared recipe revisions | +| `POST /api/v1/workbench/recipes` | append an immutable recipe revision | +| `GET /api/v1/workbench/runs?hash=…&limit=…` | recent parent runs for the caller and payload | +| `POST /api/v1/workbench/runs` | submit a saved revision or typed one-off selection | +| `GET /api/v1/workbench/runs/{id}` | reconcile and return one parent run | +| `POST /api/v1/workbench/runs/{id}/children/{analyzer_id}/{action}` | bounded deliberate retry, or cancel an exact pending marker when supported (`action` ∈ `retry`\|`cancel`) | Create, recipe-save, retry, and cancel outcomes use the existing dashboard audit sink. Audit fields name the contract fields but do not copy payload content, prompts, model replies, filenames, credentials, or tool output. ## Model-status adapter -`/var/lib/honeypot-ghidra/model-status.json` remains root-owned mode `0600`. `honeypot-model-status-adapter.service` reads and re-validates it, strips every field outside schema v1, and serves only `GET /v1/status` over `/run/honeypot-model-status/status.sock`. The dashboard mounts that runtime directory read-only and uses `MODEL_STATUS_SOCKET=/model-status/status.sock`. There is no TCP listener and no write, pull, replace, promote, prompt, or model-selection route. +`/var/lib/honeypot-ghidra/model-status.json` remains root-owned mode `0600`. `honeypot-model-status-adapter.service` reads and re-validates it, strips every field outside schema v1, and serves only `GET /v1/status` over `/run/honeypot-model-status/status.sock`. The dashboard is expected to mount that runtime directory read-only and use `MODEL_STATUS_SOCKET=/model-status/status.sock`. There is no TCP listener and no write, pull, replace, promote, prompt, or model-selection route. + +> **Undetermined:** `MODEL_STATUS_SOCKET` has no occurrence anywhere in the Rust backend, the compose files, or `.env.example` as of 2026-09-27 — the adapter half of this contract is real and installed, but the consumer side is not visible in the repository. Treat the socket wiring as intended-but-unwired rather than working, and confirm before relying on the advisory model-health field the registry returns. Re-run `sudo analysis/ghidra/install-analysis-host.sh` to install or update the adapter. Its failure only displays `unavailable`; it never disables a worker. @@ -138,7 +141,7 @@ Deploy the dashboard normally after merging. Rollback is additive and safe: 1. deploy the previous dashboard image; 2. optionally disable `honeypot-model-status-adapter.service`; -3. leave the workbench indices in Elasticsearch untouched (a rolled-back dashboard from before the #405 follow-up reads its own local `/state/analysis-workbench` copy instead and simply does not see runs created after the rollback). +3. leave the workbench indices in Elasticsearch untouched (a rolled-back pre-workbench dashboard does not see runs created after the rollback). The old `/ghidra/submit` and `/sandbox/submit` routes remain compatible. No worker or native result schema is changed by the workbench. diff --git a/docs/sandbox/README.md b/docs/sandbox/README.md index 94a28065d..28ca9ac58 100644 --- a/docs/sandbox/README.md +++ b/docs/sandbox/README.md @@ -1,12 +1,16 @@ # Hard-isolated malware sandbox plan The homeserver supports this design: Intel VT-x is enabled, `/dev/kvm` is -available, KVM is loaded, and the host exposes 84 IOMMU groups. The foundation +available, KVM is loaded, and IOMMU is on with **94 groups** (re-measured +2026-09-27). This was not free on Rocky: #1609 recorded 87 groups on the old +Ubuntu host, and the 2026-09-03 Rocky 10 rebuild came up with **zero**, so +`install-homeserver.sh`'s `step_vfio_gpu_passthrough` adds `intel_iommu=on` +to the kernel command line. The foundation installer provisions system libvirt and the dedicated isolated network. ## Route selection and evidence return across four dynamic-detonation routes -The workbench's registry (`dashboard/workbench_domain.go`) offers four +The workbench's registry (`backend-service/src/workbench_domain.rs`) offers four routes to dynamic detonation, not one sandbox with options. Each is its own guest, network, spool, and result format — this section is the canonical side-by-side comparison; each route's own internal detail lives in its own @@ -15,7 +19,7 @@ section below (Linux) or its own directory (`sandbox/windows/`, ```mermaid flowchart TB - workbench["Payload workbench —
analyst selects a route
(dashboard/workbench_domain.go)"] + workbench["Payload workbench —
analyst selects a route
(backend-service/src/workbench_domain.rs)"] subgraph linuxRoute["linux-sandbox"] direction TB @@ -41,7 +45,7 @@ flowchart TB capeGuest["Windows guest under CAPE's
own cuckoo.py orchestration
(runs on the host directly,
not in Docker) + cape-mongo"] end - sharedLock{{"honeypot-kvm-detonation.lock —
shared ONLY between windows-sandbox
and cape (#320): 16 logical CPUs total,
win11-sandbox alone already 8 vCPU.
Held only around the actual detonation
call, not the whole drain loop.
linux-sandbox and windows-ghosts are
NOT part of this lock — independent,
can run concurrently with anything."}} + sharedLock{{"honeypot-kvm-detonation.lock —
shared ONLY between windows-sandbox
and cape (#320): 48 logical CPUs total,
win11-sandbox alone already 8 vCPU.
Held only around the actual detonation
call, not the whole drain loop.
linux-sandbox and windows-ghosts are
NOT part of this lock — independent,
can run concurrently with anything."}} workbench -->|"hash-only request,
SANDBOX_REQUEST_DIR"| linuxRoute workbench -->|"hash-only request,
WINDOWS_SANDBOX_REQUEST_DIR"| winRoute @@ -72,7 +76,7 @@ credential ever crosses the dashboard/host boundary for any of the four. never wait on anything.** `sandbox/windows/run_pending.sh` and `sandbox/cape/worker/cape-worker.py` share one host-wide `honeypot-kvm-detonation.lock` (#320) — a real capacity constraint, not a -correctness one: both are KVM/QEMU domains on the same 16-logical-CPU host, +correctness one: both are KVM/QEMU domains on the same 48-logical-CPU host, and `windows-sandbox`'s own guest is already configured for 8 vCPU. The lock is held only around the actual detonation call, never the whole drain loop, so an idle worker on either side never blocks the other. `linux-sandbox` @@ -235,6 +239,12 @@ sudo bash /opt/stacks/apiary/sandbox/install-windows-forensics.sh ## Required operating controls - Reserve at most 4 vCPU and 8 GiB RAM per analysis VM; run one job initially. + For scale, the Windows analysis domain is currently defined at **8 vCPU / + 16 GiB** (`sandbox/windows/packer/win11-kvm.xml`), and + `sandbox/sandbox.env.example` ships `SANDBOX_VM_MEMORY_MB=3072` as its + default — so this line's 4 vCPU / 8 GiB guidance matches neither figure + exactly. It is operator guidance rather than a measured limit; settle it + against the real per-VM reservation before relying on it. - Enforce a 10-minute hard timeout and kill QEMU if graceful shutdown fails. - Store golden images on root-owned storage and verify SHA-256 before every job. - Sign/validate result JSON and treat all guest-produced text as untrusted.