Skip to content

merge: docs reconciliation round 3, analysis and inference slice (#3399) - #3411

Merged
Xore merged 6 commits into
docs/3395-doc-reconciliationfrom
docs/3399-r3-b
Sep 27, 2026
Merged

Xore merged 6 commits into
docs/3395-doc-reconciliationfrom
docs/3399-r3-b

Conversation

@Xore

@Xore Xore commented Sep 27, 2026

Copy link
Copy Markdown
Owner

Round 3 of the #3399 docs reconciliation, analysis and inference slice — 6 files.
Companion to #3409. Follows #3398 (r1), #3404 (map), #3405 (infra).

Gates — all six green

mermaid            40 blocks in 154 files parse cleanly
links              413 local refs resolve
paths-exist        491 tokens, 41 allowlisted
reachability       85 reachable, 35 exempt
stale-paths        passed
public-leaks       passed

docs/llm-inference-backend-comparison.md

The most drift-prone document in the set, and it had drifted in both directions:

docs/analysis/ghidra/models/README.md

Corrected the post-#2394 GPU-identity states — the doc described model
identities in terms that the numbering change invalidated.

docs/kvm-network-traffic-analysis.md

Corrected the Windows results path.

docs/persona-design.md

  • Disambiguated the T-Pot README citation, which resolved to the wrong tree.
  • Fixed a misclassified sensor.

docs/analysis/RECOVERY.md

Fixed an unscoped runbook step.

Note on process

The final dispatch was interrupted at its reporting step; the work itself was
committed and is what is in this PR. Per-file outcomes are recorded in the
issue comment.

The 2026-08-05 CPU-only research pass had two sections that later evidence
contradicts, plus one pin that has since moved. Both are now marked rather
than rewritten, so the record of what was known on that date survives.

- Status banner: the measured three-engine run lives in
  analysis/ghidra/benchmarks/engine-benchmark/README.md (2026-08-06, plus
  #832's settings tuning). It corroborates 1, 3 and 4; it overturns 6 and 7.
- Section 6: real-card throughput was in fact measured (llama.cpp ~21.4,
  Ollama ~21.5, vLLM ~23.1 tok/s) -- but on REx86 f16 7B weights, so the
  qwen3:14b comparison this doc asked for is still open.
- Section 7: vLLM's "refuses to start" failure was specific to the previous
  8GB Quadro RTX 4000 and does not reproduce on the current 20GB card. The
  stronger "the standard distribution has no CPU path at all" reading
  overstates what was shown, so it is now recorded as undetermined rather
  than restated. The image-size row and section 2's digest finding stand.
- Method: 0.32.0 was the repo's pin on 2026-08-05; production has since moved
  to ollama/ollama:0.32.13 in the ghidra compose and approved-models.json.
- Opening claim narrowed: no llama.cpp or vLLM service is deployed (39-entry
  manifest confirms), though both are now exercised in benchmark scripts.
…nbook

persona-design.md's air-gap table classified every non-Cowrie/Dionaea/Tanner
sensor as "no design reason either way". That is wrong for canarytokens: it
is a honeytoken platform whose public reachability is the product (#1487
made the switchboard's HTTP channel reachable through the VPS so planted
file/doc tokens actually fire, and that stack's compose file says inert
internal-only tokens do not serve it). The capture-vs-safety tradeoff the
doc argues does not transfer in either direction, so it gets its own row.
The catch-all enumeration was also missing nine deployed sensor networks;
they are now listed.

RECOVERY.md's numbered restore steps name SHA256SUMS,
stack-config-state.tar.gz and keycloak.sql.gz, which are
analysis/backup-honeypot.sh's on-host layout and resolve nowhere else --
yet the doc's stated purpose is restoring after the homeserver is gone,
which is precisely when that archive does not exist. The steps are now
scoped to the archive they actually describe, and the surviving
workstation archive is pointed at its own procedure, including the
install-homeserver.conf prerequisite those steps do not cover.
…esults

The 2026-08-05 research pass had drifted in four ways, all verified against
stored records rather than restated from other docs:

- Section 2 cited a function that does not exist. The drift-comparison entry
  point in analysis/ghidra/models/model-governance.py is evaluate_drift(),
  not compare_against_approved().
- The status banner claimed the later three-engine run "corroborates section
  1, 3 and 4". It does not. That run tests one engine at a time by
  construction, so it says nothing about structured-output enforcement or
  keep-alive/swap, and it explicitly leaves section 2 standing. Narrowed the
  claim to what the stored record supports: overturns 6 and 7, leaves 2,
  does not re-test 1 and 3.
- Section 4's premise -- that temperature 0 with a fixed seed is relied on for
  reproducible output today -- was overtaken by #2646. With
  OLLAMA_KEEP_ALIVE=30m a warm slot returns different text for a
  byte-identical prompt, so production is permanently in the drifting regime
  and the shipped fix is slot_generation accounting, not determinism. Noted as
  a dated revision; the per-engine measurements stand.
- Section 8 listed ghidra-worker.py as needing a client-shape change. It
  already speaks an OpenAI-compatible /v1 dialect deliberately, precisely so
  llama.cpp/vLLM/LM Studio work unchanged. The only residual coupling is the
  /api/ps slot_generation probe, which already degrades to "unavailable" on a
  non-Ollama server.

Also keeps the pending uncommitted work in this file, verified: the measured
decode throughput (llama.cpp ~21.4, Ollama ~21.5, vLLM ~23.1 tok/s), the
2026-08-06 date and RTX 4000 Ada 20GB host, and the ollama 0.32.13 production
tag were each read out of analysis/ghidra/benchmarks/engine-benchmark/README.md
and the two approved-models.json / docker-compose.ghidra.yml runtime blocks.
…s README

The section read "Two expected, honest states after #2394 (not regressions)"
and told operators that host_gpu_uuid_changed means "expected for old
snapshots, not evidence of an actual UUID change". The checker disagrees
with the second half of that.

model-governance.py builds the code as f"host_{key}_changed" over a
per-field comparison, so host_gpu_uuid_changed fires in two situations that
need opposite responses: a pre-#2394 snapshot with no recorded UUID
(benign), and a snapshot whose UUID names a different physical card (real
drift, in which every field the old schema did compare -- name, memory,
driver -- can still look correct). An operator following the old wording
would dismiss a wrong-card event. The code string stays, so
tests/docs/test_2409_fix.py still passes, but it now carries the
present-vs-absent test for telling the cases apart, and the sibling
per-field codes are named.

Also adds approved_gpu_absent, a third #2394 host-leg code the section
never mentioned: the tool ran, was pointed at the approved UUID, and no
such card exists. Unlike the other two this is not an expected rollout
state, so the heading no longer claims they all are.

Separately, "30 days is the recorded recommendation" cited a retention
record that does not exist -- the only 30-day window under
analysis/ghidra/models/ is BLOB-RETENTION-POLICY.md's unrelated #2862
dionaea-bistream rule. The advice stands, attributed to this doc.
One drift found. The doc cited "`README.md` line 296" for T-Pot's outbound
guidance without saying whose README. In this repository that pointer resolves
to the root README.md, which is 135 lines and contains no such text, so a
reader following it finds nothing.

The line number refers to T-Pot's upstream README, recorded from an unpinned
read. Outbound network access was unavailable during this reconciliation, so
the number could not be re-verified and is not restated as fact -- the
citation is now labelled as a lookup pointer rather than a stable one. The
substantive claim is unaffected: the doc already quotes T-Pot's outbound text
inline, and that quote is what the decision rests on.

Everything else in this file verified against the tree with no drift:

- All 21 sensor stacks named in the outbound table exist under arcane/home/,
  and the four named networks (cowrie_net, dionaea_net, tanner_local,
  canarytokens_net) plus the yara-scanner service's network_mode: none are as
  described.
- "no honeypot's network has ever set internal: true" still holds: exactly
  three files set it (oidc-session, honeypot-llm-data, keycloak-data) and all
  three are infrastructure networks, not sensor capture networks.
- The three air-gap flags and their default-false wiring match compose and
  .env.example in all three stacks.
- The "one org per estate" example is real: nexusai-core, nexusai-edge and
  nexusai-platform are site_ids in personas/personas.json, all under NexusAI
  Research GmbH, wired into honeypot-multipot and honeypot-http.
- cowrie.cfg's banner string, COWRIE_HOSTNAME=gpu01, and the canarytokens
  "inert (internal-only) tokens don't serve it" quote are exact.
The short-version table gave the Windows capture results location as
"sandbox/results/<run>/". Nothing tracked lives there, and the row conflated
two different directories.

The run-keyed artifact directory -- the one the dashboard actually reads
(es_results_importer.rs and sandbox_submit.rs both resolve
WINDOWS_SANDBOX_RESULTS_DIR) -- is $WINDOWS_SANDBOX_RESULTS_DIR, with one
subdirectory per sample named by its sha256, per run_sample.py's detonate().
Its own fallback when that env var is unset is reports/windows-sandbox.

sandbox/results/current is real but is only docker-compose.sandbox.yml's
bare default for a manual `up`: all five sniffer bind-mounts key off
${SANDBOX_RESULTS_DIR:-./sandbox/results/current}, and run_sample.py points
SANDBOX_RESULTS_DIR at the run's own out_dir, so the default only applies
when the stack is brought up by hand.

Everything else in the doc verified against the tree: both bridges and the
absent <forward>, the 10.10.10.x address plan, macvlan internal:true, the
sniffer NET_ADMIN/NET_RAW + network_mode: host grant versus cap_drop on
INetSim and mitmproxy, the honeypot-sandbox-strict filterref, the
controlled-mode 198.18.0.1 DNS/proxy, and all three retention figures
(cleanup.sh 30/180/7).
@github-actions

Copy link
Copy Markdown

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

@Xore

Xore commented Sep 27, 2026

Copy link
Copy Markdown
Owner Author

Closing: #3399 is explicitly out of scope for this orchestration pass (dashboard-rewrite/docs exclusion). No agent was dispatched for it; this branch is orphaned output from a terminated respawner. Reopen if it becomes wanted.

@Xore Xore closed this Sep 27, 2026
@Xore Xore reopened this Sep 27, 2026
@Xore
Xore merged commit 7b9f771 into docs/3395-doc-reconciliation Sep 27, 2026
240 of 244 checks passed
@Xore
Xore deleted the docs/3399-r3-b branch September 27, 2026 16:33
Xore added a commit that referenced this pull request Sep 27, 2026
…and gate diagram/link rot (#3398)

* ci(docs): make mermaid parse errors and whole-tree link rot fail CI

Two of the forty mermaid diagrams in the tree did not parse at all:
AI_TRIAGE.md named a node `call` (a reserved mermaid token) and
ghidra/README.md left a colon unquoted inside an edge label. Both
rendered as an error box on GitHub and nothing noticed. Separately,
check-doc-paths-exist.py (#2458) scans only README.md and docs/**, so
the component trees under arcane/**, analysis/** and sandbox/** were
never link-checked — the agent-intrusion-corpus README's dead
`../../docs/...` hop (two levels short) lived there unnoticed.

- scripts/check-mermaid.mjs parses every block in one headless browser
  (mmdc-per-block costs ~15s each, this costs seconds). mermaid.min.js is
  served over a throwaway local http server because Chromium blocks
  file:// module imports.
- scripts/check-doc-links.py resolves every non-fenced relative link and
  src=/href= in every tracked *.md. Fenced blocks are skipped: their
  paths belong to the reader's project, not this repo.
- both wired into quality.yml as home: true rows.
- the two diagrams and the one link fixed at source.

* docs(gpu-queue,ghosts,revdeck): point the three drifted stack READMEs at what ships

Continued #3399 slice 3399-legacy-core. The previous attempt was cut off
after these three edits; re-verified each against source before keeping it.

- gpu-queue: the Go dashboard deleted at #1628 is gone, so the GPU queue
  board is the Rust backend-service's gpu_queue.rs, surfaced in the
  payload workbench (not the old /ghidra page). Routes, index and
  ES-only search corrected.
- ghosts: the stack no longer builds from a remote git context -- #1506
  vendored cmu-sei/GHOSTS at v9.0.0 after Arcane's refs/heads/-only
  resolution broke both the tag and the SHA. And ghosts-api now publishes
  5000 on virbr-ghosts's own gateway 10.20.30.1, not the docker-internal
  10.90.0.2 that only ever worked host-locally. Rewrote the #325 note as
  shipped state (#325, #2444, #2257 all landed) and recorded #2257's
  accepted residual risk: still no auth on that socket.
- revdeck: .env's API_BASE/API_KEY/MODEL_NAME are overridden by the
  compose file's own environment block, so editing them does nothing.
  #568 promoted qwen3:14b to all three slots; the 7b baseline this doc
  recommended is now superseded. Corrected both.

* docs(architecture): correct stack/sensor counts and dashboard-tier stack membership

Established the authoritative inventory from
arcane/manifests/home-production.json and the compose files:

- 39 sync entries in the manifest, not 37
- 33 of them under arcane/home/ (32 honeypot-* + unsloth), not 31/32
- 6 at their own repository-root paths, unchanged
- 20 sensor stacks (decoy stacks each on its own isolated network),
  not 22 or 21
- 34 stack directories on disk under arcane/home/: the 33 manifest ones
  plus rex86-eval, which is deliberately not a deployment piece

Fixed:
- ARCHITECTURE.md prose counts and both mermaid node labels that carried
  stale sensor counts; added the sonicwall-sma decoy the second diagram
  was missing
- README.md "38 deployment pieces - 32 under arcane/home/" -> 39/33, plus
  the two stacks its table never enumerated (sonicwall-sma, unsloth)
- README.md's honeypot-dashboard-backend row was inverted: that stack
  holds the unprivileged :8081 backend-service, while the write-capable
  :8082 backend-service-mounted stayed in honeypot-dashboard. ARCHITECTURE
  .md's dashboard-tier diagram put both inside honeypot-dashboard; moved
  :8081 into its own subgraph and noted its two extra worker loops.
- NETWORK.md "zero exceptions across all 32 stacks": the HP_BIND rule does
  hold, but unsloth spells the variable UNSLOTH_BIND. Restated as the
  claim that is actually true (no 0.0.0.0 published bind) plus the one
  naming exception.

All five doc gates green.

* docs(pipelines): correct the enrichment source count, conpot classification and index names

- "16 named sources ... 22 sources in all" -> 17 + 6 conpot personas = 23.
  sonicwall-sma-honeypot joined the #1217 canonical-promotion watch list
  in #3131 and nobody updated the count. Cited discover_sources in
  ip_enrichment/mod.rs so the next reader can re-derive it.
- Conpot was listed as wholly tunnel-blind. All six personas set
  CONPOT_PROXY_PROTOCOL=1 and portbridge's RULES carry the pp flag on
  their TCP listeners, so they are PROXY-aware on TCP; only SNMP 161,
  BACnet 47808 and IPMI 623 are tunnel-blind (PROXY v1 has no UDP form).
  Moved them across, with the same per-listener nuance the cisco-asa
  entry already used.
- Index names: the payload-bytes index is dashboard-payload-bytes-v1, not
  a "-bytes-v1" suffix on dashboard-payload-inventory-v1 (which would
  read as dashboard-payload-inventory-bytes-v1). The ml score index is
  ml-anomalies. Added workbench-runs-v1, which the init stack templates
  as part of the *-analysis-v1 family but the catalog omitted.
- README's honeypot-agent-intrusion-worker row described the retired
  Python worker as the live writer of agent-intrusion-campaigns. It has
  been profiles:["legacy"] since #1649; the Rust loop writes the index.
- ROADMAP.md still described the dashboard tier as pre-cutover behind a
  `next` profile alongside the Go dashboard. Cutover completed
  2026-08-22 and no `next` profile exists.
- settings-operations.md said to revoke the admin role "in auth-backend";
  the live IdP is Keycloak (realm `apiary`, role `apiary-admin`).

* docs(personas,runner): correct the persona inventory and the Windows trigger chain

Continued #3399 slice 3399-legacy-core, finishing Group A.

docs/personas/README.md:
- The table listed 13 of the 18 entries in personas/personas.json. Added
  the five added since the table was last touched -- nexusai-directory
  (Beelzebub), nexusai-analytics-legacy (Elasticpot), meridian-legacy-web
  (Hellpot), meridian-staff-console (Galah) and harborline-pbx
  (SentryPeer) -- and gave nexusai-core its Mailoney half, retired
  multipot's own SMTP handler for in #1422. Count verified by command:
  personas.json has 18 keys.
- "The live dashboard exposes each field as a clickable pivot" overclaimed.
  persona_id/site_id/asset_id are pivots (events.rs, events.tsx's decoy
  group); organization is not. The dashboard's organization filter is
  source.as.organization_name, the attacker's network owner -- a
  different field. Nothing anywhere reads honeypot.organization.
- persona-apply is a service of arcane/home/honeypot-init/compose.yml,
  not of the root marker compose file, so the documented
  `docker compose -f compose.yml run --rm persona-apply` would fail from
  the checkout. Pointed at the honeypot-init stack's deployed copy.

docs/sandbox/windows/runner/README.md:
- The piece table claimed the path unit starts
  honeypot-windows-sandbox-worker.service. It does not: its Unit= is
  honeypot-windows-sandbox-web-requests.service, which resolves the hash
  against the capture roots and copies the sample bytes in before
  starting the worker with --no-block. The dashboard writes a bare
  {sha256}.request with no sample data, so the direct trigger dropped
  every request. Added the missing hop.

SECURITY.md verified, no drift: DECOY_ONLY is real (cowrie's
nexusai-inference .env) and is the literal allowlisted marker in
scripts/check-public-leaks.py; the default branch is main.

* docs(operations): fix stale Go-era API paths, template count and Kibana index

- "8 of the init stack's 27 index templates" -> 33. The 8 is right
  (honeypot-events-v2, suricata-events, portbridge-events, zeek-events,
  zeek-proxy-events, extracted-files, huginn-events, traefik-access); the
  total was not, because two of the script's 25 _index_template PUTs are
  loops (7 analysis-family templates + 3 dashboard-config templates).
- "/api/campaigns" and "/api/intelligence/archive" are Go-dashboard paths.
  main.rs has no intelligence route at all; the archive index
  dashboard-intelligence-archive-v1 is reachable through the generic store
  route /api/v1/store/intelligence. Campaigns is /api/v1/campaigns.
- The source-health runtime card no longer reports Go heap/goroutines --
  health.rs says so explicitly and returns uptime_seconds/rss_bytes/
  vm_bytes. There is no /api/runtime route.
- Saved objects are not in a bare ".kibana" index: the stack runs
  Elasticsearch/Kibana 9.5.3, so they live in .kibana_<n>.
- dashboard/frontend does not exist; the frontend is
  arcane/home/honeypot-dashboard/frontend-next.
- GeoIP pipeline covers three index families, not "both" templates, and the
  Arkime "db-ip country database" note contradicted this same file's own
  #2713 paragraph (GeoLite2 Country, no coordinates).
- Kibana data views are honeypot-v2-* / suricata-* / dead-letter-honeypot*.

* docs(stack-rebuild): make the full-reset runbook cover every current stack

The stop and start loops named 17 and 16 stacks. Two problems, both
silent:

- honeypot-citrix, honeypot-cisco-asa and honeypot-rdp are not
  directories. The real ones are honeypot-citrix-honeypot,
  honeypot-cisco-asa-honeypot and honeypot-rdp-honeypot, so `cd` failed
  and those three were never stopped or started.
- 15 further stacks were missing entirely, including
  honeypot-dashboard-backend, the three Go worker stacks, the nine
  post-#3131/#1418/#1424 decoys (beelzebub, canarytokens, elasticpot,
  galah, hellpot, mailoney, sentrypeer, sonicwall-sma) and unsloth. A
  "deliberate full reset" left them running.

Both loops are now the full set: 32 stacks stopped in step 2 (the
manifest's 33 minus honeypot-keycloak, handled separately) and 30
started in step 4 (those two go in step 3). Verified every name against
a real arcane/home/ directory. Left them spelled out rather than globbed
so a future manifest entry cannot be swept in unreviewed.

Also corrected the "~19 projects" figure to name it as the historical
#258 split while stating the current 33, and fixed the three wrong stack
names in that sentence.

TESTING.md's Keycloak checklist claimed the working tree greps to zero
hits for the retired auth-runtime strings. It does not: three
`forward-auth` matches remain, all of them comments/allowlist entries
saying the thing was retired. Restated the check as "no live hits" and
listed the three so the next person running it verbatim does not file a
false failure.

DASHBOARD-CUTOVER.md's step 4 points at port-tests/{backend-api,
frontend-ssr,auth-flow}.sh; the whole port-tests/ directory is gone.
Noted where the equivalent lives now.

* docs(analysis): reconcile the analysis tree docs with what ships

Corrects drifted claims across the Group B narrative docs, each verified
by command against the tree rather than by eye.

- analysis/README.md: Phase 7 env/Compose wiring has shipped (the
  GITHUB_ANALYSIS_* vars and both spool bind-mounts are in the dashboard
  service); re-point the component table at the post-#1502
  arcane/home/ locations; replace the Docker-volume log copy with the
  host bind-mount the sensors actually use; note that EveBox the
  container still runs and only its config file and SQLite store are
  gone, with packet capture now Arkime.
- ghidra/README.md, ghidra/AI_TRIAGE.md: the triage model is
  qwen3:14b at 32768 context since the #568 requalification, not
  qwen3:8b at 16384; current routes are the /api/v1 forms.
- ghidra/DASHBOARD_INTEGRATION_PLAN.md: mark as a design record, since
  every dashboard/*.go path in it describes the Go dashboard deleted at
  #1628.
- ghidra/IMPLEMENTATION_PLAN.md: five of the six scripts/ postScripts
  are gone (findcrypt.py at #136, four more at #141) and the sixth is
  unused; capa arrives through the statictools sidecar, not as one of
  the nine awesome-ghidra candidates, all of which are decided out.
- ghidra/ghidrassist/README.md: the biniamfd/ghidra-headless-rest image
  is gone, replaced at #245 by a build from analysis/ghidra/service.
- ghidra/benchmarks/README.md: the llm-worker network isolation that
  justifies the three-stage probe is a property of the Safe #66 base;
  the #1751 authorized entrypoint adds honeypot-llm-data and
  honeypot-llm, so the gates are what the probe stays out of.
- ghidra/benchmarks/corpus/README.md: 17 cases, 850 builds (14 covered
  by semantic harnesses, 3 excluded), matching manifest.json.
- yara/README.md: upstream/DROPPED is currently empty, and the auto rule
  named is AutoGen_190460923_exe.

Verified with no drift: analysis/RECOVERY.md, ghidra/models/README.md.

* docs(settings,geoip): restate settings behaviour for the Rust tier and the real mmdb consumer set

settings-operations.md still described the deleted Go dashboard's behaviour:
its in-memory config cache with a 3s poll, eight honeypot_settings_* Prometheus
gauges, POST /api/settings/config/rollback, `docker compose stop dashboard`,
and the realm role apiary-admin. Verified against the Rust backend:

- config.rs has no in-process cache at all - load_config() reads the
  dashboard-config-v1 document from Elasticsearch on every request, so the
  cross-replica visibility window is "next request", not "up to 3s".
- no honeypot_settings_* metric is emitted anywhere in the repo, and no
  prometheus file is tracked either, so there is no scrape target. /metrics on
  :8081 exposes only apiary_backend_requests_total and
  apiary_backend_request_duration_seconds from obs.rs.
- the rollback route is POST /api/v1/config/rollback; the nine-client realm
  matches the doc's "nine OIDC clients" (it said eight).
- the service is dashboard-next / hp-dashboard-next; the Go `dashboard`
  service no longer exists in the compose file.
- the admin role is a *client* role `admin` on the apiary-dashboard client,
  read from resource_access.apiary-dashboard.roles - not a realm role.
- writes carry If-Match: <revision> and answer 409 + X-Current-Revision.
- failure posture corrected: the Rust tier returns 502 on an Elasticsearch
  error rather than degrading to compiled defaults read-only. The staged
  rollout sequence is now labelled a design record, since the two metrics it
  gated on never shipped.

GEOIP-THREAT-INTEL.md said the .mmdb files are mounted into two containers.
They are mounted into three - hp-elasticsearch, hp-arkime-capture and
hp-arkime-viewer, the latter two both mounting /opt/arkime/geo - and a
fourth, hp-geoipupdate in honeypot-init, is the writer. Restarting only two
of the three consumers leaves a stale database in the third.

* docs(network-isolation): complete the honeynet membership list and name the firewall sync check

The trusted-plane paragraph listed seven members of honeynet and stopped, which
read as a complete enumeration. It is 24 services across 11 compose files.
Added the ones a reader auditing trust boundaries would most want to see: the
four Go worker stacks (all profiles: ["legacy"]), honeypot-dashboard-backend's
unprivileged backend-service, honeypot-dashboard's worker containers,
honeypot-init's three one-shot containers, honeypot-utilities' disk-space-
monitor (log-maintenance was already listed), honeypot-elk's
extracted-file-importer, and the two root-path stacks auth-events-worker and
ml-worker. Also recorded the three that are deliberately *not* on it -
backend-worker-enrichment and services-adapter are network_mode: none, and
oidc-sessions has its own single-member network - since "trusted plane" reads
as a completeness claim and its complement is the interesting half.

Line 42 said honeypot-firewall.sh was the only firewall script in the
repository; a grep finds two. Added that vps/check-firewall-portbridge-sync.sh
(#152) configures nothing and only diffs portbridge's RULES against the ufw
opens - the guard against exactly the silent drift the doc describes elsewhere
- and that it is in no workflow, so it must be run by hand.

Verified unchanged, so left alone: the NET_ADMIN/NET_RAW enumeration (exactly
8 services, 3 of them in docker-compose.sandbox.yml as stated - mitmproxy adds
only NET_BIND_SERVICE and technitium/compose.yml:68 says so explicitly); tanner
keeping tanner_local for all seven services; tanner_docker being the only
privileged container; the sandbox macvlan internal:true 10.10.10.0/24 and the
vps/suricata + sandbox README links, which resolve correctly as docs-relative
paths.

* docs(keycloak): stop calling the Keycloak image pinned

"Keycloak and PostgreSQL use pinned upstream images" was true of PostgreSQL and
false of Keycloak, and the compose file says so in a comment block right above
the line it contradicts. Postgres is a digest pin (postgres:18.6-bookworm@sha256:
1c59e2c3...). Keycloak is deliberately quay.io/keycloak/keycloak:latest with
pull_policy: always - it was pinned to 26.7.1@sha256:f1f1f01e when CVE-2026-18963
(unauthenticated account takeover through reset-credentials, CVSS 9.1) landed,
and the file's own reasoning is that a digest pin is what keeps a known-vulnerable
build running until someone edits a file. pull_policy is load-bearing, not
decoration: without it a redeploy reuses whatever :latest first resolved to.

Restated the topology bullet as the deliberate asymmetry it is, and fixed the
upgrade procedure's step 2, which told the operator to "update image tag and
digest together" - an instruction that, followed literally, would re-pin
Keycloak and reintroduce the CVE exposure.

Everything else in the file verified against the repo, so left alone: nine
confidential clients and zero public (the 6+3 gateway split, §3); no client
secret pinned in the realm template, so Keycloak generates a fresh one per
--import-realm; only *.example.invalid hostnames; hp-keycloak / hp-keycloak-postgres
container names; keycloak-data internal + keycloak-egress with PostgreSQL on
keycloak-data only and no published ports; ${HP_BIND}:${KEYCLOAK_PORT:-18080}
the sole published port; socat-hp-keycloak; the keycloak / keycloak-embedded-frames
/ keycloak-admin-console / keycloak-account-console routers; the /administrators
group; registrationAllowed false; auth-events-poller the only service-account
client; both scripts cited in the acceptance checks; and the realm-management
realm-admin grant, which is absent from the template by design because
realm-management is a built-in client.

* docs(stack-rebuild): drop the portbridge-log-rotate that no longer exists

The VPS half of the full-reset runbook named hp-portbridge-log-rotate in the
stop list and portbridge-log-rotate in the start list. Neither exists: #1779
(9fcfd108) removed the service from vps/docker-compose.yml, folding its only
job - pruning the renamed portbridge.json.* files past the retention window -
into portbridge-log-maintenance, whose script header says exactly that. The
only remaining *rotate service in the VPS compose is traefik-log-rotate.

In a runbook, `docker stop <nonexistent>` and `docker compose up -d
<nonexistent>` fail quietly, so the step looked like it had worked. Also
recorded why the list still omits portbridge-manual-blackhole-refresh (the
manual counterpart to the scheduled refresh, deliberately not auto-started) and
that suricata-update is handled in the note immediately below.

* docs(ml-worker,benchmarks): correct retention figures, manager name and check count

Three files, six claims, each checked against the repo rather than read off.

benchmarks/claim-pools/README.md counted 240 executable semantic checks.
analysis/ghidra/benchmarks/corpus/semantic_checks.json records checked: 280,
failed: 0, and a cases_covered list of 14 entries, so both the "240 executable
checks" figure and the "(240 checked, 0 failed)" parenthetical were stale by
40. Added the case count, since the rung's whole argument is about what the
harness does and does not cover.

ml-worker-plan.md stated a 180-day retention for the ml-worker-metrics index
and named the knob ML_METRICS_RETENTION. worker.py:74 is
ML_METRICS_RETENTION_DAYS with a default of 90, and worker.py:1003 applies it;
the 180 figure and the knob name were both wrong, and the sibling anomalies
index already quoted its own knob correctly two paragraphs above. Also moved
the ensure_ilm_policy() citation from worker.py:939 to :1002, where the call
actually is.

Both docs, and gpu-ml-worker-acceleration.md, described ml-worker as "its own
Dockge stack". Dockge is gone; arcane/manifests/home-production.json still
lists ml-worker/docker-compose.yml as a deployed entry, so the status line now
says Arcane-managed and the two narrative mentions say standalone, which is
what the surrounding text is actually contrasting against (a service folded
into the root docker-compose.yml).

* chore: stop tracking the agent run log and its brief

64d96b82 was staged with `git add -A` and swept in two local tooling files
that origin/main does not track: .agent-run.log (an append-only transcript of
a concurrent run) and BRIEF.md (that run's task assignment). Neither is
documentation and neither belongs in the tree.

Removed from the index only; both files remain on disk untouched.

* docs(sandbox): reconcile the Windows/CAPE/GHOSTS plans with what ships

The five sandbox design docs still described the retired Go dashboard, a
FLARE-VM provisioning pass that no longer exists, and a golden image
that was never built. All three claims are now wrong in the opposite
direction from what the docs say, so this is a status correction rather
than a restyle.

cape/IMPLEMENTATION_PLAN.md
  - File Structure tree now points at the Rust backend service
    (workbench_domain.rs, detail.rs, worker.rs) instead of the deleted
    dashboard/cape.go and dashboard/workbench_domain.go.
  - Records why #319 is only partial: workbench_orchestrator.rs
    marker_dir() has arms for ghidra, windows-sandbox, windows-ghosts,
    linux-sandbox and revdeck, but none for cape, so #317 still needs
    manual routing.

windows/packer-golden-image-guide.md
  - The golden image is built. Commit 6402f237 (#1128) records a real
    rebuild of win11-analysis.qcow2 on the homeserver, and b8741069
    (#957) confirms it live via virsh screenshot.
  - FLARE-VM was removed on 2026-08-02; 02-flarevm-start and
    03-flarevm-wait are gone, and the provisioner timeout is now the
    "60m" in 04-tools.ps1, not "5h".
  - Boot is PXE, not CD-ROM (#288/#406); win11-analysis.pkr.hcl carries
    no boot_wait or boot_command.
  - Secure Boot is deliberately off: OVMF_CODE_4M.fd paired with
    OVMF_VARS_4M.fd (#288/#419).
  - Packer defaults restated from the HCL: 16384 MB, 12 cpus, 90000 MB
    disk, 45m winrm_timeout, ide disk interface.
  - The "no Sysmon config in the tree" claim was already stale in the
    other direction: sandbox/windows/config/sysmon_config.xml does not
    exist because the config is fetched at build time and pinned to
    commit 1836897f, so a rebuild does not pick up a newer one.

windows/IMPLEMENTATION_PLAN.md
  - Status block In Progress -> Shipped, Phase 3 checklists corrected:
    only "Process Creation" is audited (a single auditpol call), so the
    4663/4657 file and registry access, "Uptime: > 3 days" and
    "> 50 processes" checks are marked not implemented rather than left
    carrying a tick.
  - libvirt-python removed from the stack, and the snapshot
    subcommand removed from the kvm_manage.sh invocation list
    (create/revert/start/stop/status are what exist).

windows-guest-risk-config-model.md
  - The GHOSTS NPC client is Ghosts.Client.Universal, built by
    sandbox/ghosts/Dockerfile.client-win and renamed to EndpointAgent,
    not a real Ghosts.Api client.
  - #904 is landed (cbb36a11): the CAPE autounattend identity is
    VPM-ENG0089 / Daniel Kowalski / Vantage Precision Manufacturing, not
    ACP-FIN0142 / Robert Tanaka / Ashford Capital Partners.
  - #901 now notes that sandbox/ghosts/loldriver-gate-test.ps1 exists as
    acceptance evidence, with no run output recorded.

ghosts/IMPLEMENTATION_PLAN.md
  - workbench_orchestrator.go -> workbench_orchestrator.rs, and the
    sandbox status route is GET /api/v1/sandbox/{job}.

* Revert "chore: stop tracking the agent run log and its brief"

This reverts commit 91d82e5f2261937f590494461cc99fe0d85b3df0.

* docs(dionaea,audits,canarytokens,benchmarks): four verified corrections

dionaea-bistreams-retention.md -- this reverses an error introduced by
a9e3086e, which claimed dedupe() never hard-links inside bistreams because
PAYLOAD_ROOTS "is the dedupe root list (/payloads/cowrie:...scripts/
script-payloads)". That three-root list is the *fallback default* in
dedupe-payloads.py:169, not the deployed value: honeypot-payload-analysis/
compose.yml:49 overrides it with a fourth root,
/payloads/dionaea/bistreams, listed last. So the shipped container dedupes and
hard-links bistreams as well as pruning it, and BISTREAMS_ROOT (compose.yml:56)
is the separate variable the pruning pass reads, not an exclusion from dedupe.
main() prunes before dedupe by design ("bounds what dedupe has to hash this
pass"), which the compose comment at :51-55 attributes to bistreams being 82%
duplicate by content in a live sample (#112).

container-writable-layer-audit-2026-09-03.md -- said rex86-eval "is not in
this repo's tracked compose files at all (it's a raw docker run/Dockge stack)".
arcane/home/rex86-eval/{compose.yml,setup.sh,.env.example} are all tracked
(restored by #847). It is absent from home-production.json, which is the
distinction that actually matters: benchmark tooling, not a deployment piece,
and so the one directory under arcane/home/ with no manifest entry. Also
records the read-only ${APIARY_REPO:-../../../../}:/repo:ro mount
(compose.yml:45), which is why 227 GB of /root/.cache still lands in the
rootfs.

canarytoken-live-fire-checklist.md -- the dashboard step queried
since=3650d; frontend-next/src/routes/canarytokens.tsx:49 issues
since=365d, with the reason in a comment at :46 (tokens fire rarely; the
events endpoint's own default is 10d). The trailing 0 was a transcription
slip, not a wider window. Also names the actual public entry point:
honeypot-canarytokens/compose.yml publishes 19427 from
canarytokens-http-router (:300-:313), which holds ADAPTER_URL and fronts the
adapter's 8083; the switchboard itself publishes no host port (:256-257).

benchmarks/claim-pools/README.md -- the ground-truth-superseded section said
this "happened once". It has happened twice: tier-a-v1-review-queue.json has 59
superseded rows, 33 for integer_overflow_alloc and 26 for
process_and_injection. #2694 retired the latter's *resistance-test* reading
rather than its ground truth (build_corpus.py:145-180: injection-gate twins),
and its 26 verdicts are placeholders like the other 33, so nothing needs
re-adjudicating.

* chore: untrack the dispatch brief and agent run log

They were dispatch scaffolding, committed by accident when the agent
staged its whole worktree. Never repo content.

* fix(docs): fail check-mermaid with a clear message when puppeteer is absent

On a fresh runner the npx cache is empty, so the require_() at the puppeteer
lookup threw a raw CJS loader stack trace instead of an actionable error. The
mermaid bundle lookup one block down already guards this exact case; give
puppeteer the same guard so the failure names the command that fixes it.

* ci: warm the npx cache so the mermaid gate can actually run

check-mermaid.mjs resolves mermaid and puppeteer from the npx cache, and the
step's own comment claimed that made it self-bootstrapping. It did not: a
fresh runner has an empty cache, so the gate exited 2 with "puppeteer not
found" and could never pass. Run the unpinned npx warm-up in the same step,
which is what the comment always assumed.

Verified from a deliberately empty HOME: 40 mermaid blocks in 154 files
parse cleanly. zizmor is unchanged (no new finding).

* merge: docs reconciliation round 2, map slice (#3399) (#3413)

* docs(readme): correct sensor directory names, retired-worker row, profile set, stack and diagram counts

- row 'more sensors' listed honeypot-citrix / honeypot-cisco-asa /
  honeypot-rdp, but the directories are honeypot-citrix-honeypot,
  honeypot-cisco-asa-honeypot and honeypot-rdp-honeypot, so the row's own
  'arcane/home/honeypot-<name>/compose.yml' pattern did not resolve
- the three worker stacks in the next row were all retired by #1649 and run
  as Rust WORKER_LOOPS now; the row still described them as live writers
- 'the only profile is geoip-update' omitted the threat-intel maintenance
  job, the four ["legacy"] rollback definitions and ghosts' ["test"] client
- '20 split home stacks' -> 26, the roster deploy-profiles/full.txt names
- ARCHITECTURE.md '6 diagrams' -> 4, its actual mermaid block count

* docs(map): drop the nonexistent archive/ tree, add design-lab/ to the exempt list

docs/archive/ has no directory and no tracked files, but the map listed it
twice (as a subdirectory, and as a dated record tree). The reachability
script does exempt docs/design-lab/ -- kept as a near-duplicate of
branding/design-lab/ -- and the map's description of that gate omitted it.
Reachability result is unchanged: 85 reachable, 35 exempt.

* docs(sensors): fix retired dashboard container, ES budget, TANNER roster, endlessh PROXY

- the investigation-UI table still named the deleted Go 'dashboard' container
  on :8090. The live view is dashboard-next, container :8080, published on
  home 19090/19092 and bridged from VPS 8090/8092
- Elasticsearch is limits.memory 12G with ES_JAVA_OPTS -Xms6g -Xmx6g, not
  8 GiB / 4 GiB; dashboard-next gets 2 CPUs, not 1
- the TANNER container list omitted tanner_docker and listed snare_clone as
  one of this stack's services; it is a honeypot-init one-shot (hp-snare-clone)
  writing the snare-pages volume. The seven tanner_local services are now listed
- endlessh parses PROXY_PROTOCOL=1 behind its :pp 2022 rule but was missing from
  the list of sensors that recover the real attacker IP that way

* docs(deception): add the missing SonicWall SMA decoy row

The 'Service-specific decoys' section states its purpose as backfilling
rows so every running sensor traces back to a decision, but
honeypot-sonicwall-sma (#3033, CVE-2026-83548 / CVE-2026-83549) had no row
anywhere in the file. Every other one of the 20 sensor stacks does. All
other claims verified against compose and the source: wordpot's directory is
gone, the 2026-08-27 retirement date matches, conpot's six personas, and
#233/#242 as cited in community-threat-intel-sharing.md.

* docs(sensors): correct the Suricata capture interface, ILM retention and payload-dedupe logging claims

* docs(roadmap,design-lab): correct CAPEv2 build status, doc-sweep scope and the design-lab harness pointer

* docs(security-fixes): date the CodeQL record and point its Go-era examples at the Rust tier

* docs(sensors): state the ILM windows at the shipped HONEYPOT_RETENTION_DAYS=21, not the code fallback

* docs(reporting,bistreams): redraw the reporter diagrams against the Go service, date the retention decision

* docs(threat-intel): correct #153 to closed-and-implemented in the reporter comparison

* docs(threat-model): resolve two follow-ups the body still reported as open

Both were already marked Done in the doc's own Follow-up scope section; the
section bodies and the applicability matrix were never updated to match.

- item 9: llm-analysis severity is wired into the Rust alert sink as
  llm_flagged_alerts (worker.rs:1277), not browse-only via /llm-analysis
- item 5: hp-autoheal no longer bind-mounts /var/run/docker.sock. #592 moved it
  onto hp-docker-socket-proxy, which holds the socket :ro scoped to
  CONTAINERS/IMAGES/POST on a private network. The residual is that
  CONTAINERS=1 stays daemon-wide, which is now what the text says
- the docker.sock grep aside claimed 'absent' from the dashboard compose, but
  that file does contain one mount (services-adapter's) -- made explicit so
  someone re-running the grep is not surprised

ROADMAP.md: verified, no drift. No 'next' profile survives anywhere, CAPE is
authored but absent from the manifest while GHOSTS is present, sandbox/cape
has the Packer file and spool worker it claims, reporter/ sits under
honeypot-utilities, and llm-worker's selftest exists.

* docs(ip-block,agent-intrusion): record the ip-block export path move and the unported sidecar default

* docs(llm-worker): name the deployed captured-data entry point from the manifest

* docs(ip-reporting): drop the Prometheus endpoint claim, make the test-coverage claim true

- 'The reporter container' still promised a /metrics Prometheus endpoint for
  Grafana. metrics.go explicitly rejects that shape and writeMetricsLoop
  overwrites dataDir/metrics.json instead; the status banner 100 lines down
  already said so, so the doc contradicted itself
- 'Every .go file has a matching _test.go' does not hold by filename:
  categorize.go is covered from event_test.go, and main.go/report.go from
  dryrun_test.go. Every function is still exercised, so the substance stands

Every other claim verified against the tree: /data/reported.db, the reporter
volumes, all thirteen .go files in the diagram, no .py, and the
metrics.json/AUDIT_BLOCKLISTDE/GREYNOISE_ENABLED env contracts.

* docs(ml-worker): correct the Tier 1 contract-check count to seven

* docs(ml-worker,writable-layer): seven contract checks, dated banner, drifted line ref

- evaluate_detectors.py runs seven check_* functions (appended at lines
  404-411), not six
- the 2026-09-03 audit had no dated banner, so its live-host measurements
  (245.8 GB, 13.18 GB reclaimable, a '45+ hours up' buildkit container) read as
  current. Added one pointing at the same do-not-mirror principle
  security-fixes.md already states, and at the open issues #2915/#2904
- its install-homeserver.sh:431-500 pointer no longer lands on the builder-GC
  discussion; the reasoning is at 392-397 and the JSON block at 463-465

Everything else verified: all four quality.yml scripts now use 'docker rm -fv',
rex86-eval is tracked but absent from the manifest and is the only such
directory under arcane/home/, and the daemon.json gc values match exactly.

* docs(design-lab): correct how the read-only seam is enforced for the mounted backend

---------

Co-authored-by: xore <xore@localhost>

* merge: docs reconciliation round 2, infra slice (#3399) (#3414)

* docs(arcane): reconcile ARCANE-GIT-SYNC.md with the 39-entry manifest

- 37/31/32 census -> 39 (33 under arcane/home/, 6 self-contained at root
  paths); document #2911 technitium and #3092 unsloth
- manifest import section: filter reaches 35 of 39, technitium replaced
  pihole's dedicated step, unsloth has no installer path at all
- manifest declares branch: production since #1943, not main; no
  refs/heads/production exists on origin (checked 2026-09-27)
- 34-of-37 build / three pullers -> 34-of-39 build / five pullers
- restart: no census: 20 hits, 13 declarations, 8 files; honeypot-elk's
  arkime-pcap-init (#3128) is a second Arcane-started one-shot
- relative-bind-mount set 9 -> 10 (ghidra); env_file claim: ghidra has one
- re-count dashboard 268->279 files, ghosts 989->962, ghosts-src 963->935
- :?required stacks are canarytokens/ghosts/technitium/unsloth

* docs(cgnat): reconcile deployment paths and the gateway chain

- 31 in-tree stacks -> 33 (add honeypot-dashboard-backend #1622 and
  honeypot-sonicwall-sma #3131, drop the retired ip-enrichment worker);
  note unsloth as the 33rd directory
- pihole -> technitium in the self-contained six and its dedicated step
- the investigation chain is a loadBalancer to each oauth2-proxy gateway,
  not a forwardAuth middleware (0 forwardAuth in dynamic.yml), and only
  five of the six gateways have a socat-hp-* bridge
- Arcane pin is v2.11.1, not v2.8.0; the confirmed limitations have not
  been re-confirmed against it
- mark the SFTP-upload/rebuild-from-terminal paragraph as pre-#1502

* docs(arcane): flag that v2.11.1 pin is unre-verified

* docs(security): document the leak gate SECURITY.md relied on

SECURITY.md names 'leak a real secret' as a reportable class but never says
the repository enforces it in CI, and never warns that the gate scans
untracked files. Both facts were load-bearing on this run: the checker
carries the home-server address as a named forbidden literal, and
RESUME.md quotes that address verbatim, so the gate fails on a clean
tracked tree purely because of dispatch artifacts sitting in the checkout.

Adds the gate's actual pattern set, its three fail-closed exemption sets
with their real sizes (3 ALLOWED_DOTENV paths, 2 ALLOWED_LITERAL_FIXTURE_FILES
paths), the four deployment-specific literals it assembles from fragments,
and the untracked-file scan. No existing prose changed.

* docs(homeserver): re-measure the live disk layout; it is not the documented one

Every row of HOMESERVER-DISK-LAYOUT.md's 'as installed' table was stale.
Re-measured read-only over ssh 2026-09-27 (lsblk/lvs/findmnt/df/du):

  - OS is Rocky Linux 10.2, not Ubuntu/subiquity. The doc's curtin
    autoinstall notes described an install that has been replaced.
  - boot disk is a different, 4x-larger NVMe (PC401 SK hynix 1TB,
    953.9G) and is now LVM-backed: rl-root 70G, rl-swap 32G, rl-home
    849.3G. The doc recorded 'no LVM' as a design decision.
  - /var is on an sdb1 partition of an 8.7T LUN, not a whole disk.
    The doc left sdb's size unstated in the table and then called it
    'the 1.7T sdb disk' in prose -- an internal contradiction, now moot.
  - sda is a USB-attached Samsung PSSD T7 at /mnt/usb-recovery, so the
    /mnt-2 bulk-storage role is gone; no sr0 optical device enumerates.
  - swap is a 32G LVM LV (14.6G in use at measurement), not an 8G
    /swap.img swapfile; /swap.img does not exist.
  - /var/lib/docker is 2.9T and /var/dockge 350G, against 103G/229G and
    332G combined in the doc. 45 stack dirs, not 23.

The autoinstall config and its manual-partitioning walkthrough are kept
as the record of the former Ubuntu layout and explicitly marked as no
longer a rebuild target for this host.

BACKUP-ESSENTIALS.md: three corrections, all verified against
scripts/backup-essentials.sh and the live host.
  - 40 stack .env files as of 2026-09-27, not 41, and phrased per-stack
    so the number is read as a measurement rather than a constant.
  - The volumes table listed short names (arcane-data, evebox-config,
    ...). Four of the five real volumes carry an Arcane project prefix,
    and backup-essentials.sh writes each archive as .tar.gz, so
    following the old table would 'restore' into volumes no stack is
    mounted against. Table and restore step 5 now carry the real names.
  - Samsung PSSD T7 is the model lsblk/udevadm report; and the dead
    Keycloak restic config is now doubly dead, since /mnt-2 itself has
    been decommissioned.

HOST-TUNING.md needed no change: all five tunings, half-of-RAM capped
8G zram, priority 100, --replace-swap, the per-class schedulers, the
cups/bluetooth/ModemManager service list and the '20 GB card' claim all
match scripts/tune-rocky10.sh, and the host's RTX 4000 Ada (20475 MiB)
confirms the card size.

* docs(ci-cd): reconcile workflow topology, counts, and VPS sync excludes

* docs(analysis): reconcile workbench and LLM-worker docs with the Rust backend

Both docs still described the Go dashboard that #1628 deleted. The
corrections were verified against backend-service/src/main.rs and the
workbench modules, and the manifest.

payload-analysis-workbench.md:
  - All 8 HTTP contract routes were wrong. The live set is
    /api/v1/workbench/{analyzers,runs,runs/{id},runs/{id}/children/
    {analyzer_id}/{action},recipes} (main.rs:456-468), so: the prefix
    moves from /api/payload-workbench to /api/v1/workbench, the hash is
    a query parameter rather than a path segment, and cancel/retry
    collapse into one route dispatched on {action} rather than being
    two paths. The cancelled child-action row is gone with the merge.
  - Registry pointer dashboard/workbench_domain.go ->
    backend-service/src/workbench_domain.rs. There are zero .go files
    under dashboard/ in this repo.
  - /payload-workbench -> /payload-workbench/results, whose
    workbench-builder section is the orchestration surface.
  - The 'closed Go schema (unknown fields are rejected)' and 64 KiB body
    cap are not present in the Rust tier (no deny_unknown_fields on
    workbench_api.rs, no body-limit constant anywhere), so the sentence
    now says not to rely on them rather than asserting a guarantee the
    server does not make.
  - Dropped the rollback note about a local /state/analysis-workbench
    copy; no such store exists, workbench_es.rs is the only one.
  - MODEL_STATUS_SOCKET appears nowhere in the backend, compose or
    .env.example. Marked undetermined rather than deleting the claim --
    the adapter is genuinely installed, and deleting would lose the
    intended contract.

gpu-llm-analysis-worker.md:
  - /api/llm/analysis does not exist; the page reads the generic store
    route /api/v1/store/llm-analysis (main.rs:438, stores.rs).
    dashboard/llm_analysis.go -> frontend-next/src/routes/llm-analysis.tsx.
  - Semantic search was marked 'Deferred' but shipped as
    /api/v1/llm-search (main.rs:341). Recorded as delivered, with a note
    that the section is historical scope rather than current state.
  - 'Managed by Dockge under /opt/stacks' is obsolete: Arcane gitops
    manages /var/dockge/stacks, and /opt/stacks is now only a
    compatibility symlink to it (created 2026-09-04).
  - The file list omitted the entrypoint the manifest actually deploys.
    arcane/manifests/home-production.json sets the llm-worker sync's
    dockerComposePath to llm-worker/docker-compose.captured-data-deploy.yml,
    so a bare docker-compose.yml bring-up does not reproduce the live
    worker (#2234). Added with that warning.
  - Host RAM/CPU 91 GiB / 16 CPUs -> 92 GiB / 48, measured on the host.

* docs(deploy-profiles): correct backbone list, structural deps, and full.txt gap

* docs(homeserver): sharpen the manifest/directory-count sentence

The previous phrasing conflated two different counts: 34 is how many
directories exist on disk under arcane/home/, while the manifest holds 39
entries of which 33 name one of those directories and 6 are root-level
stacks. Spelled out so the sentence cannot be read as '34 of the 45 stack
directories are manifest-managed'.

* docs(infra,benchmarks,gpu): fix manifest count, cohort table, and GPU device pinning

* docs(infra): fix shipped-vs-planned tense, KVM bridges, and GPU roadmap drift

ROCKY-10-MIGRATION.md: 'is moving from Ubuntu' -> 'has moved' (the rebuild
hit live 2026-09-03; install-homeserver.sh:324 says so). Also resolved the
one claim left undetermined by the previous pass: both cards really are on
the bus (P2200 17:00.0/10de:1c31, Ada 65:00.0), but only the Ada has a
driver bound, so nvidia-smi lists one GPU. Documented that explicitly,
because 'one GPU in nvidia-smi' reads like missing hardware and is the
reason the nvidia-open-vs-cuda-drivers choice still matters.

KEYCLOAK-CUTOVER.md: the hard cutover this contract specifies has shipped.
vps/forward-auth/ is gone and xore_sso/_auth/verify/strip-auth-identity
exist nowhere but this doc and docs/TESTING.md. Status line now says
SHIPPED; the contract body is untouched. Verified all 8 OIDC client IDs
still exist in arcane/home/honeypot-keycloak/keycloak/realm/apiary-realm.json
with the documented flow flags.

kvm-network-traffic-analysis.md: 'two bridges' understated it -- there are
four (virbr-hpsbx, virbr-cape, virbr-ghosts, plus 198.18.0.0/24 on
virbr-hpsbx in controlled mode). Results path was wrong: run_sample.py
overrides the compose default per run via WINDOWS_SANDBOX_RESULTS_DIR, and
sandbox/results/ does not exist. 'The Linux sandbox has no Zeek equivalent'
is false -- run-linux-sample.sh runs 'zeek -r' offline and export-result.py
reads the logs. The Phase 0 FORWARD DROP pair is a documented manual step,
not scripted, and needs re-verification against Rocky 10.2 firewalld.

kvm-snapshot-vs-golden-image.md: §5's golden-win10 and
/var/lib/libvirt/golden/ are illustrative; kvm_manage.sh uses
SANDBOX_ROOT=/var/dockge/sandbox with golden-images/win11-analysis.qcow2
and spawns via qemu-img create -b + virsh define, not virt-clone. Added a
pointer instead of rewriting §5.

ml-gpu-coordinated-roadmap.md: kept as a dated plan, added an
intent-vs-shipped banner covering the four things that moved since
2026-08-01, including that §1 decision 5 was superseded outright -- the ML
worker has no embedding code at all, and llm-worker ships 768-dim
embeddings behind LLM_EMBEDDING_ENABLED (default off).

* docs(gpu-ml): refresh torch/pyod pins, second GPU, and VRAM headroom

* docs(kvm): correct the results-path row; drop an invented compose default

The previous commit cited a 'compose default of sandbox/results/current'
for WINDOWS_SANDBOX_RESULTS_DIR. No such default exists --
docker-compose.sandbox.yml never mentions the variable. The verified facts
are that run_sample.py:124 reads the variable with a
'reports/windows-sandbox' fallback, .env.example documents
'/windows-sandbox-results', and the dashboard's es-results-importer
populates it per stack. Reworded to those, which also keeps the doc-path
lint green.

* docs(gpu-llm): correct embedding dims to 768 and note the second GPU

* docs(recovery,sandbox,ml): fix restore order, stale Go paths, and the retrain schedule

analysis/RECOVERY.md: the two backups were described as 'the same scope',
which is wrong and hides a filename trap. Scope differs -- the essentials
archive additionally carries the VPS config, WireGuard, Technitium, the
installer answers and the repo runbooks. Concretely, SHA256SUMS and
stack-config-state.tar.gz are the ON-HOST copy's only; a
backup-essentials.sh archive has neither, and this doc's own first bullet
points readers at the essentials archive. Also replaced
'docker compose -f compose.yml create' with per-volume 'docker volume
create <full name>' -- four of the five volumes carry an Arcane prefix and
live in four different stacks, so one compose create cannot make them --
and restored the real start order (honeypot-elk healthy before
honeypot-init) from STACK-REBUILD.md, which this step had lost.

sandbox/README.md: the workbench registry pointer
dashboard/workbench_domain.go -> backend-service/src/workbench_domain.rs
(the Go tree has no .go files since #1628), in prose and in the mermaid
node. Two host numbers were stale: 16 -> 48 logical CPUs, and the IOMMU
figure. The doc's 84 groups was wrong on the number AND missed the story:
the Rocky 10 rebuild came up with zero groups, which is why
install-homeserver.sh's step_vfio_gpu_passthrough now adds intel_iommu=on
to the kernel command line. Re-measured 2026-09-27: 94 groups, with
intel_iommu=on iommu=pt confirmed live in /proc/cmdline, and /dev/kvm
present. Line 237's 4 vCPU / 8 GiB guidance matches neither the real
Windows domain (8 vCPU / 16 GiB) nor sandbox.env.example's
SANDBOX_VM_MEMORY_MB=3072, so it is now flagged as operator guidance
rather than left reading like a measured limit.

ml-worker-plan.md: RETRAIN_INTERVAL is gone -- #172 replaced it with four
fixed UTC slots (worker.py:59, default 03:00,09:00,15:00,21:00). Fixed in
all four places it was asserted, including the early-retrain trigger, which
is now slot-nearest rather than interval-remaining. Corrected the status
callout: the worker runs from its own Arcane stack, not the repo-root
compose file, and docker-compose.ml-worker.gpu.yml is inert and undeployed
so the worker is CPU-only. Fixed the v0.1 paragraph, which contradicted
its own §10 and the line below it by claiming there were no tests.
Added a dated reading note separating the shipped sections from the
original-draft ones, so the remaining dashboard/*.go references read as
historical instead of as current.

Benchmark records #65-#69: dated intent-vs-shipped banners only, no
history rewritten and no recorded number touched. Added alongside the
existing /mnt-1 banners rather than replacing them. #66 is superseded by
#2694; #67's hardware table is pre-reinstall and its #2985 'missing
scripts' are now in git except gptoss_rerun.sh; #68's unsloth is the
Arcane stack, not an installer; #69's TOOLCHAIN.md 'no smoke test yet'
still stands.

* docs(ml-worker): correct Tier 1 check count to seven

* docs(ml-worker): add zeek-v1-conn source index and refresh pin record

* docs(benchmarks): land the #66 §9/§2 reconciliation banner

Uncommitted work in the tree when this run ended; committing so the
reconciliation is not lost. It sits after the 2026-08-30 supersession
banner and rewrites no measured number: §9's 'it is not merged' is stale
(the branch landed as df650a8f, the same minute as the 16:20Z report
count), §2's Tier B tally is the matrix's own rather than the shipped
64-row fixture's (15 FAIL / 15 PASS, 13 payload hits to 2, not 14/15 and
11/3), test_record_baseline.py is 54 not 50, §8's governance-gate term
list omits 'conclude benign', and §8's '59 hand-labelled answers' is the
wrong set twice over.

* docs(benchmarks): name the live corpus path in the #1947 resume-plan banner

The banner lists the #2985 scripts as now in git, but there is no
root corpus/; the tracked copies are under
analysis/ghidra/benchmarks/corpus/. The doc body already gives that
directory at the P1 section, so the banner now agrees with it.
No measured number touched.

* docs: revert five files that are outside this slice's assignment

SECURITY.md, docs/ROCKY-10-MIGRATION.md, docs/analysis/RECOVERY.md,
docs/kvm-network-traffic-analysis.md and
docs/kvm-snapshot-vs-golden-image.md are not on the round-2 r2-infra
assignment list, but earlier commits in this branch had corrected
them. Reverted to the branch point so this branch touches only its
19 assigned files and cannot collide with the slice that owns those
docs. The corrections are preserved as a handoff patch outside the
repo; they are reported in the round-2 handback.

* docs(models): reconcile the evaluation record against the stored runs and matrices

Checked every claim this file can be checked against, and recorded the result
at the top. No score in the file was changed -- the one edit is a clause that
makes an existing claim true.

Confirmed: the round-7 cold baseline reconciles cell for cell against
round7-cold-baseline.json (91 models, 182 cells, 367 records, 179 reproduced,
the same 3 escalated cells and third-run values, 11 zero-scored tags, and
every anchor including the 12.1 / 7.3 point gaps). The cold-cohort table
reconciles against 1947-cohort-cold-protocol.json down to min/run B and
Ornith's injection FAIL. The twelve-model survey reconciles against
1805c-ghidra-slot-matrix.json. The #568 approved qwen3:14b@bdbd181c33f2 and
context_tokens 32768 are confirmed against approved-models.json.

Recorded as undetermined rather than asserted either way: the archive holds no
ghidra-slot record at all (all 62 runs, 1498 records, are revdeck or sessions)
and no transcript carries VRAM or a context-probe result, so the Ghidra, 16k
and VRAM columns have no in-repo source. The round-7 "14 of 91 fully resist
(5/5)" figure likewise -- the matrix stores score and percent only.

Recorded as vintage deltas, not errors: the archive is the 2026-08-25..29 sweep
while #568 is 2026-08-05, so overlapping rows differ by design; and part 1's
sessions column sits one point low on three of eleven rows under the current
scorer, reproducibly over all three repeats, which widens the set of rows above
the incumbent named in that section's Decision.

Flagged for anyone recomputing: seven archived runs report outcome ok with
non-zero output_tokens and an empty raw, and yield 14-18/69 for qwen3:14b
against the authoritative 60-62/69.

The inline edit: the round-7 "top band by run-pooled mean total_score" list
omits the three highest means -- gemma-4-26B 90.75, Foundation-Sec Q8_0 87.25,
XORTRON LARGE 83.25 -- because those are exactly the three models whose Tier B
cell was escalated to a third run and so carry a 5-run denominator. Named the
exclusion instead of silently dropping them.

---------

Co-authored-by: docs-reconcile <docs@local>

* merge: docs reconciliation round 3, policy and host slice (#3399) (#3409)

* docs(security): document the leak gate and its one allowlisted decoy .env

* docs(knowledge-store): record that the vault worker is authored but not deployed

* docs(design-lab): mark the public copy redacted and pre-#1628 snapshot

* docs(knowledge-store): repoint stale line citations at the code as it is

The design record cited four source locations by line number, and all four
had drifted: the two-flag llm-worker gate is at worker.py:246-248 and
:314-318 (not :200-202/:254-264, which land in an IP-locality helper and
in the embedding-model config), the backup-honeypot.sh archive walk is at
:119 (not :87, which is the compose-file existence check), and the
auto_sync = 0 references in ARCANE-GIT-SYNC.md are at :425 and :478
(not :321/:374, a service list and a manual-sync endpoint).

No prose change: the claims themselves all still hold.

* docs(rocky,host-tuning): record that the Rocky migration landed and measure tuning state

The homeserver was re-provisioned from Ubuntu to Rocky Linux 10.2 on
2026-09-03, so both of these documents were describing a migration that had
already happened. Add a dated status banner to each rather than rewriting
prose that is not wrong.

ROCKY-10-MIGRATION.md: the rhel branch of install-homeserver.sh is now the
production path, not a smoke-test rehearsal. Re-checked every technical
claim against scripts/install-homeserver.sh and scripts/lib/install-common.sh
-- the pkg_update/pkg_install shims, the once-at-source-time $DISTRO_FAMILY
resolution (ID first, then ID_LIKE), the full package-name table, the verbatim
Docker centos repofile, the cuda-rhel10 repo, the container_use_devices
boolean, and the non-fatal step_preflight_rhel_platform all still match, and
the base OS is still installed by hand because no kickstart artifact exists in
the tree. The ":z/:Z label gap" note is still accurate: none of the 34 compose
files under arcane/home/ carries an SELinux relabel.

HOST-TUNING.md: read-only re-measure over `ssh homeserver` shows the script is
not fully applied -- zram-generator is not installed and no zram module is
loaded (the zram-generator.conf present is hand-written, not the file the
script writes), the udev scheduler rule is absent, and noatime is on /var but
not on the /home mount the new install added. Only the tuned profile is in
effect. The 20 GB card figure is still right (20475 MiB live).

* docs(keycloak): record that the hard cutover has shipped

The file still opened with 'Status: accepted Phase 0 decision record',
which reads as a proposal, but the cutover it specifies has been carried
out. Adds a dated status note recording what the repository shows today
and leaves the contract text untouched.

Verified against the repo, not asserted: honeypot-keycloak is a manifest
entry and a deployed stack; the matrix's six 'Isolated gateway' rows
map one-to-one onto the six oauth2-proxy services in vps/docker-compose.yml
(oidc-kibana, oidc-tanner, oidc-evebox, oidc-arkime, oidc-revdeck,
oidc-traefik) sharing one image definition; and every identifier on the
hard-cutover removal list is now absent from vps/ and arcane/, surviving
only in documentation.

The Xore/auth-backend sentence needed no change -- it claims the theme
only, never a running service, which is what SENSORS.md already says.

* docs(knowledge-store): point the backup-scope decision at the comment that already records it

Two claims in the design record had fallen behind the tree, both in the
direction of "this still has to happen" when it already has:

- Section 4 said the decision to let the vault inherit backup-honeypot.sh's
  existing scope "is a decision to record verbatim in that script's own
  comment block when #2290 lands the directory". It is already recorded, at
  backup-honeypot.sh:107-114, naming #2289, #2290 and this document.
- Section 1 described Arcane's GitOps rows as carrying `auto_sync = 0`
  "set deliberately on rows that must not auto-follow".
  ARCANE-GIT-SYNC.md:475-480 records the opposite: auto_sync is 0 on *every*
  live row, including the three the manifest flags autoSync: true, which it
  calls silently inert. The conclusion the record draws from it -- that this
  codebase's git-sync tooling defaults to manual triggers -- is still what the
  source says ("Every deploy is manual"), so the reasoning stands; only the
  characterisation of the mechanism was wrong.

Everything else verified, no change needed: the vault-worker tree
(worker.py, sanitize.py, Dockerfile, two compose files), its absence from
arcane/manifests/home-production.json, the knowledge-vault-state-v1 row in
PIPELINES.md, and every code anchor in the record
(config_history.rs:14-15,50,57-88; problem_reports.rs:32,38-74;
credentials.rs:8-11; llm-worker/worker.py:246-248,314-318;
backup-honeypot.sh:15-17,119; ARCANE-GIT-SYNC.md:425,478).

* docs(rocky): correct the compose-file count behind the :z/:Z claim

The status banner said 'none of the 34 compose files under arcane/home/'.
Counted by command, 35 compose files are tracked there: the 34
stack-level compose.yml files, plus a decoy honeyfs compose nested at
arcane/home/honeypot-cowrie/cowrie/honeyfs/ that the 34 figure had
silently excluded.

The substantive claim is unchanged and re-verified: no :z or :Z relabel
appears in any of the 35.

* docs(host-tuning): name the two numbering schemes the status note mixes

The note said the verification checklist 'fails two of its six lines' and
then listed three non-compliant tunings (#1, #2, #4), which reads as a
contradiction. Both numbers are right: the six lines are the six commands
in the code block, and only the two zram lines fail; the bullets are keyed
to the table's five tuning numbers, where #2 and #4 are inapplicable as
durable changes while the single command observing each still reads
correct today. Say so instead of leaving the reader to reconcile it.

The measurements are unchanged and were re-confirmed read-only over
ssh homeserver on 2026-09-27.

* docs(keycloak): drop an unevidenced group-to-role grant from the role note

The role paragraph said the realm's low-privilege client role is "access --
granted to the users group". The role name is right and is defined in
arcane/home/honeypot-keycloak/keycloak/realm/apiary-realm.json, but nothing in
the tree records the grant: the realm's `users` and `administrators` groups
both carry empty `roleMappings`, and no provisioning script or compose
environment assigns either group a client role. So which human holds `access`
is a live-realm fact the repository does not contain.

State the role names and the gateway enforcement that *is* in the tree
(OAUTH2_PROXY_ALLOWED_ROLES is <client>:access for kibana/tanner/evebox/arkime/
revdeck, and admin for traefik-dashboard and arcane), and mark the group
membership the matrix implies as an operator-side grant to verify.

Everything else in the file re-verified: honeypot-keycloak is a manifest entry
with a postgres and keycloak service; the six "Isolated gateway" rows map
one-to-one onto the six oidc-* oauth2-proxy services sharing one image and the
oidc-gateway-environment anchor; and every identifier on the hard-cutover
removal list is absent from vps/ and arcane/, surviving only in documentation.

* merge: docs reconciliation round 3, analysis and inference slice (#3399) (#3411)

* docs(llm): mark the #598 backend comparison's overturned claims

The 2026-08-05 CPU-only research pass had two sections that later evidence
contradicts, plus one pin that has since moved. Both are now marked rather
than rewritten, so the record of what was known on that date survives.

- Status banner: the measured three-engine run lives in
  analysis/ghidra/benchmarks/engine-benchmark/README.md (2026-08-06, plus
  #832's settings tuning). It corroborates 1, 3 and 4; it overturns 6 and 7.
- Section 6: real-card throughput was in fact measured (llama.cpp ~21.4,
  Ollama ~21.5, vLLM ~23.1 tok/s) -- but on REx86 f16 7B weights, so the
  qwen3:14b comparison this doc asked for is still open.
- Section 7: vLLM's "refuses to start" failure was specific to the previous
  8GB Quadro RTX 4000 and does not reproduce on the current 20GB card. The
  stronger "the standard distribution has no CPU path at all" reading
  overstates what was shown, so it is now recorded as undetermined rather
  than restated. The image-size row and section 2's digest finding stand.
- Method: 0.32.0 was the repo's pin on 2026-08-05; production has since moved
  to ollama/ollama:0.32.13 in the ghidra compose and approved-models.json.
- Opening claim narrowed: no llama.cpp or vLLM service is deployed (39-entry
  manifest confirms), though both are now exercised in benchmark scripts.

* docs(persona,recovery): fix a misclassified sensor and an unscoped runbook

persona-design.md's air-gap table classified every non-Cowrie/Dionaea/Tanner
sensor as "no design reason either way". That is wrong for canarytokens: it
is a honeytoken platform whose public reachability is the product (#1487
made the switchboard's HTTP channel reachable through the VPS so planted
file/doc tokens actually fire, and that stack's compose file says inert
internal-only tokens do not serve it). The capture-vs-safety tradeoff the
doc argues does not transfer in either direction, so it gets its own row.
The catch-all enumeration was also missing nine deployed sensor networks;
they are now listed.

RECOVERY.md's numbered restore steps name SHA256SUMS,
stack-config-state.tar.gz and keycloak.sql.gz, which are
analysis/backup-honeypot.sh's on-host layout and resolve nowhere else --
yet the doc's stated purpose is restoring after the homeserver is gone,
which is precisely when that archive does not exist. The steps are now
scoped to the archive they actually describe, and the surviving
workstation archive is pointed at its own procedure, including the
install-homeserver.conf prerequisite those steps do not cover.

* docs(llm): reconcile the inference-backend comparison with measured results

The 2026-08-05 research pass had drifted in four ways, all verified against
stored records rather than restated from other docs:

- Section 2 cited a function that does not exist. The drift-comparison entry
  point in analysis/ghidra/models/model-governance.py is evaluate_drift(),
  not compare_against_approved().
- The status banner claimed the later three-engine run "corroborates section
  1, 3 and 4". It does not. That run tests one engine at a time by
  construction, so it says nothing about structured-output enforcement or
  keep-alive/swap, and it explicitly leaves section 2 standing. Narrowed the
  claim to what the stored record supports: overturns 6 and 7, leaves 2,
  does not re-test 1 and 3.
- Section 4's premise -- that temperature 0 with a fixed seed is relied on for
  reproducible output today -- was overtaken by #2646. With
  OLLAMA_KEEP_ALIVE=30m a warm slot returns different text for a
  byte-identical prompt, so production is permanently in the drifting regime
  and the shipped fix is slot_generation accounting, not determinism. Noted as
  a dated revision; the per-engine measurements stand.
- Section 8 listed ghidra-worker.py as needing a client-shape change. It
  already speaks an OpenAI-compatible /v1 dialect deliberately, precisely so
  llama.cpp/vLLM/LM Studio work unchanged. The only residual coupling is the
  /api/ps slot_generation probe, which already degrades to "unavailable" on a
  non-Ollama server.

Also keeps the pending uncommitted work in this file, verified: the measured
decode throughput (llama.cpp ~21.4, Ollama ~21.5, vLLM ~23.1 tok/s), the
2026-08-06 date and RTX 4000 Ada 20GB host, and the ollama 0.32.13 production
tag were each read out of analysis/ghidra/benchmarks/engine-benchmark/README.md
and the two approved-models.json / docker-compose.ghidra.yml runtime blocks.

* docs(ghidra): correct the post-#2394 GPU-identity states in the models README

The section read "Two expected, honest states after #2394 (not regressions)"
and told operators that host_gpu_uuid_changed means "expected for old
snapshots, not evidence of an actual UUID change". The checker disagrees
with the second half of that.

model-governance.py builds the code as f"host_{key}_changed" over a
per-field comparison, so host_gpu_uuid_changed fires in two situations that
need opposite responses: a pre-#2394 snapshot with no recorded UUID
(benign), and a snapshot whose UUID names a different physical card (real
drift, in which every field the old schema did compare -- name, memory,
driver -- can still look correct). An operator following the old wording
would dismiss a wrong-card event. The code string stays, so
tests/docs/test_2409_fix.py still passes, but it now carries the
present-vs-absent test for telling the cases apart, and the sibling
per-field codes are named.

Also adds approved_gpu_absent, a third #2394 host-leg code the section
never mentioned: the tool ran, was pointed at the approved UUID, and no
such card exists. Unlike the other two this is not an expected rollout
state, so the heading no longer claims they all are.

Separately, "30 days is the recorded recommendation" cited a retention
record that does not exist -- the only 30-day w…
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant