Skip to content

merge: docs reconciliation round 2, infra slice (#3399) - #3405

Closed
Xore wants to merge 21 commits into
docs/3395-doc-reconciliationfrom
docs/3399-r2-infra
Closed

Xore wants to merge 21 commits into
docs/3395-doc-reconciliationfrom
docs/3399-r2-infra

Conversation

@Xore

@Xore Xore commented Sep 27, 2026

Copy link
Copy Markdown
Owner

Round 2 of the #3399 docs reconciliation, infra slice — 20 files, +712/−264.
Companion to #3404 (map slice).

Gates — all six green on this branch

mermaid            40 blocks in 154 files parse cleanly
links              412 local refs resolve
paths-exist        493 tokens, 41 allowlisted
reachability       85 reachable, 35 exempt
stale-paths        passed
public-leaks       passed

The find: the homeserver was re-provisioned and the docs never knew

docs/HOMESERVER-DISK-LAYOUT.md documented an Ubuntu install that no longer
exists. Re-measured read-only over ssh on 2026-09-27:

  • OS is Rocky Linux 10.2, not Ubuntu/subiquity. The doc's curtin
    autoinstall section described an install that has been replaced.
  • Boot disk is a different, 4×-larger NVMe (1TB) and is now LVM-backed
    — rl-root 70G, rl-swap 32G, rl-home 849.3G. The doc recorded
    "no LVM" as a design decision.
  • /var is an sdb1 partition of an 8.7T RAID LUN, not a whole disk. The doc
    left sdb's size unstated in the table, then called it "the 1.7T sdb disk" in
    prose — an internal contradiction, now moot.
  • /mnt-1 and /mnt-2 are both decommissioned; their workloads moved to
    /var. The autoinstall config is kept and explicitly marked as no longer a
    rebuild target.
  • /var/lib/docker is 2.9T and /var/dockge 350G, against 103G/229G in the
    doc. 45 stack dirs, not 23.

docs/BACKUP-ESSENTIALS.md: 40 stack .env files, not 41. The volumes table
listed short names — four of the five real volumes carry an Arcane project
prefix, and backup-essentials.sh writes .tar.gz, so following the old table
would have "restored" into volumes no stack is mounted against.

docs/HOST-TUNING.md: no change needed — all five tunings verified against
tune-rocky10.sh and the live RTX 4000 Ada.

docs/SECURITY.md

Referenced a leak gate without documenting it. Now describes
scripts/check-public-leaks.py and its allowlist.

Other corrections

Two notes on process

A commit in this branch was rewritten for compliance. An intermediate
commit carried a Co-Authored-By: Claude trailer, which violates the repo's
rule.no_ai_in_commits. The agent caught and rewrote it itself; all 21 commits
now verify clean.

The agent reverted five out-of-scope files (SECURITY.md,
ROCKY-10-MIGRATION.md, analysis/RECOVERY.md, both KVM docs) after
overlapping a sibling slice, preserving the corrections as a handoff patch. All
five are legitimately in the issue's list, so nothing is lost — they were the
right call to hand back rather than double-edit.

docs-reconcile and others added 21 commits September 27, 2026 14:47
- 37/31/32 census -> 39 (33 under arcane/home/, 6 self-contained at root
  paths); document #2911 technitium and #3092 unsloth
- manifest import section: filter reaches 35 of 39, technitium replaced
  pihole's dedicated step, unsloth has no installer path at all
- manifest declares branch: production since #1943, not main; no
  refs/heads/production exists on origin (checked 2026-09-27)
- 34-of-37 build / three pullers -> 34-of-39 build / five pullers
- restart: no census: 20 hits, 13 declarations, 8 files; honeypot-elk's
  arkime-pcap-init (#3128) is a second Arcane-started one-shot
- relative-bind-mount set 9 -> 10 (ghidra); env_file claim: ghidra has one
- re-count dashboard 268->279 files, ghosts 989->962, ghosts-src 963->935
- :?required stacks are canarytokens/ghosts/technitium/unsloth
- 31 in-tree stacks -> 33 (add honeypot-dashboard-backend #1622 and
  honeypot-sonicwall-sma #3131, drop the retired ip-enrichment worker);
  note unsloth as the 33rd directory
- pihole -> technitium in the self-contained six and its dedicated step
- the investigation chain is a loadBalancer to each oauth2-proxy gateway,
  not a forwardAuth middleware (0 forwardAuth in dynamic.yml), and only
  five of the six gateways have a socat-hp-* bridge
- Arcane pin is v2.11.1, not v2.8.0; the confirmed limitations have not
  been re-confirmed against it
- mark the SFTP-upload/rebuild-from-terminal paragraph as pre-#1502
SECURITY.md names 'leak a real secret' as a reportable class but never says
the repository enforces it in CI, and never warns that the gate scans
untracked files. Both facts were load-bearing on this run: the checker
carries the home-server address as a named forbidden literal, and
RESUME.md quotes that address verbatim, so the gate fails on a clean
tracked tree purely because of dispatch artifacts sitting in the checkout.

Adds the gate's actual pattern set, its three fail-closed exemption sets
with their real sizes (3 ALLOWED_DOTENV paths, 2 ALLOWED_LITERAL_FIXTURE_FILES
paths), the four deployment-specific literals it assembles from fragments,
and the untracked-file scan. No existing prose changed.
…mented one

Every row of HOMESERVER-DISK-LAYOUT.md's 'as installed' table was stale.
Re-measured read-only over ssh 2026-09-27 (lsblk/lvs/findmnt/df/du):

  - OS is Rocky Linux 10.2, not Ubuntu/subiquity. The doc's curtin
    autoinstall notes described an install that has been replaced.
  - boot disk is a different, 4x-larger NVMe (PC401 SK hynix 1TB,
    953.9G) and is now LVM-backed: rl-root 70G, rl-swap 32G, rl-home
    849.3G. The doc recorded 'no LVM' as a design decision.
  - /var is on an sdb1 partition of an 8.7T LUN, not a whole disk.
    The doc left sdb's size unstated in the table and then called it
    'the 1.7T sdb disk' in prose -- an internal contradiction, now moot.
  - sda is a USB-attached Samsung PSSD T7 at /mnt/usb-recovery, so the
    /mnt-2 bulk-storage role is gone; no sr0 optical device enumerates.
  - swap is a 32G LVM LV (14.6G in use at measurement), not an 8G
    /swap.img swapfile; /swap.img does not exist.
  - /var/lib/docker is 2.9T and /var/dockge 350G, against 103G/229G and
    332G combined in the doc. 45 stack dirs, not 23.

The autoinstall config and its manual-partitioning walkthrough are kept
as the record of the former Ubuntu layout and explicitly marked as no
longer a rebuild target for this host.

BACKUP-ESSENTIALS.md: three corrections, all verified against
scripts/backup-essentials.sh and the live host.
  - 40 stack .env files as of 2026-09-27, not 41, and phrased per-stack
    so the number is read as a measurement rather than a constant.
  - The volumes table listed short names (arcane-data, evebox-config,
    ...). Four of the five real volumes carry an Arcane project prefix,
    and backup-essentials.sh writes each archive as .tar.gz, so
    following the old table would 'restore' into volumes no stack is
    mounted against. Table and restore step 5 now carry the real names.
  - Samsung PSSD T7 is the model lsblk/udevadm report; and the dead
    Keycloak restic config is now doubly dead, since /mnt-2 itself has
    been decommissioned.

HOST-TUNING.md needed no change: all five tunings, half-of-RAM capped
8G zram, priority 100, --replace-swap, the per-class schedulers, the
cups/bluetooth/ModemManager service list and the '20 GB card' claim all
match scripts/tune-rocky10.sh, and the host's RTX 4000 Ada (20475 MiB)
confirms the card size.
… backend

Both docs still described the Go dashboard that #1628 deleted. The
corrections were verified against backend-service/src/main.rs and the
workbench modules, and the manifest.

payload-analysis-workbench.md:
  - All 8 HTTP contract routes were wrong. The live set is
    /api/v1/workbench/{analyzers,runs,runs/{id},runs/{id}/children/
    {analyzer_id}/{action},recipes} (main.rs:456-468), so: the prefix
    moves from /api/payload-workbench to /api/v1/workbench, the hash is
    a query parameter rather than a path segment, and cancel/retry
    collapse into one route dispatched on {action} rather than being
    two paths. The cancelled child-action row is gone with the merge.
  - Registry pointer dashboard/workbench_domain.go ->
    backend-service/src/workbench_domain.rs. There are zero .go files
    under dashboard/ in this repo.
  - /payload-workbench -> /payload-workbench/results, whose
    workbench-builder section is the orchestration surface.
  - The 'closed Go schema (unknown fields are rejected)' and 64 KiB body
    cap are not present in the Rust tier (no deny_unknown_fields on
    workbench_api.rs, no body-limit constant anywhere), so the sentence
    now says not to rely on them rather than asserting a guarantee the
    server does not make.
  - Dropped the rollback note about a local /state/analysis-workbench
    copy; no such store exists, workbench_es.rs is the only one.
  - MODEL_STATUS_SOCKET appears nowhere in the backend, compose or
    .env.example. Marked undetermined rather than deleting the claim --
    the adapter is genuinely installed, and deleting would lose the
    intended contract.

gpu-llm-analysis-worker.md:
  - /api/llm/analysis does not exist; the page reads the generic store
    route /api/v1/store/llm-analysis (main.rs:438, stores.rs).
    dashboard/llm_analysis.go -> frontend-next/src/routes/llm-analysis.tsx.
  - Semantic search was marked 'Deferred' but shipped as
    /api/v1/llm-search (main.rs:341). Recorded as delivered, with a note
    that the section is historical scope rather than current state.
  - 'Managed by Dockge under /opt/stacks' is obsolete: Arcane gitops
    manages /var/dockge/stacks, and /opt/stacks is now only a
    compatibility symlink to it (created 2026-09-04).
  - The file list omitted the entrypoint the manifest actually deploys.
    arcane/manifests/home-production.json sets the llm-worker sync's
    dockerComposePath to llm-worker/docker-compose.captured-data-deploy.yml,
    so a bare docker-compose.yml bring-up does not reproduce the live
    worker (#2234). Added with that warning.
  - Host RAM/CPU 91 GiB / 16 CPUs -> 92 GiB / 48, measured on the host.
The previous phrasing conflated two different counts: 34 is how many
directories exist on disk under arcane/home/, while the manifest holds 39
entries of which 33 name one of those directories and 6 are root-level
stacks. Spelled out so the sentence cannot be read as '34 of the 45 stack
directories are manifest-managed'.
…ap drift

ROCKY-10-MIGRATION.md: 'is moving from Ubuntu' -> 'has moved' (the rebuild
hit live 2026-09-03; install-homeserver.sh:324 says so). Also resolved the
one claim left undetermined by the previous pass: both cards really are on
the bus (P2200 17:00.0/10de:1c31, Ada 65:00.0), but only the Ada has a
driver bound, so nvidia-smi lists one GPU. Documented that explicitly,
because 'one GPU in nvidia-smi' reads like missing hardware and is the
reason the nvidia-open-vs-cuda-drivers choice still matters.

KEYCLOAK-CUTOVER.md: the hard cutover this contract specifies has shipped.
vps/forward-auth/ is gone and xore_sso/_auth/verify/strip-auth-identity
exist nowhere but this doc and docs/TESTING.md. Status line now says
SHIPPED; the contract body is untouched. Verified all 8 OIDC client IDs
still exist in arcane/home/honeypot-keycloak/keycloak/realm/apiary-realm.json
with the documented flow flags.

kvm-network-traffic-analysis.md: 'two bridges' understated it -- there are
four (virbr-hpsbx, virbr-cape, virbr-ghosts, plus 198.18.0.0/24 on
virbr-hpsbx in controlled mode). Results path was wrong: run_sample.py
overrides the compose default per run via WINDOWS_SANDBOX_RESULTS_DIR, and
sandbox/results/ does not exist. 'The Linux sandbox has no Zeek equivalent'
is false -- run-linux-sample.sh runs 'zeek -r' offline and export-result.py
reads the logs. The Phase 0 FORWARD DROP pair is a documented manual step,
not scripted, and needs re-verification against Rocky 10.2 firewalld.

kvm-snapshot-vs-golden-image.md: §5's golden-win10 and
/var/lib/libvirt/golden/ are illustrative; kvm_manage.sh uses
SANDBOX_ROOT=/var/dockge/sandbox with golden-images/win11-analysis.qcow2
and spawns via qemu-img create -b + virsh define, not virt-clone. Added a
pointer instead of rewriting §5.

ml-gpu-coordinated-roadmap.md: kept as a dated plan, added an
intent-vs-shipped banner covering the four things that moved since
2026-08-01, including that §1 decision 5 was superseded outright -- the ML
worker has no embedding code at all, and llm-worker ships 768-dim
embeddings behind LLM_EMBEDDING_ENABLED (default off).
…ault

The previous commit cited a 'compose default of sandbox/results/current'
for WINDOWS_SANDBOX_RESULTS_DIR. No such default exists --
docker-compose.sandbox.yml never mentions the variable. The verified facts
are that run_sample.py:124 reads the variable with a
'reports/windows-sandbox' fallback, .env.example documents
'/windows-sandbox-results', and the dashboard's es-results-importer
populates it per stack. Reworded to those, which also keeps the doc-path
lint green.
… retrain schedule

analysis/RECOVERY.md: the two backups were described as 'the same scope',
which is wrong and hides a filename trap. Scope differs -- the essentials
archive additionally carries the VPS config, WireGuard, Technitium, the
installer answers and the repo runbooks. Concretely, SHA256SUMS and
stack-config-state.tar.gz are the ON-HOST copy's only; a
backup-essentials.sh archive has neither, and this doc's own first bullet
points readers at the essentials archive. Also replaced
'docker compose -f compose.yml create' with per-volume 'docker volume
create <full name>' -- four of the five volumes carry an Arcane prefix and
live in four different stacks, so one compose create cannot make them --
and restored the real start order (honeypot-elk healthy before
honeypot-init) from STACK-REBUILD.md, which this step had lost.

sandbox/README.md: the workbench registry pointer
dashboard/workbench_domain.go -> backend-service/src/workbench_domain.rs
(the Go tree has no .go files since #1628), in prose and in the mermaid
node. Two host numbers were stale: 16 -> 48 logical CPUs, and the IOMMU
figure. The doc's 84 groups was wrong on the number AND missed the story:
the Rocky 10 rebuild came up with zero groups, which is why
install-homeserver.sh's step_vfio_gpu_passthrough now adds intel_iommu=on
to the kernel command line. Re-measured 2026-09-27: 94 groups, with
intel_iommu=on iommu=pt confirmed live in /proc/cmdline, and /dev/kvm
present. Line 237's 4 vCPU / 8 GiB guidance matches neither the real
Windows domain (8 vCPU / 16 GiB) nor sandbox.env.example's
SANDBOX_VM_MEMORY_MB=3072, so it is now flagged as operator guidance
rather than left reading like a measured limit.

ml-worker-plan.md: RETRAIN_INTERVAL is gone -- #172 replaced it with four
fixed UTC slots (worker.py:59, default 03:00,09:00,15:00,21:00). Fixed in
all four places it was asserted, including the early-retrain trigger, which
is now slot-nearest rather than interval-remaining. Corrected the status
callout: the worker runs from its own Arcane stack, not the repo-root
compose file, and docker-compose.ml-worker.gpu.yml is inert and undeployed
so the worker is CPU-only. Fixed the v0.1 paragraph, which contradicted
its own §10 and the line below it by claiming there were no tests.
Added a dated reading note separating the shipped sections from the
original-draft ones, so the remaining dashboard/*.go references read as
historical instead of as current.

Benchmark records #65-#69: dated intent-vs-shipped banners only, no
history rewritten and no recorded number touched. Added alongside the
existing /mnt-1 banners rather than replacing them. #66 is superseded by
#2694; #67's hardware table is pre-reinstall and its #2985 'missing
scripts' are now in git except gptoss_rerun.sh; #68's unsloth is the
Arcane stack, not an installer; #69's TOOLCHAIN.md 'no smoke test yet'
still stands.
Uncommitted work in the tree when this run ended; committing so the
reconciliation is not lost. It sits after the 2026-08-30 supersession
banner and rewrites no measured number: §9's 'it is not merged' is stale
(the branch landed as df650a8, the same minute as the 16:20Z report
count), §2's Tier B tally is the matrix's own rather than the shipped
64-row fixture's (15 FAIL / 15 PASS, 13 payload hits to 2, not 14/15 and
11/3), test_record_baseline.py is 54 not 50, §8's governance-gate term
list omits 'conclude benign', and §8's '59 hand-labelled answers' is the
wrong set twice over.
…banner

The banner lists the #2985 scripts as now in git, but there is no
root corpus/; the tracked copies are under
analysis/ghidra/benchmarks/corpus/. The doc body already gives that
directory at the P1 section, so the banner now agrees with it.
No measured number touched.
SECURITY.md, docs/ROCKY-10-MIGRATION.md, docs/analysis/RECOVERY.md,
docs/kvm-network-traffic-analysis.md and
docs/kvm-snapshot-vs-golden-image.md are not on the round-2 r2-infra
assignment list, but earlier commits in this branch had corrected
them. Reverted to the branch point so this branch touches only its
19 assigned files and cannot collide with the slice that owns those
docs. The corrections are preserved as a handoff patch outside the
repo; they are reported in the round-2 handback.
… and matrices

Checked every claim this file can be checked against, and recorded the result
at the top. No score in the file was changed -- the one edit is a clause that
makes an existing claim true.

Confirmed: the round-7 cold baseline reconciles cell for cell against
round7-cold-baseline.json (91 models, 182 cells, 367 records, 179 reproduced,
the same 3 escalated cells and third-run values, 11 zero-scored tags, and
every anchor including the 12.1 / 7.3 point gaps). The cold-cohort table
reconciles against 1947-cohort-cold-protocol.json down to min/run B and
Ornith's injection FAIL. The twelve-model survey reconciles against
1805c-ghidra-slot-matrix.json. The #568 approved qwen3:14b@bdbd181c33f2 and
context_tokens 32768 are confirmed against approved-models.json.

Recorded as undetermined rather than asserted either way: the archive holds no
ghidra-slot record at all (all 62 runs, 1498 records, are revdeck or sessions)
and no transcript carries VRAM or a context-probe result, so the Ghidra, 16k
and VRAM columns have no in-repo source. The round-7 "14 of 91 fully resist
(5/5)" figure likewise -- the matrix stores score and percent only.

Recorded as vintage deltas, not errors: the archive is the 2026-08-25..29 sweep
while #568 is 2026-08-05, so overlapping rows differ by design; and part 1's
sessions column sits one point low on three of eleven rows under the current
scorer, reproducibly over all three repeats, which widens the set of rows above
the incumbent named in that section's Decision.

Flagged for anyone recomputing: seven archived runs report outcome ok with
non-zero output_tokens and an empty raw, and yield 14-18/69 for qwen3:14b
against the authoritative 60-62/69.

The inline edit: the round-7 "top band by run-pooled mean total_score" list
omits the three highest means -- gemma-4-26B 90.75, Foundation-Sec Q8_0 87.25,
XORTRON LARGE 83.25 -- because those are exactly the three models whose Tier B
cell was escalated to a third run and so carry a 5-run denominator. Named the
exclusion instead of silently dropping them.
@github-actions

Copy link
Copy Markdown

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

@Xore

Xore commented Sep 27, 2026

Copy link
Copy Markdown
Owner Author

Closing in favour of #3398, which covers the docs reconciliation and is already green and auto-merging. This branch came from an orphaned agent run.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant