From bfb8874bd5dd8ae1fc429f09e302cd5614c5167f Mon Sep 17 00:00:00 2001 From: Xore Date: Sun, 27 Sep 2026 16:41:40 +0200 Subject: [PATCH 1/6] docs(llm): mark the #598 backend comparison's overturned claims The 2026-08-05 CPU-only research pass had two sections that later evidence contradicts, plus one pin that has since moved. Both are now marked rather than rewritten, so the record of what was known on that date survives. - Status banner: the measured three-engine run lives in analysis/ghidra/benchmarks/engine-benchmark/README.md (2026-08-06, plus #832's settings tuning). It corroborates 1, 3 and 4; it overturns 6 and 7. - Section 6: real-card throughput was in fact measured (llama.cpp ~21.4, Ollama ~21.5, vLLM ~23.1 tok/s) -- but on REx86 f16 7B weights, so the qwen3:14b comparison this doc asked for is still open. - Section 7: vLLM's "refuses to start" failure was specific to the previous 8GB Quadro RTX 4000 and does not reproduce on the current 20GB card. The stronger "the standard distribution has no CPU path at all" reading overstates what was shown, so it is now recorded as undetermined rather than restated. The image-size row and section 2's digest finding stand. - Method: 0.32.0 was the repo's pin on 2026-08-05; production has since moved to ollama/ollama:0.32.13 in the ghidra compose and approved-models.json. - Opening claim narrowed: no llama.cpp or vLLM service is deployed (39-entry manifest confirms), though both are now exercised in benchmark scripts. --- docs/llm-inference-backend-comparison.md | 48 ++++++++++++++++++++++-- 1 file changed, 45 insertions(+), 3 deletions(-) diff --git a/docs/llm-inference-backend-comparison.md b/docs/llm-inference-backend-comparison.md index 75bf7f917..83d09ce42 100644 --- a/docs/llm-inference-backend-comparison.md +++ b/docs/llm-inference-backend-comparison.md @@ -2,6 +2,18 @@ Status: research for [issue #598](https://github.com/Xore/APIARY/issues/598), 2026-08-05. +> **Superseded in part — read this first.** This is the CPU-only research +> pass. The *measured* three-engine comparison it asked for was run +> afterwards and is recorded in +> `analysis/ghidra/benchmarks/engine-benchmark/README.md` (2026-08-06, on the +> RTX 4000 Ada 20GB, extended by #832's settings tuning). It corroborates §1, +> §3 and §4 below and **overturns §6 and §7**: real-card performance was in +> fact measured, and vLLM's "refuses to start" failure turned out to be +> specific to the *previous* 8GB Quadro RTX 4000 rather than a property of the +> image. The sections below are left as written, as the record of what was +> known on 2026-08-05; where the later run contradicts them, the later run +> wins. + This is a task-specific decision record, matching [`local-llm-model-evaluation.md`](local-llm-model-evaluation.md)'s format for the model-selection decision it's paired with. It evaluates the inference @@ -18,6 +30,11 @@ APIs). No comparison against llama.cpp (the inference engine Ollama itself wraps) or vLLM (the throughput-oriented alternative) had been done. #598 asked for one, explicitly allowing "switch, and rearchitect" as a valid outcome. +(Still true of the *deployed* stack: `arcane/manifests/home-production.json` +deploys no llama.cpp or vLLM service. Both now appear in benchmark scripts +under `analysis/ghidra/benchmarks/` — measured, not depended on. See the +status banner above.) + ## Method Every claim below was tested directly — a real container, a real (small) @@ -33,7 +50,10 @@ Test model: `Qwen/Qwen2.5-0.5B-Instruct-GGUF` (Q4_K_M, 630M params) for generation tests; `nomic-ai/nomic-embed-text-v1.5-GGUF` (Q4_K_M) for the embeddings test. Images: `ghcr.io/ggml-org/llama.cpp:server`/`:full`, `ghcr.io/mostlygeek/llama-swap:cpu`, `ollama/ollama:0.32.0` (this repo's own -pinned version), `vllm/vllm-openai:latest` (v0.26.0). +pinned version *at the time*; production has since moved to +`ollama/ollama:0.32.13` in `analysis/ghidra/docker-compose.ghidra.yml` and +`analysis/ghidra/models/approved-models.json`), `vllm/vllm-openai:latest` +(v0.26.0). ## 1. Structured output enforcement @@ -175,14 +195,18 @@ prompt-processing (`pp`) and text-generation (`tg`) tokens/sec numbers. (`qwen3:14b`, promoted to all three slots under #568/#569 now that VRAM is confirmed ~20GB) on the real GPU, and compare against Ollama's own measured throughput for the same model (`docs/local-llm-model-evaluation.md` already -has some of these numbers). +has some of these numbers). **Partly done, on a different model** — the later +three-engine run in `analysis/ghidra/benchmarks/engine-benchmark/README.md` +did measure decode throughput on the real 20GB card (llama.cpp ~21.4, Ollama +~21.5, vLLM ~23.1 tok/s), but against the merged REx86 f16 7B weights, not +`qwen3:14b`. The `qwen3:14b`-specific comparison is still outstanding. ## 7. Operational surface | | Ollama 0.32.0 | llama.cpp `:server` | llama-swap `:cpu` | vLLM 0.26.0 | |---|---:|---:|---:|---:| | Image size | 8.06 GB | 1.21 GB | 1.24 GB | 28.1 GB | -| Runs without a GPU present | Yes (CPU fallback) | Yes (CPU fallback) | Yes | **No — confirmed** | +| Runs without a GPU present | Yes (CPU fallback) | Yes (CPU fallback) | Yes | **No — confirmed**, but hardware-specific; see below | The vLLM finding is real and concrete, not inferred: the standard `vllm/vllm-openai` image fails to even construct its own CLI argument parser @@ -201,6 +225,20 @@ less-maintained artifact, not the mainline image this stack would pull). Ollama and llama.cpp both degrade gracefully to CPU; vLLM does not degrade, it refuses to start. +> **Superseded.** That failure was hardware-specific, not a property of the +> image. On the current 20GB RTX 4000 Ada Generation, `vllm/vllm-openai:latest` +> (v0.26.0) starts cleanly, loads a 7B model, completes `torch.compile` +> warmup and serves correct completions — recorded in the "Side finding" +> section of `analysis/ghidra/benchmarks/engine-benchmark/README.md`. The +> failure above was observed on the *previous* 8GB Quadro RTX 4000. +> +> Whether the mainline image has any CPU-only path at all is therefore +> **undetermined** from this repo's evidence: the two data points are one +> failure on the old card and one success on the new one, and neither isolates +> GPU *presence* from GPU *capability*. The "no CPU path at all" reading above +> overstates what was shown. Unchanged by the later run: vLLM is by far the +> largest image of the four, and §2's digest/registry finding stands. + **Embeddings** (relevant to `ml-worker`/#151's planned semantic search, both currently spec'd around `nomic-embed-text`): llama.cpp has a native `--embedding` server mode with an OpenAI-compatible `/v1/embeddings` @@ -290,6 +328,10 @@ serving — not the current shared, bursty, multi-workload shape. - Run `llama-bench` against the currently-approved model (`qwen3:14b`, promoted under #568/#569) on the real card, side-by-side with Ollama's own numbers, to get the actual performance data this research couldn't gather. + (Partly superseded: the three-engine run in + `analysis/ghidra/benchmarks/engine-benchmark/README.md` gathered real-card + throughput for llama.cpp/Ollama/vLLM, but on REx86 f16 7B weights. The + `qwen3:14b` case is still open.) - Resolve the 768-vs-384 embedding-dimension discrepancy (§7) before #151's `dense_vector` ES mapping work begins, regardless of backend choice. - If a future re-evaluation is triggered (e.g. `model-governance.py` gets a From 42f70e53c8c06353ab6d836c361f0722c54624c9 Mon Sep 17 00:00:00 2001 From: Xore Date: Sun, 27 Sep 2026 16:45:24 +0200 Subject: [PATCH 2/6] docs(persona,recovery): fix a misclassified sensor and an unscoped runbook persona-design.md's air-gap table classified every non-Cowrie/Dionaea/Tanner sensor as "no design reason either way". That is wrong for canarytokens: it is a honeytoken platform whose public reachability is the product (#1487 made the switchboard's HTTP channel reachable through the VPS so planted file/doc tokens actually fire, and that stack's compose file says inert internal-only tokens do not serve it). The capture-vs-safety tradeoff the doc argues does not transfer in either direction, so it gets its own row. The catch-all enumeration was also missing nine deployed sensor networks; they are now listed. RECOVERY.md's numbered restore steps name SHA256SUMS, stack-config-state.tar.gz and keycloak.sql.gz, which are analysis/backup-honeypot.sh's on-host layout and resolve nowhere else -- yet the doc's stated purpose is restoring after the homeserver is gone, which is precisely when that archive does not exist. The steps are now scoped to the archive they actually describe, and the surviving workstation archive is pointed at its own procedure, including the install-homeserver.conf prerequisite those steps do not cover. --- docs/analysis/RECOVERY.md | 17 ++++++++++++++++- docs/persona-design.md | 3 ++- 2 files changed, 18 insertions(+), 2 deletions(-) diff --git a/docs/analysis/RECOVERY.md b/docs/analysis/RECOVERY.md index 00f817bb4..752372ef5 100644 --- a/docs/analysis/RECOVERY.md +++ b/docs/analysis/RECOVERY.md @@ -23,7 +23,22 @@ authenticated with an empty event history. See and the sizes behind it. Recovery is intentionally not automatic because overwriting live volumes is -destructive. On a replacement host: +destructive. + +**Which archive the steps below apply to**: an `analysis/backup-honeypot.sh` +directory — that script's on-host layout of `SHA256SUMS`, +`stack-config-state.tar.gz`, `keycloak.sql.gz` and `volumes/.tar.gz`. +Those names resolve nowhere else, and that archive only exists if the +homeserver itself survived. A restore driven by the archive that actually +survives a dead homeserver — `scripts/backup-essentials.sh`'s +`apiary-essentials-.tar.gz.gpg`, a different layout under +`homeserver/`, `vps/` and `repo/` — has its own procedure, and one +prerequisite the steps below do not cover: put +`homeserver/installer/*-install-homeserver.conf` back in place first, because +`scripts/install-homeserver.sh` will not run without it. See +[`docs/BACKUP-ESSENTIALS.md`](../BACKUP-ESSENTIALS.md). + +On a replacement host: 1. Verify `SHA256SUMS`, unpack `stack-config-state.tar.gz` into a new empty stack directory, and inspect `.env` permissions and values. diff --git a/docs/persona-design.md b/docs/persona-design.md index 4ada313c4..abae225a7 100644 --- a/docs/persona-design.md +++ b/docs/persona-design.md @@ -41,7 +41,8 @@ The tradeoff is real in both directions: | Cowrie | Allowed (flag: `COWRIE_AIR_GAPPED`, default `false`) | The one sensor in this stack designed around capturing attacker-fetched malware — its whole SSH/Telnet fake-shell premise is attackers running `wget`/`curl`/`tftp` against real URLs. See `arcane/home/honeypot-cowrie/compose.yml`'s `cowrie_net`. | | Dionaea | Allowed (flag: `DIONAEA_AIR_GAPPED`, default `false`) | Same tradeoff as Cowrie (#269/#538): captures shellcode/binaries pushed *to* it over SMB/FTP/TFTP/etc, which both ship enabled by default — this is the attacker's actual malware sample, not just the exploit attempt. See `arcane/home/honeypot-dionaea/compose.yml`'s `dionaea_net` (#541). `internal: true` still permits `tftp-relay`'s inbound forwarding on the same network — it only removes the outbound route. | | Tanner/Snare | Allowed (flag: `TANNER_AIR_GAPPED`, default `false`) | Same tradeoff as Cowrie and Dionaea: the `template_injection` emulator fetches real RFI payloads when enabled, capturing the attacker's actual payload instead of just the RFI attempt. See `arcane/home/honeypot-tanner/compose.yml`'s `tanner_local`. Setting the flag also breaks the emulator's own `REMOTE_DOCKERFILE` self-maintenance fetch (`raw.githubusercontent.com`) — a real cost, not just a capture-vs-safety tradeoff. | -| Everything else (Conpot personas, DNP3, HTTP/API honeypot, multipot, dicompot, dns-honeypot, citrix-honeypot, cisco-asa-honeypot, rdp-honeypot) | Allowed, no design reason either way | None of these protocols involve the honeypot fetching attacker-supplied URLs — outbound access is unused in the intended interaction, just never explicitly closed off. An operator who wants maximum containment can set `internal: true` directly on that sensor's network in its compose file without losing anything these honeypots actually rely on. | +| Canarytokens | Allowed (`canarytokens_net`; no air-gap flag exists) | **A different deception category, not a capture-vs-safety call.** This stack is a honeytoken *platform* — it plants fake documents, credentials and DNS names that alert when touched — not a sensor that fetches attacker-supplied content. Its reachability is the product: #1487 made the switchboard's HTTP channel publicly reachable through the VPS precisely so dashboard-created file/doc tokens fire when opened outside our own network, and that stack's compose file is explicit that "inert (internal-only) tokens don't serve it." The tradeoff argued above therefore does not transfer to it, in either direction. Note the asymmetry with the `internal: true` rule below: `internal` removes the *outbound* route, so it would not by itself break the VPS's inbound bridge — whether it is safe on `canarytokens_net` is undecided in-repo. Don't assume either way. | +| Everything else (Conpot personas, DNP3, HTTP/API honeypot, multipot, dicompot, dns-honeypot, citrix-honeypot, cisco-asa-honeypot, rdp-honeypot, sonicwall-sma, endlessh, beelzebub, hellpot, elasticpot, galah, sentrypeer, mailoney) | Allowed, no design reason either way | None of these protocols involve the honeypot fetching attacker-supplied URLs — outbound access is unused in the intended interaction, just never explicitly closed off. An operator who wants maximum containment can set `internal: true` directly on that sensor's network in its compose file without losing anything these honeypots actually rely on. | | `yara-scanner` (not a honeypot — offline payload analysis) | Blocked (`network_mode: none`) | Already air-gapped; scans captured files at rest, never needs network access at all. The one existing precedent this decision extends. | `COWRIE_AIR_GAPPED=true` (`.env`) sets `internal: true` on `cowrie_net`, From fe3b76721c0539739cccb91efb533337460245ca Mon Sep 17 00:00:00 2001 From: Xore Date: Sun, 27 Sep 2026 16:45:29 +0200 Subject: [PATCH 3/6] docs(llm): reconcile the inference-backend comparison with measured results The 2026-08-05 research pass had drifted in four ways, all verified against stored records rather than restated from other docs: - Section 2 cited a function that does not exist. The drift-comparison entry point in analysis/ghidra/models/model-governance.py is evaluate_drift(), not compare_against_approved(). - The status banner claimed the later three-engine run "corroborates section 1, 3 and 4". It does not. That run tests one engine at a time by construction, so it says nothing about structured-output enforcement or keep-alive/swap, and it explicitly leaves section 2 standing. Narrowed the claim to what the stored record supports: overturns 6 and 7, leaves 2, does not re-test 1 and 3. - Section 4's premise -- that temperature 0 with a fixed seed is relied on for reproducible output today -- was overtaken by #2646. With OLLAMA_KEEP_ALIVE=30m a warm slot returns different text for a byte-identical prompt, so production is permanently in the drifting regime and the shipped fix is slot_generation accounting, not determinism. Noted as a dated revision; the per-engine measurements stand. - Section 8 listed ghidra-worker.py as needing a client-shape change. It already speaks an OpenAI-compatible /v1 dialect deliberately, precisely so llama.cpp/vLLM/LM Studio work unchanged. The only residual coupling is the /api/ps slot_generation probe, which already degrades to "unavailable" on a non-Ollama server. Also keeps the pending uncommitted work in this file, verified: the measured decode throughput (llama.cpp ~21.4, Ollama ~21.5, vLLM ~23.1 tok/s), the 2026-08-06 date and RTX 4000 Ada 20GB host, and the ollama 0.32.13 production tag were each read out of analysis/ghidra/benchmarks/engine-benchmark/README.md and the two approved-models.json / docker-compose.ghidra.yml runtime blocks. --- docs/llm-inference-backend-comparison.md | 50 +++++++++++++++++++----- 1 file changed, 41 insertions(+), 9 deletions(-) diff --git a/docs/llm-inference-backend-comparison.md b/docs/llm-inference-backend-comparison.md index 83d09ce42..ec333b4dc 100644 --- a/docs/llm-inference-backend-comparison.md +++ b/docs/llm-inference-backend-comparison.md @@ -6,13 +6,21 @@ Status: research for [issue #598](https://github.com/Xore/APIARY/issues/598), 20 > pass. The *measured* three-engine comparison it asked for was run > afterwards and is recorded in > `analysis/ghidra/benchmarks/engine-benchmark/README.md` (2026-08-06, on the -> RTX 4000 Ada 20GB, extended by #832's settings tuning). It corroborates §1, -> §3 and §4 below and **overturns §6 and §7**: real-card performance was in -> fact measured, and vLLM's "refuses to start" failure turned out to be -> specific to the *previous* 8GB Quadro RTX 4000 rather than a property of the -> image. The sections below are left as written, as the record of what was -> known on 2026-08-05; where the later run contradicts them, the later run -> wins. +> RTX 4000 Ada 20GB, extended by #832's settings tuning). It **overturns §6 +> and §7**: real-card performance was in fact measured, and vLLM's "refuses to +> start" failure turned out to be specific to the *previous* 8GB Quadro +> RTX 4000 rather than a property of the image. It explicitly leaves §2 +> standing ("no digest/registry system", `model-governance.py`'s pipeline being +> "Ollama-shaped"), and it does not re-test §1's structured-output axis or §3's +> keep-alive/swap behaviour — it runs one engine at a time by construction, so +> it says nothing about multi-model sharing. The sections below are left as +> written, as the record of what was known on 2026-08-05; where later work +> contradicts them, the later work wins. +> +> Two later findings bear on this doc and are flagged inline where they +> apply: #2646 (a warm resident slot is *not* reproducible at temperature 0 +> with a fixed seed, which revises §4's premise) and `ghidra-worker.py`'s +> deliberate OpenAI-compatible `/v1` client, which revises §8's cost estimate. This is a task-specific decision record, matching [`local-llm-model-evaluation.md`](local-llm-model-evaluation.md)'s format for @@ -107,7 +115,7 @@ llama.cpp/vLLM lose on — Ollama's digest happens to be *exactly* what equivalent. A migration would replace "query the running server's registry digest" with "hash the GGUF file on disk directly" — simpler and arguably more auditable (no trust in a registry's own digest computation), but a real -rewrite of `collect_snapshot()`/`compare_against_approved()`'s identity model, +rewrite of `collect_snapshot()`/`evaluate_drift()`'s identity model, and every recorded `approved-models.json` entry's `ollama_repo_digest`-shaped field. @@ -153,6 +161,21 @@ instances with strict, static VRAM partitions instead of dynamic sharing). `temperature: 0, seed: 66` is relied on for reproducible output today. +> **Revised after the fact (#2646).** That reliance did not survive contact +> with the deployed configuration. `OLLAMA_KEEP_ALIVE=30m` keeps the weights +> resident between samples, and a **warm** Ollama slot returns different text +> for a byte-identical prompt at temperature 0 with a fixed seed — production +> therefore runs permanently in the drifting regime, and no setting fixes that +> without paying a reload. The fix that shipped is accounting, not determinism: +> every stored assessment now records a `slot_generation` fingerprint (read +> from Ollama's `/api/ps`) naming the resident instance that answered, so two +> results are never compared as though one instance produced both — see +> `docs/analysis/ghidra/AI_TRIAGE.md`. This is attributed to a warm **Ollama** +> slot specifically; the llama.cpp result below is not stated to have +> exercised an equivalent long-lived warm-slot regime, so it stands as +> measured. What the finding removes is this section's *premise* — that the +> deployed stack is reproducible today — not its per-engine results. + **llama.cpp**: confirmed bit-identical output across 3 repeated identical requests (`temperature: 0, seed: 66`) against the same model. No loss. @@ -269,7 +292,16 @@ contract) replaced Ollama: `/v1/chat/completions`'s `response_format` shape instead of `/api/chat`'s `format`; digest verification (`model_digest()`) rewritten to hash the local GGUF file instead of querying `/api/tags`. -- `analysis/ghidra/worker/ghidra-worker.py` — same client-shape change. +- `analysis/ghidra/worker/ghidra-worker.py` — **no client-shape change needed**, + as of the current code. It already speaks an OpenAI-compatible `/v1` chat + dialect (`GHIDRA_TRIAGE_API_BASE`, default `http://127.0.0.1:11434/v1`) and + does so deliberately: the worker records that this is what makes "llama.cpp's + server, vLLM and LM Studio work unchanged; the requirement is that it is + *local*, not that it is Ollama." The one Ollama-specific coupling left is the + `/api/ps` probe behind `GHIDRA_TRIAGE_RUNTIME_BASE` that populates + `slot_generation` (#2646); the code already handles a server that is not + Ollama, recording `unavailable`. This line was a real cost when this doc was + written and is no longer one. - `analysis/ghidra/models/model-governance.py` — the largest rewrite. Its entire `collect_snapshot()`/drift-comparison model is built around Ollama's registry APIs and Docker-image-reference identity; every field in From 6ca68803333836565b7c3cc07fb22995b282ec75 Mon Sep 17 00:00:00 2001 From: Xore Date: Sun, 27 Sep 2026 16:49:14 +0200 Subject: [PATCH 4/6] docs(ghidra): correct the post-#2394 GPU-identity states in the models README The section read "Two expected, honest states after #2394 (not regressions)" and told operators that host_gpu_uuid_changed means "expected for old snapshots, not evidence of an actual UUID change". The checker disagrees with the second half of that. model-governance.py builds the code as f"host_{key}_changed" over a per-field comparison, so host_gpu_uuid_changed fires in two situations that need opposite responses: a pre-#2394 snapshot with no recorded UUID (benign), and a snapshot whose UUID names a different physical card (real drift, in which every field the old schema did compare -- name, memory, driver -- can still look correct). An operator following the old wording would dismiss a wrong-card event. The code string stays, so tests/docs/test_2409_fix.py still passes, but it now carries the present-vs-absent test for telling the cases apart, and the sibling per-field codes are named. Also adds approved_gpu_absent, a third #2394 host-leg code the section never mentioned: the tool ran, was pointed at the approved UUID, and no such card exists. Unlike the other two this is not an expected rollout state, so the heading no longer claims they all are. Separately, "30 days is the recorded recommendation" cited a retention record that does not exist -- the only 30-day window under analysis/ghidra/models/ is BLOB-RETENTION-POLICY.md's unrelated #2862 dionaea-bistream rule. The advice stands, attributed to this doc. --- docs/analysis/ghidra/models/README.md | 13 +++++++++---- 1 file changed, 9 insertions(+), 4 deletions(-) diff --git a/docs/analysis/ghidra/models/README.md b/docs/analysis/ghidra/models/README.md index 8cb459050..087f88fcc 100644 --- a/docs/analysis/ghidra/models/README.md +++ b/docs/analysis/ghidra/models/README.md @@ -17,10 +17,15 @@ python3 /opt/honeypot-ghidra/models/model-governance.py check-runtime \ The status file contains only state and reason codes. It contains no prompts, model replies, captured data, container paths, or credentials, and is written owner-only mode `0600`. If a dashboard later needs it, expose only the sanitized object through a privileged read-only endpoint; do not mount or relax the host file. `approved`, `drift`, and `unavailable` are advisory states: the service exits successfully with `--warn-only`, so an LLM problem never stops ingestion or deterministic analysis. Omit `--warn-only` in an operator check when drift should produce a non-zero exit status. The command only reads `/api/version`, `/api/tags`, Docker inspection metadata, and `nvidia-smi` telemetry. -### Two expected, honest states after #2394 (not regressions) +### Post-#2394 GPU-identity states: which are expected, which are not -- **`approved_gpu_uuid_missing`** on the `host` leg: the deployed manifest copy at `/opt/honeypot-ghidra/models/approved-models.json` still predates the `approved_host.gpu_uuid` field until `install-analysis-host.sh` next runs end to end on that host. The checker reports this distinct, advisory-only code rather than silently comparing whichever GPU enumerates as index 0. It clears itself once the host is redeployed with the current manifest. -- **`host_gpu_uuid_changed`**: any `--snapshot` file captured before #2394 was written under the older schema and has no GPU-identity fields for the new comparison to match against. Replaying it will read as drift under the new schema even though nothing on the host changed -- expected for old snapshots, not evidence of an actual UUID change. +#2394 made the checker compare GPU identity by UUID rather than enumeration +order, and that added three host-leg codes. Only the first two are expected +during the rollout; the third is a real problem wearing the same shape. + +- **`approved_gpu_uuid_missing`** on the `host` leg — *expected.* The deployed manifest copy at `/opt/honeypot-ghidra/models/approved-models.json` still predates the `approved_host.gpu_uuid` field until `install-analysis-host.sh` next runs end to end on that host. The checker reports this distinct, advisory-only code rather than silently comparing whichever GPU enumerates as index 0. It clears itself once the host is redeployed with the current manifest. +- **`host_gpu_uuid_changed`** — *expected for a legacy `--snapshot` only, and this code is overloaded, so check before dismissing.* A snapshot captured before #2394 was written under the older schema and carries no GPU-identity field, so replaying it reads as drift even though nothing on the host changed. But the checker also emits the **same code** when `gpu_uuid` is present and names a *different physical card* than the manifest pins — that is real drift, not a schema artefact, and the field the old schema compared (name, memory, driver) can all still look correct. Tell the two apart by whether the snapshot has a `gpu_uuid` key at all: absent means a legacy replay, present-and-different means the host is not running the approved card. The same per-field construction applies to `host_gpu_changed`, `host_gpu_memory_mib_changed`, `host_driver_changed` and `host_compute_capability_changed`, so treat each of those the same way. +- **`approved_gpu_absent`** — *not expected.* The tool ran, was pointed at the approved UUID explicitly, and no such card exists on the host. Distinct from `gpu_telemetry_unavailable` (which means the telemetry could not be read at all). This one means the approved card is genuinely gone. ## When requalification is mandatory @@ -30,7 +35,7 @@ Run the complete workflow before changing any model tag or digest, Ollama image/ Use a trusted checkout on the approved analysis host. Stop unrelated GPU-heavy jobs if needed, but do not stop or modify QEMU. The benchmark uses only checked-in synthetic TEST-NET fixtures, talks only to the explicitly supplied local Ollama endpoint, records exact artifacts/settings/timing/RAM/VRAM metadata, and unloads each candidate through Ollama after its slot. It never downloads a model. -Keep verbose replies outside the repository in an operator-only directory with bounded retention (30 days is the recorded recommendation): +Keep verbose replies outside the repository in an operator-only directory with bounded retention (30 days is this doc's own recommendation; no separate retention record covers this directory): ```sh install -d -m 0700 "$HOME/model-qualification" From 7f4f86ac104969deed3a36ca46df548b5637a1ff Mon Sep 17 00:00:00 2001 From: Xore Date: Sun, 27 Sep 2026 16:49:15 +0200 Subject: [PATCH 5/6] docs(personas): disambiguate the T-Pot README citation in persona-design One drift found. The doc cited "`README.md` line 296" for T-Pot's outbound guidance without saying whose README. In this repository that pointer resolves to the root README.md, which is 135 lines and contains no such text, so a reader following it finds nothing. The line number refers to T-Pot's upstream README, recorded from an unpinned read. Outbound network access was unavailable during this reconciliation, so the number could not be re-verified and is not restated as fact -- the citation is now labelled as a lookup pointer rather than a stable one. The substantive claim is unaffected: the doc already quotes T-Pot's outbound text inline, and that quote is what the decision rests on. Everything else in this file verified against the tree with no drift: - All 21 sensor stacks named in the outbound table exist under arcane/home/, and the four named networks (cowrie_net, dionaea_net, tanner_local, canarytokens_net) plus the yara-scanner service's network_mode: none are as described. - "no honeypot's network has ever set internal: true" still holds: exactly three files set it (oidc-session, honeypot-llm-data, keycloak-data) and all three are infrastructure networks, not sensor capture networks. - The three air-gap flags and their default-false wiring match compose and .env.example in all three stacks. - The "one org per estate" example is real: nexusai-core, nexusai-edge and nexusai-platform are site_ids in personas/personas.json, all under NexusAI Research GmbH, wired into honeypot-multipot and honeypot-http. - cowrie.cfg's banner string, COWRIE_HOSTNAME=gpu01, and the canarytokens "inert (internal-only) tokens don't serve it" quote are exact. --- docs/persona-design.md | 8 ++++++-- 1 file changed, 6 insertions(+), 2 deletions(-) diff --git a/docs/persona-design.md b/docs/persona-design.md index abae225a7..dce39ce9b 100644 --- a/docs/persona-design.md +++ b/docs/persona-design.md @@ -5,8 +5,12 @@ Two related pieces of persona design that were implicit rather than documented decisions: whether a honeypot may reach the internet outbound, and how to name/place a honeypot host so it doesn't look staged. T-Pot's -own README calls both out by name (`README.md` line 296 for outbound; the -"where to place a honeypot" guidance for siting). This repo already does +own upstream README calls both out by name (its `README.md` line 296 for +outbound; the "where to place a honeypot" guidance for siting). That line +number refers to T-Pot's own repository, not to this repo's `README.md`, which +is ~135 lines; it was recorded from an unpinned upstream read and could not be +re-verified during this reconciliation, so treat the line number as a pointer +to look up rather than a stable citation. This repo already does deep, source-verified realism work for Windows personas ([#91](https://github.com/Xore/APIARY/issues/91)/[#94](https://github.com/Xore/APIARY/issues/94)/[#96](https://github.com/Xore/APIARY/issues/96)) and has a full fictional-organization inventory From f7f373f4bdce0d79764358f9988d9d0d36423865 Mon Sep 17 00:00:00 2001 From: Xore Date: Sun, 27 Sep 2026 16:54:17 +0200 Subject: [PATCH 6/6] docs(sandbox): correct the Windows results path in the KVM traffic doc The short-version table gave the Windows capture results location as "sandbox/results//". Nothing tracked lives there, and the row conflated two different directories. The run-keyed artifact directory -- the one the dashboard actually reads (es_results_importer.rs and sandbox_submit.rs both resolve WINDOWS_SANDBOX_RESULTS_DIR) -- is $WINDOWS_SANDBOX_RESULTS_DIR, with one subdirectory per sample named by its sha256, per run_sample.py's detonate(). Its own fallback when that env var is unset is reports/windows-sandbox. sandbox/results/current is real but is only docker-compose.sandbox.yml's bare default for a manual `up`: all five sniffer bind-mounts key off ${SANDBOX_RESULTS_DIR:-./sandbox/results/current}, and run_sample.py points SANDBOX_RESULTS_DIR at the run's own out_dir, so the default only applies when the stack is brought up by hand. Everything else in the doc verified against the tree: both bridges and the absent , the 10.10.10.x address plan, macvlan internal:true, the sniffer NET_ADMIN/NET_RAW + network_mode: host grant versus cap_drop on INetSim and mitmproxy, the honeypot-sandbox-strict filterref, the controlled-mode 198.18.0.1 DNS/proxy, and all three retention figures (cleanup.sh 30/180/7). --- docs/kvm-network-traffic-analysis.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/kvm-network-traffic-analysis.md b/docs/kvm-network-traffic-analysis.md index 563160ecd..a2baefddc 100644 --- a/docs/kvm-network-traffic-analysis.md +++ b/docs/kvm-network-traffic-analysis.md @@ -27,7 +27,7 @@ are **two** such bridges, because there are two sandboxes: | Network XML | `sandbox/windows/setup/sandbox-network.xml` | `sandbox/network.xml` | | Fake internet | INetSim at `10.10.10.1` | none by default; optional logged DNS + Squid allowlist (`controlled` mode) | | Capture | `docker-compose.sandbox.yml` (tcpdump, Zeek, Suricata) | root-owned `tcpdump` per job, host and guest side | -| Results | `sandbox/results//` | `/var/lib/honeypot-sandbox/results/` (root-only), sanitized export copied out | +| Results | `$WINDOWS_SANDBOX_RESULTS_DIR//` — this is the path the dashboard reads. The compose file's own `./sandbox/results/current` default only applies to a manual `docker compose -f docker-compose.sandbox.yml up`; `run_sample.py` overrides `SANDBOX_RESULTS_DIR` to the current run's out_dir, and its own fallback when the env var is unset is `reports/windows-sandbox` | `/var/lib/honeypot-sandbox/results/` (root-only), sanitized export copied out | | Orchestrator | `sandbox/windows/orchestrate/run_sample.py` | `sandbox/run-linux-sample.sh` | Neither bridge has a `` element, so neither can route anywhere. That