diff --git a/docs/analysis/RECOVERY.md b/docs/analysis/RECOVERY.md index 00f817bb4..752372ef5 100644 --- a/docs/analysis/RECOVERY.md +++ b/docs/analysis/RECOVERY.md @@ -23,7 +23,22 @@ authenticated with an empty event history. See and the sizes behind it. Recovery is intentionally not automatic because overwriting live volumes is -destructive. On a replacement host: +destructive. + +**Which archive the steps below apply to**: an `analysis/backup-honeypot.sh` +directory — that script's on-host layout of `SHA256SUMS`, +`stack-config-state.tar.gz`, `keycloak.sql.gz` and `volumes/.tar.gz`. +Those names resolve nowhere else, and that archive only exists if the +homeserver itself survived. A restore driven by the archive that actually +survives a dead homeserver — `scripts/backup-essentials.sh`'s +`apiary-essentials-.tar.gz.gpg`, a different layout under +`homeserver/`, `vps/` and `repo/` — has its own procedure, and one +prerequisite the steps below do not cover: put +`homeserver/installer/*-install-homeserver.conf` back in place first, because +`scripts/install-homeserver.sh` will not run without it. See +[`docs/BACKUP-ESSENTIALS.md`](../BACKUP-ESSENTIALS.md). + +On a replacement host: 1. Verify `SHA256SUMS`, unpack `stack-config-state.tar.gz` into a new empty stack directory, and inspect `.env` permissions and values. diff --git a/docs/analysis/ghidra/models/README.md b/docs/analysis/ghidra/models/README.md index 8cb459050..087f88fcc 100644 --- a/docs/analysis/ghidra/models/README.md +++ b/docs/analysis/ghidra/models/README.md @@ -17,10 +17,15 @@ python3 /opt/honeypot-ghidra/models/model-governance.py check-runtime \ The status file contains only state and reason codes. It contains no prompts, model replies, captured data, container paths, or credentials, and is written owner-only mode `0600`. If a dashboard later needs it, expose only the sanitized object through a privileged read-only endpoint; do not mount or relax the host file. `approved`, `drift`, and `unavailable` are advisory states: the service exits successfully with `--warn-only`, so an LLM problem never stops ingestion or deterministic analysis. Omit `--warn-only` in an operator check when drift should produce a non-zero exit status. The command only reads `/api/version`, `/api/tags`, Docker inspection metadata, and `nvidia-smi` telemetry. -### Two expected, honest states after #2394 (not regressions) +### Post-#2394 GPU-identity states: which are expected, which are not -- **`approved_gpu_uuid_missing`** on the `host` leg: the deployed manifest copy at `/opt/honeypot-ghidra/models/approved-models.json` still predates the `approved_host.gpu_uuid` field until `install-analysis-host.sh` next runs end to end on that host. The checker reports this distinct, advisory-only code rather than silently comparing whichever GPU enumerates as index 0. It clears itself once the host is redeployed with the current manifest. -- **`host_gpu_uuid_changed`**: any `--snapshot` file captured before #2394 was written under the older schema and has no GPU-identity fields for the new comparison to match against. Replaying it will read as drift under the new schema even though nothing on the host changed -- expected for old snapshots, not evidence of an actual UUID change. +#2394 made the checker compare GPU identity by UUID rather than enumeration +order, and that added three host-leg codes. Only the first two are expected +during the rollout; the third is a real problem wearing the same shape. + +- **`approved_gpu_uuid_missing`** on the `host` leg — *expected.* The deployed manifest copy at `/opt/honeypot-ghidra/models/approved-models.json` still predates the `approved_host.gpu_uuid` field until `install-analysis-host.sh` next runs end to end on that host. The checker reports this distinct, advisory-only code rather than silently comparing whichever GPU enumerates as index 0. It clears itself once the host is redeployed with the current manifest. +- **`host_gpu_uuid_changed`** — *expected for a legacy `--snapshot` only, and this code is overloaded, so check before dismissing.* A snapshot captured before #2394 was written under the older schema and carries no GPU-identity field, so replaying it reads as drift even though nothing on the host changed. But the checker also emits the **same code** when `gpu_uuid` is present and names a *different physical card* than the manifest pins — that is real drift, not a schema artefact, and the field the old schema compared (name, memory, driver) can all still look correct. Tell the two apart by whether the snapshot has a `gpu_uuid` key at all: absent means a legacy replay, present-and-different means the host is not running the approved card. The same per-field construction applies to `host_gpu_changed`, `host_gpu_memory_mib_changed`, `host_driver_changed` and `host_compute_capability_changed`, so treat each of those the same way. +- **`approved_gpu_absent`** — *not expected.* The tool ran, was pointed at the approved UUID explicitly, and no such card exists on the host. Distinct from `gpu_telemetry_unavailable` (which means the telemetry could not be read at all). This one means the approved card is genuinely gone. ## When requalification is mandatory @@ -30,7 +35,7 @@ Run the complete workflow before changing any model tag or digest, Ollama image/ Use a trusted checkout on the approved analysis host. Stop unrelated GPU-heavy jobs if needed, but do not stop or modify QEMU. The benchmark uses only checked-in synthetic TEST-NET fixtures, talks only to the explicitly supplied local Ollama endpoint, records exact artifacts/settings/timing/RAM/VRAM metadata, and unloads each candidate through Ollama after its slot. It never downloads a model. -Keep verbose replies outside the repository in an operator-only directory with bounded retention (30 days is the recorded recommendation): +Keep verbose replies outside the repository in an operator-only directory with bounded retention (30 days is this doc's own recommendation; no separate retention record covers this directory): ```sh install -d -m 0700 "$HOME/model-qualification" diff --git a/docs/kvm-network-traffic-analysis.md b/docs/kvm-network-traffic-analysis.md index 563160ecd..a2baefddc 100644 --- a/docs/kvm-network-traffic-analysis.md +++ b/docs/kvm-network-traffic-analysis.md @@ -27,7 +27,7 @@ are **two** such bridges, because there are two sandboxes: | Network XML | `sandbox/windows/setup/sandbox-network.xml` | `sandbox/network.xml` | | Fake internet | INetSim at `10.10.10.1` | none by default; optional logged DNS + Squid allowlist (`controlled` mode) | | Capture | `docker-compose.sandbox.yml` (tcpdump, Zeek, Suricata) | root-owned `tcpdump` per job, host and guest side | -| Results | `sandbox/results//` | `/var/lib/honeypot-sandbox/results/` (root-only), sanitized export copied out | +| Results | `$WINDOWS_SANDBOX_RESULTS_DIR//` — this is the path the dashboard reads. The compose file's own `./sandbox/results/current` default only applies to a manual `docker compose -f docker-compose.sandbox.yml up`; `run_sample.py` overrides `SANDBOX_RESULTS_DIR` to the current run's out_dir, and its own fallback when the env var is unset is `reports/windows-sandbox` | `/var/lib/honeypot-sandbox/results/` (root-only), sanitized export copied out | | Orchestrator | `sandbox/windows/orchestrate/run_sample.py` | `sandbox/run-linux-sample.sh` | Neither bridge has a `` element, so neither can route anywhere. That diff --git a/docs/llm-inference-backend-comparison.md b/docs/llm-inference-backend-comparison.md index 75bf7f917..ec333b4dc 100644 --- a/docs/llm-inference-backend-comparison.md +++ b/docs/llm-inference-backend-comparison.md @@ -2,6 +2,26 @@ Status: research for [issue #598](https://github.com/Xore/APIARY/issues/598), 2026-08-05. +> **Superseded in part — read this first.** This is the CPU-only research +> pass. The *measured* three-engine comparison it asked for was run +> afterwards and is recorded in +> `analysis/ghidra/benchmarks/engine-benchmark/README.md` (2026-08-06, on the +> RTX 4000 Ada 20GB, extended by #832's settings tuning). It **overturns §6 +> and §7**: real-card performance was in fact measured, and vLLM's "refuses to +> start" failure turned out to be specific to the *previous* 8GB Quadro +> RTX 4000 rather than a property of the image. It explicitly leaves §2 +> standing ("no digest/registry system", `model-governance.py`'s pipeline being +> "Ollama-shaped"), and it does not re-test §1's structured-output axis or §3's +> keep-alive/swap behaviour — it runs one engine at a time by construction, so +> it says nothing about multi-model sharing. The sections below are left as +> written, as the record of what was known on 2026-08-05; where later work +> contradicts them, the later work wins. +> +> Two later findings bear on this doc and are flagged inline where they +> apply: #2646 (a warm resident slot is *not* reproducible at temperature 0 +> with a fixed seed, which revises §4's premise) and `ghidra-worker.py`'s +> deliberate OpenAI-compatible `/v1` client, which revises §8's cost estimate. + This is a task-specific decision record, matching [`local-llm-model-evaluation.md`](local-llm-model-evaluation.md)'s format for the model-selection decision it's paired with. It evaluates the inference @@ -18,6 +38,11 @@ APIs). No comparison against llama.cpp (the inference engine Ollama itself wraps) or vLLM (the throughput-oriented alternative) had been done. #598 asked for one, explicitly allowing "switch, and rearchitect" as a valid outcome. +(Still true of the *deployed* stack: `arcane/manifests/home-production.json` +deploys no llama.cpp or vLLM service. Both now appear in benchmark scripts +under `analysis/ghidra/benchmarks/` — measured, not depended on. See the +status banner above.) + ## Method Every claim below was tested directly — a real container, a real (small) @@ -33,7 +58,10 @@ Test model: `Qwen/Qwen2.5-0.5B-Instruct-GGUF` (Q4_K_M, 630M params) for generation tests; `nomic-ai/nomic-embed-text-v1.5-GGUF` (Q4_K_M) for the embeddings test. Images: `ghcr.io/ggml-org/llama.cpp:server`/`:full`, `ghcr.io/mostlygeek/llama-swap:cpu`, `ollama/ollama:0.32.0` (this repo's own -pinned version), `vllm/vllm-openai:latest` (v0.26.0). +pinned version *at the time*; production has since moved to +`ollama/ollama:0.32.13` in `analysis/ghidra/docker-compose.ghidra.yml` and +`analysis/ghidra/models/approved-models.json`), `vllm/vllm-openai:latest` +(v0.26.0). ## 1. Structured output enforcement @@ -87,7 +115,7 @@ llama.cpp/vLLM lose on — Ollama's digest happens to be *exactly* what equivalent. A migration would replace "query the running server's registry digest" with "hash the GGUF file on disk directly" — simpler and arguably more auditable (no trust in a registry's own digest computation), but a real -rewrite of `collect_snapshot()`/`compare_against_approved()`'s identity model, +rewrite of `collect_snapshot()`/`evaluate_drift()`'s identity model, and every recorded `approved-models.json` entry's `ollama_repo_digest`-shaped field. @@ -133,6 +161,21 @@ instances with strict, static VRAM partitions instead of dynamic sharing). `temperature: 0, seed: 66` is relied on for reproducible output today. +> **Revised after the fact (#2646).** That reliance did not survive contact +> with the deployed configuration. `OLLAMA_KEEP_ALIVE=30m` keeps the weights +> resident between samples, and a **warm** Ollama slot returns different text +> for a byte-identical prompt at temperature 0 with a fixed seed — production +> therefore runs permanently in the drifting regime, and no setting fixes that +> without paying a reload. The fix that shipped is accounting, not determinism: +> every stored assessment now records a `slot_generation` fingerprint (read +> from Ollama's `/api/ps`) naming the resident instance that answered, so two +> results are never compared as though one instance produced both — see +> `docs/analysis/ghidra/AI_TRIAGE.md`. This is attributed to a warm **Ollama** +> slot specifically; the llama.cpp result below is not stated to have +> exercised an equivalent long-lived warm-slot regime, so it stands as +> measured. What the finding removes is this section's *premise* — that the +> deployed stack is reproducible today — not its per-engine results. + **llama.cpp**: confirmed bit-identical output across 3 repeated identical requests (`temperature: 0, seed: 66`) against the same model. No loss. @@ -175,14 +218,18 @@ prompt-processing (`pp`) and text-generation (`tg`) tokens/sec numbers. (`qwen3:14b`, promoted to all three slots under #568/#569 now that VRAM is confirmed ~20GB) on the real GPU, and compare against Ollama's own measured throughput for the same model (`docs/local-llm-model-evaluation.md` already -has some of these numbers). +has some of these numbers). **Partly done, on a different model** — the later +three-engine run in `analysis/ghidra/benchmarks/engine-benchmark/README.md` +did measure decode throughput on the real 20GB card (llama.cpp ~21.4, Ollama +~21.5, vLLM ~23.1 tok/s), but against the merged REx86 f16 7B weights, not +`qwen3:14b`. The `qwen3:14b`-specific comparison is still outstanding. ## 7. Operational surface | | Ollama 0.32.0 | llama.cpp `:server` | llama-swap `:cpu` | vLLM 0.26.0 | |---|---:|---:|---:|---:| | Image size | 8.06 GB | 1.21 GB | 1.24 GB | 28.1 GB | -| Runs without a GPU present | Yes (CPU fallback) | Yes (CPU fallback) | Yes | **No — confirmed** | +| Runs without a GPU present | Yes (CPU fallback) | Yes (CPU fallback) | Yes | **No — confirmed**, but hardware-specific; see below | The vLLM finding is real and concrete, not inferred: the standard `vllm/vllm-openai` image fails to even construct its own CLI argument parser @@ -201,6 +248,20 @@ less-maintained artifact, not the mainline image this stack would pull). Ollama and llama.cpp both degrade gracefully to CPU; vLLM does not degrade, it refuses to start. +> **Superseded.** That failure was hardware-specific, not a property of the +> image. On the current 20GB RTX 4000 Ada Generation, `vllm/vllm-openai:latest` +> (v0.26.0) starts cleanly, loads a 7B model, completes `torch.compile` +> warmup and serves correct completions — recorded in the "Side finding" +> section of `analysis/ghidra/benchmarks/engine-benchmark/README.md`. The +> failure above was observed on the *previous* 8GB Quadro RTX 4000. +> +> Whether the mainline image has any CPU-only path at all is therefore +> **undetermined** from this repo's evidence: the two data points are one +> failure on the old card and one success on the new one, and neither isolates +> GPU *presence* from GPU *capability*. The "no CPU path at all" reading above +> overstates what was shown. Unchanged by the later run: vLLM is by far the +> largest image of the four, and §2's digest/registry finding stands. + **Embeddings** (relevant to `ml-worker`/#151's planned semantic search, both currently spec'd around `nomic-embed-text`): llama.cpp has a native `--embedding` server mode with an OpenAI-compatible `/v1/embeddings` @@ -231,7 +292,16 @@ contract) replaced Ollama: `/v1/chat/completions`'s `response_format` shape instead of `/api/chat`'s `format`; digest verification (`model_digest()`) rewritten to hash the local GGUF file instead of querying `/api/tags`. -- `analysis/ghidra/worker/ghidra-worker.py` — same client-shape change. +- `analysis/ghidra/worker/ghidra-worker.py` — **no client-shape change needed**, + as of the current code. It already speaks an OpenAI-compatible `/v1` chat + dialect (`GHIDRA_TRIAGE_API_BASE`, default `http://127.0.0.1:11434/v1`) and + does so deliberately: the worker records that this is what makes "llama.cpp's + server, vLLM and LM Studio work unchanged; the requirement is that it is + *local*, not that it is Ollama." The one Ollama-specific coupling left is the + `/api/ps` probe behind `GHIDRA_TRIAGE_RUNTIME_BASE` that populates + `slot_generation` (#2646); the code already handles a server that is not + Ollama, recording `unavailable`. This line was a real cost when this doc was + written and is no longer one. - `analysis/ghidra/models/model-governance.py` — the largest rewrite. Its entire `collect_snapshot()`/drift-comparison model is built around Ollama's registry APIs and Docker-image-reference identity; every field in @@ -290,6 +360,10 @@ serving — not the current shared, bursty, multi-workload shape. - Run `llama-bench` against the currently-approved model (`qwen3:14b`, promoted under #568/#569) on the real card, side-by-side with Ollama's own numbers, to get the actual performance data this research couldn't gather. + (Partly superseded: the three-engine run in + `analysis/ghidra/benchmarks/engine-benchmark/README.md` gathered real-card + throughput for llama.cpp/Ollama/vLLM, but on REx86 f16 7B weights. The + `qwen3:14b` case is still open.) - Resolve the 768-vs-384 embedding-dimension discrepancy (§7) before #151's `dense_vector` ES mapping work begins, regardless of backend choice. - If a future re-evaluation is triggered (e.g. `model-governance.py` gets a diff --git a/docs/persona-design.md b/docs/persona-design.md index 4ada313c4..dce39ce9b 100644 --- a/docs/persona-design.md +++ b/docs/persona-design.md @@ -5,8 +5,12 @@ Two related pieces of persona design that were implicit rather than documented decisions: whether a honeypot may reach the internet outbound, and how to name/place a honeypot host so it doesn't look staged. T-Pot's -own README calls both out by name (`README.md` line 296 for outbound; the -"where to place a honeypot" guidance for siting). This repo already does +own upstream README calls both out by name (its `README.md` line 296 for +outbound; the "where to place a honeypot" guidance for siting). That line +number refers to T-Pot's own repository, not to this repo's `README.md`, which +is ~135 lines; it was recorded from an unpinned upstream read and could not be +re-verified during this reconciliation, so treat the line number as a pointer +to look up rather than a stable citation. This repo already does deep, source-verified realism work for Windows personas ([#91](https://github.com/Xore/APIARY/issues/91)/[#94](https://github.com/Xore/APIARY/issues/94)/[#96](https://github.com/Xore/APIARY/issues/96)) and has a full fictional-organization inventory @@ -41,7 +45,8 @@ The tradeoff is real in both directions: | Cowrie | Allowed (flag: `COWRIE_AIR_GAPPED`, default `false`) | The one sensor in this stack designed around capturing attacker-fetched malware — its whole SSH/Telnet fake-shell premise is attackers running `wget`/`curl`/`tftp` against real URLs. See `arcane/home/honeypot-cowrie/compose.yml`'s `cowrie_net`. | | Dionaea | Allowed (flag: `DIONAEA_AIR_GAPPED`, default `false`) | Same tradeoff as Cowrie (#269/#538): captures shellcode/binaries pushed *to* it over SMB/FTP/TFTP/etc, which both ship enabled by default — this is the attacker's actual malware sample, not just the exploit attempt. See `arcane/home/honeypot-dionaea/compose.yml`'s `dionaea_net` (#541). `internal: true` still permits `tftp-relay`'s inbound forwarding on the same network — it only removes the outbound route. | | Tanner/Snare | Allowed (flag: `TANNER_AIR_GAPPED`, default `false`) | Same tradeoff as Cowrie and Dionaea: the `template_injection` emulator fetches real RFI payloads when enabled, capturing the attacker's actual payload instead of just the RFI attempt. See `arcane/home/honeypot-tanner/compose.yml`'s `tanner_local`. Setting the flag also breaks the emulator's own `REMOTE_DOCKERFILE` self-maintenance fetch (`raw.githubusercontent.com`) — a real cost, not just a capture-vs-safety tradeoff. | -| Everything else (Conpot personas, DNP3, HTTP/API honeypot, multipot, dicompot, dns-honeypot, citrix-honeypot, cisco-asa-honeypot, rdp-honeypot) | Allowed, no design reason either way | None of these protocols involve the honeypot fetching attacker-supplied URLs — outbound access is unused in the intended interaction, just never explicitly closed off. An operator who wants maximum containment can set `internal: true` directly on that sensor's network in its compose file without losing anything these honeypots actually rely on. | +| Canarytokens | Allowed (`canarytokens_net`; no air-gap flag exists) | **A different deception category, not a capture-vs-safety call.** This stack is a honeytoken *platform* — it plants fake documents, credentials and DNS names that alert when touched — not a sensor that fetches attacker-supplied content. Its reachability is the product: #1487 made the switchboard's HTTP channel publicly reachable through the VPS precisely so dashboard-created file/doc tokens fire when opened outside our own network, and that stack's compose file is explicit that "inert (internal-only) tokens don't serve it." The tradeoff argued above therefore does not transfer to it, in either direction. Note the asymmetry with the `internal: true` rule below: `internal` removes the *outbound* route, so it would not by itself break the VPS's inbound bridge — whether it is safe on `canarytokens_net` is undecided in-repo. Don't assume either way. | +| Everything else (Conpot personas, DNP3, HTTP/API honeypot, multipot, dicompot, dns-honeypot, citrix-honeypot, cisco-asa-honeypot, rdp-honeypot, sonicwall-sma, endlessh, beelzebub, hellpot, elasticpot, galah, sentrypeer, mailoney) | Allowed, no design reason either way | None of these protocols involve the honeypot fetching attacker-supplied URLs — outbound access is unused in the intended interaction, just never explicitly closed off. An operator who wants maximum containment can set `internal: true` directly on that sensor's network in its compose file without losing anything these honeypots actually rely on. | | `yara-scanner` (not a honeypot — offline payload analysis) | Blocked (`network_mode: none`) | Already air-gapped; scans captured files at rest, never needs network access at all. The one existing precedent this decision extends. | `COWRIE_AIR_GAPPED=true` (`.env`) sets `internal: true` on `cowrie_net`,