Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
17 changes: 16 additions & 1 deletion docs/analysis/RECOVERY.md
Original file line number Diff line number Diff line change
Expand Up @@ -23,7 +23,22 @@ authenticated with an empty event history. See
and the sizes behind it.

Recovery is intentionally not automatic because overwriting live volumes is
destructive. On a replacement host:
destructive.

**Which archive the steps below apply to**: an `analysis/backup-honeypot.sh`
directory — that script's on-host layout of `SHA256SUMS`,
`stack-config-state.tar.gz`, `keycloak.sql.gz` and `volumes/<name>.tar.gz`.
Those names resolve nowhere else, and that archive only exists if the
homeserver itself survived. A restore driven by the archive that actually
survives a dead homeserver — `scripts/backup-essentials.sh`'s
`apiary-essentials-<stamp>.tar.gz.gpg`, a different layout under
`homeserver/`, `vps/` and `repo/` — has its own procedure, and one
prerequisite the steps below do not cover: put
`homeserver/installer/*-install-homeserver.conf` back in place first, because
`scripts/install-homeserver.sh` will not run without it. See
[`docs/BACKUP-ESSENTIALS.md`](../BACKUP-ESSENTIALS.md).

On a replacement host:

1. Verify `SHA256SUMS`, unpack `stack-config-state.tar.gz` into a new empty stack
directory, and inspect `.env` permissions and values.
Expand Down
13 changes: 9 additions & 4 deletions docs/analysis/ghidra/models/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -17,10 +17,15 @@ python3 /opt/honeypot-ghidra/models/model-governance.py check-runtime \

The status file contains only state and reason codes. It contains no prompts, model replies, captured data, container paths, or credentials, and is written owner-only mode `0600`. If a dashboard later needs it, expose only the sanitized object through a privileged read-only endpoint; do not mount or relax the host file. `approved`, `drift`, and `unavailable` are advisory states: the service exits successfully with `--warn-only`, so an LLM problem never stops ingestion or deterministic analysis. Omit `--warn-only` in an operator check when drift should produce a non-zero exit status. The command only reads `/api/version`, `/api/tags`, Docker inspection metadata, and `nvidia-smi` telemetry.

### Two expected, honest states after #2394 (not regressions)
### Post-#2394 GPU-identity states: which are expected, which are not

- **`approved_gpu_uuid_missing`** on the `host` leg: the deployed manifest copy at `/opt/honeypot-ghidra/models/approved-models.json` still predates the `approved_host.gpu_uuid` field until `install-analysis-host.sh` next runs end to end on that host. The checker reports this distinct, advisory-only code rather than silently comparing whichever GPU enumerates as index 0. It clears itself once the host is redeployed with the current manifest.
- **`host_gpu_uuid_changed`**: any `--snapshot` file captured before #2394 was written under the older schema and has no GPU-identity fields for the new comparison to match against. Replaying it will read as drift under the new schema even though nothing on the host changed -- expected for old snapshots, not evidence of an actual UUID change.
#2394 made the checker compare GPU identity by UUID rather than enumeration
order, and that added three host-leg codes. Only the first two are expected
during the rollout; the third is a real problem wearing the same shape.

- **`approved_gpu_uuid_missing`** on the `host` leg — *expected.* The deployed manifest copy at `/opt/honeypot-ghidra/models/approved-models.json` still predates the `approved_host.gpu_uuid` field until `install-analysis-host.sh` next runs end to end on that host. The checker reports this distinct, advisory-only code rather than silently comparing whichever GPU enumerates as index 0. It clears itself once the host is redeployed with the current manifest.
- **`host_gpu_uuid_changed`** — *expected for a legacy `--snapshot` only, and this code is overloaded, so check before dismissing.* A snapshot captured before #2394 was written under the older schema and carries no GPU-identity field, so replaying it reads as drift even though nothing on the host changed. But the checker also emits the **same code** when `gpu_uuid` is present and names a *different physical card* than the manifest pins — that is real drift, not a schema artefact, and the field the old schema compared (name, memory, driver) can all still look correct. Tell the two apart by whether the snapshot has a `gpu_uuid` key at all: absent means a legacy replay, present-and-different means the host is not running the approved card. The same per-field construction applies to `host_gpu_changed`, `host_gpu_memory_mib_changed`, `host_driver_changed` and `host_compute_capability_changed`, so treat each of those the same way.
- **`approved_gpu_absent`** — *not expected.* The tool ran, was pointed at the approved UUID explicitly, and no such card exists on the host. Distinct from `gpu_telemetry_unavailable` (which means the telemetry could not be read at all). This one means the approved card is genuinely gone.

## When requalification is mandatory

Expand All @@ -30,7 +35,7 @@ Run the complete workflow before changing any model tag or digest, Ollama image/

Use a trusted checkout on the approved analysis host. Stop unrelated GPU-heavy jobs if needed, but do not stop or modify QEMU. The benchmark uses only checked-in synthetic TEST-NET fixtures, talks only to the explicitly supplied local Ollama endpoint, records exact artifacts/settings/timing/RAM/VRAM metadata, and unloads each candidate through Ollama after its slot. It never downloads a model.

Keep verbose replies outside the repository in an operator-only directory with bounded retention (30 days is the recorded recommendation):
Keep verbose replies outside the repository in an operator-only directory with bounded retention (30 days is this doc's own recommendation; no separate retention record covers this directory):

```sh
install -d -m 0700 "$HOME/model-qualification"
Expand Down
2 changes: 1 addition & 1 deletion docs/kvm-network-traffic-analysis.md
Original file line number Diff line number Diff line change
Expand Up @@ -27,7 +27,7 @@ are **two** such bridges, because there are two sandboxes:
| Network XML | `sandbox/windows/setup/sandbox-network.xml` | `sandbox/network.xml` |
| Fake internet | INetSim at `10.10.10.1` | none by default; optional logged DNS + Squid allowlist (`controlled` mode) |
| Capture | `docker-compose.sandbox.yml` (tcpdump, Zeek, Suricata) | root-owned `tcpdump` per job, host and guest side |
| Results | `sandbox/results/<run>/` | `/var/lib/honeypot-sandbox/results/` (root-only), sanitized export copied out |
| Results | `$WINDOWS_SANDBOX_RESULTS_DIR/<sample sha256>/` — this is the path the dashboard reads. The compose file's own `./sandbox/results/current` default only applies to a manual `docker compose -f docker-compose.sandbox.yml up`; `run_sample.py` overrides `SANDBOX_RESULTS_DIR` to the current run's out_dir, and its own fallback when the env var is unset is `reports/windows-sandbox` | `/var/lib/honeypot-sandbox/results/` (root-only), sanitized export copied out |
| Orchestrator | `sandbox/windows/orchestrate/run_sample.py` | `sandbox/run-linux-sample.sh` |

Neither bridge has a `<forward>` element, so neither can route anywhere. That
Expand Down
84 changes: 79 additions & 5 deletions docs/llm-inference-backend-comparison.md
Original file line number Diff line number Diff line change
Expand Up @@ -2,6 +2,26 @@

Status: research for [issue #598](https://github.com/Xore/APIARY/issues/598), 2026-08-05.

> **Superseded in part — read this first.** This is the CPU-only research
> pass. The *measured* three-engine comparison it asked for was run
> afterwards and is recorded in
> `analysis/ghidra/benchmarks/engine-benchmark/README.md` (2026-08-06, on the
> RTX 4000 Ada 20GB, extended by #832's settings tuning). It **overturns §6
> and §7**: real-card performance was in fact measured, and vLLM's "refuses to
> start" failure turned out to be specific to the *previous* 8GB Quadro
> RTX 4000 rather than a property of the image. It explicitly leaves §2
> standing ("no digest/registry system", `model-governance.py`'s pipeline being
> "Ollama-shaped"), and it does not re-test §1's structured-output axis or §3's
> keep-alive/swap behaviour — it runs one engine at a time by construction, so
> it says nothing about multi-model sharing. The sections below are left as
> written, as the record of what was known on 2026-08-05; where later work
> contradicts them, the later work wins.
>
> Two later findings bear on this doc and are flagged inline where they
> apply: #2646 (a warm resident slot is *not* reproducible at temperature 0
> with a fixed seed, which revises §4's premise) and `ghidra-worker.py`'s
> deliberate OpenAI-compatible `/v1` client, which revises §8's cost estimate.

This is a task-specific decision record, matching
[`local-llm-model-evaluation.md`](local-llm-model-evaluation.md)'s format for
the model-selection decision it's paired with. It evaluates the inference
Expand All @@ -18,6 +38,11 @@ APIs). No comparison against llama.cpp (the inference engine Ollama itself
wraps) or vLLM (the throughput-oriented alternative) had been done. #598 asked
for one, explicitly allowing "switch, and rearchitect" as a valid outcome.

(Still true of the *deployed* stack: `arcane/manifests/home-production.json`
deploys no llama.cpp or vLLM service. Both now appear in benchmark scripts
under `analysis/ghidra/benchmarks/` — measured, not depended on. See the
status banner above.)

## Method

Every claim below was tested directly — a real container, a real (small)
Expand All @@ -33,7 +58,10 @@ Test model: `Qwen/Qwen2.5-0.5B-Instruct-GGUF` (Q4_K_M, 630M params) for
generation tests; `nomic-ai/nomic-embed-text-v1.5-GGUF` (Q4_K_M) for the
embeddings test. Images: `ghcr.io/ggml-org/llama.cpp:server`/`:full`,
`ghcr.io/mostlygeek/llama-swap:cpu`, `ollama/ollama:0.32.0` (this repo's own
pinned version), `vllm/vllm-openai:latest` (v0.26.0).
pinned version *at the time*; production has since moved to
`ollama/ollama:0.32.13` in `analysis/ghidra/docker-compose.ghidra.yml` and
`analysis/ghidra/models/approved-models.json`), `vllm/vllm-openai:latest`
(v0.26.0).

## 1. Structured output enforcement

Expand Down Expand Up @@ -87,7 +115,7 @@ llama.cpp/vLLM lose on — Ollama's digest happens to be *exactly* what
equivalent. A migration would replace "query the running server's registry
digest" with "hash the GGUF file on disk directly" — simpler and arguably
more auditable (no trust in a registry's own digest computation), but a real
rewrite of `collect_snapshot()`/`compare_against_approved()`'s identity model,
rewrite of `collect_snapshot()`/`evaluate_drift()`'s identity model,
and every recorded `approved-models.json` entry's `ollama_repo_digest`-shaped
field.

Expand Down Expand Up @@ -133,6 +161,21 @@ instances with strict, static VRAM partitions instead of dynamic sharing).

`temperature: 0, seed: 66` is relied on for reproducible output today.

> **Revised after the fact (#2646).** That reliance did not survive contact
> with the deployed configuration. `OLLAMA_KEEP_ALIVE=30m` keeps the weights
> resident between samples, and a **warm** Ollama slot returns different text
> for a byte-identical prompt at temperature 0 with a fixed seed — production
> therefore runs permanently in the drifting regime, and no setting fixes that
> without paying a reload. The fix that shipped is accounting, not determinism:
> every stored assessment now records a `slot_generation` fingerprint (read
> from Ollama's `/api/ps`) naming the resident instance that answered, so two
> results are never compared as though one instance produced both — see
> `docs/analysis/ghidra/AI_TRIAGE.md`. This is attributed to a warm **Ollama**
> slot specifically; the llama.cpp result below is not stated to have
> exercised an equivalent long-lived warm-slot regime, so it stands as
> measured. What the finding removes is this section's *premise* — that the
> deployed stack is reproducible today — not its per-engine results.

**llama.cpp**: confirmed bit-identical output across 3 repeated identical
requests (`temperature: 0, seed: 66`) against the same model. No loss.

Expand Down Expand Up @@ -175,14 +218,18 @@ prompt-processing (`pp`) and text-generation (`tg`) tokens/sec numbers.
(`qwen3:14b`, promoted to all three slots under #568/#569 now that VRAM is
confirmed ~20GB) on the real GPU, and compare against Ollama's own measured
throughput for the same model (`docs/local-llm-model-evaluation.md` already
has some of these numbers).
has some of these numbers). **Partly done, on a different model** — the later
three-engine run in `analysis/ghidra/benchmarks/engine-benchmark/README.md`
did measure decode throughput on the real 20GB card (llama.cpp ~21.4, Ollama
~21.5, vLLM ~23.1 tok/s), but against the merged REx86 f16 7B weights, not
`qwen3:14b`. The `qwen3:14b`-specific comparison is still outstanding.

## 7. Operational surface

| | Ollama 0.32.0 | llama.cpp `:server` | llama-swap `:cpu` | vLLM 0.26.0 |
|---|---:|---:|---:|---:|
| Image size | 8.06 GB | 1.21 GB | 1.24 GB | 28.1 GB |
| Runs without a GPU present | Yes (CPU fallback) | Yes (CPU fallback) | Yes | **No — confirmed** |
| Runs without a GPU present | Yes (CPU fallback) | Yes (CPU fallback) | Yes | **No — confirmed**, but hardware-specific; see below |

The vLLM finding is real and concrete, not inferred: the standard
`vllm/vllm-openai` image fails to even construct its own CLI argument parser
Expand All @@ -201,6 +248,20 @@ less-maintained artifact, not the mainline image this stack would pull).
Ollama and llama.cpp both degrade gracefully to CPU; vLLM does not degrade,
it refuses to start.

> **Superseded.** That failure was hardware-specific, not a property of the
> image. On the current 20GB RTX 4000 Ada Generation, `vllm/vllm-openai:latest`
> (v0.26.0) starts cleanly, loads a 7B model, completes `torch.compile`
> warmup and serves correct completions — recorded in the "Side finding"
> section of `analysis/ghidra/benchmarks/engine-benchmark/README.md`. The
> failure above was observed on the *previous* 8GB Quadro RTX 4000.
>
> Whether the mainline image has any CPU-only path at all is therefore
> **undetermined** from this repo's evidence: the two data points are one
> failure on the old card and one success on the new one, and neither isolates
> GPU *presence* from GPU *capability*. The "no CPU path at all" reading above
> overstates what was shown. Unchanged by the later run: vLLM is by far the
> largest image of the four, and §2's digest/registry finding stands.

**Embeddings** (relevant to `ml-worker`/#151's planned semantic search, both
currently spec'd around `nomic-embed-text`): llama.cpp has a native
`--embedding` server mode with an OpenAI-compatible `/v1/embeddings`
Expand Down Expand Up @@ -231,7 +292,16 @@ contract) replaced Ollama:
`/v1/chat/completions`'s `response_format` shape instead of `/api/chat`'s
`format`; digest verification (`model_digest()`) rewritten to hash the
local GGUF file instead of querying `/api/tags`.
- `analysis/ghidra/worker/ghidra-worker.py` — same client-shape change.
- `analysis/ghidra/worker/ghidra-worker.py` — **no client-shape change needed**,
as of the current code. It already speaks an OpenAI-compatible `/v1` chat
dialect (`GHIDRA_TRIAGE_API_BASE`, default `http://127.0.0.1:11434/v1`) and
does so deliberately: the worker records that this is what makes "llama.cpp's
server, vLLM and LM Studio work unchanged; the requirement is that it is
*local*, not that it is Ollama." The one Ollama-specific coupling left is the
`/api/ps` probe behind `GHIDRA_TRIAGE_RUNTIME_BASE` that populates
`slot_generation` (#2646); the code already handles a server that is not
Ollama, recording `unavailable`. This line was a real cost when this doc was
written and is no longer one.
- `analysis/ghidra/models/model-governance.py` — the largest rewrite. Its
entire `collect_snapshot()`/drift-comparison model is built around Ollama's
registry APIs and Docker-image-reference identity; every field in
Expand Down Expand Up @@ -290,6 +360,10 @@ serving — not the current shared, bursty, multi-workload shape.
- Run `llama-bench` against the currently-approved model (`qwen3:14b`,
promoted under #568/#569) on the real card, side-by-side with Ollama's own
numbers, to get the actual performance data this research couldn't gather.
(Partly superseded: the three-engine run in
`analysis/ghidra/benchmarks/engine-benchmark/README.md` gathered real-card
throughput for llama.cpp/Ollama/vLLM, but on REx86 f16 7B weights. The
`qwen3:14b` case is still open.)
- Resolve the 768-vs-384 embedding-dimension discrepancy (§7) before #151's
`dense_vector` ES mapping work begins, regardless of backend choice.
- If a future re-evaluation is triggered (e.g. `model-governance.py` gets a
Expand Down
Loading
Loading