Skip to content

docs(research): assess OllamaDrama/Ollure against this fleet's LLM-deception surface (#3394) - #3396

Merged
Xore merged 1 commit into
mainfrom
oc/3394-ollure-research
Sep 27, 2026
Merged

Xore merged 1 commit into
mainfrom
oc/3394-ollure-research

Conversation

@Xore

@Xore Xore commented Sep 27, 2026

Copy link
Copy Markdown
Owner

Summary

Research note for #3394. Docs-only: no sensor, compose file, index template, ingest pipeline, Suricata rule or dashboard file is touched. The arcane/home/honeypot-dashboard/frontend-next/ tree is untouched.

The paper verifies. The signature work it asks for cannot be built, because this fleet emulates no Ollama management API.

What checked out

Every claim the issue makes from arXiv:2609.29757 matches the primary source, including the figures most likely to be misquoted later:

  • 84 days, 4 deployments, 290,887 interactions / 2,793 unique IPs
  • "4,148 prompts / 72 unique" — exactly the System and Model Information category, not a corpus total (the corpus total is 4,378 unique prompts on /api/generate)
  • 12,288-token flooding, num_ctx=999999999999999999, num_predict=-1
  • HIVE-AI 16,682 / 1,229 / 20; Shodan ~20k; Xu et al. 152,137 over 362 days
  • CVE-2024-39722 (traversal via /api/push), CVE-2025-63389 (no auth), CVE-2026-85180 (SSRF)

Its Table 1 and Table 2 reconcile against its own published aggregates — the endpoint column sums to 290,887, the model-management endpoints to 3,983 (the paper's "3,983 attempts to execute model modifications"), and the inference endpoints to 56,063 with /api/generate at 94.53%. A transcription that was subtly wrong would not reconcile on all three.

Why the signatures don't build

There is no handler for /api/tags, /api/pull, /api/create, … anywhere in the tree.

  • galah-llm-broker is a two-route reverse proxy to the real Ollama. Its own negative test pins /api/pull and /api/tags on the 404 side.
  • http-honeypot / api-honeypot serve the OpenAI /v1 dialect. The intersection with the paper's surface is exactly one endpoint, /v1/models.
  • beelzebub has no Ollama service, deliberately.
  • The only Ollama in the tree is our own backend on an internal: true network, not internet-facing.

The KQL sketch's field paths don't exist either: event.dataset is in no template (the real discriminator is event.sensor), and the request body is a flattened leaf, which supports prefix and term queries only. Wired up as written it returns zero documents and reads exactly like "no attacks" — the same trap docs/research/2777-litellm-mcp-starlette.md documented for a previous research issue.

What the signature review actually found

Rather than re-assert the regexes "look reasonable", each was executed against the literal values the paper prints. Three results worth carrying forward:

  1. Signature 1's scope is correct and better justified than the issue argues — all 98 URL/path model-name values landed on /api/pull and /api/push, zero on the inference endpoints, so widening to /api/generate would add 52,994 requests of which 9,329 are legitimate cloud-model references. Its defects are encoding-only: case-sensitive scheme (HTTP://169.254.169.254/ matches nothing), and the %2f%2f branch is masked by the separate ^\.\. branch, so a percent-encoded form with no literal prefix also matches nothing.
  2. Signature 3 is English-only and misses the paper's own observed override. Ignore tes précedente instruction… is a MISS — the alternation (prior|previous|all) can't see précédente. It's also scoped to template, where the paper's payloads are in system (begware, Drakonchik), messages (27 begware requests) and an undocumented modelfile (the RCE/XMRig vector).
  3. Signature 4's character class is correct despite not reading the way it looks. bc1[ac-hj-np-z02-9] parses as a | c-h | j-n | p-z | 0 | 2-9 — not j and n-p — and covers 32/32 bech32 chars. The paper's own observed address matches. Worth a comment before someone "tidies" it into a broken range.

Also: signature 5's field path is wrong (options.num_ctx, not num_ctx — and /api/generate's top-level context is a deprecated field with entirely different semantics), and signature 6 is a cross-document correlation transform this stack does not run.

The one live finding

Kept in its own section because it is a different issue: galah already turns an attacker-supplied POST body into a prompt against our real model, and only a SHA-256 of that body reaches Elasticsearch. The broker caps body size (64 KiB) and bounds wall-clock (90s), but does not clamp generation options, so options.num_predict: -1 is an unmetered request against a shared single-load slot. Whether that actually burns the slot is not measured here.

Not fixed in this PR — it is about galah's exposure, not about emulating Ollama, and should be filed separately.

Recommendation

Do not build the signatures as sketched — not "later", as written they target fields that do not exist. A fake Ollama management API is a sensor decision, not a research finding, and belongs under the deception-sensor epic with its own reachability, fidelity and cost case; the field gap (honeypot.body is a flattened leaf) has to be closed first regardless, and the existing Log4Shell honeypot.*-blob processor is the template for doing it without a new mapping. No follow-up issue filed from this PR: that's an operator priority call, and filing it as an implementation task is the manufacturing-work outcome this batch is meant to avoid.

Gates

  • scripts/check-doc-paths-exist.py — pass (123 files, 485 tokens)
  • scripts/check-docs-reachable.py — pass (docs/research/ is an exempt record tree)
  • scripts/check-doc-stale-paths.py — pass
  • python -m pytest tests/docs/ -v — 497 passed, 1 xfailed
  • scripts/check-ai-attribution.py — clean on the commit message and this body

Honest limits

No Elasticsearch was reachable in this environment, so every "would we see it" judgement is read off the templates and the dashboard's own consumers, not a live query. I have not written a number anywhere that needed a measured corpus, and I did not build the honeypot.body wildcard field this would need. The paper's released dataset is not fetchable from here, so §4's hit rates are against the paper's printed examples only. Full list in §9 of the note.

…ception surface (#3394)

The paper checks out on every claim #3394 makes from it, including the
easily-mangled ones (4,148/72 prompts, 12,288-token flooding, HIVE-AI's
16,682/1,229/20, the 20k Shodan and 152,137 Xu et al. figures). Its Table 1
and Table 2 reconcile against its own published aggregates, so the taxonomy
is transcribed rather than approximated.

The signature work it asks for, though, cannot be built as written: this
fleet emulates no Ollama management API. galah-llm-broker is a two-route
proxy to the *real* Ollama (its own negative test pins /api/pull and
/api/tags on the 404 side), the HTTP decoy serves the OpenAI /v1 dialect,
beelzebub has no Ollama service by design, and the only Ollama in the tree
is our own backend on an internal-only network. The intersection with the
paper's emulated surface is one endpoint, /v1/models.

So rather than restate the sketch, measured each of the six signatures
against the literal values the paper prints. Three results worth carrying
forward: signature 1's scope is correct and better justified than the issue
argues (all 98 URL/path model names landed on /api/pull and /api/push, zero
on the inference endpoints) but it is case-sensitive on the scheme; and
signature 3 is English-only, so it misses the paper's own observed French
override, and is scoped to `template` where the paper puts payloads in
`system`, `messages` and an undocumented `modelfile`. Signature 4's
character class is correct despite not reading the way it looks — it parses
as j-n and p-z, not j and n-p — and covers 32/32 bech32 chars.

Also confirmed the KQL sketch's field paths do not exist: `event.dataset` is
in no template (the real discriminator is `event.sensor`), and the body is a
`flattened` leaf, which supports prefix and term queries only. Wired up as
written it would return zero documents and read exactly like "no attacks".

The one live finding is unrelated to signatures and is kept in its own
section rather than folded in: galah already turns an attacker-supplied
POST body into a prompt against our real model, and only a SHA-256 of that
body reaches Elasticsearch. Filed-for-later, not fixed here — it is a
different issue about galah's exposure, not about emulating Ollama.
@github-actions

Copy link
Copy Markdown

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

@Xore
Xore merged commit 8b96bf4 into main Sep 27, 2026
117 of 118 checks passed
@Xore
Xore deleted the oc/3394-ollure-research branch September 27, 2026 10:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant