Skip to content

research: OllamaDrama/Ollure (arXiv 2609.29757) — empirical attack classes against emulated Ollama management APIs, LLM-deception surface #3394

Description

@Xore

STATUS 2026-09-28 — Implementation PRs are OPEN and awaiting Niklas' merge. No agent merge authority.

Merge authorization: agents open PRs, Niklas merges

Decision (Niklas, 2026-09-28): NO agent merge authority. Open the PR and report; Niklas merges and closes the issue.

This supersedes the 2026-08-23 standing authorization (gh pr merge --squash --delete-branch once green).
Never merge, never delete a branch, and never close this issue while its implementation merely sits in an open PR.

See the issue comments for per-PR state (branch, commit, test counts).


ORIGINAL ISSUE BODY

STATUS 2026-09-28 — §5 ANSWERED: public :8888 HAS been probed with /v1/models

4 hits, 2 distinct source IPs, all 2026-09. Read-only agg, no reindex, no writes, no live traffic. Full numbers and method caveats in the analysis below; raw evidence in the issue comment.


§5 ANSWERED — empirical result (2026-09-28)

The read-only aggregation Niklas authorised has been run. Yes — public :8888 has been probed with /v1/models.

query: event.sensor:api-honeypot  AND  url.path:"/v1/models"
total hits: 4   (all 2026-09)

  2026-09-12T02:44:50Z  71.6.237.176  /v1/models
  2026-09-11T18:23:10Z  71.6.237.176  /v1/models
  2026-09-11T19:03:58Z  71.6.243.202  /v1/models
  2026-09-11T16:40:53Z  71.6.243.202  /v1/models

Two distinct source IPs, both hitting the OpenAI-compatible model-list path. This is exactly the class of behaviour the Ollure paper reports, and it is small — two IPs, four requests — which is the honest scale of the signal.

Adjacent surface, same sensor (context for the classifier)

  /v1/models             4     <- the question, answered
  /api/tags              1     <- ollama native list
  /v1/chat/completions   0
  /api/chat              0
  /api/v1/models         0

/api/tags being hit once while /api/chat never is notable: scanners are enumerating, not conversing. The Ollama-native list endpoint is being probed almost as rarely as the OpenAI-compatible one.

Method / caveats

  • Read-only _search with size:0 + a date_histogram. No reindex, no writes, no traffic sent at any endpoint.
  • The field is url.path, not event.request.path. An initial probe against the wrong field returned 0 and would have produced a false "never probed" conclusion. event.sensor is correct — it matches 11,957 api-honeypot events.
  • Index set spans 1034 shards; the agg completed on all of them.
  • Whether those 4 requests were answered as the decoy intends (200 with a fake model list vs 401) is not established by this agg — the _source returned no persona/site field for these hits. Follow-up from the event detail, not from this query.

Decision on record (Niklas, 2026-09-28)

Run the read-only aggregation. Scope limited to the read-only agg; no reindex, no writes, no live traffic. The galah prompt-injection follow-on is filed separately as #3448.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    honeypotHoneypot sensor or deception technologyresearchResearch findings or reports

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions