Personas, the seeded cowrie filesystem, GeoIP enrichment, and how to actually read the data once it's flowing.
Every sensor belongs to a stable fictional organization/site/asset identity in personas/personas.json. Native sensors include these fields in JSON; Filebeat enriches upstream logs that cannot. The dashboard makes persona, site, and asset labels clickable so investigations can span protocols without confusing defender-side identity with the attacker's ASN organization. See personas/README.md for the complete matrix and validator.
The fake shell is a believable in-use NexusAI GPU inference node, not an empty
box — realistic /etc/passwd, inference services, a credential-laden .env,
.bash_history, nginx/cron configs and logs (all fictional). Baked at build time; details in
cowrie/README-fs.md.
The custom dashboard, Arkime, and Elasticsearch all share the same local
MaxMind GeoLite2 MMDB files under analysis/geoip/, kept current by the
hp-geoipupdate container (docker compose -f compose.yml --profile geoip-update up -d geoipupdate, needs MAXMIND_ACCOUNT_ID /
MAXMIND_LICENSE_KEY in the honeypot-init stack's .env, managed by
Arcane). Dashboard lookups happen after
portbridge real-IP correlation, support IPv4/IPv6, and add country, city,
coordinates, accuracy radius, ASN, organization, and cloud/hosting/scanner
classification without sending attacker IPs to an external lookup API.
Two independent geo integrations, same source files:
- Arkime reads
analysis/geoip/GeoLite2-{Country,ASN}.mmdb, mounted read-only into both capture and viewer at/opt/arkime/geoand configured viageoLite2Country/geoLite2ASNin arkime/config.ini. Sessions get country + ASN. (#2713: this used to point at a separate, never-automatedarkime/geo/directory populated by hand from db-ip.com — retired in favor of the same files everything else already uses.) - Elasticsearch enriches every
suricata-*andhoneypot-v2-*event (and the portbridge, zeek, extracted-files, huginn and traefik families) through thegeoip-honeypotingest pipeline (set asindex.default_pipelineon 8 of the init stack's 33 index templates), writing ECSsource.geo/source.as/destination.geowith city-level lat/lon — this is what powers Kibana maps (source.geo.locationis mapped asgeo_point), fromGeoLite2-City.mmdbmounted atanalysis/geoip/ → /usr/share/elasticsearch/config/ingest-geoip, withingest.geoip.downloader.enabled=false(ES's own auto-downloader needs egress the home server doesn't have).
The .mmdb files themselves are not in git; hp-geoipupdate refreshes them
on GEOIPUPDATE_FREQUENCY (hours) automatically once credentials are set —
no manual download/copy step.
Elasticsearch field limits: Suricata's EVE output creates so many dynamic fields that the default 1000-field index cap breaks ingest — every event is 400-rejected and filebeat silently drops it (this bit us on 2026-07-19: Kibana "froze" while eve.json was fine). The
suricata-*template raises the limit to 5000; stats live in their own event-type index. Honeypot source objects use a boundedflattenedmapping.
The one-shot honeypot-kibana-setup job installs normalized honeypot,
Suricata, and dead-letter data views plus the XORE Honeypot — enriched
investigation dashboard. Panels cover recent attacks, OT personas,
commands/credentials, payloads, enriched IDS alerts, and ingest failures.
- Dashboard (built in) →
https://honeypot.<domain>, authenticated via its own native OIDC session against Keycloak directly (no gateway hop, unlike the investigation UIs below -- #1026). Live KPIs, feed-freshness states, per-sensor/protocol counts, top IPs/creds/commands, payload downloads, attack chains, and 7-day/24//64campaign correlation scored across sensors, credentials, ports, IDS alerts, and payload hashes. The interface is organized as an operations console rather than a shortcut collection: task-based navigation separates monitoring, investigation, and the Elasticsearch archive. Four top-level overview tabs group live operations, the threat landscape, attacker behavior, and evidence/campaigns so only one workflow is visible at a time; the selected tab survives live refreshes. Every KPI and ranked value pivots directly into the relevant investigation, while explanatory labels make states and metrics usable without a separate legend. The responsive navigation becomes an accessible menu on narrow displays, and light, dark, and automatic themes remain available. The frontend follows the shared Xore/theme design system (migration guide: MIGRATE-HONEYPOT-STACK.md): a semantic, server-rendered application shell (toolbar, sidebar, main canvas, command bar) styled only by the byte-identical vendoredtheme.css. APIARY does not carry a custom dashboard stylesheet; new selectors are implemented inXore/themeand re-vendored. The theme and Leaflet are served from the dashboard binary rather than a JavaScript CDN. The fixed desktop sidebar becomes a compact rail and then an off-canvas navigation panel on narrow screens, while the 32px application toolbar keeps activity, health, and theme controls (dark, light, system) available across every investigation page. The command bar also routes IP addresses, payload hashes, session IDs, ASNs, HTTP paths, and free-text input directly to the appropriate investigation. Event results use server-side pagination (25 rows by default;per_pageaccepts 25–500). Every longer table and API-fed list initially displays 25 entries, then reveals the next 25 near the end of the page or through an accessible Load 25 more control. This also covers payloads, alerts, sandbox results, commands, source lists, and Elasticsearch dead letters. Payload, event, and attack-source rows are fetched from the server in 25-row chunks, avoiding large initial HTML responses even when an inventory contains thousands of records. Attacker profiles combine network enrichment, behavior aggregates, and a chronological progression view./sessions/<id>provides an oldest-to-newest session replay, and both views add conservative MITRE ATT&CK Enterprise/ICS behavior mappings with links and explicit evidence (behavior context, never actor attribution)./clustersfinds fingerprints, payloads, ASNs, and provider classes shared by multiple source IPs. Campaign rows now explain exactly which cross-sensor, credential, payload, alert, or fingerprint factors produced their correlation score. The navbar alert badge shows unacknowledged alert state, while source health uses neutral metric tiles for feeds, Elasticsearch, Filebeat, and dead letters./api/v1/campaignsexposes the same correlation data. A balanced recent feed prevents one noisy sensor from hiding lower-volume sensors. The portbridge connection log is used only to recover real source IPs; it is not counted as a sensor or displayed as an event. The overview attack map uses the vendored Leaflet 1.9.4 client with a hardcoded OpenStreetMap raster basemap and GeoLite2 City/ASN coordinates from/api/map-points. Attack origins are geographic circles whose physical radius is weighted by event count, so their displayed size changes naturally with zoom. Hover shows IP/city/ASN/provider details and selecting a circle opens every event for that attacker. Live refreshes retain the Leaflet map DOM and update only its GeoJSON layer, preserving pan and zoom. If the map library or tiles fail, the current OpenStreetMap container remains visible with an availability message; there is no local basemap fallback. #2425:HONEYPOT_MAP_TILE_URL/HONEYPOT_MAP_ATTRIBUTIONare gone -- they fed the retired Go dashboard's server-side template; the Leaflet layer hardcodes both values outright (frontend-next/src/components/ OverviewPanels.tsx), so there is no basemap env surface to set anymore. The hourly activity chart also exposes exact counts on hover/focus. The 24-hour KPI compares activity with the preceding 24 hours and labels large changes; source health reports dashboard process uptime and memory (RSS + virtual) as the runtime card on/api/v1/source-health— the Go heap/goroutine figures that card used to show have no Rust equivalent and are gone. Event metadata is directly pivotable: sessions, HASSH/JA3/JA4/User-Agent fingerprints, exact commands and credentials, HTTP paths, IDS signatures and categories, payload hashes, ASNs, organizations, and provider classes all open their related events. Source-health tail counts open normalized events; Elasticsearch source totals open a pre-populated historical query. Management-ready A4 PDF reports are available from the overview, alert center, campaign view, attacker profiles, sessions, and every filtered Event Explorer result./export/report.pdfapplies the same filters as Event Explorer, including source IP/CIDR, ASN, organization/provider, country, sensor, signature/category, payload, session, protocol/port, HTTP path, fingerprint, and time window. Reports contain an executive risk summary, ranked sources and indicators, operational alert state, recommendations, and a bounded representative-evidence appendix. Report downloads require the dashboard's own Keycloak-derivedadminrole because they contain hostile-source telemetry. The overview also ranks fingerprints, ASNs, and provider classes with the same one-click pivots. ASN/provider pivots,/commandssession-aware command analysis, and/payload-analysis/<hash>with bounded hashes, entropy, hex, strings, PE/ELF and script classification, behavior indicators, packing likelihood, extracted URL/domain/IP indicators, real YARA rule matches from a networkless, read-only scanner, risk scoring, and Base64/hex/URL/PowerShell UTF-16 decoding. Payload inventory scans run in the background and refresh every two minutes, so walking large capture volumes cannot block/payloads; source filters and duplicate provenance are preserved. Download event rows link directly to the matching static report./historyadds Elasticsearch search/export;/source-healthshows Filebeat/Elasticsearch diagnostics. Alerts have persistent cooldown/acknowledgment, live refresh uses SSE, and events pivot directly to Kibana, EveBox, Arkime, and VirusTotal. Event tables support keyboard-accessible sorting, selectable columns, and an expandable normalized-row JSON view; live events on investigation pages raise a transient notification. Browser API contracts live inarcane/home/honeypot-dashboard/frontend-nextas strict TypeScript and compile to the committed, dependency-free production bundle, so Node.js is only a development tool and never part of the container. - Operational APIs —
/metricsexposes Prometheus text metrics for event, sensor, ingestion, Filebeat, runtime, dead-letter, and YARA health./dead-lettersinvestigates rejected Elasticsearch documents, and durable campaign/cluster snapshots indashboard-intelligence-archive-v1are readable through the generic index-store route/api/v1/store/intelligence(there is no dedicated intelligence route in the Rust router). Alert acknowledgements and captured-malware downloads require the dashboard's own Keycloak-derivedadminrole. - Safe payload triage —
yara-scannerinventories all mounted Dionaea, Cowrie, and script captures without network access or execution. Its results enrich/payloads, static-analysis reports, risk scores, alerts, and health. - Backups — run
sudo analysis/backup-honeypot.sh; Elasticsearch uses its snapshot API and other named volumes are archived separately. Test and restore procedures are indocs/analysis/RECOVERY.md. - Kibana saved objects (dashboards, visualizations, data views you build
by hand) live only in Elasticsearch's Kibana saved-objects index — on this
stack's Kibana 9.5.3 that is
.kibana_<n>, not a bare.kibana— so an ES reset, migration, or upgrade loses them with no recovery path unless you've exported first. Runanalysis/kibana-export.shbefore any ES-affecting change (matchingKIBANA_URLto how you reach Kibana — defaults tohttp://kibana:5601, the in-cluster address); restore withanalysis/kibana-import.sh. Export first, the same way you'd back up anything else you'd be upset to lose —backup-honeypot.shabove doesn't cover these, only the raw Elasticsearch data. - Hard-isolated sandbox — the optional
sandbox/host setup installs KVM/libvirt beside Docker, defines a non-forwarding network, disables libvirt's default NAT network, and uses disposable qcow2 overlays. It never mounts the Docker socket, host folders, or payloads into a running guest. The root-only hash resolver and serial systemd queue accept existing captures, deduplicate them by SHA-256, enforce guest/host deadlines, and export only bounded escaped JSON summaries to the dashboard's/sandboxinvestigation page. The view includes queue health, search, risk, timeout/duration, static versus dynamic evidence, ATT&CK behavior, Windows PE forensics, DNS query/response evidence, host and guest packet summaries, and sanitized JSON export. Content-based routing distinguishes PE executables, DLLs, VBS, JScript, batch, PowerShell, shell, Python, PHP, Node.js, ELF, documents, archives, and unknown data before selecting a static-only or type-specific detonation path. Optional headless Wine execution and allowlisted real-DNS/HTTP(S) forensic retrieval remain inside a freshly recreated VM and a host-enforced proxy boundary. Authenticated administrators can queue an existing captured hash with the payload Analyze button through a narrow host-owned request spool. Bounded host/guest PCAPs can be downloaded by administrators for Wireshark; complete traces, oversize captures, and direct queue control remain outside Docker. A libvirt NIC filter prevents MAC spoofing and unwanted L2 traffic. - analyze.py — summarize the bind-mounted logs directly on the home server:
python3 analysis/analyze.py /opt/stacks/apiary/logs --top 20
- Kibana →
https://kibana.<domain>(Keycloak via the oauth2-proxy gateway). Data views already exist:honeypot-v2-*andsuricata-*(time field@timestamp) plusdead-letter-honeypot*, alongside Arkime Sessions (arkime_sessions3-*, time fieldlastPacket). All suricata and honeypot events carrysource.geo/source.as— build maps onsource.geo.location. Arkime sessions have country + ASN only (GeoLite2 Country has no coordinates). - Arkime →
http://<HP_BIND>:19080— full-packet session search over everything Suricata captured on the VPS. - TANNER dashboard →
https://tanner.<domain>(Keycloak via the oauth2-proxy gateway) — web-attack analysis. - Dionaea/Conpot write their own JSON into the shared volume for jq/ELK; the live dashboard ingests them alongside Cowrie, multipot, HTTP, and Suricata.
apiary-backend (the Rust tier behind every /api/v1 route) answers two
different questions on two different endpoints. They used to be one endpoint
with a constant answer: /healthz returned {"ok": true, "es": <bool>},
where ok was literally hardcoded to true, so a backend that could not
reach Elasticsearch at all still told its own healthcheck, and anything else
that probed it, that it was healthy. The #3283 ingest outage would not have
appeared there. Split into:
| Endpoint | Question | Touches Elasticsearch | Non-200 |
|---|---|---|---|
/livez |
Is the process up and serving? | No | never |
/healthz |
Same handler as /livez, historical name |
No | never |
/readyz |
Can it actually do its job? | Yes (2 probes) | 503 + reason |
$ curl -s http://backend-service:8081/livez
{"live":true,"built":"2026-09-27T01:46:34+00:00","revision":"3dca4457f1b2c0d4e5a69788796a5b4c3d2e1f0ab"}
$ curl -s http://backend-service:8081/readyz
{"ready":true,"cluster":"green","write_blocked":[]}
$ curl -s -o /dev/null -w '%{http_code}\n' http://backend-service:8081/readyz # during an ES outage
503/readyz deliberately carries no revision: readiness is about whether
this process can do its job, which is a property of the running container and
not of a build, and the two questions have different answers during a rollback.
The revision belongs to the liveness body, which is the one every probe and
every ops script already reads.
/livez is what the container HEALTHCHECK curls, and it must stay that
way. A probe that can block on Elasticsearch converts that dependency's
outage into a restart loop of a container that was never the problem. The
image's own comment on the HEALTHCHECK line says this too. /healthz is
kept as an alias because the port-test harness
(arcane/home/honeypot-dashboard/port-tests/lib.sh) and the ops scripts
already use that name; it is a rename-with-a-twist, not a rename.
- Reachable —
GET /_cluster/healthanswers. This is the check the oldes: <bool>field gestured at; the difference is that failing it now produces a 503. - Not red — red means unassigned primaries, against which reads and writes both fail. Yellow stays ready: it means unassigned replicas, which is the ordinary shape of a replicated cluster during a rolling restart, and gating on it would mark the backend not-ready on every deploy.
- Writable — no
index.blocks.writeset on any of the index families this tier writes to (the list ises::WRITE_TARGET_FAMILIESinbackend-service/src/es.rs,dashboard-*plus the bundled worker loops' own families). This is what catches the flood-stage disk watermark, which sets the block on every index at once, and an operator'sPUT /<index>/_block/write.
The response names the blocked indices rather than counting them, and the
reason string points at _cat/allocation because the endpoint cannot
tell those two causes apart — an honest limit, stated rather than papered
over. A missing index is not a block: every dashboard-owned index is created
lazily on its first write, so a fresh cluster correctly reads ready.
Both probes run concurrently under a 5s deadline
(READINESS_TIMEOUT), shorter than the shared client's 30s, because a probe
that blocks for 30s is indistinguishable from the outage it exists to report.
- Docker/compose healthcheck, Traefik, uptime pings →
/livez(or/healthz). Never/readyz; see above. - Deploy verification, diagnostics, "is ingest actually working?" →
/readyz, and treat 503 as a real answer, not a transport error. This is the endpoint that would have shown#3283. - "Why is the dashboard empty?" →
/api/v1/source-health(the per-sensor freshness page)./readyzsays the backend cannot write; only source-health says whether events are arriving. - "Which commit is actually deployed?" →
revisionon/livez, plus the image label. See below.
The failure this exists for is the one docs/ARCANE-GIT-SYNC.md names
directly: a content-change redeploy leaves the dashboard "green, healthy,
running the old code." A fresh container id, a passing healthcheck and a
recent image build all say the machinery ran — none of them say this code is
what is running. So both dashboard images carry the git revision they were
built from, and scripts/verify-deploy.sh compares it against what you
expect.
The contract, in one table. Four surfaces, one value:
| Surface | Where | Read it with |
|---|---|---|
| Backend health body | {"live":…,"built":…,"revision":…} on /livez and /healthz |
curl -s …/healthz |
| Backend boot log | revision=… on the apiary-backend listening line |
docker logs … | grep listening |
| Frontend static file | /build.json → {"built":…,"revision":…,"source":…} |
curl -s …/build.json |
| Both images | org.opencontainers.image.revision (with .source, .created) |
docker image inspect --format '{{ index .Config.Labels "org.opencontainers.image.revision" }}' apiary-backend:latest |
The health body and build.json are small enough to read unfiltered, so
there is nothing to parse here and nothing that needs jq — which is
deliberately absent from these hosts (verify-deploy.sh uses python3 for
the same reason).
Where the value comes from, and what "unknown" means. GIT_SHA is a
Docker build arg, declared in both Dockerfiles and passed by the stacks'
compose files as build.args.GIT_SHA: ${GIT_SHA:-}. The backend compiles it
in (build.rs re-exports GIT_SHA as APIARY_GIT_SHA); the frontend writes
it into public/build.json before npm run build. Unset is normal and not
an error — you get the literal string unknown, never a fabricated revision,
because a plausible-looking wrong answer is worse than no answer. Both tiers
normalize to a bare lowercase hex object name (optional sha256: prefix
stripped, 7–64 hex chars) or unknown, and both are held to one shared table,
backend-service/src/revision-corpus.json, so the two normalizers — different
languages, different CI lanes — cannot drift.
build.rs also emits cargo:rerun-if-env-changed=GIT_SHA. That line is
load-bearing, not hygiene: cargo caches a build script's output by its inputs
and an environment variable is not one of them unless declared, so without it
the second build of an identical tree silently reports the first build's
revision — the exact failure this section exists to end, reproduced by the
stamp itself.
Turning it on. Nothing repo-tracked writes GIT_SHA; the stacks' .env
files are root-owned and provisioned outside this repository, and there is no
.git on the host to ask (deploy.yml rsyncs with --exclude .git/). So
the first run after this landed correctly reports not stamped. From a
clone you are deploying, one line per stack:
echo "GIT_SHA=$(git rev-parse HEAD)" >> /var/dockge/stacks/honeypot-dashboard-backend/.env
echo "GIT_SHA=$(git rev-parse HEAD)" >> /var/dockge/stacks/honeypot-dashboard/.envfollowed by the POST /projects/{id}/build that a content-change sync does
not do. CI (.github/workflows/containers.yml) passes github.sha for the
two Dockerfiles that declare the arg.
Running the check. scripts/verify-deploy.sh compares the live revision,
the image label, and an expected revision — which it takes from origin/main
in a real clone, or from an argument:
# On a host with a clone of apiary, the full check:
scripts/verify-deploy.sh --healthz-exec \
'docker exec hp-apiary-backend curl -sf http://127.0.0.1:8081/healthz' \
--image apiary-backend:latest
# Without one, name the expected revision yourself:
scripts/verify-deploy.sh --behind-days 0 "$(git rev-parse origin/main)" \
--healthz-exec 'docker exec hp-apiary-backend curl -sf http://127.0.0.1:8081/healthz'
# The frontend's form:
scripts/verify-deploy.sh --healthz-url http://host:19090/build.jsonExit codes are the point, and are deliberately three-valued: 0 the
deployed revision matches; 1 a finding (missing, unstamped, malformed, or
stale past --behind-days); 2 could not tell — an unreadable body or an
image that is not on this host. An unreadable answer must never read as a
stale deploy, which is why 2 is not 1. Add --image to check a label against
the running container, which catches the commoner variant where a
compose up recreated from an older image than the one whose label you just
read. The lag check measures the age of the commit, not of the container, so
it survives a container that has been up since before the merge it is missing
— but it needs a clone to measure against, hence --behind-days 0 for the
clone-less form.
diagnostics.yml runs it in --warn-only mode in its own "Deployed
revision" section, against $GITHUB_SHA: the home runner has no clone to
measure against, so the section degrades to "is the running revision the
current tip of main" and says so in the report. Its exit 2 is still
reported as could not tell, not as a pass.
hp-disk-space-monitor (arcane/home/honeypot-utilities/analysis/disk-space-check.sh)
polls df on the bind-mounted host paths in DISK_CHECK_PATHS (default:
honeypot-logs=/logs:honeypot-state=/state:dionaea-payloads=/dionaea-lib)
plus Elasticsearch's own data volume over HTTP (_cat/allocation), and
writes a warning line to /logs/diagnostics/disk-space.json whenever a
checked filesystem's free space drops below DISK_WARN_PERCENT_FREE
(default 15%). Filebeat ships those lines under
event.module:disk-space-check.
#2707: alerts are grouped by physical filesystem, not by bind-mounted
path. honeypot-logs (/logs) and honeypot-state (/state) are both
bind-mounts of the same host directory tree
(/opt/stacks/apiary on the home server), so a low-free-space condition on
that one filesystem used to fire one near-identical WARNING line per bind
mount. The monitor now reads the backing device from df -Pk's first
column, groups every checked path that shares a device into a single
alert, and names the largest of the grouped paths (by du) as
top_contributor so a real spike points straight at a directory instead
of leaving an operator to compare N identical percentages by hand. Shape:
{
"@timestamp": "2026-08-30T12:00:00Z",
"event": {"module": "disk-space-check", "category": "host"},
"disk": {
"source": "/dev/sda1",
"labels": "honeypot-logs,honeypot-state",
"percent_free": 8,
"available_kb": 1234567,
"total_kb": 20000000,
"top_contributor": {"label": "honeypot-logs", "path": "/logs", "used_kb": 15000000}
},
"level": "warning"
}elasticsearch-data alerts (from _cat/allocation, since es-data is a
stack-private volume never bind-mounted cross-stack) are unaffected --
there is only ever one of them per check.
When a benchmark or training leg needs more headroom than the homeserver has,
pause the decoy stacks, run the leg, resume them afterwards. The
classification and the command list both live in
scripts/honeypot-pause.sh — this section is
the reasoning behind the table it carries, and the two measured facts that
make the procedure different from a one-liner.
scripts/honeypot-pause.sh list # the table, with reasons
scripts/honeypot-pause.sh pause honeypot-elasticpot # pause (gate first)
scripts/honeypot-pause.sh status # what is paused, in order
scripts/honeypot-pause.sh resume # reverse order, verifieddocker pause frees CPU, not RAM. The issue says "free RAM/CPU"; only
the CPU half is true. Measured on 2026-09-27, cgroup v2 memory.current for
hp-elasticpot was 36,339,712 B running and 36,360,192 B while paused —
it went up 20 KiB. docker pause freezes the cgroup (the kernel stops
scheduling its tasks) but keeps the memory charged to it. For real RAM you
want docker stop, which is a different and more destructive operation: it
tears down the container, so in-container state and live sessions are gone
unless they are on a volume. This script does not substitute stop for
pause — the issue specifies pause, and silently substituting a more
destructive call is not a call this script gets to make on its own. If a leg
is genuinely RAM-bound, that is a separate decision with its own rollback.
So the honest summary of what a pause buys you: CPU scheduling, and the appearance of a dark decoy. On a CPU-starved leg that is real. On a RAM-starved leg it is close to nothing.
This is the part that will bite you if you improvise. Every decoy on this
host carries the label autoheal=true, and hp-autoheal polls every
AUTOHEAL_INTERVAL=30 seconds. A paused container's healthcheck exec cannot
complete, so Docker marks it unhealthy and autoheal docker restarts it —
which unpauses it. Measured on 2026-09-27, pausing hp-elasticpot at
18:12:21Z produced this in docker logs hp-autoheal:
18:12:38 Container /hp-elasticpot (4c07ab37c137) found to be unhealthy - Restarting container now with 10s timeout
Seventeen seconds later the decoy was running again, docker ps showed a
normal Up, and RestartCount was still 0 — a docker restart is not a
policy restart, so nothing in the container's own bookkeeping says anything
happened. A naive pause is a ~30-second no-op that leaves the operator
believing a decoy is dark when it is not. That is worse than an honest
outage, so pause shuts hp-autoheal down first and resume brings it back
last, after every decoy has been probed healthy again.
Default-deny: a stack not in this table cannot be paused at all, and the refusal prints the reason. The wrong "safe" does not fail loudly — it silently breaks a decoy that was carrying the honeypot surface.
The test applied to each container, in order:
- Is it a decoy? A sensor whose only job is to answer an attacker. A pipeline stage, a datastore, an orchestrator or an operator surface is not a decoy, however much RAM it holds.
- Does it hold a listener or capture path? Pausing a listener does not refuse connections — the kernel keeps the socket and the SYN is queued, so a client hangs until it times out. For a decoy that is the intended dark. For anything in front of a capture path it is silent data loss.
- Does anything depend on it? Shared state, a FUSE mount, an internal datastore. A partial pause hangs a live session instead of ending it cleanly, so these go as a whole stack.
- Is it a worker anything waits on, or an active test's fixture? No.
- Is it a GPU, training or eval leg? Never — those are the reason for pausing, not a target of it.
| Stack | Containers | Verdict | Reason |
|---|---|---|---|
honeypot-elasticpot |
hp-elasticpot |
safe | Standalone Elasticsearch decoy, no dependents, no shared state. Single container, so no in-flight session to strand. |
honeypot-multipot |
hp-multipot |
safe | pop3/imap/socks/dockerv2/tls/adb/redis/elastic decoy, no dependents, no upstream. |
honeypot-dicompot |
hp-dicompot |
safe | Standalone DICOM decoy, no dependents. |
honeypot-dnp3 |
hp-dnp3 |
safe | Standalone DNP3 decoy, no dependents. |
honeypot-sentrypeer |
hp-sentrypeer |
safe | Standalone SIP decoy, no dependents. |
honeypot-hellpot |
hp-hellpot |
safe | Standalone SMTP/HTTP/FTP decoy, no dependents. |
honeypot-endlessh |
hp-endlessh |
safe | Standalone SSH tarpit. Long-lived tarpit sessions are abandoned, which is the intended dark. |
honeypot-rdp-honeypot |
hp-rdp-honeypot |
safe | Standalone RDP decoy, no dependents. |
honeypot-cisco-asa-honeypot |
hp-cisco-asa-honeypot |
safe | Standalone ASA decoy, no dependents. |
honeypot-sonicwall-sma-honeypot |
hp-sonicwall-sma-honeypot |
safe | Standalone SMA decoy, no dependents. |
honeypot-citrix-honeypot |
hp-citrix-honeypot |
safe | Standalone Citrix decoy, no dependents. |
honeypot-mailoney |
hp-mailoney |
safe | Standalone SMTP decoy, no dependents. |
honeypot-beelzebub |
hp-beelzebub |
safe | Standalone adaptive SSH/HTTP/MCP decoy. It hosts an MCP endpoint; pausing it is a deliberate dark, not a defect. |
These have a real dependency on each other. Pausing half of one hangs a live session rather than ending it cleanly, so the script refuses to pause a partial set — it takes the named stacks whole, in dependency order.
| Stack | Containers (pause order) | Reason |
|---|---|---|
honeypot-conpot |
hp-conpot, hp-conpot-guardian, hp-conpot-s7-1500, hp-conpot-s7-1200, hp-conpot-iec104, hp-conpot-kamstrup |
Six OT decoys sharing an image and a project. hp-conpot is the shared modbus/s7 front and the rest are peers on it, so a partial pause leaves listeners up with no engine behind them. |
honeypot-cowrie |
hp-cowrie, hp-honeyfs-implant |
hp-honeyfs-implant is the FUSE server serving cowrie's fake filesystem. Freezing the implant alone hangs every filesystem operation inside a live session. Resume implant first, cowrie second. |
honeypot-dionaea |
hp-dionaea, hp-tftp-relay |
The largest single decoy (~369 MiB, ~14% CPU) and the one most worth pausing. hp-tftp-relay is stack-internal, and dionaea holds an internal MySQL plus live sessions. Never dionaea alone. |
honeypot-galah |
hp-galah, hp-galah-llm-broker |
hp-galah proxies its LLM calls through the broker. Freezing the broker alone makes galah's own decoy paths fail in a way that looks like a broken decoy rather than a stand-down. |
honeypot-canarytokens |
hp-canarytokens-http-router, -frontend, -adapter, -switchboard, -redis |
Internal chain router → adapter → switchboard → redis. A partial pause leaves the switchboard blocked on a frozen redis. |
| Stack | Containers | Reason |
|---|---|---|
honeypot-elk |
hp-elasticsearch, hp-filebeat, hp-zeek-proxy, hp-arkime-capture, hp-arkime-viewer, hp-kibana, hp-evebox, hp-pcap-sync, hp-extracted-file-importer |
The capture pipeline. Freezing arkime-capture/zeek-proxy/pcap-sync stops packet capture at the point of arrival — silent data loss. Freezing hp-elasticsearch stalls every sensor's event write. Not a decoy; it is what makes the decoys worth running. Highest single allocation on the host (~10.6 GiB) and the least safe to interrupt. |
honeypot-dashboard |
hp-dashboard-next, hp-dashboard-oidc-sessions, hp-apiary-worker, -worker-enrichment, -worker-importer, -worker-payload-inventory, hp-apiary-backend-mounted, hp-services-adapter |
The operator surface and the ES write consumers. Freezing the workers buffers Elasticsearch bulk queues, and a frozen dashboard means you cannot observe the stand-down you are performing — which defeats the point of recording it. |
honeypot-dashboard-backend |
hp-apiary-backend |
Auth and API surface. A frozen backend fails every dashboard and CLI call with a hang, not a clean error. |
honeypot-keycloak |
hp-keycloak, hp-keycloak-postgres |
Identity tier. Dashboard, Arcane and the OIDC login tests all authenticate through it. |
honeypot-arcane |
hp-arcane |
GitOps control plane — the container that would re-create a paused decoy on its next sync, so freezing it is arguably necessary. But it is also what reconciles the fleet, and a control plane held frozen across a long leg cannot report or repair drift. Reconcile first, pause decoys, do not freeze the orchestrator. |
honeypot-utilities |
hp-docker-socket-proxy, hp-disk-space-monitor, hp-docker-hygiene, hp-log-maintenance, hp-reporter |
Host watchdogs. Freezing the disk and hygiene monitors during exactly the memory-hungry leg they exist to catch removes the guard against the failure you are creating. hp-docker-socket-proxy is also the socket autoheal drives. |
honeypot-tanner |
hp-tanner, hp-tanner-api, hp-tanner-web, hp-tanner-redis, hp-tanner-docker, hp-tanner-phpox, hp-snare |
Event sink, not a decoy. cowrie and dionaea POST events here, so freezing it makes live sensors error on their own event path, and hp-tanner-redis holds session state in flight. Analysis of collected events is exactly what a leg must not interrupt. |
honeypot-payload-analysis |
hp-yara-scanner, hp-payload-dedupe |
Downstream payload analysis fed from Elasticsearch. Not a decoy and holds no listener, but it is a pipeline stage and freezing it mid-file leaves partial state. |
honeypot-init |
hp-geoipupdate, hp-threat-cidrs-refresh |
Periodic updaters. Freezing hp-geoipupdate mid-write can leave a truncated GeoIP database, which then fails every enrichment silently. |
ml-worker |
hp-ml-worker |
Named in the cold protocol's LIVE_WORKERS. The cold-run mechanism governs this container via STOP_WORKERS=1 + trap; pausing it here would create a second, unreconciled source of truth for whether it is running. |
auth-events-worker |
hp-auth-events-worker |
Consumes Keycloak auth-failure events. Freezing it drops the signal that tier exists to capture. |
ghidra |
ghidra-ollama-1, ghidra-revdeck-1, ghidra-statictools-1, ghidra-ghidra-1 |
Hard prohibition. ghidra-ollama-1 is the GPU slot holder, named explicitly in #3135: never pause it while a benchmark holds the GPU. ghidra-revdeck-1 is in the cold protocol's LIVE_WORKERS. These are also the containers a leg runs to make room for. |
unsloth |
hp-unsloth-studio |
The training leg itself — the reason for pausing, not a target of it. |
rex86-eval |
rex86-eval |
The eval leg itself, same reasoning. |
technitium |
technitium-dns |
Real recursive DNS for 192.168.42.50, not a decoy. A frozen resolver takes the host's name resolution with it. |
pentagi |
graphiti, neo4j, pentagi, pgvector, pentagi-ollama-embedding, pgexporter, scraper |
Unrelated product stack, not part of the honeypot. Out of scope. |
ghosts |
ghosts-ghosts-api-1, ghosts-ghosts-postgres-1 |
Belongs to the sandbox isolation stack that #3312 audits. Standing that down is a separate, declared act (scripts/sandbox-standdown.sh), not a decoy pause. |
dashkcnext-dashkcchaos |
dashkcnext-{pg,kc,redis}-414734, dashkcchaos-{pg,kc,redis}-404812 |
OIDC chaos-test fixtures — these are an active test. "Depended on by an active test" is an explicit not-pause-safe condition in #3135. |
The pentagi-terminal-* containers and the buildkit builder carry no compose
project label and are out of scope for this table; none is a decoy.
Pause
- Confirm nothing in the cold-run protocol is running.
pauserefuses on its own (pgrepoversweep_extra.sh,record_baseline.py,round7_sweep.sh,coldrun.sh,round7_coldrun.sh— the same guardcoldrun.shuses to avoid double-booking the GPU), but the refusal is worth reading rather than working around. scripts/honeypot-pause.sh pause <stack> [stack...]- The script pauses
hp-autohealfirst, then each stack in dependency order, and records every container inAPIARY_PAUSE_DIR/inventory(default/var/lib/apiary/honeypot-pause/) plus an append-only what-and-when line inrecord.lognext to it. scripts/honeypot-pause.sh status— what is paused, in pause order.
Resume
scripts/honeypot-pause.sh resumewalks the inventory in strict reverse order. For each container it unpauses, waits for the healthcheck to return tohealthybefore releasinghp-autoheal(so autoheal cannot restart it mid-verification), then makes one real probe request at the decoy's own port: a completed TCP connect, or an HTTP request that must come back200with a body.- Any container that does not verify keeps the inventory on disk and exits non-zero. Do not treat the leg as clean until they answer.
Verify by hand (the two things the script does, if you are checking):
docker ps --filter status=paused # should be empty when resumed
docker ps --filter name=hp-elasticpot # status, including (healthy)
curl -s -o /dev/null -w '%{http_code}\n' http://10.8.0.2:9201/ # expect 200
docker logs --since 5m hp-autoheal # expect no "found to be unhealthy"A paused decoy hangs rather than refusing: the kernel keeps the listening
socket and queues the SYN, so curl exits 28 (timeout), not 7
(connection refused). Expect the timeout, and do not read it as a decoy that
is wedged differently from how you paused it.
Recorded here rather than only in a PR, because the next person to run this will want to know what has actually been executed against the live host.
Proven on the live homeserver, on the real honeypot-elasticpot stack —
pause → probe stops answering → unpause → probe answers again:
| Step | Observed |
|---|---|
| Baseline | Up (healthy), probe HTTP 200, 339 bytes of decoy ES ("name": "Green Goblin") |
| Naive pause (no gate) | Paused, probe HTTP 000 / curl exit 28; autoheal restored it 17s later — see above |
| Pause with the autoheal gate | Held Paused through 100s (3+ AUTOHEAL_INTERVALs), State.Health=unhealthy, autoheal silent |
| cgroup memory across the pause | 36,339,712 B running → 36,360,192 B paused — no memory released |
| Reverse-order resume | Unpaused, health unhealthy → healthy at t+35s |
| Probe after resume | HTTP 200, 339 bytes, byte-identical to baseline (cmp clean) |
| After autoheal released last | Both healthy, no restart in docker logs hp-autoheal |
The resume path was additionally exercised for real against throwaway
containers (never decoys, never autoheal, never the GPU leg): an inventory of
a,b was walked as b,a; a genuinely paused container was unpaused and
returned to healthy; and a container that could not return to healthy
produced WARN, exit 1, and kept its inventory for retry rather than
reporting a clean leg.
Not yet proven end to end: a full script-driven pause → resume on a
live decoy. The attempt was correctly refused — a benchmark was holding the
GPU slot at the time (record_baseline.py, 8h in, 18.4 GB resident in
ghidra-ollama-1), and the interlock fired:
ABORT: cold-run protocol is active (matched 'record_baseline.py').
That refusal is the interlock working, not a gap in it. The consequence to be
honest about is narrower: the script's own pause path has not yet been run
end to end against a decoy, so the first operator to use it after the GPU slot
frees should expect to watch it once rather than treat it as battle-tested.
The measurements above were taken with the same docker pause/unpause calls
the script makes, on the same host, against the same stack.
Stop-workers-during-cold-run stays governed by STOP_WORKERS=1 and the
restore trap in analysis/ghidra/benchmarks/corpus/sweep_extra.sh. That
mechanism is unchanged by #3135, and honeypot-pause.sh does not stop,
start, or otherwise manage a worker. It only refuses to run while a cold-run
pattern is live, so that ad-hoc pausing cannot become a second, competing
record of which workers are down.