What happened
Every sensor in honeypot-v2-* stopped indexing at about 2026-09-21 18:19 UTC (cowrie, multipot, cisco-asa, hellpot, sentrypeer, beelzebub, http, endlessh, citrix, rdp, api, dnp3, dns, elasticpot, dicompot, mailoney, galah all show the same last @timestamp). Dionaea incidents are affected the same way.
Cause (observed on the homeserver, 2026-09-24)
_cluster/health: 991 active shards plus 9 unassigned, which is the 1000 default of cluster.max_shards_per_node on this single-node cluster. _cluster/settings has no override.
- The newest dead-letter documents all carry the same error:
validation_exception: this action would add [1] shards, but this cluster currently has [1000]/[1000] maximum normal shards open.
Each new daily index or data-stream backing index is refused, so Filebeat routes every event to the dead-letter index.
dead-letter-honeypot now holds ~1.73 billion documents, 600 GB. Those documents contain the missing events, so they can be replayed and should not be deleted before the gap is backfilled.
hp-filebeat shows a RestartCount of 1401 and logs Failed to index N events in last 10s: tried dead letter index every 10 seconds.
- Disk is not the cause: 68% used, 2.7 TB free.
Where the shards go
The biggest consumers are many tiny daily indices: about 30 zeek-proxy-v1-<log type>-* and zeek-v1-* families with 11 to 14 days each (several of them hold under 100 docs in total), arkime_sessions3-*, dashboard-backend-v1-* and dionaea-incidents-v1-*.
Fix options
- Immediate unblock (reversible): raise the limit, e.g. set
cluster.max_shards_per_node to 1500 as a persistent setting. This buys headroom, not a fix.
- Real fix: consolidate the zeek per-log-type daily indices (one data stream, or weekly/monthly rollover), add ILM or retention for
arkime_sessions3-*, dashboard-backend-v1-* and dashboard-bff-v1-*, and set replicas to 0 on the single node (the 9 unassigned shards).
- Backfill: replay the dead-letter documents from 2026-09-21 18:19 onward through the
honeypot pipeline, then trim dead-letter-honeypot.
- Alerting: the source-health page and the live toasts should flag "all sensors silent" and a growing dead-letter rate; this went unnoticed for three days.
Found incidentally while auditing the dashboard mock-data coverage.
What happened
Every sensor in
honeypot-v2-*stopped indexing at about 2026-09-21 18:19 UTC (cowrie, multipot, cisco-asa, hellpot, sentrypeer, beelzebub, http, endlessh, citrix, rdp, api, dnp3, dns, elasticpot, dicompot, mailoney, galah all show the same last@timestamp). Dionaea incidents are affected the same way.Cause (observed on the homeserver, 2026-09-24)
_cluster/health: 991 active shards plus 9 unassigned, which is the 1000 default ofcluster.max_shards_per_nodeon this single-node cluster._cluster/settingshas no override.validation_exception: this action would add [1] shards, but this cluster currently has [1000]/[1000] maximum normal shards open.Each new daily index or data-stream backing index is refused, so Filebeat routes every event to the dead-letter index.
dead-letter-honeypotnow holds ~1.73 billion documents, 600 GB. Those documents contain the missing events, so they can be replayed and should not be deleted before the gap is backfilled.hp-filebeatshows a RestartCount of 1401 and logsFailed to index N events in last 10s: tried dead letter indexevery 10 seconds.Where the shards go
The biggest consumers are many tiny daily indices: about 30
zeek-proxy-v1-<log type>-*andzeek-v1-*families with 11 to 14 days each (several of them hold under 100 docs in total),arkime_sessions3-*,dashboard-backend-v1-*anddionaea-incidents-v1-*.Fix options
cluster.max_shards_per_nodeto 1500 as a persistent setting. This buys headroom, not a fix.arkime_sessions3-*,dashboard-backend-v1-*anddashboard-bff-v1-*, and set replicas to 0 on the single node (the 9 unassigned shards).honeypotpipeline, then trimdead-letter-honeypot.Found incidentally while auditing the dashboard mock-data coverage.