Skip to content

ops: Elasticsearch at 1000/1000 shards, all sensor ingest dead-lettered since 2026-09-21 #3283

Description

@Xore

What happened

Every sensor in honeypot-v2-* stopped indexing at about 2026-09-21 18:19 UTC (cowrie, multipot, cisco-asa, hellpot, sentrypeer, beelzebub, http, endlessh, citrix, rdp, api, dnp3, dns, elasticpot, dicompot, mailoney, galah all show the same last @timestamp). Dionaea incidents are affected the same way.

Cause (observed on the homeserver, 2026-09-24)

  • _cluster/health: 991 active shards plus 9 unassigned, which is the 1000 default of cluster.max_shards_per_node on this single-node cluster. _cluster/settings has no override.
  • The newest dead-letter documents all carry the same error:
    validation_exception: this action would add [1] shards, but this cluster currently has [1000]/[1000] maximum normal shards open.
    Each new daily index or data-stream backing index is refused, so Filebeat routes every event to the dead-letter index.
  • dead-letter-honeypot now holds ~1.73 billion documents, 600 GB. Those documents contain the missing events, so they can be replayed and should not be deleted before the gap is backfilled.
  • hp-filebeat shows a RestartCount of 1401 and logs Failed to index N events in last 10s: tried dead letter index every 10 seconds.
  • Disk is not the cause: 68% used, 2.7 TB free.

Where the shards go

The biggest consumers are many tiny daily indices: about 30 zeek-proxy-v1-<log type>-* and zeek-v1-* families with 11 to 14 days each (several of them hold under 100 docs in total), arkime_sessions3-*, dashboard-backend-v1-* and dionaea-incidents-v1-*.

Fix options

  1. Immediate unblock (reversible): raise the limit, e.g. set cluster.max_shards_per_node to 1500 as a persistent setting. This buys headroom, not a fix.
  2. Real fix: consolidate the zeek per-log-type daily indices (one data stream, or weekly/monthly rollover), add ILM or retention for arkime_sessions3-*, dashboard-backend-v1-* and dashboard-bff-v1-*, and set replicas to 0 on the single node (the 9 unassigned shards).
  3. Backfill: replay the dead-letter documents from 2026-09-21 18:19 onward through the honeypot pipeline, then trim dead-letter-honeypot.
  4. Alerting: the source-health page and the live toasts should flag "all sensors silent" and a growing dead-letter rate; this went unnoticed for three days.

Found incidentally while auditing the dashboard mock-data coverage.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't workingin-progressActively being worked onopsDeployment, runners, observability, host access

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions