You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
hp-filebeat crashes every one to two minutes with a Go runtime fatal error:
runtime: program exceeds 10000-thread limit
fatal error: thread exhaustion
Docker restarts it each time. The restart count went from 1401 (2026-09-24 20:26 UTC) to 1774 (2026-09-25 05:38 UTC), about 40 an hour.
Observed (2026-09-25, homeserver)
Between crashes the process is ordinary: 12 threads, 41 open fds, about 250 MB RSS. It runs 70–90 s, then the thread count explodes past 10,000 within the runtime's own check (checkmcount via sysmon/handoffp in the stack). That looks like a burst of goroutines blocked in syscalls, not a slow leak.
The last log line before a crash is often an elasticsearch.deadLetterIndexForPolicy warning.
No OOM kill (OOMKilled=false). The container has no healthcheck and no autoheal label, so these are crashes, not health restarts.
This predates #3283's fix. It was already restarting (count 1401) while the shard limit was the visible problem.
Leads, not yet checked
Which input's harvesters spike: /logs/zeek*/*.log, /logs/enriched/*.json, suricata eve-*.json and the others are all filestream inputs.
Whether writes to the event_data/dead-letter log (logging.event_data) block on disk I/O during the dead-letter storm and pile up threads.
Setting GODEBUG / runtime/debug.SetMaxThreads only hides it. Better: harvester_limit per input, or close.on_state_change.inactive, and checking the file count under /logs/zeek*.
What happens
hp-filebeatcrashes every one to two minutes with a Go runtime fatal error:Docker restarts it each time. The restart count went from 1401 (2026-09-24 20:26 UTC) to 1774 (2026-09-25 05:38 UTC), about 40 an hour.
Observed (2026-09-25, homeserver)
checkmcountviasysmon/handoffpin the stack). That looks like a burst of goroutines blocked in syscalls, not a slow leak.elasticsearch.deadLetterIndexForPolicywarning.OOMKilled=false). The container has no healthcheck and no autoheal label, so these are crashes, not health restarts.honeypot-v2-*in 30 min, with noFailed to indexsince the shard limit was raised (ops: Elasticsearch at 1000/1000 shards, all sensor ingest dead-lettered since 2026-09-21 #3283). So nothing is lost, but events arrive in bursts, and every restart re-scans all inputs.This predates #3283's fix. It was already restarting (count 1401) while the shard limit was the visible problem.
Leads, not yet checked
/logs/zeek*/*.log,/logs/enriched/*.json, suricataeve-*.jsonand the others are all filestream inputs.logging.event_data) block on disk I/O during the dead-letter storm and pile up threads.GODEBUG/runtime/debug.SetMaxThreadsonly hides it. Better:harvester_limitper input, orclose.on_state_change.inactive, and checking the file count under/logs/zeek*.