Skip to content

fix(ops): prune zeek rotations on the VPS, where they are writable (#3284, #3283) - #3335

Merged
Xore merged 1 commit into
mainfrom
fix/ops-zeek-harvester-limit
Sep 26, 2026
Merged

Xore merged 1 commit into
mainfrom
fix/ops-zeek-harvester-limit

Conversation

@Xore

@Xore Xore commented Sep 26, 2026

Copy link
Copy Markdown
Owner

Root cause of both open ops issues, and it was invisible from the homeserver.

The catch

/opt/stacks/apiary/logs/zeek is reached over sshfs and is deliberately read-only:

root@10.8.0.1:/opt/stacks/apiary/logs/zeek  /var/dockge/stacks/apiary/logs/zeek  fuse.sshfs ro

(same for suricata, portbridge, huginn). log-maintenance.sh on the homeserver therefore cannot delete these files, and its find carries 2>/dev/null || true, so the failure is silent. Measured live: 12,390 paths printed by the script, file count unchanged, both touch and find -delete returning Read-only file system.

Nothing on the VPS pruned them either — no logrotate rule, no cron, no maintenance script. So 13,717 hourly rotations (16 GB) had accumulated back to 2026-09-03.

How that drove both issues

The fix

  • vps/zeek-log-maintenance.sh (new) — follows the existing suricata-log-maintenance.sh / portbridge-log-maintenance.sh pattern and the Unify log/ILM/artifact retention behind a single top-level env var #261 HONEYPOT_RETENTION_DAYS ratio, and runs where the files are actually written. Deployed and verified: 13,717 → 1,874 files, live generations untouched, wired to @reboot plus a running instance.
  • filebeat.yml — harvester_limit: 64 and ignore_older: 72h on both zeek inputs, so a restart cannot re-open thousands of harvesters even before the first prune pass lands. These are the two settings Filebeat offers for exactly this; GODEBUG/SetMaxThreads would only hide the symptom.
  • log-maintenance.sh — keeps its /logs/zeek-proxy line (zeek-proxy logs are never pruned despite compose claiming 'the pruner owns deletion' (~400MB/day growth) #2323; that directory is local and writable, 834 → 822 files confirmed) and documents why no /logs/zeek line can work there, so nobody re-adds a find that fails quietly.

Verified live

  • hp-filebeat: RestartCount 0, zero thread-exhaustion errors after 5 minutes, against a 70–90s crash cycle before.
  • Filebeat config validated with filebeat test config → Config OK before restart. Backups kept as *.bak-3284.

Not done here

#3283's consolidation of the existing 836 indices is unchanged — this stops new ones accumulating. Replicas=0 on the single node (12 unassigned shards) and the dead-letter backfill from 2026-09-21 18:19 are also still open. Stopping new growth is a prerequisite for those, not a substitute.

Refs #3283, #3284

…3284, #3283)

Root cause of both open ops issues, and it was invisible from the
homeserver.

/opt/stacks/apiary/logs/zeek is reached over sshfs and is deliberately
read-only: root@10.8.0.1:/opt/stacks/apiary/logs/zeek ->
/var/dockge/stacks/apiary/logs/zeek, fuse.sshfs ro in /etc/fstab (same
for suricata, portbridge, huginn). log-maintenance.sh on the homeserver
therefore cannot delete these files at all, and its find carries
2>/dev/null || true, so the failure is invisible: measured live, 12,390
paths printed by the script, file count unchanged, both touch and
find -delete returning EROFS. Nothing on the VPS pruned them either --
no logrotate rule, no cron, no maintenance script -- so 13,717 hourly
rotations (16 GB) had accumulated back to 2026-09-03.

That drove both issues:
- #3284: filestream opens a harvester per file, so each hp-filebeat
  restart created ~13,700 goroutines until the Go runtime hit its
  10000-thread limit and the process died. Roughly every 70-90s.
- #3283: one ES index per day per log type, 836 of the cluster's 1091
  shards, several holding a single document.

vps/zeek-log-maintenance.sh follows the existing
suricata-log-maintenance.sh / portbridge-log-maintenance.sh pattern and
the #261 HONEYPOT_RETENTION_DAYS ratio, and runs where the files are
actually written. Deployed and verified: 13,717 -> 1,874 files, live
generations untouched, wired to @reboot plus a running instance.

log-maintenance.sh keeps its /logs/zeek-proxy line (#2323, that
directory is local and writable -- 834 -> 822 files confirmed) and now
documents why no /logs/zeek line can work there, so the next person
does not re-add a find that fails quietly.

filebeat.yml also gains harvester_limit 64 and ignore_older 72h on both
zeek inputs, so a restart cannot re-open thousands of harvesters again
even before the first prune pass lands. Verified live: RestartCount 0
and zero thread-exhaustion errors after 5 minutes, against a 70-90s
crash cycle before.

Refs #3283, #3284
@Xore
Xore force-pushed the fix/ops-zeek-harvester-limit branch from fcb3a6c to 11e8bb5 Compare September 26, 2026 11:32
@github-actions

Copy link
Copy Markdown

Dependency Review

✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.

Scanned Files

None

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant