fix(ops): prune zeek rotations on the VPS, where they are writable (#3284, #3283) - #3335
Merged
Merged
Conversation
…3284, #3283) Root cause of both open ops issues, and it was invisible from the homeserver. /opt/stacks/apiary/logs/zeek is reached over sshfs and is deliberately read-only: root@10.8.0.1:/opt/stacks/apiary/logs/zeek -> /var/dockge/stacks/apiary/logs/zeek, fuse.sshfs ro in /etc/fstab (same for suricata, portbridge, huginn). log-maintenance.sh on the homeserver therefore cannot delete these files at all, and its find carries 2>/dev/null || true, so the failure is invisible: measured live, 12,390 paths printed by the script, file count unchanged, both touch and find -delete returning EROFS. Nothing on the VPS pruned them either -- no logrotate rule, no cron, no maintenance script -- so 13,717 hourly rotations (16 GB) had accumulated back to 2026-09-03. That drove both issues: - #3284: filestream opens a harvester per file, so each hp-filebeat restart created ~13,700 goroutines until the Go runtime hit its 10000-thread limit and the process died. Roughly every 70-90s. - #3283: one ES index per day per log type, 836 of the cluster's 1091 shards, several holding a single document. vps/zeek-log-maintenance.sh follows the existing suricata-log-maintenance.sh / portbridge-log-maintenance.sh pattern and the #261 HONEYPOT_RETENTION_DAYS ratio, and runs where the files are actually written. Deployed and verified: 13,717 -> 1,874 files, live generations untouched, wired to @reboot plus a running instance. log-maintenance.sh keeps its /logs/zeek-proxy line (#2323, that directory is local and writable -- 834 -> 822 files confirmed) and now documents why no /logs/zeek line can work there, so the next person does not re-add a find that fails quietly. filebeat.yml also gains harvester_limit 64 and ignore_older 72h on both zeek inputs, so a restart cannot re-open thousands of harvesters again even before the first prune pass lands. Verified live: RestartCount 0 and zero thread-exhaustion errors after 5 minutes, against a 70-90s crash cycle before. Refs #3283, #3284
Xore
force-pushed
the
fix/ops-zeek-harvester-limit
branch
from
September 26, 2026 11:32
fcb3a6c to
11e8bb5
Compare
Dependency Review✅ No vulnerabilities or license issues or OpenSSF Scorecard issues found.Scanned FilesNone |
This was referenced Sep 26, 2026
Closed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Root cause of both open ops issues, and it was invisible from the homeserver.
The catch
/opt/stacks/apiary/logs/zeekis reached over sshfs and is deliberately read-only:(same for suricata, portbridge, huginn).
log-maintenance.shon the homeserver therefore cannot delete these files, and itsfindcarries2>/dev/null || true, so the failure is silent. Measured live: 12,390 paths printed by the script, file count unchanged, bothtouchandfind -deletereturningRead-only file system.Nothing on the VPS pruned them either — no logrotate rule, no cron, no maintenance script. So 13,717 hourly rotations (16 GB) had accumulated back to 2026-09-03.
How that drove both issues
hp-filebeatrestart created ~13,700 goroutines until Go hit its 10000-thread limit. Crash cycle was 70–90s.The fix
vps/zeek-log-maintenance.sh(new) — follows the existingsuricata-log-maintenance.sh/portbridge-log-maintenance.shpattern and the Unify log/ILM/artifact retention behind a single top-level env var #261HONEYPOT_RETENTION_DAYSratio, and runs where the files are actually written. Deployed and verified: 13,717 → 1,874 files, live generations untouched, wired to@rebootplus a running instance.filebeat.yml—harvester_limit: 64andignore_older: 72hon both zeek inputs, so a restart cannot re-open thousands of harvesters even before the first prune pass lands. These are the two settings Filebeat offers for exactly this;GODEBUG/SetMaxThreadswould only hide the symptom.log-maintenance.sh— keeps its/logs/zeek-proxyline (zeek-proxy logs are never pruned despite compose claiming 'the pruner owns deletion' (~400MB/day growth) #2323; that directory is local and writable, 834 → 822 files confirmed) and documents why no/logs/zeekline can work there, so nobody re-adds a find that fails quietly.Verified live
hp-filebeat: RestartCount 0, zero thread-exhaustion errors after 5 minutes, against a 70–90s crash cycle before.filebeat test config→Config OKbefore restart. Backups kept as*.bak-3284.Not done here
#3283's consolidation of the existing 836 indices is unchanged — this stops new ones accumulating. Replicas=0 on the single node (12 unassigned shards) and the dead-letter backfill from 2026-09-21 18:19 are also still open. Stopping new growth is a prerequisite for those, not a substitute.
Refs #3283, #3284