Repository navigation
rendezvous: stop holding the global write lock across the whole reap - #133
Merged
Merged
Conversation
reapStaleNodes took s.mu.Lock() and held it while building + sorting a ~200k-entry id slice and emitting audit/log lines — the top mutex-contention site after the beacon relay (mutex profile: ~160k s cumulative delay). Now: snapshot ids under RLock and sort WITHOUT the lock; scan one chunk under RLock collecting stale candidates (no mutation/logging); delete candidates in short write-locked batches (re-checking staleness); advance the cursor and do audit/logging outside the lock. Behaviour is unchanged (same deletions, cursor semantics, save trigger); the global stall is gone.
Codecov Report✅ All modified and coverable lines are covered by tests. 📢 Thoughts on this report? Let us know! |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
reapStaleNodes(runs every 10s) took the global write lock and held it while:s.nodes,slog.Info+s.auditlines,…then kept holding it for the whole delete pass.
The mutex profile shows this as the #3 contention site (
reapStaleNodes.deferwrap1, ~160,000 s cumulative wait — low count, huge delay = long holds), stalling every concurrent heartbeat/register.Fix
RLock(cheap) and sort outside the lock.RLock, collecting stale candidates without mutating or logging.reapDeleteBatch = 128), re-checking staleness (a node may have heartbeated between scan and delete).reapCursorand do the slog/audit outside the lock.Behaviour is unchanged: same deletions (backbone membership, hostname index, pubkey/owner index preserved), same cursor semantics, same
save()trigger.Verification
go test ./...(20 packages) green; reap tests green;-raceon reap/stale tests green.Context: after the merged encoder + a 1 TB pd-ssd, the live registry is ~88% CPU-bound on 16 vCPUs (
pilot-rendezvous~11.4 cores); the profile is now dominated by network syscalls + lock contention (beacon relay, then this reap, then report_trust).