Skip to content

rendezvous: stop holding the global write lock across the whole reap - #133

Merged
TeoSlayer merged 1 commit into
mainfrom
perf/shorten-reap-lock
Oct 1, 2026
Merged

TeoSlayer merged 1 commit into
mainfrom
perf/shorten-reap-lock

Conversation

@TeoSlayer

Copy link
Copy Markdown
Contributor

Problem

reapStaleNodes (runs every 10s) took the global write lock and held it while:

  • building a ~200k-entry id slice from s.nodes,
  • sorting it (O(n log n)),
  • emitting slog.Info + s.audit lines,

…then kept holding it for the whole delete pass.

The mutex profile shows this as the #3 contention site (reapStaleNodes.deferwrap1, ~160,000 s cumulative wait — low count, huge delay = long holds), stalling every concurrent heartbeat/register.

Fix

  • Snapshot node ids under RLock (cheap) and sort outside the lock.
  • Scan one chunk under RLock, collecting stale candidates without mutating or logging.
  • Delete candidates in short write-locked batches (reapDeleteBatch = 128), re-checking staleness (a node may have heartbeated between scan and delete).
  • Advance reapCursor and do the slog/audit outside the lock.

Behaviour is unchanged: same deletions (backbone membership, hostname index, pubkey/owner index preserved), same cursor semantics, same save() trigger.

Verification

  • go test ./... (20 packages) green; reap tests green; -race on reap/stale tests green.

Context: after the merged encoder + a 1 TB pd-ssd, the live registry is ~88% CPU-bound on 16 vCPUs (pilot-rendezvous ~11.4 cores); the profile is now dominated by network syscalls + lock contention (beacon relay, then this reap, then report_trust).

reapStaleNodes took s.mu.Lock() and held it while building + sorting a
~200k-entry id slice and emitting audit/log lines — the top mutex-contention
site after the beacon relay (mutex profile: ~160k s cumulative delay).

Now: snapshot ids under RLock and sort WITHOUT the lock; scan one chunk under
RLock collecting stale candidates (no mutation/logging); delete candidates in
short write-locked batches (re-checking staleness); advance the cursor and do
audit/logging outside the lock. Behaviour is unchanged (same deletions,
cursor semantics, save trigger); the global stall is gone.
@codecov

codecov Bot commented Oct 1, 2026

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.

📢 Thoughts on this report? Let us know!

@TeoSlayer
TeoSlayer merged commit b85bc76 into main Oct 1, 2026
9 checks passed
@TeoSlayer
TeoSlayer deleted the perf/shorten-reap-lock branch October 1, 2026 22:24
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant