Skip to content

daemon: stop maintaining paths to peers nobody is using - #513

Open
TeoSlayer wants to merge 2 commits into
mainfrom
perf/idle-peer-upkeep
Open

TeoSlayer wants to merge 2 commits into
mainfrom
perf/idle-peer-upkeep

Conversation

@TeoSlayer

Copy link
Copy Markdown
Collaborator

Summary

The daemon kept every peer it had ever exchanged keys with warm, forever:

  • a NAT keepalive every 25 s,
  • path-watchdog probes and resets whenever the peer went quiet,
  • for each relayed peer, a direct-upgrade attempt every 15 s (relayProbeLoop → tryDirectUpgrade: a beacon punch request, which makes the beacon send a punch command to both sides, plus 5 probes) and a registry Resolve every 60 s (resolve-cache TTL).

None of it stopped when the peer stopped being used. It also kept the stale-peer reaper from ever firing: reapStalePeers drops a peer with no open connection and no contact for 5 minutes, and the keepalives counted as contact.

This change tracks application traffic per peer: anything sent through SendTo, a new key exchange, and anything delivered to a port other than path probes. A peer with no application traffic for 2 minutes and no open connection gets no keepalives, no path-watchdog probes and no upgrade attempts. Five minutes later the existing reaper removes it, as it was meant to.

Why: the service-agent fleet (read-only measurements, 2026-10-06)

Daemons on the agent VM 432, 22–24 cores in total, flat: median 5.3% of a core each, the 10 busiest only 6% of the total
Peers per agent 3,500–10,900, mostly relayed (catfact-ninja: 4,025 peers, 3,683 relayed)
Upkeep traffic, agent with no open connection ~120 KB/s out (zippopotam-zip 114 KB/s, catfact-ninja 128 KB/s)
catfact-ninja daemon log, 56 min 358 inbox messages; ~17,000 path events (4,204 path recoveries, 3,237 relay flips, 1,909 abandoned rekeys, 1,851 desync resets, 1,608 watchdog resets)
Public rendezvous (pilot-rendezvous, registry + beacon) 9.8 of 16 cores, 147k UDP pkt/s in, 62k out

At ~3,600 relayed peers per agent, the 15 s upgrade loop alone is ~240 attempts/s per daemon. Across 432 agents that is on the order of 100k beacon punch requests/s and 26k registry resolves/s, almost all for peers that have sent nothing in hours. The rendezvous load is an estimate from the code and the counters; a canary agent would measure it.

Functionality

  • A peer with an open connection is never idle, however quiet the connection.
  • Sending to an idle peer works as before. If a NAT mapping expired, the first packet goes through the usual fallbacks (address learning, blackhole detection, relay). If the peer was reaped, the next contact runs a fresh key exchange, as first contact does. A peer that still holds the old session gets the "no key" rekey reply, the same recovery used after a daemon restart (test_tunnel_desync_recovery, test_peer_restarted_*).
  • Behaviour to know: the first message to a peer unused for over ~7 minutes pays a key exchange again (one round trip, plus a registry lookup).

Tests

  • Unit: TestKeepaliveSweepSkipsIdlePeers, TestOnlyApplicationTrafficCountsAsActivity, TestPathWatchSkipsIdlePeers, TestReaperDropsIdlePeersButKeepsTalkingOnes (the last fails on main: a peer sending over the relay was reaped).
  • go test ./pkg/... ./cmd/... ./internal/... -short and go test ./tests/ pass.
  • Docker (harness from tests: make the Docker integration suite build and boot again #506, image built from this branch):
    • New test_nat_idle_peer_resume.sh: cone NAT with a 10 s UDP conntrack timeout, idle 160 s (keepalives stopped, mapping expired), then an echo each way: pass (also passes on main).
    • New test_nat_idle_peer_reaped.sh: same NAT, idle 8 minutes. On this branch both daemons reaped each other (a-forgot-b b-forgot-a) and the echo each way passes after a fresh key exchange. On main: reaped: none, echoes pass.
    • All 109 of 109 tests that pass on main pass on this branch. The image was built from a checkout of this branch with tests: make the Docker integration suite build and boot again #506 merged, and the tests ran from that checkout, so the NAT tests' compose up --build rebuilt the same code.

🤖 Generated with Claude Code

Teo Calin and others added 2 commits October 6, 2026 02:45
Every peer the daemon had exchanged keys with got a NAT keepalive every
25 s, path-watchdog probes and resets when quiet, and, when relayed, a
direct-upgrade attempt every 15 s (beacon punch, five probes, a registry
resolve each minute), forever. The keepalives also counted as contact,
so the stale-peer reaper never fired. A service agent with ~4,000 peers
and no open connection sent ~120 KB/s of this.

Track application traffic per peer. A peer with none for 2 minutes and
no open connection gets no keepalives, path probes or upgrade attempts;
the reaper then drops it five minutes after its last frame. The reaper
now counts relayed inbound frames as contact, so a peer still sending
to us is kept.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Two Docker scenarios on a cone NAT with a 10 s UDP conntrack timeout:
idle 160 s (keepalives stopped, mapping expired) and idle 8 minutes
(both sides reap each other), then an echo in each direction.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant