Skip to content

daemon: forget peers nobody uses, so agents stop growing until they run out of memory - #525

Merged
TeoSlayer merged 10 commits into
mainfrom
fix/daemon-memory-growth
Oct 8, 2026
Merged

TeoSlayer merged 10 commits into
mainfrom
fix/daemon-memory-growth

Conversation

@TeoSlayer

Copy link
Copy Markdown
Collaborator

Problem

Every pilot-daemon keeps every peer that ever talked to it, so its memory grows with every client it has ever seen. On the service-agent VM, the ~430 agent daemons grew from ~26 MB to ~330 MB each in two days (~142 GB in total). On 2026-10-08 the VM ran out of memory system-wide and had to be reset.

Root cause

  1. The reaper never fires. reapStalePeers drops a peer after 5 min without contact, but the daemon's own 25 s keepalive to every connected peer counts as contact. So every peer stays, along with its upkeep: keepalives, path-watchdog probes and, for a relayed peer, a registry lookup, punch and 5 probes every 15 s.
  2. Dead peers loop forever. When a rekey to a vanished client gives up, onRekeyGaveUp resets the path, which starts another rekey that gives up. The 64-retransmits-per-4-s cap stretches each round past the 30 s reset cooldown, so the loop never stops. Each round spawns a registry lookup, and its frames count as contact.
  3. Goroutines pile up. When the registry is slow, those lookups pile up as goroutines. Their descriptors are never returned to the heap.

Production evidence (read-only, since the 10:47 UTC VM reset)

  • Fleet average peers went 186 → 781 between 11:02 and 12:58 UTC, and only up. RSS ≈ 29 MB + 11–14 KB per peer (r = 0.83–0.87).
  • github-public: 2,235 → 8,396 peers, 52 → 111 MB, about 94 MB/h at first.
  • catfact-ninja logged 25,147 rekey-gave-up resets in its first 90 minutes.

Change

  • daemon: stop maintaining paths to peers nobody is using #513 included, unchanged: stop maintaining paths to peers nobody is using. A path-probe reply no longer counts as activity.
  • Reaping: a reaped peer leaves nothing behind. Frames waiting on a rekey that never completes expire after 2 minutes; before, after 256 dead peers, every first send to a new peer failed with "too many pending key exchanges".
  • No reset loop: a rekey that gives up no longer resets a peer nobody is using.
  • Leftover state: per-peer state left for nodes with no tunnel entry is reaped, including relay flags, which count against the 4,096 relay-peer cap.
  • Direct upgrades: at most one direct-upgrade attempt per relayed peer at a time.
  • Profiling: an opt-in -pprof listener, loopback addresses only, off by default, with no cmdline page.
  • Tests:
    • a bounded-churn test: 12 rounds of 400 clients stay flat at 400 peers and ~1.5 MB, while the old behaviour grows 400 → 4,800
    • a soak test that runs only when PILOT_SOAK_AGENT_BIN is set

Soak

One agent with the fleet's flags, 4 new clients per second for 20 minutes, then quiet. Figures at 30 minutes:

peers relay flags live heap peak goroutines gave-up resets
main 2,427 (3,904 at 20 min) 2,329 22.1 MB 4,237 5,177
#513 alone 126 142 9.0 MB 315 3,021
this branch 0 0 8.1 MB 256 48

go build, go vet and go test pass across ./pkg/... ./cmd/... ./internal/.... -race passes on ./pkg/daemon/... and ./cmd/daemon. The full ./tests -parallel 4 suite passes in 292 s. v1.15.0 and main leak the same way: none of the code involved changed in between.

Supersedes #513.

🤖 Generated with Claude Code

Teo Calin and others added 10 commits October 8, 2026 15:45
The service-agent fleet's daemons grow from ~26 MB to ~330 MB RSS over two
days, and nothing in the binary could show where: v1.15.0 exposes no
runtime metrics and no profiles, so the only evidence was RSS from /proc.

-pprof 127.0.0.1:<port> serves Go's runtime profiles (heap, allocs,
goroutine, threadcreate, block, mutex, CPU profile, execution trace,
symbol) on a private mux. It is off by default, and refuses any address
that is not loopback ("localhost" binds 127.0.0.1). /debug/pprof/cmdline
is not served, since the command line can carry -admin-token. Any local
user can reach a loopback TCP port, unlike the IPC socket, so the flag's
help says to leave it off on shared hosts.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every peer the daemon had exchanged keys with got a NAT keepalive every
25 s, path-watchdog probes and resets when quiet, and, when relayed, a
direct-upgrade attempt every 15 s (beacon punch, five probes, a registry
resolve each minute), forever. The keepalives also counted as contact,
so the stale-peer reaper never fired. A service agent with ~4,000 peers
and no open connection sent ~120 KB/s of this.

Track application traffic per peer. A peer with none for 2 minutes and
no open connection gets no keepalives, path probes or upgrade attempts;
the reaper then drops it five minutes after its last frame. The reaper
now counts relayed inbound frames as contact, so a peer still sending
to us is kept.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit 0b1135d)
Two Docker scenarios on a cone NAT with a 10 s UDP conntrack timeout:
idle 160 s (keepalives stopped, mapping expired) and idle 8 minutes
(both sides reap each other), then an echo in each direction.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit a6da8d1)
Review fix. SendTo counted every packet it sent as application traffic,
including the pong to a peer's path probe. Any idle peer with a path
watchdog (v1.13.10 and later) probes after 55 s of quiet, so the pong made
it active again for PeerIdleAfter and the upkeep this change stops came
back for those peers. Outbound pings and pongs are now treated as they
already were inbound.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
(cherry picked from commit ec10ffd)
… out

The stale-peer reaper (now able to fire, see the idle-peer change before
this) removes a peer with RemovePeer, which keeps the frames queued for the
peer's key exchange: a path reset calls it too, and re-keys at once. For a
peer that is being forgotten those frames are never sent. They were dropped
only when a key exchange completed, so for a peer that never answered they
stayed for the daemon's lifetime; once maxPendingPeers (256) of them had
piled up, every first send to a new peer without a session failed with
"too many pending key exchanges".

- reapStalePeers drops a peer through forgetPeer, which also discards its
  pending queue.
- The idle sweep drops any pending queue older than pendingMaxAge (2 min;
  a key exchange is given up after 20 s), whoever still holds it.
- routing.Manager and keyexchange.Manager report how many per-peer entries
  they hold (PeerStateEntries), and TunnelManager.peerStateEntries sums
  every per-peer table, for tests and diagnostics.

TestPeerChurnStateStaysBounded runs the service-agent workload in-process:
12 rounds of 400 new clients that each make one request and go quiet, with
the daemon's keepalive sweep and reaper run between rounds. The tables hold
one round (400 peers, 2,800 entries) throughout and the live heap stays at
1.2-1.5 MB. Its control subtest turns idleness tracking off, the keepalive
behaviour before the idle-peer change, and shows the leak: 400 -> 4,800
peers and 1.2 -> 8.0 MB, one round's worth per round.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
onRekeyGaveUp resets the peer's path (resolve, punch, a fresh PILA). For a
peer that went away the new key exchange gives up too, and the next reset
follows once gaveUpResetCooldown (30 s) has passed. With a few hundred such
peers in rekey, the keyexchange loop's cap of 64 retransmits per 4 s tick
stretched each peer's five attempts past that cooldown, so the cycle never
ended: on 2026-10-08, 90 minutes after a restart, catfact-ninja had logged
25,147 "rekey gave up (session desync) — reset peer path" and
hn-algolia-search 8,267. Each round is a registry Resolve in a goroutine,
a punch and five PILAs, and those PILAs stamp lastOutboundSend, so the
stale-peer reaper never dropped these peers either. The idle-peer change
stopped keepalives, path probes and upgrades to unused peers, but not this.

- onRekeyGaveUp resets only a peer in use: an open connection, or
  application traffic within PeerIdleAfter (peerInUse). A peer with no
  activity record is not in use here, unlike for the keepalive and upgrade
  loops (idleFilter), where only ready sessions are considered and every
  new session records activity: one it only ever got a key request from,
  or never heard from, is exactly the case to leave alone. Left alone, it
  goes quiet and the reaper drops it; whoever uses it next re-keys it.
- resetPeerPath keeps the peer's activity record across its RemovePeer:
  a reset re-keys the same peer. Dropping the record made a reset peer look
  untracked, which the keepalive and upgrade loops treat as active.

TestOnRekeyGaveUpResetsWithCooldown now marks its peer in use. New tests:
TestRekeyGaveUpLeavesUnusedPeerToReaper (fails without the gate) and
TestResetPeerPathKeepsActivity.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
reapStalePeers walks the tunnel's peer table, so state kept for a node
that has no entry there is never reaped. A path reset produces it:
resetPeerPath removes the peer, puts its relay flag back so the recovery
PILA travels by relay, and re-adds the peer only if ensureTunnel's
registry resolve succeeds. When it fails (the peer went away, or the
registry is slow), the relay flag stays with no tunnel behind it. The
relay probe loop then attempted a direct upgrade for it every 15 s, and
it counted against MaxRelayPeers (4,096) for the daemon's lifetime; a
service agent holds ~3,600 relayed peers, so the cap is within reach,
and past it new relayed peers are dropped ("relay peers cap reached").
In a local soak run of this branch (silent clients, 23 minutes), health
showed 875 relay flags for 381 peers.

- reapStalePeers also sweeps every node that routing, key exchange or
  the session store holds state for (new PeerIDs listings) but that has
  no tunnel entry and no open connection: it is forgotten once its last
  contact is older than peerReapIdleTimeout, or, with no contact ever
  recorded, peerReapIdleTimeout after the sweep first finds it
  (orphanSeen).
- The last-contact computation is shared (lastPeerContact) and the
  timeout is a package constant.
- relayProbeLoop skips relay flags with no tunnel peer: there is no
  session to upgrade.

TestReaperForgetsOrphanedPeerState covers both kinds of orphan and that
state with recent contact stays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ault)

TestSoakSwarmAgentMemory runs one agent daemon as its own process, from
PILOT_SOAK_AGENT_BIN with the service-agent fleet's flags (-encrypt
-public -trust-auto-approve -idle-timeout 3600s -keepalive 300s) and
-pprof, against an in-process registry and beacon. A swarm of in-process
clients churns through it the way the fleet's traffic does: each client
trust-handshakes, sends a message and gets a reply the way the fleet's
responder sends one (a handshake, then a dial back to port 1001). Some
stay up idle for the whole run; some register an endpoint nobody listens
on, so the agent reaches them over the relay.

Every sample records the agent's live heap after a forced GC, goroutines,
RSS, peers, relayed peers and connections to a CSV; it saves a heap
profile at peak RSS, a goroutine dump the first time goroutines pass
2,000, and a heap profile and goroutine dump at the end. Knobs are in the
doc comment. It skips unless PILOT_SOAK_AGENT_BIN is set.

Before and after the idle-peer and reaping changes on this branch (30
minutes, 4 new clients per second for the first 20, 60 s client lifetime):

  main: peers grew one for one with clients ever seen; rekey give-ups
  reset dead peers' paths ~5 times a second; the registry calls behind
  them piled up to 21,149 goroutines; RSS 238 MB at 11 minutes and
  436 MB at 23.
  this branch: peers stayed at the clients of the last few minutes and
  fell after churn stopped (126 at the end); goroutines 38-120; RSS
  48-57 MB.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
relayProbeLoop started tryDirectUpgrade for every relayed peer every
RelayProbeInterval (15 s). An attempt without a cached resolve asks the
registry through withRegistryDeadline, which stops waiting after 8 s but
leaves the call running: the registry client then waits for a free pooled
connection with no deadline. While the registry was slow, every tick
added one more waiting call per relayed peer, so goroutines grew with
the length of the slowdown rather than with the number of peers. In the
local soak test on main this reached 21,149 goroutines at one sample (the
gave-up path resets of the previous commits feed the same registry);
afterwards 30 MB of the agent's live heap was runtime.malg, i.e. about
70,000 goroutine descriptors that Go keeps for the life of the process.
On the fleet the registry connections do stall: since the 10:47 restart
catfact-ninja logged 79 "registry pool conn reconnect failed" and 9
half-open reconnects.

- relayProbeTick (the loop's body, now testable) skips a peer whose
  previous attempt is still running (upgradeInFlight).
- tryDirectUpgrade's resolve is waited for in full instead of through
  withRegistryDeadline, so the attempt, and the guard, last as long as
  the registry call. It already runs in its own goroutine.

TestRelayProbeTickRunsOneUpgradePerPeer: with the registry stuck, 8 ticks
over 25 relayed peers start 25 attempts (178 without the guard), and the
next tick after they finish starts 25 again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@TeoSlayer
TeoSlayer merged commit 8ed371f into main Oct 8, 2026
14 checks passed
@TeoSlayer
TeoSlayer deleted the fix/daemon-memory-growth branch October 8, 2026 13:15
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant