Repository navigation
daemon: forget peers nobody uses, so agents stop growing until they run out of memory - #525
Merged
Merged
Conversation
The service-agent fleet's daemons grow from ~26 MB to ~330 MB RSS over two
days, and nothing in the binary could show where: v1.15.0 exposes no
runtime metrics and no profiles, so the only evidence was RSS from /proc.
-pprof 127.0.0.1:<port> serves Go's runtime profiles (heap, allocs,
goroutine, threadcreate, block, mutex, CPU profile, execution trace,
symbol) on a private mux. It is off by default, and refuses any address
that is not loopback ("localhost" binds 127.0.0.1). /debug/pprof/cmdline
is not served, since the command line can carry -admin-token. Any local
user can reach a loopback TCP port, unlike the IPC socket, so the flag's
help says to leave it off on shared hosts.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Every peer the daemon had exchanged keys with got a NAT keepalive every 25 s, path-watchdog probes and resets when quiet, and, when relayed, a direct-upgrade attempt every 15 s (beacon punch, five probes, a registry resolve each minute), forever. The keepalives also counted as contact, so the stale-peer reaper never fired. A service agent with ~4,000 peers and no open connection sent ~120 KB/s of this. Track application traffic per peer. A peer with none for 2 minutes and no open connection gets no keepalives, path probes or upgrade attempts; the reaper then drops it five minutes after its last frame. The reaper now counts relayed inbound frames as contact, so a peer still sending to us is kept. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit 0b1135d)
Two Docker scenarios on a cone NAT with a 10 s UDP conntrack timeout: idle 160 s (keepalives stopped, mapping expired) and idle 8 minutes (both sides reap each other), then an echo in each direction. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit a6da8d1)
Review fix. SendTo counted every packet it sent as application traffic, including the pong to a peer's path probe. Any idle peer with a path watchdog (v1.13.10 and later) probes after 55 s of quiet, so the pong made it active again for PeerIdleAfter and the upkeep this change stops came back for those peers. Outbound pings and pongs are now treated as they already were inbound. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> (cherry picked from commit ec10ffd)
… out The stale-peer reaper (now able to fire, see the idle-peer change before this) removes a peer with RemovePeer, which keeps the frames queued for the peer's key exchange: a path reset calls it too, and re-keys at once. For a peer that is being forgotten those frames are never sent. They were dropped only when a key exchange completed, so for a peer that never answered they stayed for the daemon's lifetime; once maxPendingPeers (256) of them had piled up, every first send to a new peer without a session failed with "too many pending key exchanges". - reapStalePeers drops a peer through forgetPeer, which also discards its pending queue. - The idle sweep drops any pending queue older than pendingMaxAge (2 min; a key exchange is given up after 20 s), whoever still holds it. - routing.Manager and keyexchange.Manager report how many per-peer entries they hold (PeerStateEntries), and TunnelManager.peerStateEntries sums every per-peer table, for tests and diagnostics. TestPeerChurnStateStaysBounded runs the service-agent workload in-process: 12 rounds of 400 new clients that each make one request and go quiet, with the daemon's keepalive sweep and reaper run between rounds. The tables hold one round (400 peers, 2,800 entries) throughout and the live heap stays at 1.2-1.5 MB. Its control subtest turns idleness tracking off, the keepalive behaviour before the idle-peer change, and shows the leak: 400 -> 4,800 peers and 1.2 -> 8.0 MB, one round's worth per round. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
onRekeyGaveUp resets the peer's path (resolve, punch, a fresh PILA). For a peer that went away the new key exchange gives up too, and the next reset follows once gaveUpResetCooldown (30 s) has passed. With a few hundred such peers in rekey, the keyexchange loop's cap of 64 retransmits per 4 s tick stretched each peer's five attempts past that cooldown, so the cycle never ended: on 2026-10-08, 90 minutes after a restart, catfact-ninja had logged 25,147 "rekey gave up (session desync) — reset peer path" and hn-algolia-search 8,267. Each round is a registry Resolve in a goroutine, a punch and five PILAs, and those PILAs stamp lastOutboundSend, so the stale-peer reaper never dropped these peers either. The idle-peer change stopped keepalives, path probes and upgrades to unused peers, but not this. - onRekeyGaveUp resets only a peer in use: an open connection, or application traffic within PeerIdleAfter (peerInUse). A peer with no activity record is not in use here, unlike for the keepalive and upgrade loops (idleFilter), where only ready sessions are considered and every new session records activity: one it only ever got a key request from, or never heard from, is exactly the case to leave alone. Left alone, it goes quiet and the reaper drops it; whoever uses it next re-keys it. - resetPeerPath keeps the peer's activity record across its RemovePeer: a reset re-keys the same peer. Dropping the record made a reset peer look untracked, which the keepalive and upgrade loops treat as active. TestOnRekeyGaveUpResetsWithCooldown now marks its peer in use. New tests: TestRekeyGaveUpLeavesUnusedPeerToReaper (fails without the gate) and TestResetPeerPathKeepsActivity. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
reapStalePeers walks the tunnel's peer table, so state kept for a node
that has no entry there is never reaped. A path reset produces it:
resetPeerPath removes the peer, puts its relay flag back so the recovery
PILA travels by relay, and re-adds the peer only if ensureTunnel's
registry resolve succeeds. When it fails (the peer went away, or the
registry is slow), the relay flag stays with no tunnel behind it. The
relay probe loop then attempted a direct upgrade for it every 15 s, and
it counted against MaxRelayPeers (4,096) for the daemon's lifetime; a
service agent holds ~3,600 relayed peers, so the cap is within reach,
and past it new relayed peers are dropped ("relay peers cap reached").
In a local soak run of this branch (silent clients, 23 minutes), health
showed 875 relay flags for 381 peers.
- reapStalePeers also sweeps every node that routing, key exchange or
the session store holds state for (new PeerIDs listings) but that has
no tunnel entry and no open connection: it is forgotten once its last
contact is older than peerReapIdleTimeout, or, with no contact ever
recorded, peerReapIdleTimeout after the sweep first finds it
(orphanSeen).
- The last-contact computation is shared (lastPeerContact) and the
timeout is a package constant.
- relayProbeLoop skips relay flags with no tunnel peer: there is no
session to upgrade.
TestReaperForgetsOrphanedPeerState covers both kinds of orphan and that
state with recent contact stays.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ault) TestSoakSwarmAgentMemory runs one agent daemon as its own process, from PILOT_SOAK_AGENT_BIN with the service-agent fleet's flags (-encrypt -public -trust-auto-approve -idle-timeout 3600s -keepalive 300s) and -pprof, against an in-process registry and beacon. A swarm of in-process clients churns through it the way the fleet's traffic does: each client trust-handshakes, sends a message and gets a reply the way the fleet's responder sends one (a handshake, then a dial back to port 1001). Some stay up idle for the whole run; some register an endpoint nobody listens on, so the agent reaches them over the relay. Every sample records the agent's live heap after a forced GC, goroutines, RSS, peers, relayed peers and connections to a CSV; it saves a heap profile at peak RSS, a goroutine dump the first time goroutines pass 2,000, and a heap profile and goroutine dump at the end. Knobs are in the doc comment. It skips unless PILOT_SOAK_AGENT_BIN is set. Before and after the idle-peer and reaping changes on this branch (30 minutes, 4 new clients per second for the first 20, 60 s client lifetime): main: peers grew one for one with clients ever seen; rekey give-ups reset dead peers' paths ~5 times a second; the registry calls behind them piled up to 21,149 goroutines; RSS 238 MB at 11 minutes and 436 MB at 23. this branch: peers stayed at the clients of the last few minutes and fell after churn stopped (126 at the end); goroutines 38-120; RSS 48-57 MB. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
relayProbeLoop started tryDirectUpgrade for every relayed peer every RelayProbeInterval (15 s). An attempt without a cached resolve asks the registry through withRegistryDeadline, which stops waiting after 8 s but leaves the call running: the registry client then waits for a free pooled connection with no deadline. While the registry was slow, every tick added one more waiting call per relayed peer, so goroutines grew with the length of the slowdown rather than with the number of peers. In the local soak test on main this reached 21,149 goroutines at one sample (the gave-up path resets of the previous commits feed the same registry); afterwards 30 MB of the agent's live heap was runtime.malg, i.e. about 70,000 goroutine descriptors that Go keeps for the life of the process. On the fleet the registry connections do stall: since the 10:47 restart catfact-ninja logged 79 "registry pool conn reconnect failed" and 9 half-open reconnects. - relayProbeTick (the loop's body, now testable) skips a peer whose previous attempt is still running (upgradeInFlight). - tryDirectUpgrade's resolve is waited for in full instead of through withRegistryDeadline, so the attempt, and the guard, last as long as the registry call. It already runs in its own goroutine. TestRelayProbeTickRunsOneUpgradePerPeer: with the registry stuck, 8 ticks over 25 relayed peers start 25 attempts (178 without the guard), and the next tick after they finish starts 25 again. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
Every pilot-daemon keeps every peer that ever talked to it, so its memory grows with every client it has ever seen. On the service-agent VM, the ~430 agent daemons grew from ~26 MB to ~330 MB each in two days (~142 GB in total). On 2026-10-08 the VM ran out of memory system-wide and had to be reset.
Root cause
reapStalePeersdrops a peer after 5 min without contact, but the daemon's own 25 s keepalive to every connected peer counts as contact. So every peer stays, along with its upkeep: keepalives, path-watchdog probes and, for a relayed peer, a registry lookup, punch and 5 probes every 15 s.onRekeyGaveUpresets the path, which starts another rekey that gives up. The 64-retransmits-per-4-s cap stretches each round past the 30 s reset cooldown, so the loop never stops. Each round spawns a registry lookup, and its frames count as contact.Production evidence (read-only, since the 10:47 UTC VM reset)
Change
-pproflistener, loopback addresses only, off by default, with nocmdlinepage.PILOT_SOAK_AGENT_BINis setSoak
One agent with the fleet's flags, 4 new clients per second for 20 minutes, then quiet. Figures at 30 minutes:
go build,go vetandgo testpass across./pkg/... ./cmd/... ./internal/....-racepasses on./pkg/daemon/...and./cmd/daemon. The full./tests -parallel 4suite passes in 292 s. v1.15.0 and main leak the same way: none of the code involved changed in between.Supersedes #513.
🤖 Generated with Claude Code