Repository navigation
daemon: recover at once when this node's IP address changes - #495
Merged
Merged
Conversation
TeoSlayer
force-pushed
the
fix/ip-change-recovery
branch
2 times, most recently
from
October 7, 2026 13:59
af63bfe to
db5277b
Compare
When the address under a running daemon changed, nothing noticed. Peers kept sending to the old address until their own timeouts moved them to the relay or our 25s keepalive reached them, our pooled registry connections stayed bound to the old address until a 30s read timeout (so our own dials hung too), and the registry kept the old endpoint. Add an address watcher (pkg/daemon/addrwatch.go). Once a second it asks the kernel which source address it would use toward the beacon (a connected UDP socket, nothing sent). When that changes it re-registers with the beacon, sends one authenticated probe straight to every tunnel peer so each learns the new source from it, and re-registers with the registry over a fresh connection, refreshing the advertised LAN addresses. The beacon's discover reply on the tunnel socket is now parsed as a second input, so a NATed node whose public IP changes runs the same recovery (unit-tested only). Recoveries are at least 10s apart, doubling to 2 minutes while changes keep coming; a change seen during the gap runs when it ends. An interface returning with the same address triggers nothing. The beacon+registry re-registration that the rx watchdog's soft recovery and the host-resume handler each spelled out is now one routine, reestablishTransport, used by all three. Docker lab (3 nodes, one moved to a new address on the same network): peer to moved node 32s -> under 1s; moved node to peer 30s of failed sends -> none; node with no tunnel: 30s timeout -> direct in 0.2s. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…recovery at a time - A beacon switch no longer reads as a move: both baselines (local route source and beacon-observed IP) belong to the beacon they were taken against, and start over when the beacon refresh picks another one. Discover replies carry the beacon that sent them. - The beacon-observed IP is damped: a reply counts only within 2 s of a discover this node sent, and a new IP only once two replies in a row agree (one differing reply triggers one confirming discover, at most one per 30 s). Its backoff starts at 1 minute and caps at 1 hour, separate from the local-address backoff (10 s to 2 minutes). A->B->A inside the cooldown cancels, as it does for the local address. - Address-change recoveries re-register the endpoint only (reRegisterEndpoint): no SetVisibility, SetHostname or per-peer ReportTrust. The full restore still runs when the registry has not answered for 5 minutes, answers under another node ID, or drops the configured hostname. Resume and rx-silence keep the full re-register. - The recovery registers with the beacon again 31 s after it started, past the beacon's 30 s per-node limit on endpoint updates. - reestablishTransport runs one at a time, and a resume or rx-silence run is skipped when one the registry accepted finished under 5 s ago (wall clock). Address changes always run. - reRegister returns an error; registry_ok reports it instead of inferring success from lastRegistryOKNano, which heartbeats also move. - tunnel.addr_changed carries the observed IPs for observed_endpoint. - RecoverFromAddrChange is no longer a Daemon method; RecoverFromAddrChangeForTest serves ./tests. - The end-to-end test heals within a deadline computed from A's last inbound frame and the key-exchange timers, with the recovery run concurrently, instead of fixed sleeps with ~2.5 s of margin. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… keeps a waiting move, wall-clock cooldowns - -no-addr-watch (config.json no_addr_watch, Config.DisableAddrWatch) turns the watcher off, wired like -no-rx-watchdog and -no-path-watch; pilotctl's config key table knows the key. - A beacon switch no longer drops a local change still waiting to be announced. noteBeacon keeps the local baselines when a change is pending, and the switch tick takes one last sample on the old route first, so a move in the same second as the switch is seen. Only the observed-path baselines always start over. - The watcher runs on the wall clock (addrWatchNow, time.Now().Round(0)): the monotonic clock does not advance while a Linux host is suspended, so a cooldown or backoff would not count time asleep. A last recovery that lies ahead of a stepped-back clock does not hold a change back. - The heartbeat's own registry reconnect and re-registration take reestablishMu too (reconnectRegistrySerialised, reRegisterSerialised), so every reconnect and re-registration runs one at a time. - A full caller (resume, rx-silence) is skipped only after a full successful run, not after an endpoint-only one. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
TeoSlayer
force-pushed
the
fix/ip-change-recovery
branch
from
October 7, 2026 14:29
db5277b to
a323b44
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Pull Request
Summary
When a running node's IP address changed, nothing in the daemon noticed. Recovery waited for other nodes' timeouts: measured on v1.14.0-rc.1, messages failed in both directions for 23–32s, an open stream stalled ~40s, and the registry held the old address for minutes. This makes the node detect the change and announce itself.
Changes
reestablishTransport, shared with the watcher.tunnel.addr_changed.Peers need no update: the measurements below used a patched moved node with unpatched peers.
Measured (three nodes and a registry on one Docker network, one node moved to a new address)
Detection fired 0.16–0.8s after the address appeared. Three disconnect/reconnect cycles on the same address triggered nothing; eight moves in 26s produced two recoveries.
Limits
Daemon.RecoverFromAddrChange()exists for that test and is not rate-limited.Test Plan
go build ./...,go vet ./..., unit suite withGOWORK=offpkg/daemon/zz_addrwatch_test.go,tests/zz_addr_change_recovery_test.gogo test -parallel 4 -count=1 ./tests/: one failure,TestManualSnapshotTrigger, which binds fixed port 127.0.0.1:18080 held by another process on the test machine; it fails the same way onmainthereChecklist
go.mod/go.sumunchanged🤖 Generated with Claude Code