Skip to content

daemon: recover at once when this node's IP address changes - #495

Merged
TeoSlayer merged 3 commits into
mainfrom
fix/ip-change-recovery
Oct 7, 2026
Merged

TeoSlayer merged 3 commits into
mainfrom
fix/ip-change-recovery

Conversation

@TeoSlayer

Copy link
Copy Markdown
Collaborator

Pull Request

Summary

When a running node's IP address changed, nothing in the daemon noticed. Recovery waited for other nodes' timeouts: measured on v1.14.0-rc.1, messages failed in both directions for 23–32s, an open stream stalled ~40s, and the registry held the old address for minutes. This makes the node detect the change and announce itself.

Changes

  • Detection. Once a second the daemon asks the kernel which source address it would use toward the beacon (a connected UDP socket; nothing is sent). That changes exactly when the daemon's own path changes and ignores unrelated interfaces (Docker bridges, VPNs that do not take the route, link-local addresses). No route reads as "offline" and triggers nothing; the same address coming back triggers nothing.
  • Recovery, in latency order: re-register with the beacon; send one authenticated probe to every tunnel peer so each learns the new source address at once (repeated at +1s and +3s only for peers that have not answered); then a fresh registry connection — the pooled ones are bound to the old address, which is why the moved node's own sends failed for ~30s — and a re-registration.
  • NAT. A NATed node cannot see its public address change locally. The beacon's discover reply, previously discarded, is now parsed; a changed observed IP runs the same recovery (bounded by the 60s beacon keepalive). Port-only changes are ignored.
  • Rate limit. Recoveries are at least 10s apart, doubling to a 2-minute cap and resetting after 10 quiet minutes, so a flapping interface costs a handful of re-registrations and then one every two minutes.
  • The beacon+registry re-registration that soft recovery and resume each had inline is now one routine, reestablishTransport, shared with the watcher.
  • New event tunnel.addr_changed.

Peers need no update: the measurements below used a patched moved node with unpatched peers.

Measured (three nodes and a registry on one Docker network, one node moved to a new address)

Before After
Peer → moved node first success +32s under 1s, no failures
Moved node → peer failures until ~+30s no failures
Node with no tunnel, 10s after the move 30s timeout (registry had the old address) 0.2s, direct
Subscription on the moved node, 1 publish/s 2 lost, first delivery +37s 60/60, gap ~2s
Registry learns the new address 2m46s milliseconds after detection

Detection fired 0.16–0.8s after the address appeared. Three disconnect/reconnect cycles on the same address triggered nothing; eight moves in 26s produced two recoveries.

Limits

  • The NAT path has unit tests only; it was not run against a real NAT. A source-port-only change and compat (WSS) mode are untested; the watcher is off in compat mode.
  • Only the route toward the beacon is watched; a change affecting only peers reached through another interface is not detected.
  • Several real moves within ten minutes raise the gap, so a fourth or fifth can wait up to two minutes and behaves as before meanwhile.
  • A recovery re-reports every trust pair: roughly 3 + N registry calls for N trusted peers.
  • The integration harness runs on loopback and cannot change an address; the integration test points a peer's entry at a dead address and checks the recovery corrects it. The move itself was reproduced in Docker.
  • An exported Daemon.RecoverFromAddrChange() exists for that test and is not rate-limited.

Test Plan

  • go build ./..., go vet ./..., unit suite with GOWORK=off
  • pkg/daemon/zz_addrwatch_test.go, tests/zz_addr_change_recovery_test.go
  • Full go test -parallel 4 -count=1 ./tests/: one failure, TestManualSnapshotTrigger, which binds fixed port 127.0.0.1:18080 held by another process on the test machine; it fails the same way on main there

Checklist

  • New code includes the SPDX license header
  • go.mod / go.sum unchanged
  • CHANGELOG updated

🤖 Generated with Claude Code

@TeoSlayer
TeoSlayer force-pushed the fix/ip-change-recovery branch 2 times, most recently from af63bfe to db5277b Compare October 7, 2026 13:59
Teo Calin and others added 3 commits October 7, 2026 17:28
When the address under a running daemon changed, nothing noticed. Peers
kept sending to the old address until their own timeouts moved them to
the relay or our 25s keepalive reached them, our pooled registry
connections stayed bound to the old address until a 30s read timeout
(so our own dials hung too), and the registry kept the old endpoint.

Add an address watcher (pkg/daemon/addrwatch.go). Once a second it asks
the kernel which source address it would use toward the beacon (a
connected UDP socket, nothing sent). When that changes it re-registers
with the beacon, sends one authenticated probe straight to every tunnel
peer so each learns the new source from it, and re-registers with the
registry over a fresh connection, refreshing the advertised LAN
addresses. The beacon's discover reply on the tunnel socket is now
parsed as a second input, so a NATed node whose public IP changes runs
the same recovery (unit-tested only).

Recoveries are at least 10s apart, doubling to 2 minutes while changes
keep coming; a change seen during the gap runs when it ends. An
interface returning with the same address triggers nothing.

The beacon+registry re-registration that the rx watchdog's soft
recovery and the host-resume handler each spelled out is now one
routine, reestablishTransport, used by all three.

Docker lab (3 nodes, one moved to a new address on the same network):
peer to moved node 32s -> under 1s; moved node to peer 30s of failed
sends -> none; node with no tunnel: 30s timeout -> direct in 0.2s.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…recovery at a time

- A beacon switch no longer reads as a move: both baselines (local route
  source and beacon-observed IP) belong to the beacon they were taken
  against, and start over when the beacon refresh picks another one.
  Discover replies carry the beacon that sent them.
- The beacon-observed IP is damped: a reply counts only within 2 s of a
  discover this node sent, and a new IP only once two replies in a row
  agree (one differing reply triggers one confirming discover, at most one
  per 30 s). Its backoff starts at 1 minute and caps at 1 hour, separate
  from the local-address backoff (10 s to 2 minutes). A->B->A inside the
  cooldown cancels, as it does for the local address.
- Address-change recoveries re-register the endpoint only
  (reRegisterEndpoint): no SetVisibility, SetHostname or per-peer
  ReportTrust. The full restore still runs when the registry has not
  answered for 5 minutes, answers under another node ID, or drops the
  configured hostname. Resume and rx-silence keep the full re-register.
- The recovery registers with the beacon again 31 s after it started, past
  the beacon's 30 s per-node limit on endpoint updates.
- reestablishTransport runs one at a time, and a resume or rx-silence run
  is skipped when one the registry accepted finished under 5 s ago (wall
  clock). Address changes always run.
- reRegister returns an error; registry_ok reports it instead of inferring
  success from lastRegistryOKNano, which heartbeats also move.
- tunnel.addr_changed carries the observed IPs for observed_endpoint.
- RecoverFromAddrChange is no longer a Daemon method;
  RecoverFromAddrChangeForTest serves ./tests.
- The end-to-end test heals within a deadline computed from A's last
  inbound frame and the key-exchange timers, with the recovery run
  concurrently, instead of fixed sleeps with ~2.5 s of margin.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
… keeps a waiting move, wall-clock cooldowns

- -no-addr-watch (config.json no_addr_watch, Config.DisableAddrWatch)
  turns the watcher off, wired like -no-rx-watchdog and -no-path-watch;
  pilotctl's config key table knows the key.
- A beacon switch no longer drops a local change still waiting to be
  announced. noteBeacon keeps the local baselines when a change is
  pending, and the switch tick takes one last sample on the old route
  first, so a move in the same second as the switch is seen. Only the
  observed-path baselines always start over.
- The watcher runs on the wall clock (addrWatchNow, time.Now().Round(0)):
  the monotonic clock does not advance while a Linux host is suspended,
  so a cooldown or backoff would not count time asleep. A last recovery
  that lies ahead of a stepped-back clock does not hold a change back.
- The heartbeat's own registry reconnect and re-registration take
  reestablishMu too (reconnectRegistrySerialised, reRegisterSerialised),
  so every reconnect and re-registration runs one at a time.
- A full caller (resume, rx-silence) is skipped only after a full
  successful run, not after an endpoint-only one.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@TeoSlayer
TeoSlayer force-pushed the fix/ip-change-recovery branch from db5277b to a323b44 Compare October 7, 2026 14:29
@TeoSlayer
TeoSlayer merged commit 12dc7a8 into main Oct 7, 2026
15 checks passed
@TeoSlayer
TeoSlayer deleted the fix/ip-change-recovery branch October 7, 2026 14:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant