Skip to content

redial disconnected cluster peers and remove dead ones - #148

Merged
fabracht merged 2 commits into
mainfrom
fix/146-redial
Sep 22, 2026
Merged

fabracht merged 2 commits into
mainfrom
fix/146-redial

Conversation

@fabracht

@fabracht fabracht commented Sep 22, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Test plan

  • cargo make clippy (pedantic, all feature combos)
  • cargo make test (1203 tests, 0 failures)
  • dead_peer_is_removed_and_redialled integration test (send-side failure removes the peer; redial reconnects it)
  • TLA+ specs/ClusterRedial.tla: safe cfg holds safety + liveness; unguarded / reusable-id / no-redial each produce the expected counterexample; overlive cfg confirms InvNoLiveDrop holds when redial fires over a still-live connection
  • live 5-node cluster E2E, fresh cluster per victim (nodes 3, 1-leader, 5): each crashed and restarted accept-only, all 4 survivors re-dialled it, mesh healed to 5/5, cross-node op through the rejoined node recovered (3/3)

Follow-ups (non-blocking, #140/robustness track — see companion issue)

  • reconnection latency up to ~75s (15s heartbeat dead-detect + 60s mesh tick)
  • redial keys on strict Alive, so a briefly-Suspected but healthy peer is eligible for a bounded, safe re-dial once per tick; could gate on Dead/Unknown only
  • hard-crash peer-map removal waits on quinn's ~30s idle timeout (removal is non-load-bearing for reconnection)
  • no per-dial timeout, so a black-hole peer serializes a tick's remaining redials
  • no address-change handling (static --peers; address gossip / discovery is cluster: nodes never dial peers they discover, never retry failed dials, and never drop dead peers #140)

closes #146

@fabracht
fabracht merged commit 2f1b09d into main Sep 22, 2026
9 checks passed
@fabracht
fabracht deleted the fix/146-redial branch September 22, 2026 15:14
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

cluster: congestion-drop reconciliation and stale-peer redial after writer-task death

1 participant