You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
Non-blocking robustness follow-ups from the redial half (#146 / PR #148), on the #140 mesh/reconnection track.
reconnection latency: a crashed peer leaves alive_nodes after the ~15s heartbeat timeout and is re-dialled only on the 60s mesh-check tick, so worst-case reconnection is ~75s. consider driving a re-dial directly off death-detection, or a shorter/adaptive tick.
suspect-window churn: redial_unlinked keys on strict Alive, so a peer briefly in Suspected (healthy but ~7.5s late) is eligible for re-dial, which replaces a working connection with an equivalent fresh one. bounded (once per 60s tick) and safe, but could gate on Dead/Unknown only.
hard-crash removal latency: remove_peer_generation fires on a send-side write_all error, which for a hard crash (no clean close) waits on quinn's ~30s idle timeout. removal is not load-bearing for reconnection (redial keys on liveness and connect_to_peer replaces stale entries), but the map/direct_peers view lags for that window.
no per-dial timeout: connect_to_peer has no dial timeout and redial_unlinked dials sequentially, so a black-hole peer stalls that tick's remaining redials up to quinn's handshake timeout (the redial runs detached, so the event loop is unaffected).
Non-blocking robustness follow-ups from the redial half (#146 / PR #148), on the #140 mesh/reconnection track.
alive_nodesafter the ~15s heartbeat timeout and is re-dialled only on the 60s mesh-check tick, so worst-case reconnection is ~75s. consider driving a re-dial directly off death-detection, or a shorter/adaptive tick.redial_unlinkedkeys on strictAlive, so a peer briefly inSuspected(healthy but ~7.5s late) is eligible for re-dial, which replaces a working connection with an equivalent fresh one. bounded (once per 60s tick) and safe, but could gate onDead/Unknownonly.remove_peer_generationfires on a send-sidewrite_allerror, which for a hard crash (no clean close) waits on quinn's ~30s idle timeout. removal is not load-bearing for reconnection (redial keys on liveness andconnect_to_peerreplaces stale entries), but the map/direct_peersview lags for that window.connect_to_peerhas no dial timeout andredial_unlinkeddials sequentially, so a black-hole peer stalls that tick's remaining redials up to quinn's handshake timeout (the redial runs detached, so the event loop is unaffected).peer_addrsis populated only from--peers; a peer that restarts at a new address cannot be re-dialled to it. proper discovery needs address gossip, which is cluster: nodes never dial peers they discover, never retry failed dials, and never drop dead peers #140.