Skip to content

cluster: congestion-drop reconciliation and stale-peer redial after writer-task death #146

Description

@fabracht

Follow-ups from #143 / #145 (per-peer QUIC writer task), out of scope there. Both belong to the #140 full-peer-mesh / reconnection track.

  • congestion-drop reconciliation: send_to_peer now drops a whole frame with SendQueueFull once the 1024-slot per-peer queue fills, instead of the old blocking send. this removes event-loop head-of-line blocking but changes overload semantics — under a bursty-but-recoverable stall the old path eventually delivered with latency, the new path drops once the queue backs up. needs a reconciliation / adaptive-queue story so cluster traffic dropped under congestion is reconciled.
  • stale-peer redial after writer-task death: when a peer_writer_task exits on a stream write error, its peer-map entry is left in place and nothing re-dials outward, so later sends return SendFailed/Disconnected until an incoming connection replaces the entry. same reconnection gap as cluster: nodes never dial peers they discover, never retry failed dials, and never drop dead peers #140.

neither affects framing safety, which is fixed in #145.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions