Skip to content

bug(Rust SDK): TCP, QUIC and WebSocket clients block after 1000 diagnostic events #4327

Description

@TomPlanche

Bug description

The TCP, QUIC and WebSocket clients of the Rust SDK stop working after 1000 diagnostic events. The 1001st call to publish_event blocks forever, and the connect(), disconnect() or shutdown() call that publishes this event never returns. @hubcio found this problem in the review of #4288. I checked it against the code on master at commit 9aff4424a, and a program reproduces it with the SDK at commit 5d8129e95.

Each client creates a bounded async_broadcast channel with a capacity of 1000, and keeps the original receiver in its events field:

  • TCP, QUIC and WebSocket call broadcast(1000).
  • subscribe_events() returns a clone of this receiver, but nothing reads the original receiver.
  • publish_event calls self.events.0.broadcast(event).await. Overflow mode is off, thus broadcast waits for a free slot when the channel is full.

The channel frees the slot of a message only when every active receiver has received it. The original receiver never receives a message, thus the channel is full after 1000 events and stays full. A subscriber that reads from subscribe_events() does not free the slots.

The TCP client publishes Connected in each connect and Disconnected in each disconnect. For a client without auto-login, the 1001st event is the Connected event of the 501st connect. A long-running client that loses its connection often can reach this limit, and then its reconnect blocks forever.

  • Expected: a client publishes any number of diagnostic events, and connect() and disconnect() always return.
  • Actual: the 501st connect() of a client without auto-login never returns.

The HTTP client also creates this channel, but it does not publish events, thus it does not block.

Possible fix

@hubcio suggests set_overflow(true) on the sender. In overflow mode, broadcast drops the oldest message when the channel is full and does not wait. A subscriber that reads too slowly then loses old events, but no client blocks.

Affected area / component

Rust SDK

Deployment

Compiled from source

Versions

  • Code checked: master at commit 9aff4424a7908a0fdf8a16c1f39e68dd061c1729
  • Reproduction: Rust SDK at commit 5d8129e95de1d1fdf2ed42af015340b82af1c59a, the commit that the repros/sdk-lifecycle crate uses

Sample code

The program is event_channel_blocks.rs in repros/sdk-lifecycle on my fork. The crate uses iggy from apache/iggy at commit 5d8129e95. The program connects and disconnects a TCP client without auto-login, up to 600 times. It sets reestablish_after to zero so that each connect starts at once, and it stops at the first connect() or disconnect() that takes more than 5 s.

Logs

$ cargo run --bin event_channel_blocks
connect 501 blocked

Iggy server config

No response

Reproduction

  1. From the root of an apache/iggy checkout on Linux, start a server with the default root credentials:

    cargo run --bin iggy-server -- --with-default-root-credentials --fresh
  2. Clone the tomplanche/repros branch of my fork, then go to repros/sdk-lifecycle.

  3. Run cargo run --bin event_channel_blocks. The program shows connect 501 blocked.

Contribution

  • I'm willing to submit a pull request to fix this bug

Good first issue

  • I think this could be a good first issue for a new contributor

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

rustPull requests that update Rust code

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions