Bug description
The TCP, QUIC and WebSocket clients of the Rust SDK stop working after 1000 diagnostic events. The 1001st call to publish_event blocks forever, and the connect(), disconnect() or shutdown() call that publishes this event never returns. @hubcio found this problem in the review of #4288. I checked it against the code on master at commit 9aff4424a, and a program reproduces it with the SDK at commit 5d8129e95.
Each client creates a bounded async_broadcast channel with a capacity of 1000, and keeps the original receiver in its events field:
- TCP, QUIC and WebSocket call
broadcast(1000).
subscribe_events() returns a clone of this receiver, but nothing reads the original receiver.
publish_event calls self.events.0.broadcast(event).await. Overflow mode is off, thus broadcast waits for a free slot when the channel is full.
The channel frees the slot of a message only when every active receiver has received it. The original receiver never receives a message, thus the channel is full after 1000 events and stays full. A subscriber that reads from subscribe_events() does not free the slots.
The TCP client publishes Connected in each connect and Disconnected in each disconnect. For a client without auto-login, the 1001st event is the Connected event of the 501st connect. A long-running client that loses its connection often can reach this limit, and then its reconnect blocks forever.
- Expected: a client publishes any number of diagnostic events, and
connect() and disconnect() always return.
- Actual: the 501st
connect() of a client without auto-login never returns.
The HTTP client also creates this channel, but it does not publish events, thus it does not block.
Possible fix
@hubcio suggests set_overflow(true) on the sender. In overflow mode, broadcast drops the oldest message when the channel is full and does not wait. A subscriber that reads too slowly then loses old events, but no client blocks.
Affected area / component
Rust SDK
Deployment
Compiled from source
Versions
- Code checked: master at commit
9aff4424a7908a0fdf8a16c1f39e68dd061c1729
- Reproduction: Rust SDK at commit
5d8129e95de1d1fdf2ed42af015340b82af1c59a, the commit that the repros/sdk-lifecycle crate uses
Sample code
The program is event_channel_blocks.rs in repros/sdk-lifecycle on my fork. The crate uses iggy from apache/iggy at commit 5d8129e95. The program connects and disconnects a TCP client without auto-login, up to 600 times. It sets reestablish_after to zero so that each connect starts at once, and it stops at the first connect() or disconnect() that takes more than 5 s.
Logs
$ cargo run --bin event_channel_blocks
connect 501 blocked
Iggy server config
No response
Reproduction
-
From the root of an apache/iggy checkout on Linux, start a server with the default root credentials:
cargo run --bin iggy-server -- --with-default-root-credentials --fresh
-
Clone the tomplanche/repros branch of my fork, then go to repros/sdk-lifecycle.
-
Run cargo run --bin event_channel_blocks. The program shows connect 501 blocked.
Contribution
Good first issue
Bug description
The TCP, QUIC and WebSocket clients of the Rust SDK stop working after 1000 diagnostic events. The 1001st call to
publish_eventblocks forever, and theconnect(),disconnect()orshutdown()call that publishes this event never returns. @hubcio found this problem in the review of #4288. I checked it against the code on master at commit9aff4424a, and a program reproduces it with the SDK at commit5d8129e95.Each client creates a bounded
async_broadcastchannel with a capacity of 1000, and keeps the original receiver in itseventsfield:broadcast(1000).subscribe_events()returns a clone of this receiver, but nothing reads the original receiver.publish_eventcallsself.events.0.broadcast(event).await. Overflow mode is off, thusbroadcastwaits for a free slot when the channel is full.The channel frees the slot of a message only when every active receiver has received it. The original receiver never receives a message, thus the channel is full after 1000 events and stays full. A subscriber that reads from
subscribe_events()does not free the slots.The TCP client publishes
Connectedin each connect andDisconnectedin each disconnect. For a client without auto-login, the 1001st event is theConnectedevent of the 501st connect. A long-running client that loses its connection often can reach this limit, and then its reconnect blocks forever.connect()anddisconnect()always return.connect()of a client without auto-login never returns.The HTTP client also creates this channel, but it does not publish events, thus it does not block.
Possible fix
@hubcio suggests
set_overflow(true)on the sender. In overflow mode,broadcastdrops the oldest message when the channel is full and does not wait. A subscriber that reads too slowly then loses old events, but no client blocks.Affected area / component
Rust SDK
Deployment
Compiled from source
Versions
9aff4424a7908a0fdf8a16c1f39e68dd061c17295d8129e95de1d1fdf2ed42af015340b82af1c59a, the commit that therepros/sdk-lifecyclecrate usesSample code
The program is
event_channel_blocks.rsinrepros/sdk-lifecycleon my fork. The crate usesiggyfromapache/iggyat commit5d8129e95. The program connects and disconnects a TCP client without auto-login, up to 600 times. It setsreestablish_afterto zero so that each connect starts at once, and it stops at the firstconnect()ordisconnect()that takes more than 5 s.Logs
Iggy server config
No response
Reproduction
From the root of an
apache/iggycheckout on Linux, start a server with the default root credentials:Clone the
tomplanche/reprosbranch of my fork, then go torepros/sdk-lifecycle.Run
cargo run --bin event_channel_blocks. The program showsconnect 501 blocked.Contribution
Good first issue