Repository navigation
Stream segments fit one packet; loss recovery and short writes that work with them - #509
Merged
Merged
Conversation
…s that work with them
A full stream segment was 4096 bytes: about 4.2 KB as a UDP datagram and
three IP fragments on a 1500-byte path. NATs, firewalls and some virtual
networks drop fragments, so on those paths handshakes, pings and messages
under ~1.4 KB worked while every full segment was lost — larger replies
never arrived and send-file failed.
Segments are now at most SendSegmentSize (1152) bytes: a 1222-byte
datagram direct, 1231 relayed, within the 1232 that fit IPv6's minimum
MTU. Sender-side only; receivers accept up to MaxSegmentSize as before, so
old and new nodes interoperate in both directions.
Smaller segments exposed faults the large ones had hidden. They are on
main today and are fixed here because the change is not safe without them.
Window accounting
- The peer's window, advertised in segments, is enforced in segments.
Counting bytes only, short segments overran the peer's receive slots.
- The congestion window's growth and floors stay in MaxSegmentSize units;
only adjustments that stand for one segment leaving the network use
SendSegmentSize. Scaling everything down made a lossy path markedly
slower (5 MB at 30ms RTT, 0.5% loss: 17s against 10s).
Loss recovery (on main, 2 of 5 5 MB transfers failed at 2% loss)
- In fast recovery SACKed segments were left out of the amount in flight
while the window also grew per duplicate ACK, counting each twice; the
amount outstanding doubled every round trip.
- The peer buffers at most MaxOOOBuf (128) out-of-order segments and
drops the rest silently. The sender now keeps at most that many
unacknowledged (MaxSegmentsOutstanding).
- After a retransmission timeout each further missing segment waited for
its own timeout, which doubles each time. A partial ACK in timeout
recovery now retransmits the next missing segment at once.
Short writes
Nagle held a short segment until everything before it was ACKed and
blocked the writer until it had gone out. Nothing could join a held write,
a stream of short writes moved at one write per round trip, and because a
client's IPC sends are handled in order, its other requests waited too.
With 1152-byte segments every 4 KB write ends in a short tail.
- The tail of a write of a segment or more leaves with it.
- A write shorter than a segment waits only for an earlier short segment.
- A held write does not block: it stays buffered, later writes join it,
and flushHeldTail sends it when that segment is ACKed, after NagleTimeout
at the latest, or CloseConnection sends it ahead of the FIN. SendMu
keeps segments in write order whichever goroutine sends them.
- A receiver ACKs a lone segment at once for the first QuickACKBudget (32)
segments of a connection and again after each time its delayed-ACK
timer has had to fire, so a peer that waits for an ACK per small write
(any peer up to v1.14.1) is not stalled for the timer each time.
Measured in a container lab (details in the pull request):
5 MB over a 30ms path, five runs each, before -> after
2% loss: 9.7-12.4s with 2 of 5 failed -> 10.3-11.8s, none failed
0.5% loss: 9.4-10.5s -> 7.3-7.6s
no loss: 8.2-8.4s -> 7.6-7.7s
60 MB on one host: 4.8s -> 1.2s; into v1.13.5: 1.8s (v1.13.5 to itself
22.7s)
with non-first fragments dropped at the receiver: messages over one
packet and file transfers failed -> all delivered, no fragments sent
Interop was checked both ways against v1.13.5, v1.14.1 and main. Two
costs: a node that is not upgraded and echoes back what it reads now
sends one 1152-byte piece per round trip (its own Nagle rule); and the
current beacon reorders relayed packets and caps packets per source,
which costs smaller segments more (5 MB relayed: 5.8s against 3.6s)
until pilot-protocol/beacon#61 is deployed (1.8s).
Twelve tests pinned the old behaviour and are updated: three segmentation
tests (1152-byte segments), four fast-recovery tests (window grows by
1152 per duplicate ACK), the window-update wake test (two free slots, one
taken by the segment in flight), the delayed-ACK scheduling test (starts
with the quick-ACK budget spent), and three Nagle tests that asserted the
writer blocks on a held write.
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What was wrong
Every full stream segment was fragmented by IP. A segment carried up to 4096 bytes, about 4.2 KB as a UDP datagram, which is three IP fragments on a 1500-byte path. NATs, firewalls and some virtual networks drop fragments. On those paths the handshake, pings and messages under about 1.4 KB worked and every full segment was lost: larger replies never arrived,
send-filefailed,benchreported a few kbit/s until its retransmissions gave up. This is the cause behind "request/response breaks behind NAT" inhosted/FINDINGS.md; reply connections behind NAT themselves work.Making segments smaller exposed faults that the large segments had been hiding, in loss recovery and in how short writes are sent. Both are on main today; they are fixed here because the change is not safe without them.
What changes
Segment size
SendSegmentSize = 1152. A full segment is a 1222-byte datagram direct and 1231 relayed, within the 1232 bytes that fit IPv6's minimum MTU. Sender-side only: receivers accept segments up toMaxSegmentSize(4096) as before.Loss recovery (on main, 2 of 5 transfers of 5 MB failed at 2% loss after 214 s and 326 s)
Short writes
Nagle's algorithm held a short segment until everything before it was acknowledged, and blocked the writer until it had gone out. That ACK is delayed by the peer when the full segments before it are an odd number (5 ms; 40 ms up to v1.14.1). Nothing could join a held write because its writer could not write again, so a stream of short writes moved at one write per round trip. And because a client's sends are handled in order on its IPC connection, every other request from that client waited too. With 1152-byte segments every 4 KB write ends in a short tail, so:
SendMukeeps segments in write order whichever goroutine sends them.QuickACKBudget). This is what keeps a v1.14.1-or-older peer moving when it waits for an ACK per small write.Interop
Container lab: two nodes and a registry,
netemfor delay and loss,nftablesto drop fragments and to block the direct path. "fix" is this branch; v1.13.5 is what the fleet runs; v1.14.1 is the latest release. Every check below was run in both directions; messages are byte-counted in the receiver's inbox and files are compared by SHA-256.Result: every check passes in both directions with v1.13.5, v1.14.1 and main, except where the sender is not upgraded and the path drops fragments (the bug itself). Two costs of a part-upgraded network are listed under "What to know".
Clean network, 1 CPU per node. Times in seconds; bench is a 1 MB echo.
Same, with the kernel's socket buffer limit at the stock 212992 bytes (an earlier attempt to remove the pause between writes collapsed here: four concurrent 20 MB files took 34–40 s):
30 ms round trip with 1% loss in each direction:
The slow main → fix runs are the client's own loss recovery, traced: it has sent everything and retransmits one segment per timeout (after 1.8 s, 3.6 s, 6.3 s, 14 s). That is the fault this PR repairs for senders that run it.
5 MB file from the fix into main over the same 30 ms path, five runs at each loss rate, against main into main:
Non-first IP fragments dropped at the receiver:
Relayed (direct path blocked):
pub/sub over the relay timed out in every pairing with beacon v0.2.10, baselines included, and passed in every pairing with beacon#61.
What to know before merging
pilotctl benchfrom an upgraded node against one takes about 35 s for 1 MB at 30 ms (0.2 s on a LAN). Requests, replies, files and pub/sub are not affected; upgrading the echoing node removes it. I found no way to avoid this from the upgraded side.nagleWritePiece = 16 * MaxSegmentSize. Whichever of the two merges second should make that a multiple ofSendSegmentSize.Tests
go vet ./...is clean. Unit tests pass, also under-race. The integration package (./tests, local only) passes on a quiet machine. Run while the lab was loading the machine it failed one test in each of two runs,TestWaitForTrustBlocksUntilApprovedandTestHandshakeTrustLoadVerify; the second fails the same way on main under that load and is fixed in tests: wait for the trust file instead of reading it the moment trust appears #508, and the first was a registry port collision when I reran it.🤖 Generated with Claude Code