Skip to content

Stream segments fit one packet; loss recovery and short writes that work with them - #509

Merged
TeoSlayer merged 1 commit into
mainfrom
fix/segment-fits-one-packet
Oct 5, 2026
Merged

TeoSlayer merged 1 commit into
mainfrom
fix/segment-fits-one-packet

Conversation

@TeoSlayer

Copy link
Copy Markdown
Collaborator

What was wrong

Every full stream segment was fragmented by IP. A segment carried up to 4096 bytes, about 4.2 KB as a UDP datagram, which is three IP fragments on a 1500-byte path. NATs, firewalls and some virtual networks drop fragments. On those paths the handshake, pings and messages under about 1.4 KB worked and every full segment was lost: larger replies never arrived, send-file failed, bench reported a few kbit/s until its retransmissions gave up. This is the cause behind "request/response breaks behind NAT" in hosted/FINDINGS.md; reply connections behind NAT themselves work.

Making segments smaller exposed faults that the large segments had been hiding, in loss recovery and in how short writes are sent. Both are on main today; they are fixed here because the change is not safe without them.

What changes

Segment size

  • SendSegmentSize = 1152. A full segment is a 1222-byte datagram direct and 1231 relayed, within the 1232 bytes that fit IPv6's minimum MTU. Sender-side only: receivers accept segments up to MaxSegmentSize (4096) as before.
  • The peer's window is advertised in segments (free receive slots) and is now enforced in segments. Counting bytes only, short segments overran the peer's slots.
  • The congestion window's growth and floors stay in 4096-byte units, so a connection ramps up and backs off over the same bytes as before. Only the adjustments that stand for one segment leaving the network use the new size. Scaling everything down was measurably slower under loss (17 s against 10 s for 5 MB at 0.5% loss).

Loss recovery (on main, 2 of 5 transfers of 5 MB failed at 2% loss after 214 s and 326 s)

  1. In fast recovery, SACKed segments were left out of the amount in flight while the window also grew by one segment per duplicate ACK. Each was counted twice, and the amount outstanding doubled every round trip while the first loss stayed unrepaired.
  2. The peer buffers at most 128 out-of-order segments and silently drops the rest. The sender ran past that (traced: receiver's buffer at 128, sender with 166 outstanding). It now keeps at most 128 segments unacknowledged.
  3. After a retransmission timeout, each further missing segment waited for a timeout of its own, and the timeout doubles each time (0.8 s, 1.6 s, 3.3 s, 6.4 s, 9.4 s in the trace). A partial ACK in timeout recovery now retransmits the next missing segment at once.

Short writes

Nagle's algorithm held a short segment until everything before it was acknowledged, and blocked the writer until it had gone out. That ACK is delayed by the peer when the full segments before it are an odd number (5 ms; 40 ms up to v1.14.1). Nothing could join a held write because its writer could not write again, so a stream of short writes moved at one write per round trip. And because a client's sends are handled in order on its IPC connection, every other request from that client waited too. With 1152-byte segments every 4 KB write ends in a short tail, so:

  • The tail of a write of a segment or more leaves with it.
  • A write shorter than a segment waits only for an earlier short segment.
  • A held write does not block. It stays buffered, later writes join it, and it is sent when that earlier segment is acknowledged, after 40 ms at the latest, or ahead of the FIN on close. SendMu keeps segments in write order whichever goroutine sends them.
  • A receiver acknowledges a lone segment at once for the first 32 segments of a connection and again after each time its delayed-ACK timer has had to fire (QuickACKBudget). This is what keeps a v1.14.1-or-older peer moving when it waits for an ACK per small write.

Interop

Container lab: two nodes and a registry, netem for delay and loss, nftables to drop fragments and to block the direct path. "fix" is this branch; v1.13.5 is what the fleet runs; v1.14.1 is the latest release. Every check below was run in both directions; messages are byte-counted in the receiver's inbox and files are compared by SHA-256.

Result: every check passes in both directions with v1.13.5, v1.14.1 and main, except where the sender is not upgraded and the path drops fragments (the bug itself). Two costs of a part-upgraded network are listed under "What to know".

Clean network, 1 CPU per node. Times in seconds; bench is a 1 MB echo.

Sender → receiver 60 MB file four 20 MB files at once bench
fix → v1.13.5 1.8 2.6 0.19
v1.13.5 → fix 3.6 1.6 0.08
fix → v1.14.1 2.1 2.3 0.20
v1.14.1 → fix 3.1 1.5 0.02
fix → main 2.1 1.4 0.18
main → fix 3.1 1.6 0.03
fix → fix 1.2 1.1 0.03
main → main (baseline) 4.8 2.8 0.02
v1.13.5 → v1.13.5 (baseline) 22.7 10.1 0.02

Same, with the kernel's socket buffer limit at the stock 212992 bytes (an earlier attempt to remove the pause between writes collapsed here: four concurrent 20 MB files took 34–40 s):

Sender → receiver 60 MB file four 20 MB files at once
fix → v1.13.5 1.8 1.9
v1.13.5 → fix 3.6 1.6
fix → fix 1.2 1.1
main → main (baseline) 4.8 2.8
v1.13.5 → v1.13.5 (baseline) 23.0 13.3

30 ms round trip with 1% loss in each direction:

Sender → receiver 5 MB file bench (sender is the client, receiver echoes)
fix → v1.13.5 8.6 35 (see below)
v1.13.5 → fix 9.4 2.0
fix → main 8.0 34 (see below)
main → fix 10.5 23 of 30 runs in 0.9–2.9; 7 of 30 over 20, 4 of them timing out at 120
fix → fix 8.4 30 of 30 runs in 1.1–2.2
main → main (baseline) 11.4 2.4
v1.13.5 → v1.13.5 (baseline) 12.8 failed: 925 KB of 1 MB back after 120

The slow main → fix runs are the client's own loss recovery, traced: it has sent everything and retransmits one segment per timeout (after 1.8 s, 3.6 s, 6.3 s, 14 s). That is the fault this PR repairs for senders that run it.

5 MB file from the fix into main over the same 30 ms path, five runs at each loss rate, against main into main:

Loss main → main fix → main
2% 9.7–12.4, and 2 of 5 failed (214 and 326) 10.3–11.8, none failed
0.5% 9.4–10.5 7.3–7.6
0% 8.2–8.4 7.6–7.7

Non-first IP fragments dropped at the receiver:

Sender → receiver 2 KB / 3.9 KB / 60 KB messages 64 KB / 5 MB files fragments dropped
fix → v1.13.5 delivered delivered 0
fix → main delivered delivered 0
v1.13.5 → fix lost failed 217
main → fix lost failed 215
v1.13.5 → v1.13.5 (baseline) lost failed —

Relayed (direct path blocked):

Sender → receiver 5 MB file, beacon v0.2.10 5 MB file, beacon#61 bench, v0.2.10 bench, beacon#61
fix → v1.13.5 6.6 1.7 8.1 8.8 (the old node echoing, see below)
v1.13.5 → fix 2.3 2.5 1.6 0.32
fix → fix 5.8 1.8 2.4 0.04
main → main (baseline) 3.6 2.7 0.26 0.04
v1.13.5 → v1.13.5 (baseline) 4.7 not run 0.32 not run

pub/sub over the relay timed out in every pairing with beacon v0.2.10, baselines included, and passed in every pairing with beacon#61.

What to know before merging

  • Only an upgraded sender stops fragmenting. A reply of more than a packet from a node that is not upgraded is still lost on a path that drops fragments. Clients behind such a path get large replies from the fleet only once the fleet is upgraded.
  • A not-yet-upgraded node that echoes what it reads is slow. It reads 1152-byte pieces, writes each back as a write below its own 4096-byte segment size, and its own Nagle rule sends one per round trip. pilotctl bench from an upgraded node against one takes about 35 s for 1 MB at 30 ms (0.2 s on a LAN). Requests, replies, files and pub/sub are not affected; upgrading the echoing node removes it. I found no way to avoid this from the upgraded side.
  • The 128-segment limit caps the window at 147 KB, where it could reach 1 MB before (512 KB before the peer started dropping). On a clean path with a long round trip that is the ceiling: about 1.5 MB/s at 100 ms. Lifting it needs receivers with a larger reorder buffer and a way for the sender to know.
  • Relayed traffic is 3.5 times as many packets, and with the current beacon the fix is slower over the relay than main (5 MB in 5.8–6.6 s against 3.6 s). The relay reorders packets within a flow and limits each source to 1000 packets a second; smaller segments suffer more from both. relay: keep a short batch's remainder; per-source cap to 4000/s beacon#61 fixes them (1.7–1.8 s for the same transfer) and should be deployed before or with this.
  • An upgraded sender overruns an old receiver's socket buffer more. v1.14.1 and older use the default 229 KB socket buffer; the lab counted 1600–2000 dropped datagrams at the receiver across the clean-network checks. The transfers above include that and recover from it.
  • pub/sub is unreliable under loss in every version, this one included: at 5% loss 27 of 30 exchanges succeeded fix-to-fix and 26 of 30 main-to-main. Not changed here.
  • send-message: payload from stdin or a file; messages over 256 KB are delivered; an unacknowledged send fails #501 adds nagleWritePiece = 16 * MaxSegmentSize. Whichever of the two merges second should make that a multiple of SendSegmentSize.

Tests

  • New unit tests: datagram sizes read off a real socket; the segment-counted window; SACK accounting in fast recovery; the 128-segment limit; ACK-clocked timeout recovery; the tail rules; coalescing behind a held write; a held write going out ahead of the FIN; the quick-ACK budget.
  • New end-to-end tests: first write of a few segments on a connection; a stream of 4 KB writes.
  • Twelve existing tests changed. Three segmentation tests expect 1152-byte segments; four fast-recovery tests expect the window to grow by 1152 per duplicate ACK; the window-update wake test advertises two free slots (one is taken by the segment in flight); the delayed-ACK scheduling test starts with the quick-ACK budget spent; three Nagle tests asserted that the writer blocks on a held write and now assert that it does not.
  • go vet ./... is clean. Unit tests pass, also under -race. The integration package (./tests, local only) passes on a quiet machine. Run while the lab was loading the machine it failed one test in each of two runs, TestWaitForTrustBlocksUntilApproved and TestHandshakeTrustLoadVerify; the second fails the same way on main under that load and is fixed in tests: wait for the trust file instead of reading it the moment trust appears #508, and the first was a registry port collision when I reran it.

🤖 Generated with Claude Code

…s that work with them

A full stream segment was 4096 bytes: about 4.2 KB as a UDP datagram and
three IP fragments on a 1500-byte path. NATs, firewalls and some virtual
networks drop fragments, so on those paths handshakes, pings and messages
under ~1.4 KB worked while every full segment was lost — larger replies
never arrived and send-file failed.

Segments are now at most SendSegmentSize (1152) bytes: a 1222-byte
datagram direct, 1231 relayed, within the 1232 that fit IPv6's minimum
MTU. Sender-side only; receivers accept up to MaxSegmentSize as before, so
old and new nodes interoperate in both directions.

Smaller segments exposed faults the large ones had hidden. They are on
main today and are fixed here because the change is not safe without them.

Window accounting
- The peer's window, advertised in segments, is enforced in segments.
  Counting bytes only, short segments overran the peer's receive slots.
- The congestion window's growth and floors stay in MaxSegmentSize units;
  only adjustments that stand for one segment leaving the network use
  SendSegmentSize. Scaling everything down made a lossy path markedly
  slower (5 MB at 30ms RTT, 0.5% loss: 17s against 10s).

Loss recovery (on main, 2 of 5 5 MB transfers failed at 2% loss)
- In fast recovery SACKed segments were left out of the amount in flight
  while the window also grew per duplicate ACK, counting each twice; the
  amount outstanding doubled every round trip.
- The peer buffers at most MaxOOOBuf (128) out-of-order segments and
  drops the rest silently. The sender now keeps at most that many
  unacknowledged (MaxSegmentsOutstanding).
- After a retransmission timeout each further missing segment waited for
  its own timeout, which doubles each time. A partial ACK in timeout
  recovery now retransmits the next missing segment at once.

Short writes
Nagle held a short segment until everything before it was ACKed and
blocked the writer until it had gone out. Nothing could join a held write,
a stream of short writes moved at one write per round trip, and because a
client's IPC sends are handled in order, its other requests waited too.
With 1152-byte segments every 4 KB write ends in a short tail.
- The tail of a write of a segment or more leaves with it.
- A write shorter than a segment waits only for an earlier short segment.
- A held write does not block: it stays buffered, later writes join it,
  and flushHeldTail sends it when that segment is ACKed, after NagleTimeout
  at the latest, or CloseConnection sends it ahead of the FIN. SendMu
  keeps segments in write order whichever goroutine sends them.
- A receiver ACKs a lone segment at once for the first QuickACKBudget (32)
  segments of a connection and again after each time its delayed-ACK
  timer has had to fire, so a peer that waits for an ACK per small write
  (any peer up to v1.14.1) is not stalled for the timer each time.

Measured in a container lab (details in the pull request):
  5 MB over a 30ms path, five runs each, before -> after
    2% loss:   9.7-12.4s with 2 of 5 failed -> 10.3-11.8s, none failed
    0.5% loss: 9.4-10.5s -> 7.3-7.6s
    no loss:   8.2-8.4s -> 7.6-7.7s
  60 MB on one host: 4.8s -> 1.2s; into v1.13.5: 1.8s (v1.13.5 to itself
  22.7s)
  with non-first fragments dropped at the receiver: messages over one
  packet and file transfers failed -> all delivered, no fragments sent
Interop was checked both ways against v1.13.5, v1.14.1 and main. Two
costs: a node that is not upgraded and echoes back what it reads now
sends one 1152-byte piece per round trip (its own Nagle rule); and the
current beacon reorders relayed packets and caps packets per source,
which costs smaller segments more (5 MB relayed: 5.8s against 3.6s)
until pilot-protocol/beacon#61 is deployed (1.8s).

Twelve tests pinned the old behaviour and are updated: three segmentation
tests (1152-byte segments), four fast-recovery tests (window grows by
1152 per duplicate ACK), the window-update wake test (two free slots, one
taken by the segment in flight), the delayed-ACK scheduling test (starts
with the quick-ACK budget spent), and three Nagle tests that asserted the
writer blocks on a held write.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@TeoSlayer
TeoSlayer merged commit 1fbd7e3 into main Oct 5, 2026
14 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant