Repository navigation
docs(benchmarks): re-measure the draft block width after the #1939 fix - #1945
Merged
Merged
Conversation
…1939 prefill fix Issue #1797 seeded a measured default of 4 for `(12, 1, Affine)` on a DFlash drafter from a sweep taken before PR #1939 fixed the burst's prompt prefill. That fix changes the state every round starts from, which changes acceptance, which changes throughput per width, and nobody had looked since. The same harness, the same pairing, widths 2, 3, 4 and 6 at n = 3 with the classic brackets at both ends, on a binary built from `origin/main` and verified to carry none of the issue #1935 branch's gate. The ordering moved. Width 3 runs at 1.19x classic and its slowest run clears width 4's fastest by 0.80 tok/s, twice the session's full bracket spread, with the drift working against width 3 rather than for it since width 4 ran later. Width 2 and width 4 overlap and this sweep does not order them. Width 6 is below every classic run. Acceptance falls monotonically as the block widens (0.754, 0.599, 0.511, 0.432) against a per-round device-sync cost that rises (19.9, 24.8, 30.3, 47.4 ms), and the product of those peaks at 3 rather than at 4. The record states the drift rather than printing a clean table over it: the two classic brackets do not overlap, by about half a percent, and that is the resolution floor for everything between them. Nothing is differenced against the #1797 table, which ran on a different binary and a different tree. This record does not change the default; it files the evidence so the number can be re-examined against 3. Refs #1797, #1935
…ill matters The reading section quoted the classic brackets' full spread (0.40 tok/s) in one sentence and the shift between their means (0.27) in the next without naming either, which invites a reader to add them. Both now say which they are. And the closing paragraph said the width was academic for this pairing because PR #1944's gate declines the burst on CUDA. It is not: `MLXCEL_MTP_ALLOW_INEXACT=1` is a documented escape and an operator who sets it gets exactly this curve, so the default still decides what they run at. The paragraph now says plainly that on this evidence 4 is not the right number for this pairing on this host and 3 is. Refs #1797, #1935
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Issue #1797's measured draft-block-width default of 4 for
(12, 1, Affine)on a DFlash drafter came from a sweep taken before PR #1939 fixed the DFlash burst's prompt prefill. That fix changes the KV and gated-delta state every round starts from, which changes acceptance, which changes throughput per width. Re-measured on current main with the same harness and the same pairing: the ordering moved, and width 3 is faster than the shipped 4.What the sweep says
Width 3 runs at 1.19x classic and its slowest run clears width 4's fastest by 0.80 tok/s, which is twice the session's full classic-bracket spread, and width 4 ran after width 3 so the drift works against width 3 rather than for it. Width 2 and width 4 overlap and this sweep does not order them. Width 6 is below every classic run. Acceptance falls monotonically as the block widens (0.754, 0.599, 0.511, 0.432) against a per-round device-sync cost that rises (19.9, 24.8, 30.3, 47.4 ms), and the product of those peaks at 3.
The two classic brackets do not overlap, by about half a percent, and the record states that as the resolution floor rather than printing a clean table over it. Nothing is differenced against the #1797 table, which ran on a different binary and a different tree.
This PR changes no default. It files the evidence so the number can be re-examined against 3, which belongs in its own issue.
Test plan
sweep_server_widths.pyfrom the perf(speculative): the default draft block width of 16 loses on GB10 #1797 harness, widths 2, 3, 4 and 6, n = 3, classic brackets at both ends, on a binary built fromorigin/mainand verified to carry none of the issue fix(speculative): Qwen 3.5 DFlash verify diverges from classic decode #1935 branch's gate.NV_ERR_NO_MEMORYcount 0 before and after, per arm.Refs #1797, #1935