[Core][SM70] Decouple the Mamba state grid from the KV block size - #645
Merged
Merged
Conversation
The long-prefill chunk is min(max_num_scheduled_tokens, mamba_state_block_size), so the recurrent-state grid -- not the token budget -- decided whether the 75T Q8000 prefill route (chunk in [8000, 8192]) was reachable at all. Page-size unification scaled that grid together with the KV block size, which tied the fast prefill path to a 4096 block. Honour an explicit --mamba-block-size in align mode, and stop multiplying a MambaSpec's block_size by the page-unification ratio when one is given. The physical page is still padded to max_page_size, and align-mode state memory is page_size_bytes * (2 + num_speculative_blocks), which does not depend on block_size, so the change is memory-neutral. Both edits are inert unless the flag is passed, so existing deployments are unaffected: --block-size 4096 still yields an 8192 grid, an 8192 chunk and an 858,310-token pool. --block-size 2048 --mamba-block-size 8192 now gives an 8192 chunk with the 75T route and a 1,058,133-token pool, 23-29% more than the 4096 block, at equal prefill (3336 versus a 3295-3426 tok/s band on a unique-salt 100K prompt). The grid must remain a multiple of the block size. Otherwise no prefix length is simultaneously block-aligned (KV blocks) and grid-aligned (recurrent state), and prefix caching silently drops to zero hits: measured 812,040 queries against 0 hits at --block-size 1648 with an 8192 grid, with an identical repeated 232K-token prompt still taking 86.5 s. The assertion is retained for that reason. Prefix-cache reuse is quantised by the grid, not the block, so the smaller block buys KV capacity only and not finer reuse; that is recorded in the design note along with the measured probes. Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
yangzhuxinyzx
force-pushed
the
codex/v100-75t-prefix-bridge-20260916
branch
from
September 16, 2026 17:13
2153e4f to
f5dbb36
Compare
SabaTech-dev
added a commit
to SabaTech-dev/1Cat-vLLM
that referenced
this pull request
Sep 18, 2026
…rid GDN), strategy
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What
The long-prefill chunk is
min(max_num_scheduled_tokens, mamba_state_block_size), and the 75T Q8000 dense prefill route only dispatches for a chunk in[8000, 8192]. So the Mamba recurrent-state grid -- not the token budget -- decided whether the fast prefill path was reachable at all. Page-size unification scaled that grid together with the KV block size, which tied the fast path to a 4096 block and forced a choice between fast prefill and KV capacity.This honours an explicit
--mamba-block-sizein align mode, and stopsunify_kv_cache_spec_page_sizemultiplying aMambaSpec'sblock_sizeby the page-unification ratio when one is given. The physical page is still padded tomax_page_size, and align-mode state memory ispage_size_bytes * (2 + num_speculative_blocks), which does not depend onblock_size-- so the change is memory-neutral. Both edits are inert unless the flag is passed.Result
Qwen3.8-27B NVFP4, TP4, fp8_e4m3 KV, DFlash2, unique-salt 100K prompt:
--block-size--mamba-block-sizeThe default path is unchanged, re-measured with
--block-size 4096and no flag: grid 8192, chunk 8192, 75T dispatch, 858,310 KV tokens, 3272 tok/s.Rejected:
--block-size 16481648 is the smallest legal block and measured the largest pool (1,112,091) and the fastest prefill (3551 tok/s), but it does not divide 8192, so no prefix length is simultaneously block-aligned (KV blocks) and grid-aligned (recurrent state). Result:
prefix_cache_queries_total = 812,040againsthits_total = 0.0, with an identical repeated 232K-token prompt still taking 86.5 s. The alignment assertion is retained for that reason; the in-tree example inscheduler.pyobeys the same rule (816 = 51 x 16).Note on prefix-cache reuse
Reuse is quantised by the grid, not the block: at grid 8192 a byte-identical prompt below 8192 tokens gets no reuse at all. Since
chunk <= grid, the reuse quantisation, the chunk size and the grid are one and the same knob, so the smaller block buys KV capacity only and not finer reuse. That is inherent to align-mode hybrid scheduling, and is documented rather than worked around.Evidence
docs/design/sm70_mamba_state_grid_decoupling.md-- derivation, raw probe output, reuse-granularity table, adopted default.docs/design/sm70_v100_migration_control.md-- control entry.scripts/serve_qwen38_27b_nvfp4_v100.sh-- adopted default launcher for this model.Verification
Re-tested on the committed revision, launched through the canonical script: 100K prompt at 3318 tok/s, 75T dispatch, 1,108,650 KV tokens,
prefix_cache_hits_total = 32,768at 8192 granularity, and cache-hit greedy outputs identical to the cold outputs at every size tested.