Skip to content

[Core][SM70] Decouple the Mamba state grid from the KV block size - #645

Merged
yangzhuxinyzx merged 1 commit into
mainfrom
codex/v100-75t-prefix-bridge-20260916
Sep 16, 2026
Merged

yangzhuxinyzx merged 1 commit into
mainfrom
codex/v100-75t-prefix-bridge-20260916

Conversation

@yangzhuxinyzx

Copy link
Copy Markdown
Contributor

What

The long-prefill chunk is min(max_num_scheduled_tokens, mamba_state_block_size), and the 75T Q8000 dense prefill route only dispatches for a chunk in [8000, 8192]. So the Mamba recurrent-state grid -- not the token budget -- decided whether the fast prefill path was reachable at all. Page-size unification scaled that grid together with the KV block size, which tied the fast path to a 4096 block and forced a choice between fast prefill and KV capacity.

This honours an explicit --mamba-block-size in align mode, and stops unify_kv_cache_spec_page_size multiplying a MambaSpec's block_size by the page-unification ratio when one is given. The physical page is still padded to max_page_size, and align-mode state memory is page_size_bytes * (2 + num_speculative_blocks), which does not depend on block_size -- so the change is memory-neutral. Both edits are inert unless the flag is passed.

Result

Qwen3.8-27B NVFP4, TP4, fp8_e4m3 KV, DFlash2, unique-salt 100K prompt:

--block-size --mamba-block-size grid chunk 75T @100k KV pool
4096 (shipped) - 8192 8192 yes 3426 818,142
2048 - 4096 4096 no 2954 1,058,133
2048 8192 8192 8192 yes 3318-3336 1,058,133-1,108,650

The default path is unchanged, re-measured with --block-size 4096 and no flag: grid 8192, chunk 8192, 75T dispatch, 858,310 KV tokens, 3272 tok/s.

Rejected: --block-size 1648

1648 is the smallest legal block and measured the largest pool (1,112,091) and the fastest prefill (3551 tok/s), but it does not divide 8192, so no prefix length is simultaneously block-aligned (KV blocks) and grid-aligned (recurrent state). Result: prefix_cache_queries_total = 812,040 against hits_total = 0.0, with an identical repeated 232K-token prompt still taking 86.5 s. The alignment assertion is retained for that reason; the in-tree example in scheduler.py obeys the same rule (816 = 51 x 16).

Note on prefix-cache reuse

Reuse is quantised by the grid, not the block: at grid 8192 a byte-identical prompt below 8192 tokens gets no reuse at all. Since chunk <= grid, the reuse quantisation, the chunk size and the grid are one and the same knob, so the smaller block buys KV capacity only and not finer reuse. That is inherent to align-mode hybrid scheduling, and is documented rather than worked around.

Evidence

  • docs/design/sm70_mamba_state_grid_decoupling.md -- derivation, raw probe output, reuse-granularity table, adopted default.
  • docs/design/sm70_v100_migration_control.md -- control entry.
  • scripts/serve_qwen38_27b_nvfp4_v100.sh -- adopted default launcher for this model.

Verification

Re-tested on the committed revision, launched through the canonical script: 100K prompt at 3318 tok/s, 75T dispatch, 1,108,650 KV tokens, prefix_cache_hits_total = 32,768 at 8192 granularity, and cache-hit greedy outputs identical to the cold outputs at every size tested.

The long-prefill chunk is min(max_num_scheduled_tokens,
mamba_state_block_size), so the recurrent-state grid -- not the token budget --
decided whether the 75T Q8000 prefill route (chunk in [8000, 8192]) was
reachable at all. Page-size unification scaled that grid together with the KV
block size, which tied the fast prefill path to a 4096 block.

Honour an explicit --mamba-block-size in align mode, and stop multiplying a
MambaSpec's block_size by the page-unification ratio when one is given. The
physical page is still padded to max_page_size, and align-mode state memory is
page_size_bytes * (2 + num_speculative_blocks), which does not depend on
block_size, so the change is memory-neutral. Both edits are inert unless the
flag is passed, so existing deployments are unaffected: --block-size 4096 still
yields an 8192 grid, an 8192 chunk and an 858,310-token pool.

--block-size 2048 --mamba-block-size 8192 now gives an 8192 chunk with the 75T
route and a 1,058,133-token pool, 23-29% more than the 4096 block, at equal
prefill (3336 versus a 3295-3426 tok/s band on a unique-salt 100K prompt).

The grid must remain a multiple of the block size. Otherwise no prefix length is
simultaneously block-aligned (KV blocks) and grid-aligned (recurrent state), and
prefix caching silently drops to zero hits: measured 812,040 queries against 0
hits at --block-size 1648 with an 8192 grid, with an identical repeated
232K-token prompt still taking 86.5 s. The assertion is retained for that
reason.

Prefix-cache reuse is quantised by the grid, not the block, so the smaller block
buys KV capacity only and not finer reuse; that is recorded in the design note
along with the measured probes.

Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
@yangzhuxinyzx
yangzhuxinyzx force-pushed the codex/v100-75t-prefix-bridge-20260916 branch from 2153e4f to f5dbb36 Compare September 16, 2026 17:13
@yangzhuxinyzx
yangzhuxinyzx merged commit b711d53 into main Sep 16, 2026
SabaTech-dev added a commit to SabaTech-dev/1Cat-vLLM that referenced this pull request Sep 18, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant