Skip to content

[Memory] Bound SM70 prefill scratch and compact hybrid cache pages - #690

Draft
yangzhuxinyzx wants to merge 2 commits into
mainfrom
codex/v100-memory-budget-20260925-132848
Draft

yangzhuxinyzx wants to merge 2 commits into
mainfrom
codex/v100-memory-budget-20260925-132848

Conversation

@yangzhuxinyzx

@yangzhuxinyzx yangzhuxinyzx commented Sep 25, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

Reduce SM70 warmup fragmentation and large-chunk NVFP4 prefill scratch, and avoid padding hybrid Mamba states to the larger FP16 draft-KV page. The 64K TP1/C2 NVFP4+DFlash2 configuration currently fails normal GUM 0.93 admission; this Draft does not claim that capacity or model-quality gate is solved.

The first warmup uses a scoped allocator setting, restored before profiling and serving. Gated prefill tiles independent output columns while preserving the full K reduction, and large non-gated projections reuse smaller dense-weight allocations. Compatible Flash-V100 sliding-window pages shrink without changing KV dtype or window. No TP-count or model-name gate is introduced; small prefill blocks keep the measured faster existing implementation.

Test Plan

  • Source-build native extensions without wheels, external kernel DSOs or preloads.
  • Check effective weights, changed-input CUDA Graph replay, allocator restoration and hybrid page capacities.
  • Compare exact-shape Graph time and allocation peaks with the unmodified source control.
  • Run normal single-V100 startup and cold long requests with Graph enabled. Require matched serving quality/speed and TP2/TP4 checks before promotion.

Test Result

  • 29 focused CUDA/allocator tests passed; five focused KV tests passed. Full KV utility suite: 69 passed and one existing DeepSeek fixture failure, reproduced on the base. All commit hooks passed.
  • TP1 M8192 gate/up: 676 MiB less scratch, 35.213 -> 34.242 ms. TP2/TP4 local shapes: +0.38%/-0.28% latency. TP1 down projection: 84 MiB less scratch, +0.72% latency. These are operator results, not serving throughput.
  • Measured model activation peak: 2.254 -> 1.594 GiB. Required 64K KV capacity: 6.000 -> 4.28125 GiB. GUM 0.93 still lacks about 1.4 GiB.
  • GUM 0.98 starts with a calibrated 512 MiB Graph reserve (actual capture 414 MiB), but ordinary-allocator cold requests OOM. Keeping max_split_size_mb:20 during serving at GUM 0.98 then passes cold 32768/64512-token retrievals with correct complete answers and natural EOS. This is an explicit runtime configuration, not a new source default. A 32768-input/256-output C1 vllm bench serve request completes with TTFT 35.049 s, TPOT 9.042 ms and zero failures; prefix cache was reset. It is not a 256K result.
  • Numerical replay checks pass, but tiled cuBLAS output is not bitwise identical. Dataset-level quality and a matched end-to-end <=3% speed gate remain pending.

See docs/design/sm70_memory_budget.md and the added standalone benchmark for details and rejected variants. Measured base: d49e32b3587d4d34ffccb0ffd376e63974b06c88; implementation: f39f7099fb. Main advanced to fcf59f8e9ae50c186333e98e5cf6aae705f320de during the task and has been merged into the owned branch (664e657b2a), preserving both worklog sections. The successful serving measurement remains tied to f39f7099fb; serving with the new DFlash defaults needs separate validation. After integration, the owned native extensions rebuild successfully, the 29 focused CUDA/allocator and five KV tests pass again, and no private runtime DSO dependency appears in readelf. Task-owned serving processes are stopped.

AI assistance: implementation, review and measurements were performed with Codex. Human review and promotion gates remain pending.

AI-assisted implementation and validation with Codex. Keep serving promotion pending the documented capacity and quality gates.

Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
@Leonccaa

Copy link
Copy Markdown
Contributor

Adding measurements from Qwen3.8-Flash-Next that bear on the SM70 prefill scratch budget, in case they are useful here.

Setup: 4x V100 32 GB, TP4 with --decode-context-parallel-size 2 (#696), AWQ W4A16 target, E4M3 main/draft KV, native MTP3 draft with FP8 experts, max_num_seqs=4, max_model_len=262144, max_num_batched_tokens=8192, expandable_segments:True. Code is main 1e90d17f2c + #705/#664/#696 + #707.

The profiling run never sees a 4-way batched prefill. With 1,600-token state blocks, four long prompts prefill 4 x 1,600 tokens per step. At gpu_memory_utilization=0.96 this killed the EngineCore twice:

  1. In the PLE short-conv batched prefill, whose dense padded copies peaked at 758 MiB for the op. [Perf][Qwen4Exp] Hold two batch-sized buffers in the batched PLE short-conv prefill #707 cuts that to 254 MiB with bitwise-identical outputs.
  2. With [Perf][Qwen4Exp] Hold two batch-sized buffers in the batched PLE short-conv prefill #707 applied, in the vision encoder: a 144 MiB all-reduce buffer for a 4096x4096 image that arrived while three ~60K-token prompts were prefilling.

What is live at the peak now. This is torch CUDA memory history on one rank, counting only allocations made during a 4 x 1,600-token prefill step (#707 applied):

allocation site MiB
MTP draft FP8 MoE prefill scratch, fp8_sm70_moe.py _get_buffers: 312.5 + 312.5 + 62.5 + 31.2 + 31.2 ~750
TP all_gather output 125
hc.py _hc_combine_norm 125
rest (_hc_gate_mix, shared experts, QSA) ~130
total 1,130

_get_buffers keeps persistent buffers only up to _DEFAULT_PERSISTENT_MAX_TOKENS = 32. Every prefill therefore allocates [num_tokens * top_k, hidden]-sized temporaries; at 6,400 tokens the permuted input and the sorted output are 312.5 MiB each. Running this path in token chunks, like VLLM_FUSED_MOE_CHUNK_SIZE does for the modular fused-MoE kernel, would take it to ~120 MiB at 1K-token chunks. Draft numerics only change the proposals, not the verified output.

The reservation grows more than the live peak. From nvidia-smi at gmu 0.93:

  • idle: 30,074 MiB;
  • after the 4-way text prefill: 31,436 MiB (+1,362);
  • with the 4096x4096 image on top: 31,696 MiB (+1,622).

The live peak is 1,130 MiB, so the difference looks like the fragmentation you are addressing with max_split_size_mb.

For this configuration the measured-safe setting today is 0.945, which leaves about 295 MiB under the image + prefill worst case. Reaching 0.96 needs roughly the FP8 MoE scratch above removed. I am happy to send the SM70 FP8 MoE chunking as a separate PR if that does not overlap with your plans for this one.

@yangzhuxinyzx

Copy link
Copy Markdown
Contributor Author

I have extracted the independent sliding-window cache-page compaction into #829 and continued validation on current main. All 71 KV-cache CPU tests and the shared SM70 gate passed. Dtype, window size and recurrent-state capacity are preserved; exact-versus-partial-page native comparisons are pending an available test slot.

The tiled GEMM and allocator parts remain here for separate review and validation. This is a collaborative continuation rather than waiting for a new revision from the original author.

yangzhuxinyzx added a commit that referenced this pull request Oct 3, 2026
Extract the exact cache-geometry change from #690 while preserving dtype, attention window and recurrent-state capacity.

Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
yangzhuxinyzx added a commit that referenced this pull request Oct 3, 2026
Compact compatible Flash-V100 draft sliding-window cache pages before hybrid unification. This avoids growing recurrent-state padding solely to match larger FP16 draft pages, preserving dtype, attention window and state capacity.

Validation: 130 CPU merge-gate controls, scoped pre-commit and PR CI passed on latest main. Six V100 native comparisons produced bit-identical outputs for 1024/2048 token pages at exact and partial-page boundaries across 1/2/4 KV heads. No kernel arithmetic changes are included.

Unverified: full-model serving throughput and capacity for this isolated change. The tiled GEMM and allocator portions of #690 remain separate.
carrey-feng pushed a commit to carrey-feng/1Cat-vLLM that referenced this pull request Oct 3, 2026
Extract the exact cache-geometry change from 1CatAI#690 while preserving dtype, attention window and recurrent-state capacity.

Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
@yangzhuxinyzx

Copy link
Copy Markdown
Contributor Author

The compile-warmup allocator fix is being integrated independently in #840. The existing scoped setting is restored before memory measurement, on explicit KV budgets and on exceptions. The 95 CPU checks and scoped pre-commit pass; hardware-dependent peak-memory tests were skipped for this revision. Together with the hybrid-page portion already merged through #829, this preserves the changes that do not alter projection arithmetic. The tiled cuBLAS portion remains here for separate numerical and quality validation.

yxuef71 pushed a commit to yxuef71/1Cat-vLLM that referenced this pull request Oct 3, 2026
Scope compile-time allocation settings and restore the complete configuration before steady-state profiling or returning an explicit KV budget. Extracts the allocator lifetime fix from 1CatAI#690 without changing projection arithmetic.

Co-authored-by: OpenAI Codex <noreply@openai.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants