[Memory] Bound SM70 prefill scratch and compact hybrid cache pages - #690
yangzhuxinyzx wants to merge 2 commits into
Conversation
AI-assisted implementation and validation with Codex. Keep serving promotion pending the documented capacity and quality gates. Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
|
Adding measurements from Qwen3.8-Flash-Next that bear on the SM70 prefill scratch budget, in case they are useful here. Setup: 4x V100 32 GB, TP4 with The profiling run never sees a 4-way batched prefill. With 1,600-token state blocks, four long prompts prefill 4 x 1,600 tokens per step. At
What is live at the peak now. This is torch CUDA memory history on one rank, counting only allocations made during a 4 x 1,600-token prefill step (#707 applied):
The reservation grows more than the live peak. From nvidia-smi at gmu 0.93:
The live peak is 1,130 MiB, so the difference looks like the fragmentation you are addressing with For this configuration the measured-safe setting today is 0.945, which leaves about 295 MiB under the image + prefill worst case. Reaching 0.96 needs roughly the FP8 MoE scratch above removed. I am happy to send the SM70 FP8 MoE chunking as a separate PR if that does not overlap with your plans for this one. |
|
I have extracted the independent sliding-window cache-page compaction into #829 and continued validation on current main. All 71 KV-cache CPU tests and the shared SM70 gate passed. Dtype, window size and recurrent-state capacity are preserved; exact-versus-partial-page native comparisons are pending an available test slot. The tiled GEMM and allocator parts remain here for separate review and validation. This is a collaborative continuation rather than waiting for a new revision from the original author. |
Extract the exact cache-geometry change from #690 while preserving dtype, attention window and recurrent-state capacity. Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Compact compatible Flash-V100 draft sliding-window cache pages before hybrid unification. This avoids growing recurrent-state padding solely to match larger FP16 draft pages, preserving dtype, attention window and state capacity. Validation: 130 CPU merge-gate controls, scoped pre-commit and PR CI passed on latest main. Six V100 native comparisons produced bit-identical outputs for 1024/2048 token pages at exact and partial-page boundaries across 1/2/4 KV heads. No kernel arithmetic changes are included. Unverified: full-model serving throughput and capacity for this isolated change. The tiled GEMM and allocator portions of #690 remain separate.
Extract the exact cache-geometry change from 1CatAI#690 while preserving dtype, attention window and recurrent-state capacity. Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
|
The compile-warmup allocator fix is being integrated independently in #840. The existing scoped setting is restored before memory measurement, on explicit KV budgets and on exceptions. The 95 CPU checks and scoped pre-commit pass; hardware-dependent peak-memory tests were skipped for this revision. Together with the hybrid-page portion already merged through #829, this preserves the changes that do not alter projection arithmetic. The tiled cuBLAS portion remains here for separate numerical and quality validation. |
Scope compile-time allocation settings and restore the complete configuration before steady-state profiling or returning an explicit KV budget. Extracts the allocator lifetime fix from 1CatAI#690 without changing projection arithmetic. Co-authored-by: OpenAI Codex <noreply@openai.com> Signed-off-by: yangzhuxinyzx <153831768+yangzhuxinyzx@users.noreply.github.com>
Purpose
Reduce SM70 warmup fragmentation and large-chunk NVFP4 prefill scratch, and avoid padding hybrid Mamba states to the larger FP16 draft-KV page. The 64K TP1/C2 NVFP4+DFlash2 configuration currently fails normal GUM 0.93 admission; this Draft does not claim that capacity or model-quality gate is solved.
The first warmup uses a scoped allocator setting, restored before profiling and serving. Gated prefill tiles independent output columns while preserving the full K reduction, and large non-gated projections reuse smaller dense-weight allocations. Compatible Flash-V100 sliding-window pages shrink without changing KV dtype or window. No TP-count or model-name gate is introduced; small prefill blocks keep the measured faster existing implementation.
Test Plan
Test Result
max_split_size_mb:20during serving at GUM 0.98 then passes cold 32768/64512-token retrievals with correct complete answers and natural EOS. This is an explicit runtime configuration, not a new source default. A 32768-input/256-output C1vllm bench serverequest completes with TTFT 35.049 s, TPOT 9.042 ms and zero failures; prefix cache was reset. It is not a 256K result.See
docs/design/sm70_memory_budget.mdand the added standalone benchmark for details and rejected variants. Measured base:d49e32b3587d4d34ffccb0ffd376e63974b06c88; implementation:f39f7099fb. Main advanced tofcf59f8e9ae50c186333e98e5cf6aae705f320deduring the task and has been merged into the owned branch (664e657b2a), preserving both worklog sections. The successful serving measurement remains tied tof39f7099fb; serving with the new DFlash defaults needs separate validation. After integration, the owned native extensions rebuild successfully, the 29 focused CUDA/allocator and five KV tests pass again, and no private runtime DSO dependency appears in readelf. Task-owned serving processes are stopped.AI assistance: implementation, review and measurements were performed with Codex. Human review and promotion gates remain pending.