Conversation
Add per-layer calibrated E4M3 K/V storage to Qwen4Exp QSA for the fixed SM70, TP4, MTP0, and FP16 activation contract. The main QSA cache uses E4M3 while the raw and compressed indexer caches remain FP16. Design decisions: - Keep checkpoint loading slots at an invalid -1.0 sentinel, separate from runtime scale buffers whose upstream default remains 1.0. Finalization validates and deletes the loading slots. - Fail closed if any of the 12 layers' K/V scales is missing, non-finite, or non-positive. This intentionally replaces the upstream KV-cache warning and unit-scale fallback, and is confined to 1Cat-owned Qwen4Exp files. - Keep indexer caches FP16 so K/V quantization cannot change QSA block selection. - Use P256 for E4M3 on generic G6 XQA virtual-page4. The P1024 specialization assumes the page1568 layout and is not valid here. - Do not bundle a scale overlay. Generate an exact-checkpoint-revision- bound model-kvscales overlay with tools/qwen4_exp/qsa_kv_calibration.py; E4M3 startup rejects unless all 24 scales are explicitly loaded. Also add offline normal-inference scale collection, report and overlay generation, and matched route/selected-block comparison. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
Apply the pinned CUDA formatter, add the required SPDX copyright header, and express the reference attention with equivalent matmul operations so the tensor subscripts do not trigger the spelling hook. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
E4M3 XQA accepts at most 16 query rows, while chunked scheduling can form non-grouped catch-up/decode batches above that limit. Route the divisible bulk through grouped page4 and split the remainder into supported XQA chunks instead of falling back to the larger Triton split-K workspace. Add routing coverage and an SM70 numerical regression for the observed 46+1+1+1 request layout.
This was referenced Sep 2, 2026
The dedicated PLE offload process builds the full model structure on meta but intentionally loads only the PLE subtree. Skip QSA KV scale finalization in that process; it never runs QSA forward, while every GPU worker retains the existing 24/24 fail-closed gate.
Leonccaa
marked this pull request as ready for review
September 2, 2026 21:20
Contributor
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
This PR adds calibrated per-layer E4M3 K/V storage to the main Qwen4Exp
QSA cache for the first validated contract:
Overlay/checkpoint revision binding and calibration provenance are offline
artifact responsibilities. The runtime gate deliberately answers only whether
the 24 expected scales were explicitly loaded and are valid.
Relationship to vLLM #54846
vllm-project/vllm#54846
has the same broad goal but a different hardware and deployment contract. It
targets SM120 with BF16 activations, adds generic Triton E4M3 and NVFP4 KV
reads, and explicitly leaves calibrated per-layer scales out of scope.
This PR targets 1Cat's SM70 path. Its E4M3 decode is not interchangeable with
the SM120 implementation: SM70 does not have the native Triton
fp8e4nvconversion used there, so this path reads byte storage and performs software
E4M3 bit decoding. It also covers the fork-specific Flash-V100 grouped-page4
and XQA-page4 paths. NVFP4 KV storage is intentionally outside this first
phase. Both implementations keep the indexer cache unquantized.
Diff anatomy
The corpus builders and frozen-token quality runner were split into a
stacked draft PR in the submitter's fork
so this PR can review the runtime and scale contract directly. That follow-up
is optional reproducibility tooling and is not a build, runtime, or overlay
dependency of this PR.
+509/-97+1084/-4+815/-0+105/-9+2513/-110, 17 filesNon-obvious design decisions
at the invalid
-1.0sentinel, while runtime_k_scale/_v_scalebuffers retain the upstream default
1.0. This follows the upstreamkv_cache.pyshape: finalization validates the loading slots, copies theminto runtime state, then deletes the slots.
1.0remains a legal explicitscale and is not a missing-scale signal.
continues with unit scale when scales are absent. This path raises instead.
The intentional divergence is contained in the Qwen4Exp/Flash-V100-owned
path rather than changing the generic loader contract.
and compressed indexer caches stay FP16 so quantization is not introduced
directly into the block-selection cache.
assumes the page1568 layout and is not valid for the generic E4M3 layout.
checkpoint-revision-bound overlay with
qsa_kv_calibration.py; otherwisethe startup gate rejects E4M3.
Why calibrated scale is part of this contract
Unit scale did not saturate on either the upstream SM120 report or this SM70
calibration. The difference in this deployment contract is low-magnitude
coverage. A conservative upper-bin analysis of 30,953,961,970 nonzero K/V
elements from the formal long-context page4 collection found, at scale 1.0:
0.2040%below E4M3's minimum subnormal magnitude (2^-9), and1.4235%in the subnormal range below2^-6.The calibrated scales (
0.0171247218to0.0841238871) move these values intoa better-resolved range while retaining zero observed saturation for all 24
tensors. This is a measurable mechanism for calibration under this exact
checkpoint/activation/topology contract; it is not a claim that unit scale
must cause a large downstream quality loss on every checkpoint.
Implementation scope
Triton reads, Flash-V100 grouped page4, and XQA reads.
never executes QSA forward; all GPU workers retain the 24/24 fail-closed
gate.
grouped-page4 bulk and XQA tail chunks of at most 16 rows, avoiding both an
unsupported E4M3 XQA launch and the larger Triton split-K workspace.
calculate_kv_scalesin calibrated E4M3 mode.engine initialization, excluding dummy/profile/graph-capture forwards.
p99.9, p99.99, saturation, and per-shard stability.
No private corpus or scale overlay is included in the repository. Generic
long/tool corpus builders and the frozen-token API quality runner are isolated
in Leonccaa/1Cat-vLLM#6.
Test Plan
Static checks over the PR's Python diff:
Public regression set, run in the SM70 CUDA build on V100:
End-to-end quality used the same checkpoint, exact prompt token IDs, and
deterministic sampling for FP16 KV and calibrated E4M3 KV. An explicitly
loaded unit-scale overlay was used only as a negative control. Calibration and
quality inputs were disjoint and covered Chinese, English, code, tool calls,
multi-turn data, and 1K/16K/64K/128K context lengths.
Test Result
Automated checks
CI-format fixes.
clang-format,typos, andcheck-spdx-header: passed for the files thatfailed the first pre-commit run.
git diff --check upstream/main: passed.pre-commit over the three changed files, including mypy, passed.
production startup matrix separately confirmed missing overlay
0/24rejection, calibrated
24/24loading on all GPU workers, and successful PLEworker startup.
(
46+1+1+1requests): grouped 48 + XQA 1 matched the E4M3 Tritonreference with max absolute error
1.52587890625e-05and mean absoluteerror
1.5917149767e-06.Startup and calibration gates
Loaded 0/24, startup rejected.1.0remains alegal value, not an error sentinel.
calculate_kv_scalesis rejected for this calibrated E4M3 contract.because it loads only PLE tensors and never runs QSA forward; the GPU-worker
startup gate remains unchanged.
0.0171247218to0.0841238871;observed saturation ratio was zero for every K/V tensor.
TP4 / MTP0 quality evidence
704,741 E4M3 tokens (
1.83x). A 128,001-token request's reportedconcurrency increased from
3.01xto5.51x.all 96 expected per-layer dumps were paired.
FP16 and calibrated E4M3 both scored 6/6, with token-identical paired
outputs.
10/12, with zero correctness regressions, identical first tokens for
12/12 comparisons, and exact repeat stability for 6/6 cases. Both modes
failed the same baseline case.
0.0324132, cosine0.999311,RMSE
0.0160753.0.990783, mean FP16 recall0.995301, exact-row ratio0.570984, and micro Jaccard0.987579.The first QSA layer (layer 3) remained exactly identical on all four ranks;
later differences are downstream of earlier quantized attention outputs.
on the six single-run tool cases. This small sample is supporting evidence,
not the primary justification for calibration.
With
max_num_batched_tokens=4096, complete chunks in these quality runs usedthe grouped Flash-V100 page4 path, while tails/decode used generic Triton. The
logs showed no XQA entry, so the FP16 P1024 versus E4M3 P256 XQA partition
difference did not confound the reported A/B. A future experiment that enters
XQA must add an FP16-at-P256 control arm or explicitly report reduction order
as a second variable.
Targeted mixed-batch scheduler regression
At
gpu_memory_utilization=0.89, C4 at 64K produced the scheduler layout thatexposed the issue: a 49-row batch consisting of a 46-row catch-up chunk plus
three decode rows. The fixed path routed this as grouped 48 + XQA 1. Warmup
and all 3/3 scored repeats (12/12 requests) completed without CUDA OOM or
engine failure. This closes the observed mixed-batch correctness/capacity bug;
it is not the formal matched performance matrix.
NVFP4 formal performance update (2026-09-02)
The matched SM70 / TP4 / MTP0 / GMU 0.89 matrix is complete on this PR plus the two isolated Triton optimizations now split across #452 and #453: 17 E4M3 cells passed, C8 x 128K was the sole expected capacity skip, and all 87 scored requests completed. At the same memory budget, KV capacity increased from 507,093 FP16 tokens to 931,100 E4M3 tokens (1.836x). Across 15 matched cells, E4M3 decode throughput was 1.92% lower and server E2E elapsed time was 1.26% higher by geometric mean; at 16K and longer the corresponding figures were -2.94% and +3.38%. This validates the fixed first-phase performance contract as a capacity win with a modest long-context speed cost, not a global speedup. The sibling PR descriptions contain the complete matrix, caveats, and split-specific diffs.
The full matrix used the source tree of #447 plus the two optimizations before their review-only split into #452 and #453; that exact tree is preserved by the fork tag
evidence/qsa-e4m3-perf-combined-20260902. The split combined tree binds the same bitcast decoder and changes only its local imported symbol name. The core-only final route was separately checked at C4 x 64K and was effectively neutral versus the optimized tree (-0.32% decode, +0.21% E2E), but the core-only tree was not rerun across all 18 cells.AWQ formal-r6 validation update (2026-09-02)
A second end-to-end campaign validated the same QSA FP8 contract with a
per-expert asymmetric AWQ W4A16 g32 checkpoint (zero point enabled, GEMM) and
an FP8 E4M3 PLE splice. The exact checkpoint index SHA-256 was
da62ba7bfef57f8117f35fc78bf5daf14c0ab5eb1b8173b4d712c1f4d294819b.The AWQ-specific MoE TP-shard alignment patch used by that checkpoint is
outside this PR; these results validate this PR's QSA E4M3 path under the AWQ
runtime composition rather than claiming that #447 alone adds AWQ loading.
Calibration used normal FP16-KV inference forwards over 4,888,771 processed
tokens spanning Chinese, English, code, multi-turn dialogue, tool use, and
1K/16K/64K/128K contexts. The 24 scales ranged from
0.0171072837to0.0806361660, with zero observed saturation. The AWQ/NVFP4 per-tensor scaleratio ranged from about
0.955to1.028, confirming that the NVFP4 overlaymust not be reused for AWQ.
Matched route-neutral quality results were:
0.995906, minimum0.991822;0.998597, maximum relative L20.052962.The fixed GMU 0.89 performance matrix covered concurrency 1/4/8 and
1K/4K/16K/32K/64K/128K with 256 output tokens and three scored repeats per
cell. FP16 completed 15/18 cells and E4M3 completed 17/18, with no failed cells.
Measured KV capacity increased from 310,480 to 569,579 tokens
(1.8345x). Over the nine matched cells at 16K and longer, E4M3 prefill was
3.42% lower, decode was 2.54% lower, and server E2E elapsed time was 3.12%
higher.
C4 x 128KandC8 x 64Kwere E4M3-only capacity cells; both modesskipped
C8 x 128K.This AWQ matrix used #447 plus both optional optimizations. FP16 P1024 versus
E4M3 P256 XQA and mixed-batch splitting remain a second variable, so these
figures characterize the deployed bundle rather than pure converter overhead.
The supported conclusion is a 1.8345x KV-capacity gain at about a 3%
long-context latency cost under the exact tested contract.
Pending before non-draft / broader enablement
production acceptance.
AI assistance disclosure
This PR was developed with OpenAI Codex assistance. Design and validation were
supervised by the human submitter. The PR remains draft pending maintainer review. TP8, MTP, prefix caching,
standalone graph-versus-eager parity, and 256K remain outside first-phase acceptance.
Checklist
tools/qwen4_exp/README.md.