Skip to content

[Model][SM70][FP8] Add calibrated E4M3 QSA KV cache - #447

Closed
Leonccaa wants to merge 4 commits into
1CatAI:mainfrom
Leonccaa:feat/qwen4exp-qsa-e4m3-kv-core
Closed

Leonccaa wants to merge 4 commits into
1CatAI:mainfrom
Leonccaa:feat/qwen4exp-qsa-e4m3-kv-core

Conversation

@Leonccaa

@Leonccaa Leonccaa commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

This PR adds calibrated per-layer E4M3 K/V storage to the main Qwen4Exp
QSA cache for the first validated contract:

  • NVIDIA SM70 / V100
  • TP4, MTP0
  • FP16 activations; validated with both NVFP4 and per-expert AWQ W4A16 g32 checkpoints
  • E4M3 main QSA K/V cache
  • FP16 raw and compressed QSA indexer caches

Required external artifact: the code in this PR does not contain a
KV-scale overlay. fp8_e4m3 requires a model-kvscales.safetensors
overlay generated for the exact checkpoint revision with
tools/qwen4_exp/qsa_kv_calibration.py. The generator adds all 24 scale
keys to the derived checkpoint's standard model.safetensors.index.json
without modifying the base checkpoint. Starting E4M3 without all 24
explicitly loaded, finite, positive K/V scales fails closed.

Published validated scale pack: leoncca/Qwen3.8-Flash-Next-NVFP4-QSA-FP8-E4M3-KV-Scales at tag v1.1-base-7b719225 (the scale payload is unchanged from v1-base-7b719225; v1.1 refreshes the split performance-PR references). The repository contains only the 24 scale scalars, a public manifest/checksums, and a revision-checking materializer; it does not contain the base weights or a pre-materialized checkpoint. The materializer requires base revision 7b719225242aacd3dbd3f9407468c2ee9a9d2594 and merges the scale keys into the standard index locally.

Overlay/checkpoint revision binding and calibration provenance are offline
artifact responsibilities. The runtime gate deliberately answers only whether
the 24 expected scales were explicitly loaded and are valid.

Relationship to vLLM #54846

vllm-project/vllm#54846
has the same broad goal but a different hardware and deployment contract. It
targets SM120 with BF16 activations, adds generic Triton E4M3 and NVFP4 KV
reads, and explicitly leaves calibrated per-layer scales out of scope.

This PR targets 1Cat's SM70 path. Its E4M3 decode is not interchangeable with
the SM120 implementation: SM70 does not have the native Triton fp8e4nv
conversion used there, so this path reads byte storage and performs software
E4M3 bit decoding. It also covers the fork-specific Flash-V100 grouped-page4
and XQA-page4 paths. NVFP4 KV storage is intentionally outside this first
phase. Both implementations keep the indexer cache unquantized.

Diff anatomy

The corpus builders and frozen-token quality runner were split into a
stacked draft PR in the submitter's fork
so this PR can review the runtime and scale contract directly. That follow-up
is optional reproducibility tooling and is not a build, runtime, or overlay
dependency of this PR.

Area Diff
Runtime +509/-97
Tests +1084/-4
Calibration tooling and documentation +815/-0
Route/quality comparison benchmark +105/-9
Total +2513/-110, 17 files

Non-obvious design decisions

  • Separate loading and runtime scale state. Checkpoint loading slots start
    at the invalid -1.0 sentinel, while runtime _k_scale / _v_scale
    buffers retain the upstream default 1.0. This follows the upstream
    kv_cache.py shape: finalization validates the loading slots, copies them
    into runtime state, then deletes the slots. 1.0 remains a legal explicit
    scale and is not a missing-scale signal.
  • Fail closed only on missing/invalid scale loading. Upstream warns and
    continues with unit scale when scales are absent. This path raises instead.
    The intentional divergence is contained in the Qwen4Exp/Flash-V100-owned
    path rather than changing the generic loader contract.
  • Keep the indexer cache FP16. Only the main K/V cache is quantized. Raw
    and compressed indexer caches stay FP16 so quantization is not introduced
    directly into the block-selection cache.
  • Use P256 for generic E4M3 G6 XQA virtual-page4. The P1024 specialization
    assumes the page1568 layout and is not valid for the generic E4M3 layout.
  • Do not bundle an overlay. Operators must generate and distribute the
    checkpoint-revision-bound overlay with qsa_kv_calibration.py; otherwise
    the startup gate rejects E4M3.

Why calibrated scale is part of this contract

Unit scale did not saturate on either the upstream SM120 report or this SM70
calibration. The difference in this deployment contract is low-magnitude
coverage. A conservative upper-bin analysis of 30,953,961,970 nonzero K/V
elements from the formal long-context page4 collection found, at scale 1.0:

  • 0.2040% below E4M3's minimum subnormal magnitude (2^-9), and
  • an additional 1.4235% in the subnormal range below 2^-6.

The calibrated scales (0.0171247218 to 0.0841238871) move these values into
a better-resolved range while retaining zero observed saturation for all 24
tensors. This is a measurable mechanism for calibration under this exact
checkpoint/activation/topology contract; it is not a claim that unit scale
must cause a large downstream quality loss on every checkpoint.

Implementation scope

  • Propagate independent per-layer K/V scales through cache writes, generic
    Triton reads, Flash-V100 grouped page4, and XQA reads.
  • Skip scale finalization only in the dedicated PLE offload process, which
    never executes QSA forward; all GPU workers retain the 24/24 fail-closed
    gate.
  • Split non-grouped E4M3 mixed catch-up/decode batches across the divisible
    grouped-page4 bulk and XQA tail chunks of at most 16 rows, avoiding both an
    unsupported E4M3 XQA launch and the larger Triton split-K workspace.
  • Reject unsupported first-phase configurations and disable
    calculate_kv_scales in calibrated E4M3 mode.
  • Collect K/V statistics only through normal FP16-KV inference forwards after
    engine initialization, excluding dummy/profile/graph-capture forwards.
  • Generate independent per-layer, per-tensor K/V scales and report max-abs,
    p99.9, p99.99, saturation, and per-shard stability.
  • Generate a derived overlay without modifying the original checkpoint.
  • Add route and selected-block comparison support for matched quality traces.

No private corpus or scale overlay is included in the repository. Generic
long/tool corpus builders and the frozen-token API quality runner are isolated
in Leonccaa/1Cat-vLLM#6.

Test Plan

Static checks over the PR's Python diff:

mapfile -t files < <(
  git diff --name-only --diff-filter=ACMR upstream/main | rg '\.py$'
)
uv tool run ruff check "${files[@]}"
uv tool run ruff format --check "${files[@]}"
git diff --check upstream/main

Public regression set, run in the SM70 CUDA build on V100:

python -m pytest -q \
  tests/benchmarks/test_compare_sm70_quality_trace_qsa.py \
  tests/kernels/attention/test_sm70_qsa_grouped_page4.py \
  tests/models/qwen4_exp/test_config.py \
  tests/models/qwen4_exp/test_qsa_cache.py \
  tests/models/qwen4_exp/test_qsa_e4m3.py \
  tests/models/qwen4_exp/test_qsa_kv_calibration.py \
  tests/models/qwen4_exp/test_qsa_ops.py \
  tests/models/qwen4_exp/test_weight_loading.py

End-to-end quality used the same checkpoint, exact prompt token IDs, and
deterministic sampling for FP16 KV and calibrated E4M3 KV. An explicitly
loaded unit-scale overlay was used only as a negative control. Calibration and
quality inputs were disjoint and covered Chinese, English, code, tool calls,
multi-turn data, and 1K/16K/64K/128K context lengths.

Test Result

TP4 / MTP0 / 128K quality and fixed-contract performance are validated. The two optional Triton optimizations remain isolated in sibling drafts #452 and #453.

Automated checks

  • Core regression set: 60 passed, 14 warnings in 17.51s on V100 after the
    CI-format fixes.
  • Ruff check: passed.
  • Ruff format check: 13 Python files already formatted.
  • clang-format, typos, and check-spdx-header: passed for the files that
    failed the first pre-commit run.
  • git diff --check upstream/main: passed.
  • Post-fix focused CPU/static regression: 10 passed, 7 skipped; full
    pre-commit over the three changed files, including mypy, passed.
  • Dedicated PLE-worker scale-gate regression: passed in the CUDA image. The
    production startup matrix separately confirmed missing overlay 0/24
    rejection, calibrated 24/24 loading on all GPU workers, and successful PLE
    worker startup.
  • SM70 numerical regression for the observed 49-row mixed layout
    (46+1+1+1 requests): grouped 48 + XQA 1 matched the E4M3 Triton
    reference with max absolute error 1.52587890625e-05 and mean absolute
    error 1.5917149767e-06.

Startup and calibration gates

  • Base checkpoint without an overlay: Loaded 0/24, startup rejected.
  • Calibrated overlay: 24/24 scales loaded before ready.
  • Explicit unit-scale negative-control overlay: 24/24 loaded; 1.0 remains a
    legal value, not an error sentinel.
  • calculate_kv_scales is rejected for this calibrated E4M3 contract.
  • The dedicated PLE offload worker is excluded from QSA scale finalization
    because it loads only PLE tensors and never runs QSA forward; the GPU-worker
    startup gate remains unchanged.
  • Final 24 calibrated scales ranged from 0.0171247218 to 0.0841238871;
    observed saturation ratio was zero for every K/V tensor.

TP4 / MTP0 quality evidence

  • KV capacity at the same memory budget: 384,760 FP16 tokens versus
    704,741 E4M3 tokens (1.83x). A 128,001-token request's reported
    concurrency increased from 3.01x to 5.51x.
  • 4K trace: first token and complete output token sequence were identical;
    all 96 expected per-layer dumps were paired.
  • 128K retrieval at 0%, 50%, and 100% insertion depth, two repeats each:
    FP16 and calibrated E4M3 both scored 6/6, with token-identical paired
    outputs.
  • Held-out real tool-selection cases: FP16 and calibrated E4M3 both scored
    10/12, with zero correctness regressions, identical first tokens for
    12/12 comparisons, and exact repeat stability for 6/6 cases. Both modes
    failed the same baseline case.
  • Mean QSA core-output error: relative L2 0.0324132, cosine 0.999311,
    RMSE 0.0160753.
  • Selected-block overlap: mean row Jaccard 0.990783, mean FP16 recall
    0.995301, exact-row ratio 0.570984, and micro Jaccard 0.987579.
    The first QSA layer (layer 3) remained exactly identical on all four ranks;
    later differences are downstream of earlier quantized attention outputs.
  • The explicit unit-scale negative control regressed from FP16's 5/6 to 4/6
    on the six single-run tool cases. This small sample is supporting evidence,
    not the primary justification for calibration.

With max_num_batched_tokens=4096, complete chunks in these quality runs used
the grouped Flash-V100 page4 path, while tails/decode used generic Triton. The
logs showed no XQA entry, so the FP16 P1024 versus E4M3 P256 XQA partition
difference did not confound the reported A/B. A future experiment that enters
XQA must add an FP16-at-P256 control arm or explicitly report reduction order
as a second variable.

Targeted mixed-batch scheduler regression

At gpu_memory_utilization=0.89, C4 at 64K produced the scheduler layout that
exposed the issue: a 49-row batch consisting of a 46-row catch-up chunk plus
three decode rows. The fixed path routed this as grouped 48 + XQA 1. Warmup
and all 3/3 scored repeats (12/12 requests) completed without CUDA OOM or
engine failure. This closes the observed mixed-batch correctness/capacity bug;
it is not the formal matched performance matrix.

NVFP4 formal performance update (2026-09-02)

The matched SM70 / TP4 / MTP0 / GMU 0.89 matrix is complete on this PR plus the two isolated Triton optimizations now split across #452 and #453: 17 E4M3 cells passed, C8 x 128K was the sole expected capacity skip, and all 87 scored requests completed. At the same memory budget, KV capacity increased from 507,093 FP16 tokens to 931,100 E4M3 tokens (1.836x). Across 15 matched cells, E4M3 decode throughput was 1.92% lower and server E2E elapsed time was 1.26% higher by geometric mean; at 16K and longer the corresponding figures were -2.94% and +3.38%. This validates the fixed first-phase performance contract as a capacity win with a modest long-context speed cost, not a global speedup. The sibling PR descriptions contain the complete matrix, caveats, and split-specific diffs.

The full matrix used the source tree of #447 plus the two optimizations before their review-only split into #452 and #453; that exact tree is preserved by the fork tag evidence/qsa-e4m3-perf-combined-20260902. The split combined tree binds the same bitcast decoder and changes only its local imported symbol name. The core-only final route was separately checked at C4 x 64K and was effectively neutral versus the optimized tree (-0.32% decode, +0.21% E2E), but the core-only tree was not rerun across all 18 cells.

AWQ formal-r6 validation update (2026-09-02)

A second end-to-end campaign validated the same QSA FP8 contract with a
per-expert asymmetric AWQ W4A16 g32 checkpoint (zero point enabled, GEMM) and
an FP8 E4M3 PLE splice. The exact checkpoint index SHA-256 was
da62ba7bfef57f8117f35fc78bf5daf14c0ab5eb1b8173b4d712c1f4d294819b.
The AWQ-specific MoE TP-shard alignment patch used by that checkpoint is
outside this PR; these results validate this PR's QSA E4M3 path under the AWQ
runtime composition rather than claiming that #447 alone adds AWQ loading.

Calibration used normal FP16-KV inference forwards over 4,888,771 processed
tokens spanning Chinese, English, code, multi-turn dialogue, tool use, and
1K/16K/64K/128K contexts. The 24 scales ranged from 0.0171072837 to
0.0806361660, with zero observed saturation. The AWQ/NVFP4 per-tensor scale
ratio ranged from about 0.955 to 1.028, confirming that the NVFP4 overlay
must not be reused for AWQ.

Matched route-neutral quality results were:

  • 128K retrieval: 6/6 for FP16 and calibrated E4M3, zero regressions;
  • held-out tool selection: 10/12 for both, zero regressions;
  • first token: 18/18 equal across the matched tool and 128K comparisons;
  • GSM8K five-shot: 29/32 for both, zero flips;
  • HumanEval/MBPP functional subset: 9/10 for both;
  • IFEval strict and loose: 3/5 prompts and 9/12 instructions for both;
  • selected-block FP16 recall: mean 0.995906, minimum 0.991822;
  • QSA output: minimum cosine 0.998597, maximum relative L2 0.052962.

The fixed GMU 0.89 performance matrix covered concurrency 1/4/8 and
1K/4K/16K/32K/64K/128K with 256 output tokens and three scored repeats per
cell. FP16 completed 15/18 cells and E4M3 completed 17/18, with no failed cells.
Measured KV capacity increased from 310,480 to 569,579 tokens
(1.8345x). Over the nine matched cells at 16K and longer, E4M3 prefill was
3.42% lower, decode was 2.54% lower, and server E2E elapsed time was 3.12%
higher. C4 x 128K and C8 x 64K were E4M3-only capacity cells; both modes
skipped C8 x 128K.

This AWQ matrix used #447 plus both optional optimizations. FP16 P1024 versus
E4M3 P256 XQA and mixed-batch splitting remain a second variable, so these
figures characterize the deployed bundle rather than pure converter overhead.
The supported conclusion is a 1.8345x KV-capacity gain at about a 3%
long-context latency cost under the exact tested contract.

Pending before non-draft / broader enablement

AI assistance disclosure

This PR was developed with OpenAI Codex assistance. Design and validation were
supervised by the human submitter. The PR remains draft pending maintainer review. TP8, MTP, prefix caching,
standalone graph-versus-eager parity, and 256K remain outside first-phase acceptance.

Checklist

Leonccaa and others added 2 commits September 1, 2026 18:41
Add per-layer calibrated E4M3 K/V storage to Qwen4Exp QSA for the
fixed SM70, TP4, MTP0, and FP16 activation contract. The main QSA
cache uses E4M3 while the raw and compressed indexer caches remain FP16.

Design decisions:

- Keep checkpoint loading slots at an invalid -1.0 sentinel, separate
  from runtime scale buffers whose upstream default remains 1.0.
  Finalization validates and deletes the loading slots.
- Fail closed if any of the 12 layers' K/V scales is missing, non-finite,
  or non-positive. This intentionally replaces the upstream KV-cache
  warning and unit-scale fallback, and is confined to 1Cat-owned
  Qwen4Exp files.
- Keep indexer caches FP16 so K/V quantization cannot change QSA block
  selection.
- Use P256 for E4M3 on generic G6 XQA virtual-page4. The P1024
  specialization assumes the page1568 layout and is not valid here.
- Do not bundle a scale overlay. Generate an exact-checkpoint-revision-
  bound model-kvscales overlay with
  tools/qwen4_exp/qsa_kv_calibration.py; E4M3 startup rejects unless all
  24 scales are explicitly loaded.

Also add offline normal-inference scale collection, report and overlay
generation, and matched route/selected-block comparison.

Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
Apply the pinned CUDA formatter, add the required SPDX copyright header, and express the reference attention with equivalent matmul operations so the tensor subscripts do not trigger the spelling hook.

Co-authored-by: OpenAI Codex <codex@openai.com>

Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
@Leonccaa Leonccaa changed the title [Model][SM70] Add calibrated E4M3 QSA KV cache [Model][SM70] Add calibrated FP8 E4M3 QSA KV cache Sep 2, 2026
E4M3 XQA accepts at most 16 query rows, while chunked scheduling can form non-grouped catch-up/decode batches above that limit. Route the divisible bulk through grouped page4 and split the remainder into supported XQA chunks instead of falling back to the larger Triton split-K workspace.

Add routing coverage and an SM70 numerical regression for the observed 46+1+1+1 request layout.
The dedicated PLE offload process builds the full model structure on meta but intentionally loads only the PLE subtree. Skip QSA KV scale finalization in that process; it never runs QSA forward, while every GPU worker retains the existing 24/24 fail-closed gate.
@Leonccaa Leonccaa changed the title [Model][SM70] Add calibrated FP8 E4M3 QSA KV cache [Model][SM70][FP8] Add calibrated E4M3 QSA KV cache Sep 2, 2026
@Leonccaa
Leonccaa marked this pull request as ready for review September 2, 2026 21:20
@yangzhuxinyzx

Copy link
Copy Markdown
Contributor

Superseded by merged #462, which replays this runtime/test work onto current main, preserves #460 QSA workspace sharing, carries both E4M3 hot-loop follow-ups, and records the narrow TP4/MTP0 capacity contract and validation boundary.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants