Skip to content

[Perf][SM70][FP8] Hoist QSA E4M3 KV scales - #453

Closed
Leonccaa wants to merge 6 commits into
1CatAI:mainfrom
Leonccaa:perf/qwen4exp-qsa-e4m3-scale-hoist
Closed

Leonccaa wants to merge 6 commits into
1CatAI:mainfrom
Leonccaa:perf/qwen4exp-qsa-e4m3-scale-hoist

Conversation

@Leonccaa

@Leonccaa Leonccaa commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

This stacked follow-up hoists calibrated E4M3 K/V scaling out of the
per-element generic Triton Qwen4Exp QSA load loop on SM70.

Depends on #447. Until #447 merges, GitHub also shows its four parent
commits in this PR. This follow-up itself is one commit changing two files
(+32/-5).

This does not change the KV-scale overlay, loader gate, cache-write path,
indexer cache, Flash-V100 grouped path, XQA path, or routing policy. E4M3 still
requires the checkpoint-revision-bound model-kvscales.safetensors overlay
described in #447 and fails closed unless all 24 calibrated values are loaded.

Changes

  • Fold the per-layer K scale into the post-dot softmax factor.
  • Apply the per-layer V scale once after normalized accumulation; split-K
    applies it once after merge.
  • Add calibrated E4M3 regression coverage for both single-split and split-K
    merge shapes.

The FP16/BF16 path remains behind the existing compile-time KV_E4M3=false
branch and is unchanged.

Split and merge independence

The bitcast-decoder optimization is isolated in sibling draft #452. Neither
performance PR depends on the other: both use #447 as their only functional
base, and both merge orders were checked locally without conflicts. #452 keeps
the decoder call sites unchanged through an import alias so this PR can own the
scale-placement lines cleanly.

Test plan

Static and CPU collection checks:

uv tool run pre-commit run --files \
  vllm/models/qwen4_exp/nvidia/ops/qsa.py \
  tests/models/qwen4_exp/test_qsa_e4m3.py
git diff --check 562bd404df000a6bc4f8410dd6e7e2ceffab302b..HEAD
PYTHONPATH=. python -m pytest -q tests/models/qwen4_exp/test_qsa_e4m3.py

SM70 numerical coverage exercises calibrated E4M3 single-split and split-K
merge shapes, including a long topk=2051 case in the combined validated
tree.

Test results

  • Full pre-commit over both changed files: passed, including ruff check/format,
    mypy, SPDX, forbidden-import, and attention-config checks.
  • git diff --check: passed.
  • CPU-only pytest collection: 7 CUDA tests skipped as expected.
  • Both merge orders with [Perf][SM70][FP8] Use bitcast E4M3 decoder in QSA #452: conflict-free.

Combined SM70 validation provenance

The GPU numerical checks and formal performance matrix were run on the source
tree of #447 plus both optimizations before this review-only split
(d3a2c62c9a). The newly split combined tree
binds the same exact bitcast decoder and differs only in its local imported
symbol name; no kernel arithmetic or dispatch changed.

On that combined V100 / TP4 / MTP0 / FP16-activation tree:

  • E4M3 long split-K versus FP32 dequant reference: max absolute error
    8.262693881988525e-06.
  • Mixed grouped/XQA route versus optimized Triton reference: max absolute
    error 1.52587890625e-05, mean absolute error
    1.2406309224388679e-06.
  • The formal matrix covered concurrency 1/4/8 x context
    1K/4K/16K/32K/64K/128K, with 17 E4M3 cells passing, one expected capacity
    skip (C8 x 128K), and all 87 scored requests completing.
  • At the same memory budget, KV capacity was 507,093 FP16 tokens versus
    931,100 E4M3 tokens (1.836x). Across 15 matched cells, E4M3 decode was
    1.92% lower and server E2E elapsed time was 1.26% higher by geometric mean.

The measured optimization effect combines this PR and #452 and is near run
noise in aggregate; it must not be attributed independently to scale hoisting.
The result supports a hot-loop simplification, not a global speedup claim.

AWQ formal-r6 replication (2026-09-02)

The same combined #447 + #452 + #453 source composition was then validated
with a per-expert asymmetric AWQ W4A16 g32 checkpoint and FP8 E4M3 PLE splice.
The updated stacked branches also carry #447's dedicated-PLE-worker gate scope
fix; every GPU worker retains the 24/24 fail-closed scale gate.

  • The SM70 numerical suite passed, including all 256 E4M3 encodings, FP16
    split-K parity, calibrated E4M3 single/split/long shapes, and the mixed
    grouped/XQA route.
  • The fixed GMU 0.89 matrix covered concurrency 1/4/8 and
    1K/4K/16K/32K/64K/128K with three scored repeats per cell.
  • FP16 completed 15/18 cells and E4M3 completed 17/18, with no failed cells;
    all 174 FP16 and 210 E4M3 scored requests produced 256 tokens.
  • KV capacity increased from 310,480 FP16 tokens to 569,579 E4M3 tokens
    (1.8345x). C4 x 128K and C8 x 64K were E4M3-only capacity
    cells; both modes skipped C8 x 128K.
  • Over the nine matched cells at 16K and longer, E4M3 prefill was 3.42% lower,
    decode was 2.54% lower, and server E2E elapsed time was 3.12% higher.

This replication again measures #452 and #453 together and does not attribute
the aggregate result to either optimization independently. The matching AWQ
quality acceptance and scale-calibration details are reported in #447.

Measurement caveats

  • FP16 XQA uses P1024 while generic E4M3 G6 XQA uses P256, so reduction order
    is a second variable in FP16/E4M3 A/B runs that enter XQA.
  • Synthetic performance output hashes are not quality evidence. The separate
    TP4/MTP0/128K quality, retrieval, held-out tool-task, selected-block, and
    attention-output acceptance remains in [Model][SM70][FP8] Add calibrated E4M3 QSA KV cache #447.

AI assistance disclosure

This PR was developed with OpenAI Codex assistance. The human submitter
reviewed the scope and supervised the numerical and end-to-end validation.

Leonccaa and others added 4 commits September 1, 2026 18:41
Add per-layer calibrated E4M3 K/V storage to Qwen4Exp QSA for the
fixed SM70, TP4, MTP0, and FP16 activation contract. The main QSA
cache uses E4M3 while the raw and compressed indexer caches remain FP16.

Design decisions:

- Keep checkpoint loading slots at an invalid -1.0 sentinel, separate
  from runtime scale buffers whose upstream default remains 1.0.
  Finalization validates and deletes the loading slots.
- Fail closed if any of the 12 layers' K/V scales is missing, non-finite,
  or non-positive. This intentionally replaces the upstream KV-cache
  warning and unit-scale fallback, and is confined to 1Cat-owned
  Qwen4Exp files.
- Keep indexer caches FP16 so K/V quantization cannot change QSA block
  selection.
- Use P256 for E4M3 on generic G6 XQA virtual-page4. The P1024
  specialization assumes the page1568 layout and is not valid here.
- Do not bundle a scale overlay. Generate an exact-checkpoint-revision-
  bound model-kvscales overlay with
  tools/qwen4_exp/qsa_kv_calibration.py; E4M3 startup rejects unless all
  24 scales are explicitly loaded.

Also add offline normal-inference scale collection, report and overlay
generation, and matched route/selected-block comparison.

Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
Apply the pinned CUDA formatter, add the required SPDX copyright header, and express the reference attention with equivalent matmul operations so the tensor subscripts do not trigger the spelling hook.

Co-authored-by: OpenAI Codex <codex@openai.com>

Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
E4M3 XQA accepts at most 16 query rows, while chunked scheduling can form non-grouped catch-up/decode batches above that limit. Route the divisible bulk through grouped page4 and split the remainder into supported XQA chunks instead of falling back to the larger Triton split-K workspace.

Add routing coverage and an SM70 numerical regression for the observed 46+1+1+1 request layout.
Fold the per-layer K scale into the post-dot softmax factor and apply the V scale after normalized accumulation. Split-K applies V scaling once after merge, while the single-split path applies it directly before the final store.

This independent follow-up keeps the existing scalar E4M3 decoder unchanged, removes per-element K/V scale multiplies, and preserves the FP16 constexpr branch and calibrated scale contract.

Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
The dedicated PLE offload process builds the full model structure on meta but intentionally loads only the PLE subtree. Skip QSA KV scale finalization in that process; it never runs QSA forward, while every GPU worker retains the existing 24/24 fail-closed gate.
@yangzhuxinyzx

Copy link
Copy Markdown
Contributor

Superseded by merged #462, which replays this runtime/test work onto current main, preserves #460 QSA workspace sharing, carries both E4M3 hot-loop follow-ups, and records the narrow TP4/MTP0 capacity contract and validation boundary.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants