Skip to content

[Perf][SM70][FP8] Use bitcast E4M3 decoder in QSA - #452

Closed
Leonccaa wants to merge 6 commits into
1CatAI:mainfrom
Leonccaa:perf/qwen4exp-qsa-e4m3-triton-decode
Closed

Leonccaa wants to merge 6 commits into
1CatAI:mainfrom
Leonccaa:perf/qwen4exp-qsa-e4m3-triton-decode

Conversation

@Leonccaa

@Leonccaa Leonccaa commented Sep 2, 2026 •

Copy link
Copy Markdown
Contributor

Purpose

This stacked follow-up routes the generic Triton Qwen4Exp QSA E4M3 read path
to the existing exact bitcast software decoder on SM70.

Depends on #447. Until #447 merges, GitHub also shows its four parent
commits in this PR. This follow-up itself is one commit changing one import
(+1/-1).

This does not change the KV-scale overlay, loader gate, scale placement,
cache-write path, indexer cache, Flash-V100 grouped path, XQA path, or routing
policy. E4M3 still requires the checkpoint-revision-bound
model-kvscales.safetensors overlay described in #447 and fails closed unless
all 24 calibrated values are loaded.

Change

Bind the existing fp8_e4m3fn_bits_to_fp32 call-site name to
fp8_e4m3fn_bits_to_fp32_bitcast. The exact decoded FP32 values are unchanged,
while the generic Triton path avoids the scalar decoder's per-element exp2
work on SM70.

The call sites are deliberately left unchanged so the scale-hoisting change in
sibling draft #453 can be reviewed and merged independently.

Split and merge independence

Neither performance PR depends on the other: #452 and #453 both use #447 as
their only functional base, and both merge orders were checked locally without
conflicts. The combined tree calls the same bitcast decoder that was used for
the V100 numerical and performance validation before this review-only split.

Test plan

uv tool run pre-commit run --files \
  vllm/models/qwen4_exp/nvidia/ops/qsa.py
git diff --check 562bd404df000a6bc4f8410dd6e7e2ceffab302b..HEAD
PYTHONPATH=. python -m pytest -q tests/models/qwen4_exp/test_qsa_e4m3.py

The decoder regression in the #447 test set covers all 256 E4M3 byte
encodings, including finite-value equality and identical NaN masks.

Test results

  • Full pre-commit over the changed file: passed, including ruff check/format,
    mypy, SPDX, forbidden-import, and attention-config checks.
  • git diff --check: passed.
  • CPU-only pytest collection: 6 CUDA tests skipped as expected.
  • Both merge orders with [Perf][SM70][FP8] Hoist QSA E4M3 KV scales #453: conflict-free.
  • All 256 decoder inputs on the combined validated tree: finite results were
    bit-exact with the scalar decoder and NaN masks were identical.

Combined SM70 validation provenance

The GPU numerical checks and formal performance matrix were run on the source
tree of #447 plus both optimizations before this review-only split
(d3a2c62c9a). The newly split combined tree
binds the same exact bitcast decoder and differs only in its local imported
symbol name; no kernel arithmetic or dispatch changed.

On that combined V100 / TP4 / MTP0 / FP16-activation tree:

  • E4M3 long split-K versus FP32 dequant reference: max absolute error
    8.262693881988525e-06.
  • Mixed grouped/XQA route versus optimized Triton reference: max absolute
    error 1.52587890625e-05, mean absolute error
    1.2406309224388679e-06.
  • The formal matrix covered concurrency 1/4/8 x context
    1K/4K/16K/32K/64K/128K, with 17 E4M3 cells passing, one expected capacity
    skip (C8 x 128K), and all 87 scored requests completing.
  • At the same memory budget, KV capacity was 507,093 FP16 tokens versus
    931,100 E4M3 tokens (1.836x). Across 15 matched cells, E4M3 decode was
    1.92% lower and server E2E elapsed time was 1.26% higher by geometric mean.

The measured optimization effect combines this PR and #453 and is near run
noise in aggregate; it must not be attributed independently to the decoder
switch. The result supports a small hot-loop simplification, not a global
speedup claim.

AWQ formal-r6 replication (2026-09-02)

The same combined #447 + #452 + #453 source composition was then validated
with a per-expert asymmetric AWQ W4A16 g32 checkpoint and FP8 E4M3 PLE splice.
The updated stacked branches also carry #447's dedicated-PLE-worker gate scope
fix; every GPU worker retains the 24/24 fail-closed scale gate.

  • The SM70 numerical suite passed, including all 256 E4M3 encodings, FP16
    split-K parity, calibrated E4M3 single/split/long shapes, and the mixed
    grouped/XQA route.
  • The fixed GMU 0.89 matrix covered concurrency 1/4/8 and
    1K/4K/16K/32K/64K/128K with three scored repeats per cell.
  • FP16 completed 15/18 cells and E4M3 completed 17/18, with no failed cells;
    all 174 FP16 and 210 E4M3 scored requests produced 256 tokens.
  • KV capacity increased from 310,480 FP16 tokens to 569,579 E4M3 tokens
    (1.8345x). C4 x 128K and C8 x 64K were E4M3-only capacity
    cells; both modes skipped C8 x 128K.
  • Over the nine matched cells at 16K and longer, E4M3 prefill was 3.42% lower,
    decode was 2.54% lower, and server E2E elapsed time was 3.12% higher.

This replication again measures #452 and #453 together and does not attribute
the aggregate result to either optimization independently. The matching AWQ
quality acceptance and scale-calibration details are reported in #447.

Measurement caveats

  • FP16 XQA uses P1024 while generic E4M3 G6 XQA uses P256, so reduction order
    is a second variable in FP16/E4M3 A/B runs that enter XQA.
  • Synthetic performance output hashes are not quality evidence. The separate
    TP4/MTP0/128K quality, retrieval, held-out tool-task, selected-block, and
    attention-output acceptance remains in [Model][SM70][FP8] Add calibrated E4M3 QSA KV cache #447.

AI assistance disclosure

This PR was developed with OpenAI Codex assistance. The human submitter
reviewed the scope and supervised the numerical and end-to-end validation.

Leonccaa and others added 3 commits September 1, 2026 18:41
Add per-layer calibrated E4M3 K/V storage to Qwen4Exp QSA for the
fixed SM70, TP4, MTP0, and FP16 activation contract. The main QSA
cache uses E4M3 while the raw and compressed indexer caches remain FP16.

Design decisions:

- Keep checkpoint loading slots at an invalid -1.0 sentinel, separate
  from runtime scale buffers whose upstream default remains 1.0.
  Finalization validates and deletes the loading slots.
- Fail closed if any of the 12 layers' K/V scales is missing, non-finite,
  or non-positive. This intentionally replaces the upstream KV-cache
  warning and unit-scale fallback, and is confined to 1Cat-owned
  Qwen4Exp files.
- Keep indexer caches FP16 so K/V quantization cannot change QSA block
  selection.
- Use P256 for E4M3 on generic G6 XQA virtual-page4. The P1024
  specialization assumes the page1568 layout and is not valid here.
- Do not bundle a scale overlay. Generate an exact-checkpoint-revision-
  bound model-kvscales overlay with
  tools/qwen4_exp/qsa_kv_calibration.py; E4M3 startup rejects unless all
  24 scales are explicitly loaded.

Also add offline normal-inference scale collection, report and overlay
generation, and matched route/selected-block comparison.

Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
Apply the pinned CUDA formatter, add the required SPDX copyright header, and express the reference attention with equivalent matmul operations so the tensor subscripts do not trigger the spelling hook.

Co-authored-by: OpenAI Codex <codex@openai.com>

Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
E4M3 XQA accepts at most 16 query rows, while chunked scheduling can form non-grouped catch-up/decode batches above that limit. Route the divisible bulk through grouped page4 and split the remainder into supported XQA chunks instead of falling back to the larger Triton split-K workspace.

Add routing coverage and an SM70 numerical regression for the observed 46+1+1+1 request layout.
Route the existing QSA decoder symbol to the exact bitcast implementation. Keeping the call sites unchanged makes this optimization independent from the sibling scale-hoisting change while preserving the same Triton specialization and generated behavior validated in the combined performance tree.

Co-authored-by: OpenAI Codex <codex@openai.com>
Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
@Leonccaa
Leonccaa force-pushed the perf/qwen4exp-qsa-e4m3-triton-decode branch from d3a2c62 to 6e42fb8 Compare September 2, 2026 08:21
@Leonccaa Leonccaa changed the title [Perf][SM70][FP8] Optimize E4M3 QSA Triton decode [Perf][SM70][FP8] Use bitcast E4M3 decoder in QSA Sep 2, 2026
The dedicated PLE offload process builds the full model structure on meta but intentionally loads only the PLE subtree. Skip QSA KV scale finalization in that process; it never runs QSA forward, while every GPU worker retains the existing 24/24 fail-closed gate.
@yangzhuxinyzx

Copy link
Copy Markdown
Contributor

Superseded by merged #462, which replays this runtime/test work onto current main, preserves #460 QSA workspace sharing, carries both E4M3 hot-loop follow-ups, and records the narrow TP4/MTP0 capacity contract and validation boundary.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants