Conversation
Add per-layer calibrated E4M3 K/V storage to Qwen4Exp QSA for the fixed SM70, TP4, MTP0, and FP16 activation contract. The main QSA cache uses E4M3 while the raw and compressed indexer caches remain FP16. Design decisions: - Keep checkpoint loading slots at an invalid -1.0 sentinel, separate from runtime scale buffers whose upstream default remains 1.0. Finalization validates and deletes the loading slots. - Fail closed if any of the 12 layers' K/V scales is missing, non-finite, or non-positive. This intentionally replaces the upstream KV-cache warning and unit-scale fallback, and is confined to 1Cat-owned Qwen4Exp files. - Keep indexer caches FP16 so K/V quantization cannot change QSA block selection. - Use P256 for E4M3 on generic G6 XQA virtual-page4. The P1024 specialization assumes the page1568 layout and is not valid here. - Do not bundle a scale overlay. Generate an exact-checkpoint-revision- bound model-kvscales overlay with tools/qwen4_exp/qsa_kv_calibration.py; E4M3 startup rejects unless all 24 scales are explicitly loaded. Also add offline normal-inference scale collection, report and overlay generation, and matched route/selected-block comparison. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
Apply the pinned CUDA formatter, add the required SPDX copyright header, and express the reference attention with equivalent matmul operations so the tensor subscripts do not trigger the spelling hook. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
E4M3 XQA accepts at most 16 query rows, while chunked scheduling can form non-grouped catch-up/decode batches above that limit. Route the divisible bulk through grouped page4 and split the remainder into supported XQA chunks instead of falling back to the larger Triton split-K workspace. Add routing coverage and an SM70 numerical regression for the observed 46+1+1+1 request layout.
Fold the per-layer K scale into the post-dot softmax factor and apply the V scale after normalized accumulation. Split-K applies V scaling once after merge, while the single-split path applies it directly before the final store. This independent follow-up keeps the existing scalar E4M3 decoder unchanged, removes per-element K/V scale multiplies, and preserves the FP16 constexpr branch and calibrated scale contract. Co-authored-by: OpenAI Codex <codex@openai.com> Signed-off-by: Leonccaa <166551845+Leonccaa@users.noreply.github.com>
This was referenced Sep 2, 2026
The dedicated PLE offload process builds the full model structure on meta but intentionally loads only the PLE subtree. Skip QSA KV scale finalization in that process; it never runs QSA forward, while every GPU worker retains the existing 24/24 fail-closed gate.
Leonccaa
marked this pull request as ready for review
September 2, 2026 21:20
Contributor
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Purpose
This stacked follow-up hoists calibrated E4M3 K/V scaling out of the
per-element generic Triton Qwen4Exp QSA load loop on SM70.
This does not change the KV-scale overlay, loader gate, cache-write path,
indexer cache, Flash-V100 grouped path, XQA path, or routing policy. E4M3 still
requires the checkpoint-revision-bound
model-kvscales.safetensorsoverlaydescribed in #447 and fails closed unless all 24 calibrated values are loaded.
Changes
applies it once after merge.
merge shapes.
The FP16/BF16 path remains behind the existing compile-time
KV_E4M3=falsebranch and is unchanged.
Split and merge independence
The bitcast-decoder optimization is isolated in sibling draft #452. Neither
performance PR depends on the other: both use #447 as their only functional
base, and both merge orders were checked locally without conflicts. #452 keeps
the decoder call sites unchanged through an import alias so this PR can own the
scale-placement lines cleanly.
Test plan
Static and CPU collection checks:
SM70 numerical coverage exercises calibrated E4M3 single-split and split-K
merge shapes, including a long
topk=2051case in the combined validatedtree.
Test results
mypy, SPDX, forbidden-import, and attention-config checks.
git diff --check: passed.Combined SM70 validation provenance
The GPU numerical checks and formal performance matrix were run on the source
tree of #447 plus both optimizations before this review-only split
(
d3a2c62c9a). The newly split combined treebinds the same exact bitcast decoder and differs only in its local imported
symbol name; no kernel arithmetic or dispatch changed.
On that combined V100 / TP4 / MTP0 / FP16-activation tree:
8.262693881988525e-06.error
1.52587890625e-05, mean absolute error1.2406309224388679e-06.1K/4K/16K/32K/64K/128K, with 17 E4M3 cells passing, one expected capacity
skip (
C8 x 128K), and all 87 scored requests completing.931,100 E4M3 tokens (
1.836x). Across 15 matched cells, E4M3 decode was1.92% lower and server E2E elapsed time was 1.26% higher by geometric mean.
The measured optimization effect combines this PR and #452 and is near run
noise in aggregate; it must not be attributed independently to scale hoisting.
The result supports a hot-loop simplification, not a global speedup claim.
AWQ formal-r6 replication (2026-09-02)
The same combined #447 + #452 + #453 source composition was then validated
with a per-expert asymmetric AWQ W4A16 g32 checkpoint and FP8 E4M3 PLE splice.
The updated stacked branches also carry #447's dedicated-PLE-worker gate scope
fix; every GPU worker retains the 24/24 fail-closed scale gate.
split-K parity, calibrated E4M3 single/split/long shapes, and the mixed
grouped/XQA route.
1K/4K/16K/32K/64K/128K with three scored repeats per cell.
all 174 FP16 and 210 E4M3 scored requests produced 256 tokens.
(
1.8345x).C4 x 128KandC8 x 64Kwere E4M3-only capacitycells; both modes skipped
C8 x 128K.decode was 2.54% lower, and server E2E elapsed time was 3.12% higher.
This replication again measures #452 and #453 together and does not attribute
the aggregate result to either optimization independently. The matching AWQ
quality acceptance and scale-calibration details are reported in #447.
Measurement caveats
is a second variable in FP16/E4M3 A/B runs that enter XQA.
TP4/MTP0/128K quality, retrieval, held-out tool-task, selected-block, and
attention-output acceptance remains in [Model][SM70][FP8] Add calibrated E4M3 QSA KV cache #447.
AI assistance disclosure
This PR was developed with OpenAI Codex assistance. The human submitter
reviewed the scope and supervised the numerical and end-to-end validation.