Repository navigation
cuda: emit native MXFP4 images from HC normalization - #153
Draft
GenerelSchwerz wants to merge 26 commits into
Draft
GenerelSchwerz wants to merge 26 commits into
GenerelSchwerz wants to merge 26 commits into
Conversation
Carry the reviewed kernel sources from PR125, PR129 and PR132 as the matching control for the separate tensor-affine HC follow-up. Assisted-by: Codex
Preserve gate intermediates, original rounding and live HC outputs. Retain strict input alias checks and standard allocation dependencies. Assisted-by: Codex
Use the exact tested control sources from published PR130, PR129 and PR132. Assisted-by: Codex
Reuse the allocated affine match to find the planned HC_POST and retain prepared emission priority. Reject expanded normalization cuts whose terminal is not written by this kernel. Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
Assisted-by: Codex
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Extend #152's HC normalization path to emit native MXFP4 activation images directly, removing a separate conversion before unchanged MMQ readers. Reuse the existing native float4 packing and bank stores; retain the original HC and normalized F32 outputs and optional F16/BF16 images.
This covers eligible HC_POST -> RMS_NORM with identity/combination, optional MUL or pure SCALE, and preceding tensor-affine gates. Bank/sample axes and canonical regroupings use the existing shape, precision, capability and physical-range checks. No model rules, new build option or public API are added. NVFP4 keeps its original converter; mixed Q8 readers retain the conversion they need. Live intermediate normalization outputs, biased SCALE, noncontiguous/offset readers and unsupported cases retain their existing behavior.
Measured on RTX 5070 Ti, SM120a, against the literal #152 head
5aab048625d3039ea35e0ba73dc38c8e7cd00a95:The two useful cases won all six alternating control/candidate pairs, with 2,000 unprofiled graph replays each. NVFP4 and forced-Q8 controls were flat. These are component results; no model-serving, MTP, older-GPU or combined-series speedup is claimed.
Validation:
Each hardware suite held the ordered shared build/GPU leases throughout, under a verified 4 GiB host limit, zero swap, finite timeouts and whole-process-tree teardown. Private dispatcher/allocator integration remains separate.
Leaf follows #152; the fixed review base remains
reference/upstream-kernels-20261001. Evidence:/home/gencoolpc/moe-cache-tests/results/upstream-kernel-series-20261001/HYPERCONNECTION/HC-NATIVE-IMAGE-ASSESSMENT.AI assisted implementation and validation. This is an owner-authorized fork review; no upstream submission.