Skip to content

cuda: emit native MXFP4 images from HC normalization - #153

Draft
GenerelSchwerz wants to merge 26 commits into
reference/upstream-kernels-20261001from
kernel/hc-native-mxfp4-emission-20261008
Draft

GenerelSchwerz wants to merge 26 commits into
reference/upstream-kernels-20261001from
kernel/hc-native-mxfp4-emission-20261008

Conversation

@GenerelSchwerz

Copy link
Copy Markdown
Owner

Extend #152's HC normalization path to emit native MXFP4 activation images directly, removing a separate conversion before unchanged MMQ readers. Reuse the existing native float4 packing and bank stores; retain the original HC and normalized F32 outputs and optional F16/BF16 images.

This covers eligible HC_POST -> RMS_NORM with identity/combination, optional MUL or pure SCALE, and preceding tensor-affine gates. Bank/sample axes and canonical regroupings use the existing shape, precision, capability and physical-range checks. No model rules, new build option or public API are added. NVFP4 keeps its original converter; mixed Q8 readers retain the conversion they need. Live intermediate normalization outputs, biased SCALE, noncontiguous/offset readers and unsupported cases retain their existing behavior.

Measured on RTX 5070 Ti, SM120a, against the literal #152 head 5aab048625d3039ea35e0ba73dc38c8e7cd00a95:

Component Control Candidate Less time
Native MX, N512/H16/T3 12.2861 us 10.2487 us 16.58%
Combination + weighted native MX, N1024/H16/T3 18.4270 us 16.3959 us 11.02%
Sample-axis native MX, N512/H17/T3 12.2975 us 12.2858 us Flat

The two useful cases won all six alternating control/candidate pairs, with 2,000 unprofiled graph replays each. NVFP4 and forced-Q8 controls were flat. These are component results; no model-serving, MTP, older-GPU or combined-series speedup is claimed.

Validation:

  • Paired GPU output/reference/upload bytes are exact, with unchanged matrix geometry, resources and arenas. Actual native precision-Q4 dispatch and converter removal are verified in normal upstream scheduler/capture profiles.
  • 776 per-stage CPU checks preserve 5e-4 for HC/RMS/typed/Q8 outputs and the existing 2e-2 tolerance only for actual native matrix outputs. The native approximation also occurs in the control; this change adds no matched GPU-output drift.
  • 56 distinct raw-image datasets match the original standalone quantizer, including headers, packed values, padding, typed rounding and guards. Generic norm groups of 256, 768 and 2304 exercise both 256/1024-thread variants.
  • 40 backend checks and nine memcheck/racecheck/synccheck runs pass. Existing test infrastructure gains four registrations for native MX and unchanged NV bank/sample cases.
  • All 747 prior host/norm device globals retain identical resources and instructions on this build; 72 native variants add 418,904 bytes. New variants have no local spills, with stack 0-24 bytes and registers 38-64.

Each hardware suite held the ordered shared build/GPU leases throughout, under a verified 4 GiB host limit, zero swap, finite timeouts and whole-process-tree teardown. Private dispatcher/allocator integration remains separate.

Leaf follows #152; the fixed review base remains reference/upstream-kernels-20261001. Evidence: /home/gencoolpc/moe-cache-tests/results/upstream-kernel-series-20261001/HYPERCONNECTION/HC-NATIVE-IMAGE-ASSESSMENT.

AI assisted implementation and validation. This is an owner-authorized fork review; no upstream submission.

Carry the reviewed kernel sources from PR125, PR129 and PR132 as the
matching control for the separate tensor-affine HC follow-up.

Assisted-by: Codex
Preserve gate intermediates, original rounding and live HC outputs.
Retain strict input alias checks and standard allocation dependencies.

Assisted-by: Codex
Use the exact tested control sources from published PR130, PR129 and PR132.

Assisted-by: Codex
Reuse the allocated affine match to find the planned HC_POST and retain prepared emission priority. Reject expanded normalization cuts whose terminal is not written by this kernel.

Assisted-by: Codex
Assisted-by: Codex
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant