Skip to content

perf(gfx1100): add opt-in MQ4V2 K5120 gate/up decode specialization - #816

Open
HUSRCF wants to merge 5 commits into
warpfront:betafrom
HUSRCF:perf/mq4xt-decode-20260930
Open

HUSRCF wants to merge 5 commits into
warpfront:betafrom
HUSRCF:perf/mq4xt-decode-20260930

Conversation

@HUSRCF

@HUSRCF HUSRCF commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

perf(gfx1100): add opt-in MQ4V2 K5120 gate/up decode specialization

Summary

Add an opt-in specialization of the existing MQ4V2 gate/up decode kernel for exact gfx1100 and (gate_m, up_m, K) = (17408, 17408, 5120). It fixes the group count at compile time while preserving the existing weight representation and arithmetic. Other architectures and shapes retain their existing routes; there is no additional resident weight copy.

The setting is kernel.mq4v2_gateup_k5120, with compatibility override HIPFIRE_MQ4V2_GATEUP_K5120=0/1. It defaults off and is resolved through ProcessConfig/FeatureFlags at initialization. The new symbol is registered for kernel packaging and replay; compiler-less installations need an updated pack to opt in.

Scope

  • Tested base: beta 45f1abd1d. Current integration includes beta a193e4491 through merge commit 3123f3865; original commits/history are preserved.
  • Commits: cb2f84ae3, 671e88e02, f689f559c.
  • Surfaces: kernel dispatch/registration/replay, configuration, documentation.
  • No Architecture trait, attention, quantization, or speculative-decoding changes.
  • Independent of the split-KV attention verifier work.

Measured Result

Primary final check: W7900 Dual Slot GPU1 / gfx1100 / HIP 7.15, same pinned model, Q8/VMM, AR, explicit max_think_tokens=1, TG4096. ABBA off/on/on/off = 38.9 / 39.6 / 39.6 / 38.8 tok/s. Median 38.85 -> 39.60 (+1.93%). All four requests generated 4096 visible-output tokens with no reasoning or cached prompt tokens; request and decoded-text hashes match. This is separate from the earlier GPU0 reasoning run below; do not pool the two devices or modes.

Reproducer and summary: benchmarks/scripts/mq4v2_k5120_abba.sh and benchmarks/results/mq4v2-k5120-20261005/README.md.

Raw evidence: mq4v2-k5120-evidence-20261005.tar.gz (337179 bytes), SHA256 38e70e8d1bbf1540920482b77f5cec432ea248f3c8804623f37974ce19e954f8.

Latest-beta integration boundary

After the full validation below, beta advanced to a193e4491. I merged it without rewriting the measured commits. The latest conflicts were confined to generated crate maps and the environment inventory count, regenerated with the official scripts; no K5120 kernel or dispatch conflict required manual resolution. On the merged tree, cargo test --locked -p rdna-compute --lib k5120 passes all four tests (config lifecycle, exact arch/shape guard, inventory and replay contract), cargo test --locked -p hipfire-config --lib passes 107 tests, and lifecycle, all 48 crate maps and whitespace checks pass. Full CI and GPU ABBA results below belong to the pre-merge candidate, not a newly benchmarked latest-beta binary. This PR remains a draft pending full integration validation and remote CI.

Earlier GPU0 Reasoning Run

W7900 gfx1100, physical GPU0, HIP 7.15; Qwen3.8-27B MQ4-XT; Q8 KV/VMM; speculative decoding off. Each fresh process had a 10-second DPM warmup and a distinct 128-token warmup request, followed by the same TG4096 request. Arms ran A/B/B/A with 15-second idle intervals. Rates are daemon-reported decode throughput, excluding TTFT/load.

Arm Setting Decode tok/s
A1 off 39.0
B1 on 39.6
B2 on 39.6
A2 off 38.9

Median: 38.95 -> 39.60 tok/s, +1.67%. This is a small workload-specific observation from two samples per condition, not a universal speedup claim. All requests generated 4096 tokens with zero cached prompt tokens; decoded reasoning/content hashes matched. Token-ID equality is not inferred from text hashes.

The model contract ignored the named thinking-off budget: these measured requests produced reasoning, not final answers. This benchmark is separate from the successful functional serving battery. No per-arm profiler symbol trace was collected during timing; dispatch/replay registration has separate validation.

Artifact SHA256: 80e7c624424fd1d363ba86681d3dc1e5ac5534e0e064306a32be204c4843d0f3.
CLI MD5: 894739f626fe3bf5cd4e077bf7743a6d.
Daemon MD5: 529a278af991815057690ff384cd5846.

Verification Status

  • Product CLI/daemon release build.
  • Configuration tests: 107 passed.
  • K5120 config roundtrip and arch/shape guard tests: 2 passed.
  • Registry/replay and source-fingerprint checks from rebase validation.
  • Official serve battery: off 5/5, on 5/5; matching requests, token counts and decoded responses.
  • Official Redline harness: four consecutive positions with bit-exact HIP/blob/PM4 logits, KV and recurrent/GDN state; 659 launches, 16 families.
  • Lifecycle documentation checker and whitespace checks.
  • Independent narrow code review (gpt-6-sol high).
  • Complete official no-GPU CI with PYTHONSAFEPATH=1 PYTHONPATH="$PWD"; workspace lib tests: 4336 passed, 67 ignored. The initial Python import-shadowing failure and successful rerun are both retained.
  • Refreshed crate maps, scratch-growth and ratchet-diff checks pass.
  • All structural CI green: daemon_lines=5392 exceeds 5377, and two historical changelog sections differ. Both failures reproduce unchanged on clean beta 45f1abd1d; no policy/threshold edits included. Remote CI and maintainer review remain required.
  • Locked global speed-gate: not claimed by this targeted ABBA experiment.
  • Commit portable reproducer, original prompt and summary; prepare raw-evidence archive with file hashes.
  • Publish the evidence archive and link it in the PR.

Unrelated local research and unsuccessful optimization experiments are excluded from this submission.

HUSRCF added 5 commits October 4, 2026 22:34
Cover canonical and legacy keys through ProcessConfig serialization and exact-arch FeatureFlags guards. Document persistent/env opt-in, daemon restart, and compiler-less kernel-pack requirements.

Validation: two K5120 flag tests and lifecycle check pass. Prior controlled TG4096 AR ABBA on W7900: off 39.0/38.9, on 39.6/39.6 tok/s (+1.67% median); decoded reasoning/content identical. Reasoning remained enabled on the current model contract; this is not an answer-quality claim.
Record GPU1 no-thinking TG4096 ABBA: 38.85 to 39.60 tok/s median (+1.93%), identical decoded text. Add official-harness warmup wrapper and fixed prompt; refresh generated crate maps.

Full no-gpu-ci passes with safe Python path; workspace lib tests: 4336 passed, 67 ignored. Existing daemon-line ratchet and released-changelog failures reproduce on clean beta 45f1abd; no thresholds or policy changed. Raw logs remain attachment-only.
@HUSRCF

HUSRCF commented Oct 5, 2026

Copy link
Copy Markdown
Contributor Author

Updated against beta a193e44 in 3123f38, preserving the existing history. The conflicts were generated crate maps and the environment inventory count; regenerated them with the official scripts. GitHub now reports mergeable=true. Rechecked: 4 K5120 tests, 107 config tests, lifecycle inventory, all 48 crate maps and whitespace checks pass. No K5120 arithmetic/dispatch changes were needed. The attached GPU ABBA evidence remains explicitly tied to the previously tested commit; this remains draft pending full latest-beta integration validation.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant