Conversation
Cover canonical and legacy keys through ProcessConfig serialization and exact-arch FeatureFlags guards. Document persistent/env opt-in, daemon restart, and compiler-less kernel-pack requirements. Validation: two K5120 flag tests and lifecycle check pass. Prior controlled TG4096 AR ABBA on W7900: off 39.0/38.9, on 39.6/39.6 tok/s (+1.67% median); decoded reasoning/content identical. Reasoning remained enabled on the current model contract; this is not an answer-quality claim.
Record GPU1 no-thinking TG4096 ABBA: 38.85 to 39.60 tok/s median (+1.93%), identical decoded text. Add official-harness warmup wrapper and fixed prompt; refresh generated crate maps. Full no-gpu-ci passes with safe Python path; workspace lib tests: 4336 passed, 67 ignored. Existing daemon-line ratchet and released-changelog failures reproduce on clean beta 45f1abd; no thresholds or policy changed. Raw logs remain attachment-only.
Contributor
Author
|
Updated against beta a193e44 in 3123f38, preserving the existing history. The conflicts were generated crate maps and the environment inventory count; regenerated them with the official scripts. GitHub now reports mergeable=true. Rechecked: 4 K5120 tests, 107 config tests, lifecycle inventory, all 48 crate maps and whitespace checks pass. No K5120 arithmetic/dispatch changes were needed. The attached GPU ABBA evidence remains explicitly tied to the previously tested commit; this remains draft pending full latest-beta integration validation. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
perf(gfx1100): add opt-in MQ4V2 K5120 gate/up decode specialization
Summary
Add an opt-in specialization of the existing MQ4V2 gate/up decode kernel for exact gfx1100 and
(gate_m, up_m, K) = (17408, 17408, 5120). It fixes the group count at compile time while preserving the existing weight representation and arithmetic. Other architectures and shapes retain their existing routes; there is no additional resident weight copy.The setting is
kernel.mq4v2_gateup_k5120, with compatibility overrideHIPFIRE_MQ4V2_GATEUP_K5120=0/1. It defaults off and is resolved through ProcessConfig/FeatureFlags at initialization. The new symbol is registered for kernel packaging and replay; compiler-less installations need an updated pack to opt in.Scope
45f1abd1d. Current integration includes betaa193e4491through merge commit3123f3865; original commits/history are preserved.cb2f84ae3,671e88e02,f689f559c.Measured Result
Primary final check: W7900 Dual Slot GPU1 / gfx1100 / HIP 7.15, same pinned model, Q8/VMM, AR, explicit
max_think_tokens=1, TG4096. ABBA off/on/on/off = 38.9 / 39.6 / 39.6 / 38.8 tok/s. Median 38.85 -> 39.60 (+1.93%). All four requests generated 4096 visible-output tokens with no reasoning or cached prompt tokens; request and decoded-text hashes match. This is separate from the earlier GPU0 reasoning run below; do not pool the two devices or modes.Reproducer and summary:
benchmarks/scripts/mq4v2_k5120_abba.shandbenchmarks/results/mq4v2-k5120-20261005/README.md.Raw evidence: mq4v2-k5120-evidence-20261005.tar.gz (337179 bytes), SHA256
38e70e8d1bbf1540920482b77f5cec432ea248f3c8804623f37974ce19e954f8.Latest-beta integration boundary
After the full validation below, beta advanced to
a193e4491. I merged it without rewriting the measured commits. The latest conflicts were confined to generated crate maps and the environment inventory count, regenerated with the official scripts; no K5120 kernel or dispatch conflict required manual resolution. On the merged tree,cargo test --locked -p rdna-compute --lib k5120passes all four tests (config lifecycle, exact arch/shape guard, inventory and replay contract),cargo test --locked -p hipfire-config --libpasses 107 tests, and lifecycle, all 48 crate maps and whitespace checks pass. Full CI and GPU ABBA results below belong to the pre-merge candidate, not a newly benchmarked latest-beta binary. This PR remains a draft pending full integration validation and remote CI.Earlier GPU0 Reasoning Run
W7900 gfx1100, physical GPU0, HIP 7.15; Qwen3.8-27B MQ4-XT; Q8 KV/VMM; speculative decoding off. Each fresh process had a 10-second DPM warmup and a distinct 128-token warmup request, followed by the same TG4096 request. Arms ran A/B/B/A with 15-second idle intervals. Rates are daemon-reported decode throughput, excluding TTFT/load.
Median: 38.95 -> 39.60 tok/s, +1.67%. This is a small workload-specific observation from two samples per condition, not a universal speedup claim. All requests generated 4096 tokens with zero cached prompt tokens; decoded reasoning/content hashes matched. Token-ID equality is not inferred from text hashes.
The model contract ignored the named thinking-off budget: these measured requests produced reasoning, not final answers. This benchmark is separate from the successful functional serving battery. No per-arm profiler symbol trace was collected during timing; dispatch/replay registration has separate validation.
Artifact SHA256:
80e7c624424fd1d363ba86681d3dc1e5ac5534e0e064306a32be204c4843d0f3.CLI MD5:
894739f626fe3bf5cd4e077bf7743a6d.Daemon MD5:
529a278af991815057690ff384cd5846.Verification Status
PYTHONSAFEPATH=1 PYTHONPATH="$PWD"; workspace lib tests: 4336 passed, 67 ignored. The initial Python import-shadowing failure and successful rerun are both retained.45f1abd1d; no policy/threshold edits included. Remote CI and maintainer review remain required.Unrelated local research and unsuccessful optimization experiments are excluded from this submission.