Conversation
added 26 commits
October 3, 2026 18:14
Qwen3.5 single-GPU decode now binds each layer as one mixer Step (DeltaNetMixer or GatedAttention) plus one FFN Step (SwigluFfn or the sealed Moe call), executed by execute_steps. The executor bodies and every arch/shape/dtype-gated route predicate move verbatim into hipfire_dispatch::pipeline::hybrid, keyed by HybridDims instead of the model config. qwen35/program.rs only binds weights, state, KV and scratch. The EP super-op bindings delegate to the same op stages, so the handler bodies exist once. Hand arms (hidden-state capture, mrope) are unchanged. Verified on gfx1151 against the pre-change daemon (d61ba753): - greedy serve battery + chain, qwen35-4b.mq4: 10/10 turns byte-identical - greedy battery, ornith-1.5-35b-a3b.mq4r: 5/5 byte-identical - Redline --pm4: tape hash unchanged (4B 8b44d5083a01b746, 426 launches; Ornith b26bea769274b161, 843 launches), PM4 shadow exact
The two byte-identical prefill FFN bodies (DeltaNet and full-attention layers, ~1250 lines) become one rows>1 executor of SwigluFfnOp. The batched projection library they depend on (IU4/A8/FP8 producers, key-selected GEMM runners, residual/partial epilogues, fast-route gates) moves into hipfire_dispatch::pipeline::batched and takes WeightRef; qwen35 callers pass dispatch_ref(). The runtime's batched rotation helpers and GEMM/fused-QKV family singletons delegate to dispatch so each exists once. Verified on gfx1151 against the pre-change daemon: - qwen35-4b.mq4 greedy battery and 1131-token prompt: byte-identical - emulated TP2 vs single (qwen3.8-27b.mq4-xt, Partial epilogue): PASS, worst logit diff 2.968e-1, unchanged from the warpfront#755 record Rebased onto beta: the moved FFN carries beta's packed-MQ4 gate/up and down routes, the gfx1151 A4 gate/up epilogue (FfnGateOutput::Iu4A4) and the widened s4_residual_fast gate (supports_dflash_f16_residual_fusions).
The full-attention prefill sublayer (input projection, Q/K prep, KV write + flash attention, A4 epilogue, gated output projection, ~1700 lines with its producer routes) moves verbatim into hipfire_dispatch::pipeline::batched_attention. BatchSemantics, TreeVerifyCtx, valid_lane_mask and WIDENED_COMMIT_ROWS move with it; the arch crate re-exports them and binds its weights, batch scratch, flash scratch and KV cache through field-compatible views. The dense layer, the F2 chunk pair and the MoE attention prep/finish all call the shared stages. The TriAttention tap is one arch closure shared by decode and prefill; the fused qkv/qkvza key selectors join gate/up in dispatch. Verified on gfx1151 against the pre-port daemon: - qwen35-4b.mq4 greedy battery + 1131-token prompt: byte-identical - ornith-1.5-35b-a3b.mq4r greedy battery: byte-identical - emulated EP2/EP4 oracles: exact (worst abs diff 0); TP2: 2.968e-1, unchanged Rebased onto beta: the moved attend step keeps beta's FA2 context gate (fa2_gfx11_ctx_admitted) and the gfx1151 multirow assert; the fused qkv/qkvza key selectors moved to hipfire-dispatch keep beta's MQ6/HFQ6 arms.
The dense DeltaNet prefill sublayer (qkvza projection with its FP8/A4/IU4 producer routes, conv + gates, sequential/chunk-scan/tree/tape recurrence, gated norm and output projection, ~1300 lines) moves verbatim into hipfire_dispatch::pipeline::batched_deltanet. StateQuant becomes the dispatch type (re-exported); the arch crate binds its layer weights, batch scratch, recurrent state and DFlash GDN tape through field-compatible views. HybridDims gains the conv kernel width. Verified on gfx1151 against the pre-port daemon: - qwen35-4b.mq4 greedy battery, chain and 1131-token prompt: byte-identical - ornith-1.5-35b-a3b.mq4r greedy battery: byte-identical - emulated EP2/EP4 oracles exact; TP2 worst diff 2.968e-1, unchanged
The DeltaNet-MoE linear-attention body becomes batched_deltanet::execute_deltanet_moe_layer_batched and the full-attention MoE prep/output projection become batched_attention::attention_moe_prepare_batched and attention_moe_output_projection_batched. The arch crate keeps binders (deltanet_moe_layer_view, fa_moe_attention_weights) and the routed FFN step. Code moved verbatim; the MoE executors stay separate from the dense ones because their input producers, PARO Givens arms and always-residual output projection differ numerically. Verified on gfx1151 against the pre-port daemon: 4B battery/chain/1131-token prompt and Ornith battery byte-identical; PARO A3B (z-lab Qwen3.5, shisa Qwen3.6; default and HIPFIRE_PARO_BATCHED=1) greedy text byte-identical; EP2/EP4 oracles exact, TP2 2.968e-1 unchanged.
mtp_head_forward_block_only_with_pos_buf now executes [Step::GatedAttention, Step::SwigluFfn | Step::Moe] instead of a hand-written attention/FFN body. MoE MTP experts load through the trunk MoE loader (qwen35::load::load_moe_ffn under the 'mtp_head' prefix, sidecar names mapped in qwen35_tensor_name_candidates) as a sealed MoeFfnWeights, so the MoE FFN is the shared sealed decode recipe. Deletes Qwen35MtpMoe*Weights, load_mtp_moe_ffn, mtp_moe_ffn_decode and the orphaned o/ffn_out scratch. The shared steps take the trunk's certified MQ4 fusions (fused FA prep, flash epilogue gate, fused QKV, residual GEMV), so draft logits move in their last bits. Ornith MTP battery (gfx1151, greedy, --mtp on): tau per turn 1.46/1.36/1.21/0.89/1.13 vs 1.46/1.36/1.20/0.89/1.13 before, 1 of 5 turns diverges at token 3 (MoE spec decode is not AR-invariant), all coherent. The sealed MoE step matches the old hand MoE decode byte for byte. Redline PM4 retained run matches the HIP run on both builds; trunk tape 843 launches. Rebased onto beta: the MTP MoE layer config sets beta's i_gpu_start (0, fully resident), and qwen35_tensor_name_candidates stays pub for beta's host_offload_smoke example.
…ep program forward_scratch_layers now always runs the per-layer Step program; the 650-line hand arms that served hidden-state capture (DFlash) and three-section RoPE (VL) are deleted along with the qwen35 HIPFIRE_FORWARD_LOWERED=0 escape hatch and the dead triattn_tap helper. GatedAttentionOp gains an mrope operand (rope_mrope_halfsplit_f32 over pos_buf3; fused FA prep off for it); the ring-buffer capture runs where a layer's hidden state is final (after the deferred MoE combine). lower_variant stays as the EP executor's super-op program. Verified on gfx1151 against the previous build: 4B battery/chain, Ornith battery and Ornith MTP battery byte-identical; Redline decode tape hashes unchanged (4B 8b44d5083a01b746/426, Ornith b26bea769274b161/843); Qwen3.5-9B + DFlash battery byte-identical, chain 1 of 5 turns diverges late (tau 1.68->1.61, coherent); Qwen3.8-27B + vision sidecar (mrope rope_delta=-272) answers byte-identical.
EAGLE verify target) is one typed step list
`[SandwichAttention, SandwichMlp, PerLayerInput?, Scale?]` built by
`hipfire-arch-gemma4/src/program.rs`; the head is
`[RmsnormAutomatic, Gemv, Softcap]`. Executors live in the new model-neutral
`hipfire-dispatch/src/pipeline/sandwich.rs` with `rows == 1` (decode) and
`rows > 1` (batched prefill / EAGLE verify) bodies moved verbatim from
forward.rs, so every route issues the same launches. RoPE is an explicit typed
`Rope { kind: RotateHalf | PartialHalved { rot_pairs }, theta }`. KV-shared
layers and the K=V full layers are op fields, not hand branches.
`AttentionFamily::run_attend_only` serves read-only (shared-KV) attention.
The fused-route developer knobs (HIPFIRE_GEMMA4_FUSED_*, EAGLE strict) are
read once in dispatch with unchanged names and defaults.
Deleted from forward.rs: sliding/full layer arms, attn_q8_swa,
finish_attn_and_ffn, both PLE branches, batch_attn_block, batch_ffn_block,
proj_gemm_batched and the fused-key helpers (-1.6k lines). Dispatch-bypass
ledger: gemma4 17 -> 6.
Gate (gfx1151, pre-port daemon 8a6c4710 vs this build, greedy, 5 prompts x
128 tokens, HIPFIRE_GEMMA4_LOGIT_TRACE_DIR top-16 logits compared with cmp):
- gemma4-12b.mq4 decode (graph): tokens + logit trace byte-identical
- gemma4-12b.mq4 HIPFIRE_GEMMA4_EAGLE=1 (unfused route): byte-identical
- gemma4-12b.mq4 HIPFIRE_GEMMA4_PREFILL_BATCH=64 (rows>1 prefill): identical
- gemma4-e2b / e4b q8 decode and PREFILL_BATCH=64 (PLE, KV sharing): identical
- gemma4-12b.mq4 + gemma-4-12B-it-assistant drafter, EAGLE spec 3: identical
text, tau per prompt unchanged (3.368/3.459/2.667/3.447/3.514)
gfx1100/gfx1201-only batched fusions (Q8 fused prefill, wide PLE GEMM, fused
PLE activation) moved verbatim; not exercised on this host.
…; drop super-op path arch-22 EAGLE/MTP draft head now run the same typed step program as the eager stack: - `ParallelMoeMlp` (dispatch `sandwich.rs`): dense GeGLU MLP in parallel with the routed experts (router norm/scale, softmax top-8 renorm, indexed gate/up + scaled indexed down, per-expert scale, triple post-norms), moved verbatim from lowered.rs `apply_moe_branch`. Expert formats without an indexed kernel pair refuse at load instead of taking a host D2H loop. - `program.rs` binds either weight stack (eager, lowered, drafter) through one `Geometry` / `LayerRefs` / `ProgramBinding`; K/V projection is a typed `Option<KvProjection>` (query-only layers read a cache they do not write). - Draft step = `[Gemv(pre_projection), query-only [SandwichAttention, SandwichMlp, Scale?] x n, RmsnormAutomatic, Gemv(lm_head), Gemv(post_projection)]` over the target's last slot of each type. - Deleted: lowered.rs super-op facade (`Gemma4Variant`, `lower_variant`, `Gemma4Bindings`, `HIPFIRE_FORWARD_LOWERED` gate), the lowered decode hand arms, dead v1 prefill, `drafter_layer`, and the Gemma-only dead super-op vocabulary in superop.rs (`AttnFlavor`, `ActFlavor`, `RopeFlavor`, `EscapeKind::GemmaLogitSoftcap`). lowered.rs -2.5k lines. - Residual / K=V copies are `copy_f32_buffer` launches (visible to hipGraph capture and the Redline recorder); a D2D memcpy made PM4 replay inexact. Gate (gfx1151, pre-port daemon 8a6c4710, greedy 5 prompts x 128 tokens): - 12B, 12B EAGLE-strict, 12B PREFILL_BATCH=64, E2B/E4B decode and PREFILL_BATCH=64: tokens + top-16 logit traces byte-identical. - 12B + gemma-4-12B-it-assistant (q8f16) EAGLE spec 3: text identical, tau per prompt unchanged 3.368/3.459/2.667/3.447/3.514. - gemma4-26b-a4b.q8-experts-reference: with HIPFIRE_GEMMA4_FUSED_QK_ROPE=0 HIPFIRE_GEMMA4_FUSED_POSTNORM=0 byte-identical (faithful move); default takes the eager fused routes: 2/5 identical, 3/5 diverge late, coherent. - 12B HIPFIRE_BATCHED_PREFILL=1 (lowered opt-in): byte-identical with every HIPFIRE_GEMMA4_FUSED_* at 0. - 26B redline_daemon_harness --pm4 --skip-prefill: shadow exact, 1083 launches (was 1263), tape hash 249f24e5f39d425d (was 0909f3792962c37d). - gemma4-26b-a4b.mq4 (MQ6G256 gate_up / HFQ4G128 down) produced garbage pre-port via the host expert loop; it now refuses at load. Not ported: lowered::forward_prefill_batch (tool-only batched prefill with calibration taps); see docs/design/gemma4-declarative-program.md. Rebased onto beta: beta's lowered hipGraph capture/replay (lowered_graph_binding, gfx1201 default) now captures the program body; the fingerprint drops the deleted HIPFIRE_FORWARD_LOWERED switch.
The fused Gemma 4 routes (attn-norm+FWHT, Q8 q+k, qk-norm+RoPE, post-norm+residual, gate+up) are always on; only greedy EAGLE (HIPFIRE_GEMMA4_EAGLE=1) still turns off the two non-bitwise fusions. Default behaviour unchanged: 12B decode tokens and top-16 logits bitwise vs the pre-port daemon, EAGLE text and tau (3.368/3.459/2.667/3.447/3.514) identical. Daemon md5 d02826078f1f343392fed7517defbfea.
through lowered::forward_prefill_batch, now a thin driver over the shared layer program (rows <= 64 per chunk, Q8 KV). The hand-written batched prefill (v1/v2 layer bodies, run_prefill_gemm, batched MoE, pb_* scratch) is deleted. - Calibration taps live in dispatch: the sandwich projection helpers (gemv, gemm_rows) and the fused-QKV family call maybe_capture_activation with the unrotated input, so any architecture on these steps calibrates without per-architecture taps. While a collector is armed, rows == 1 fusions that never materialize the unrotated input step aside. - gemm_rows stages BF16 teacher inputs for the MFMA GEMM and runs formats without a batched kernel as one GEMV per row. - ParallelMoeMlp gains a rows > 1 executor (batched dense half, routed experts row by row); rows == 1 issues the same launches as before. - HIPFIRE_BATCHED_PREFILL / HIPFIRE_WMMA_PREFILL are removed: dense Gemma 4 always loads eager, MoE on the lowered stack. - eval_hipfire's Gemma arm sat inside the llama branch, so Gemma models scored 0 tokens; it is its own branch again. Evidence (gfx1151, daemon 21384cc30cb4df9580cb6a057fd91649): - prefill_parity_gemma4, 725-token prompt: batched vs per-token same argmax and identical 24-token continuation; 12B mq4 35.2s -> 4.4s, 26B-A4B q8-experts 19.1s -> 7.9s. - calib_sweep 12B mq4 (2048 tokens, seq 512): coverage 328/328, consistency 0, q/k/(v) and gate/up identity PASS. - eval_hipfire 12B mq4, synthetic ref: prefill KLD 21.170 vs per-token 21.122. - Serve gates vs pre-port daemon 8a6c4710: 12B, 12B pb64, E2B pb64, E4B tokens and top-16 logits bitwise; EAGLE text and tau 3.368/3.459/2.667/3.447/3.514 unchanged; 26B text identical to the first port; 26B Redline PM4 shadow exact, 1083 launches, tape 249f24e5f39d425d. Rebased onto beta: lowered.rs keeps beta's gemma4_decode_flash_partials_len (and its test), the loader files keep beta's formatting with only the gemma4_source_uses_lowered / opt-in removal applied, and the carrier log line reports beta's full-tier KV mode.
…s in the hybrid SwiGLU op swiglu_down_residual replaced llama::weight_gemv_swiglu_residual for Qwen3.5 dense decode but kept only the MQ, HFQ and unrotated arms: MQ8G256, MQ4G128, ParoQ4G128, MFP4/MFP3/MFP2 and the other rotated formats refused with UnsupportedVariant on the first decoded token. They now run a plain GEMV with the weight's own rotation and AWQ sidecar plus an in-place add, as before. Also carries beta's memory.offload_exec=cpu arm: the fused GPU arms hide the down GEMV from the execute_steps cpu_exec seam, so host-mapped layers split here (SiLU on the GPU, down projection and residual on the CPU).
…beta - CHANGELOG: the port's entries move from the released v0.4.0 section into v0.4.1 (check-changelog), and list every removed Gemma 4 switch. - debt-dispatch-bypass: qwen35 184 -> 90 and gemma4 17 -> 2 (registry 0, now debt) as measured; bypass_slack is 0 again. - env-vars inventory and crate maps regenerated.
…edicates The declarative port moved mq_f16_projection_fast_route, the batched attend step and the decode attend into hipfire-dispatch (batched.rs, batched_attention.rs, hybrid.rs) and deleted the Gemma 4 lowered hand arms, so their SITES rows move with them. The sandwich executor's calibration-collector check was named `capturing`, which the inventory scanner reads as the graph-slot capture predicate; it is renamed `calibrating`.
…e ports; correct stale design-doc vocabulary
fivetide
force-pushed
the
feat/gemma-declarative
branch
from
October 3, 2026 17:44
f7ad7a0 to
ae9a02f
Compare
…HFQ4-G128 The local gemma4-26b-a4b.mq4 produced garbage on beta for two reasons: - Its routed-expert down_proj is HFQ4-G128 with K = 704 and was quantized before b484628, which packed groups across rows. Every per-row kernel reads such a tensor as garbage. The lowered loader now refuses HFQ4-G128 tensors whose byte size is the cross-row layout, naming the cause. - 4 of its 30 layers carry MQ6G256 expert gate/up, which had no indexed route. MQ6G256 is the HFQ6G256 container over a FWHT-rotated input. The current quantizer emits MQ4G256V2 (qt 44) for the remaining layers, which the lowered loader did not map at all. Routed experts now accept MQ4G256V2 and MQ6G256 gate/up (rotate, then the shared indexed kernel), the loader maps qt 44, and the hipGraph eligibility check uses RoutedExperts::supports instead of a second format list. Verified on gfx1151 with a requantization of google/gemma-4-26B-A4B-it: coherent greedy text, serve battery clean, Redline --pm4 shadow exact (1202 launches), hipGraph decode equal to direct decode, prefill_parity same continuation. Q8-expert 26B and 12B token streams are unchanged.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Targets
beta. Rebased ontobetae268a0798; carries the #795 Qwen3.5 declarative port (its slices are the first 14 commits). Tracking: #666 G6.Design and per-slice ledgers:
docs/design/gemma4-declarative-program.md,docs/design/qwen35-declarative-program.md.1. Gemma 4 declarative refactor
Ports #795's execution model (typed
Steplists, executors and route choice inhipfire_dispatch, arch crate = declarations + bindings) tohipfire-arch-gemma4. Every Gemma 4 layer is[SandwichAttention, SandwichMlp | ParallelMoeMlp, PerLayerInput?, Scale?]declared in
program.rs; the executors live inhipfire-dispatch/src/pipeline/sandwich.rs(model-neutral sandwich-norm ops). One program serves decode (rows == 1), batched prefill and EAGLE verify (rows > 1).forward.rsParallelMoeMlp), lowered decode on the shared program, arch-22 EAGLE/MTP draft head as one step list (query-only attention over the target cache);lower_variant/Gemma4Bindingssuper-op path andHIPFIRE_FORWARD_LOWEREDhand arms deletedcalib_sweep,eval_hipfire,prefill_parity_gemma4) on the program; dispatch-owned calibration taps; old batched prefill deletedBehaviour changes:
gemma4-26b-a4b.mq4produced garbage on beta because its HFQ4-G128 expertdown_proj(K = 704) was packed across rows by the quantizer beforeb4846285e; such files, and expert formats without an indexed kernel pair, now refuse at load. A fresh requantization ofgoogle/gemma-4-26B-A4B-itruns coherently (serve battery clean, Redline--pm4shadow exact with 1202 launches, hipGraph decode equal to direct decode,prefill_parity_gemma4same continuation).HIPFIRE_GEMMA4_FUSED_{FFN,QK,QK_ROPE,POSTNORM,ATTN_NORM,PROJ}(fused routes always on; greedy EAGLEHIPFIRE_GEMMA4_EAGLE=1still turns off the two non-bitwise fusions),HIPFIRE_BATCHED_PREFILL/HIPFIRE_WMMA_PREFILL,HIPFIRE_GEMMA4_{BASELINE_ATTN,ATTN_VERIFY,GEMM_VERIFY},HIPFIRE_MOE_{BYPASS,BUCKETED}.eval_hipfire's Gemma arm sat inside the llama branch, so Gemma models scored 0 tokens.2. Rebase onto beta
Beta changed code this branch moves into new dispatch files, so those edits produced no textual conflict. They are carried into the moved copies, each in the commit that moves the code:
pipeline/batched.rs(shared SwiGLU FFN): packed-MQ4 gate/up and down routes, the gfx1151 A4 gate/up epilogue (FfnGateOutput::Iu4A4,HIPFIRE_V2B_A4_EPI),s4_residual_fastonsupports_dflash_f16_residual_fusions().pipeline/batched_attention.rs:fa2_gfx11_ctx_admittedfor the wide Q8 attend; gfx1151 in the multirow assert.families/fused_qkv.rs: the MQ6G256/HFQ6G256 arms offused_qkv(za)_key_for.mtp_head.rs:i_gpu_start: 0in the MTP MoE config (build break on beta otherwise);qwen35_tensor_name_candidatesstayspub.gemma4/lowered.rs: beta'sgemma4_decode_flash_partials_len, lowered hipGraph capture (lowered_graph_binding, now capturing the program body) and Q8 full tier are kept.Fixes on top:
hybrid::swiglu_down_residual(Qwen3.5 dense decode) had dropped the generic arm ofllama::weight_gemv_swiglu_residual: MFP4/MQ8/MQ4G128/ParoQ4G128w_downrefused withUnsupportedVarianton the first token. Restored, together with beta'smemory.offload_exec=cpuarm.railgun-certrecording inventory follows the moved predicates; the sandwich calibration check is renamedcalibrating(the scanner readcapturingas graph capture).check-changelog), dispatch-bypass rows qwen35 184 → 90 and gemma4 17 → 2 (bypass_slack0), env-vars inventory and crate maps regenerated, changed files rustfmt'd. The beta-side formatting churn in the loader is dropped.3. Evidence (gfx1151, against beta)
Beta daemon
94070f1ce9ce722fcbf4dcf7cc87da58(e268a0798) vs this branchea7b9af9f33b242a81a400486f169f02.Gemma 4:
HIPFIRE_GEMMA4_EAGLE=1(strict), E2B, E4B, E2B/E4B prefill batch 64.gemma-4-12B-it-assistantq8f16; opt-in on beta viaHIPFIRE_GEMMA4_EAGLE=1): text identical, per-prompt tau 3.368/3.459/2.667/3.447/3.514 on both.serve_harness.py --thinking off --sampling greedy --compare-transcript: 12B battery + chain, E4B battery, 26B batterytranscript_byte_identical=true.--pm4 --skip-prefill: shadow exact on both; 1083 launches (beta 1263), tape hashf2675ac01238613c(beta9f57d0ca48889bda); decode 38.2 vs 38.1 tok/s.prefill_parity_gemma4(725 tokens): batched vs per-token same argmax, identical 24-token continuation; 12B 33.8 s → 4.2 s, 26B 19.2 s → 8.2 s.calib_sweep12B: coverage 328/328, Hessian consistency 0, q/k/(v) and gate/up identity PASS.eval_hipfirevs the BF16 reference (ea5e297378c7ff692013af63d8411323),gemma4-12b.mq4, Q8 KV: per-token 0.733489, prefill 0.824203; both.kldseqbyte-identical to the pre-rebase branch. The absolute level is the-itcheckpoint on raw WT2 and the MQ4 recipe, not the engine (HF transformers cross-check in the previous revision of this description: hipfire F32 oracle vs HF F32 KL < 1e-4).Qwen3.5 family (
--compare-transcript, alltranscript_byte_identical=true):qwen35-4b.mq4battery, chain, 2k-token prompt;qwen3.8-27b.mq4-xts2k-token prompt (MQ4V2 + AWQ, A4 epilogue route);ornith-1.5-35b-a3b.mq4rAR and MTP;qwen3.6-35b-a3b.mq4rMTP;qwen3.5-9b.mq4DFlash and a 40k-token prompt (wide FA2 prefill past 32K). MFP4Qwen3.5-4Bgreedy decode byte-identical to beta (the pre-rebase head failed it).qwen3.5-9b.mq4with 20 resident layers andoffload_exec=cpubyte-identical to beta, every host-mapped step on the CPU.qwen35-4b.mq4Redline--pm4shadow exact (426 dispatches).Not measurable on this box: the packed-MQ4 FFN (gfx1100 opt-in),
s4_residual_fast(gfx1100/gfx1201), MQ6/HFQ6 Qwen3.5 checkpoints (key mapping covered by theforward_slotsunit test), gfx1201 graph capture of the lowered program.Known gaps
gemma4.rseager,lowered.rsMoE) with duplicate config/loaders; both bind the same program.e268a0798:leanup-ratchetsdaemon_lines(5392 > 5377),cargo test --libinpeacemaker-ir(two codec tests) andhipfire-arch-qwen4(state_parityGPU test),cargo clippyinhipfire-xdna.railgun-cert/src/recording.rsis not rustfmt-clean on beta; formatting it would rewrite ~1.8k lines, so this branch leaves it as is.