Skip to content

refactor(gemma4): run Gemma 4 layers as declarative Steps; calibration tools on the step program (carries #795) - #813

Open
fivetide wants to merge 27 commits into
warpfront:betafrom
fivetide:feat/gemma-declarative
Open

fivetide wants to merge 27 commits into
warpfront:betafrom
fivetide:feat/gemma-declarative

Conversation

@fivetide

@fivetide fivetide commented Oct 3, 2026 •

Copy link
Copy Markdown
Collaborator

Targets beta. Rebased onto beta e268a0798; carries the #795 Qwen3.5 declarative port (its slices are the first 14 commits). Tracking: #666 G6.

Design and per-slice ledgers: docs/design/gemma4-declarative-program.md, docs/design/qwen35-declarative-program.md.

1. Gemma 4 declarative refactor

Ports #795's execution model (typed Step lists, executors and route choice in hipfire_dispatch, arch crate = declarations + bindings) to hipfire-arch-gemma4. Every Gemma 4 layer is

[SandwichAttention, SandwichMlp | ParallelMoeMlp, PerLayerInput?, Scale?]

declared in program.rs; the executors live in hipfire-dispatch/src/pipeline/sandwich.rs (model-neutral sandwich-norm ops). One program serves decode (rows == 1), batched prefill and EAGLE verify (rows > 1).

Slice
12B dense and E2B/E4B (PLE, KV sharing, double-wide FFN): decode, batched prefill, EAGLE verify; ~1.6k lines of hand layer code out of forward.rs
26B-A4B MoE (ParallelMoeMlp), lowered decode on the shared program, arch-22 EAGLE/MTP draft head as one step list (query-only attention over the target cache); lower_variant / Gemma4Bindings super-op path and HIPFIRE_FORWARD_LOWERED hand arms deleted
Calibration tools (calib_sweep, eval_hipfire, prefill_parity_gemma4) on the program; dispatch-owned calibration taps; old batched prefill deleted

Behaviour changes:

  • 26B-A4B: the lowered stack loads MQ4G256V2 (qt 44) and runs MQ4G256V2 and MQ6G256 routed experts on the indexed kernels, which is what the current quantizer emits. The old local gemma4-26b-a4b.mq4 produced garbage on beta because its HFQ4-G128 expert down_proj (K = 704) was packed across rows by the quantizer before b4846285e; such files, and expert formats without an indexed kernel pair, now refuse at load. A fresh requantization of google/gemma-4-26B-A4B-it runs coherently (serve battery clean, Redline --pm4 shadow exact with 1202 launches, hipGraph decode equal to direct decode, prefill_parity_gemma4 same continuation).
  • Removed developer switches: HIPFIRE_GEMMA4_FUSED_{FFN,QK,QK_ROPE,POSTNORM,ATTN_NORM,PROJ} (fused routes always on; greedy EAGLE HIPFIRE_GEMMA4_EAGLE=1 still turns off the two non-bitwise fusions), HIPFIRE_BATCHED_PREFILL / HIPFIRE_WMMA_PREFILL, HIPFIRE_GEMMA4_{BASELINE_ATTN,ATTN_VERIFY,GEMM_VERIFY}, HIPFIRE_MOE_{BYPASS,BUCKETED}.
  • Fix: eval_hipfire's Gemma arm sat inside the llama branch, so Gemma models scored 0 tokens.

2. Rebase onto beta

Beta changed code this branch moves into new dispatch files, so those edits produced no textual conflict. They are carried into the moved copies, each in the commit that moves the code:

  • pipeline/batched.rs (shared SwiGLU FFN): packed-MQ4 gate/up and down routes, the gfx1151 A4 gate/up epilogue (FfnGateOutput::Iu4A4, HIPFIRE_V2B_A4_EPI), s4_residual_fast on supports_dflash_f16_residual_fusions().
  • pipeline/batched_attention.rs: fa2_gfx11_ctx_admitted for the wide Q8 attend; gfx1151 in the multirow assert.
  • families/fused_qkv.rs: the MQ6G256/HFQ6G256 arms of fused_qkv(za)_key_for.
  • mtp_head.rs: i_gpu_start: 0 in the MTP MoE config (build break on beta otherwise); qwen35_tensor_name_candidates stays pub.
  • gemma4/lowered.rs: beta's gemma4_decode_flash_partials_len, lowered hipGraph capture (lowered_graph_binding, now capturing the program body) and Q8 full tier are kept.

Fixes on top:

  • hybrid::swiglu_down_residual (Qwen3.5 dense decode) had dropped the generic arm of llama::weight_gemv_swiglu_residual: MFP4/MQ8/MQ4G128/ParoQ4G128 w_down refused with UnsupportedVariant on the first token. Restored, together with beta's memory.offload_exec=cpu arm.
  • railgun-cert recording inventory follows the moved predicates; the sandwich calibration check is renamed calibrating (the scanner read capturing as graph capture).
  • Ledger/doc bookkeeping: CHANGELOG entries move into v0.4.1 (check-changelog), dispatch-bypass rows qwen35 184 → 90 and gemma4 17 → 2 (bypass_slack 0), env-vars inventory and crate maps regenerated, changed files rustfmt'd. The beta-side formatting churn in the loader is dropped.

3. Evidence (gfx1151, against beta)

Beta daemon 94070f1ce9ce722fcbf4dcf7cc87da58 (e268a0798) vs this branch ea7b9af9f33b242a81a400486f169f02.

Gemma 4:

  • Greedy daemon, 5 prompts per row, token stream and per-step top-16 logits bitwise: 12B AR, 12B prefill batch 64, 12B HIPFIRE_GEMMA4_EAGLE=1 (strict), E2B, E4B, E2B/E4B prefill batch 64.
  • EAGLE (gemma-4-12B-it-assistant q8f16; opt-in on beta via HIPFIRE_GEMMA4_EAGLE=1): text identical, per-prompt tau 3.368/3.459/2.667/3.447/3.514 on both.
  • 26B-A4B q8-experts: 3/5 identical, 2 diverge late (chars 295, 322), coherent (shared fused qk-norm+RoPE / post-norm routes).
  • serve_harness.py --thinking off --sampling greedy --compare-transcript: 12B battery + chain, E4B battery, 26B battery transcript_byte_identical=true.
  • 26B Redline --pm4 --skip-prefill: shadow exact on both; 1083 launches (beta 1263), tape hash f2675ac01238613c (beta 9f57d0ca48889bda); decode 38.2 vs 38.1 tok/s.
  • prefill_parity_gemma4 (725 tokens): batched vs per-token same argmax, identical 24-token continuation; 12B 33.8 s → 4.2 s, 26B 19.2 s → 8.2 s.
  • calib_sweep 12B: coverage 328/328, Hessian consistency 0, q/k/(v) and gate/up identity PASS.
  • eval_hipfire vs the BF16 reference (ea5e297378c7ff692013af63d8411323), gemma4-12b.mq4, Q8 KV: per-token 0.733489, prefill 0.824203; both .kldseq byte-identical to the pre-rebase branch. The absolute level is the -it checkpoint on raw WT2 and the MQ4 recipe, not the engine (HF transformers cross-check in the previous revision of this description: hipfire F32 oracle vs HF F32 KL < 1e-4).

Qwen3.5 family (--compare-transcript, all transcript_byte_identical=true): qwen35-4b.mq4 battery, chain, 2k-token prompt; qwen3.8-27b.mq4-xts 2k-token prompt (MQ4V2 + AWQ, A4 epilogue route); ornith-1.5-35b-a3b.mq4r AR and MTP; qwen3.6-35b-a3b.mq4r MTP; qwen3.5-9b.mq4 DFlash and a 40k-token prompt (wide FA2 prefill past 32K). MFP4 Qwen3.5-4B greedy decode byte-identical to beta (the pre-rebase head failed it). qwen3.5-9b.mq4 with 20 resident layers and offload_exec=cpu byte-identical to beta, every host-mapped step on the CPU. qwen35-4b.mq4 Redline --pm4 shadow exact (426 dispatches).

Not measurable on this box: the packed-MQ4 FFN (gfx1100 opt-in), s4_residual_fast (gfx1100/gfx1201), MQ6/HFQ6 Qwen3.5 checkpoints (key mapping covered by the forward_slots unit test), gfx1201 graph capture of the lowered program.

Known gaps

  • Two Gemma 4 weight stacks remain (gemma4.rs eager, lowered.rs MoE) with duplicate config/loaders; both bind the same program.
  • The MoE expert-format refusal runs at carrier load, after the daemon unloads the prior model; admission does not probe expert dtypes yet.
  • CI checks that also fail on beta e268a0798: leanup-ratchets daemon_lines (5392 > 5377), cargo test --lib in peacemaker-ir (two codec tests) and hipfire-arch-qwen4 (state_parity GPU test), cargo clippy in hipfire-xdna. railgun-cert/src/recording.rs is not rustfmt-clean on beta; formatting it would rewrite ~1.8k lines, so this branch leaves it as is.

@fivetide
fivetide requested a review from Kaden-Schutt as a code owner October 3, 2026 13:53
@fivetide
fivetide changed the base branch from feat/qwen35-declarative to beta October 3, 2026 15:42
Bjoern Agent added 26 commits October 3, 2026 18:14
Qwen3.5 single-GPU decode now binds each layer as one mixer Step
(DeltaNetMixer or GatedAttention) plus one FFN Step (SwigluFfn or the
sealed Moe call), executed by execute_steps. The executor bodies and every
arch/shape/dtype-gated route predicate move verbatim into
hipfire_dispatch::pipeline::hybrid, keyed by HybridDims instead of the
model config. qwen35/program.rs only binds weights, state, KV and scratch.

The EP super-op bindings delegate to the same op stages, so the handler
bodies exist once. Hand arms (hidden-state capture, mrope) are unchanged.

Verified on gfx1151 against the pre-change daemon (d61ba753):
- greedy serve battery + chain, qwen35-4b.mq4: 10/10 turns byte-identical
- greedy battery, ornith-1.5-35b-a3b.mq4r: 5/5 byte-identical
- Redline --pm4: tape hash unchanged (4B 8b44d5083a01b746, 426 launches;
  Ornith b26bea769274b161, 843 launches), PM4 shadow exact
The two byte-identical prefill FFN bodies (DeltaNet and full-attention
layers, ~1250 lines) become one rows>1 executor of SwigluFfnOp. The
batched projection library they depend on (IU4/A8/FP8 producers,
key-selected GEMM runners, residual/partial epilogues, fast-route gates)
moves into hipfire_dispatch::pipeline::batched and takes WeightRef; qwen35
callers pass dispatch_ref(). The runtime's batched rotation helpers and
GEMM/fused-QKV family singletons delegate to dispatch so each exists once.

Verified on gfx1151 against the pre-change daemon:
- qwen35-4b.mq4 greedy battery and 1131-token prompt: byte-identical
- emulated TP2 vs single (qwen3.8-27b.mq4-xt, Partial epilogue): PASS,
  worst logit diff 2.968e-1, unchanged from the warpfront#755 record

Rebased onto beta: the moved FFN carries beta's packed-MQ4 gate/up and
down routes, the gfx1151 A4 gate/up epilogue (FfnGateOutput::Iu4A4) and the
widened s4_residual_fast gate (supports_dflash_f16_residual_fusions).
The full-attention prefill sublayer (input projection, Q/K prep, KV write
+ flash attention, A4 epilogue, gated output projection, ~1700 lines with
its producer routes) moves verbatim into
hipfire_dispatch::pipeline::batched_attention. BatchSemantics,
TreeVerifyCtx, valid_lane_mask and WIDENED_COMMIT_ROWS move with it; the
arch crate re-exports them and binds its weights, batch scratch, flash
scratch and KV cache through field-compatible views. The dense layer, the
F2 chunk pair and the MoE attention prep/finish all call the shared stages.
The TriAttention tap is one arch closure shared by decode and prefill; the
fused qkv/qkvza key selectors join gate/up in dispatch.

Verified on gfx1151 against the pre-port daemon:
- qwen35-4b.mq4 greedy battery + 1131-token prompt: byte-identical
- ornith-1.5-35b-a3b.mq4r greedy battery: byte-identical
- emulated EP2/EP4 oracles: exact (worst abs diff 0); TP2: 2.968e-1, unchanged

Rebased onto beta: the moved attend step keeps beta's FA2 context gate
(fa2_gfx11_ctx_admitted) and the gfx1151 multirow assert; the fused
qkv/qkvza key selectors moved to hipfire-dispatch keep beta's MQ6/HFQ6 arms.
The dense DeltaNet prefill sublayer (qkvza projection with its FP8/A4/IU4
producer routes, conv + gates, sequential/chunk-scan/tree/tape recurrence,
gated norm and output projection, ~1300 lines) moves verbatim into
hipfire_dispatch::pipeline::batched_deltanet. StateQuant becomes the
dispatch type (re-exported); the arch crate binds its layer weights, batch
scratch, recurrent state and DFlash GDN tape through field-compatible views.
HybridDims gains the conv kernel width.

Verified on gfx1151 against the pre-port daemon:
- qwen35-4b.mq4 greedy battery, chain and 1131-token prompt: byte-identical
- ornith-1.5-35b-a3b.mq4r greedy battery: byte-identical
- emulated EP2/EP4 oracles exact; TP2 worst diff 2.968e-1, unchanged
The DeltaNet-MoE linear-attention body becomes
batched_deltanet::execute_deltanet_moe_layer_batched and the full-attention
MoE prep/output projection become batched_attention::attention_moe_prepare_batched
and attention_moe_output_projection_batched. The arch crate keeps binders
(deltanet_moe_layer_view, fa_moe_attention_weights) and the routed FFN step.
Code moved verbatim; the MoE executors stay separate from the dense ones
because their input producers, PARO Givens arms and always-residual output
projection differ numerically.

Verified on gfx1151 against the pre-port daemon: 4B battery/chain/1131-token
prompt and Ornith battery byte-identical; PARO A3B (z-lab Qwen3.5, shisa
Qwen3.6; default and HIPFIRE_PARO_BATCHED=1) greedy text byte-identical;
EP2/EP4 oracles exact, TP2 2.968e-1 unchanged.
mtp_head_forward_block_only_with_pos_buf now executes
[Step::GatedAttention, Step::SwigluFfn | Step::Moe] instead of a hand-written
attention/FFN body. MoE MTP experts load through the trunk MoE loader
(qwen35::load::load_moe_ffn under the 'mtp_head' prefix, sidecar names mapped
in qwen35_tensor_name_candidates) as a sealed MoeFfnWeights, so the MoE FFN is
the shared sealed decode recipe. Deletes Qwen35MtpMoe*Weights,
load_mtp_moe_ffn, mtp_moe_ffn_decode and the orphaned o/ffn_out scratch.

The shared steps take the trunk's certified MQ4 fusions (fused FA prep, flash
epilogue gate, fused QKV, residual GEMV), so draft logits move in their last
bits. Ornith MTP battery (gfx1151, greedy, --mtp on): tau per turn
1.46/1.36/1.21/0.89/1.13 vs 1.46/1.36/1.20/0.89/1.13 before, 1 of 5 turns
diverges at token 3 (MoE spec decode is not AR-invariant), all coherent. The
sealed MoE step matches the old hand MoE decode byte for byte. Redline PM4
retained run matches the HIP run on both builds; trunk tape 843 launches.

Rebased onto beta: the MTP MoE layer config sets beta's i_gpu_start (0,
fully resident), and qwen35_tensor_name_candidates stays pub for beta's
host_offload_smoke example.
…ep program

forward_scratch_layers now always runs the per-layer Step program;
the 650-line hand arms that served hidden-state capture (DFlash) and
three-section RoPE (VL) are deleted along with the qwen35
HIPFIRE_FORWARD_LOWERED=0 escape hatch and the dead triattn_tap helper.
GatedAttentionOp gains an mrope operand (rope_mrope_halfsplit_f32 over
pos_buf3; fused FA prep off for it); the ring-buffer capture runs where a
layer's hidden state is final (after the deferred MoE combine).
lower_variant stays as the EP executor's super-op program.

Verified on gfx1151 against the previous build: 4B battery/chain, Ornith
battery and Ornith MTP battery byte-identical; Redline decode tape hashes
unchanged (4B 8b44d5083a01b746/426, Ornith b26bea769274b161/843);
Qwen3.5-9B + DFlash battery byte-identical, chain 1 of 5 turns diverges
late (tau 1.68->1.61, coherent); Qwen3.8-27B + vision sidecar (mrope
rope_delta=-272) answers byte-identical.
EAGLE verify target) is one typed step list
`[SandwichAttention, SandwichMlp, PerLayerInput?, Scale?]` built by
`hipfire-arch-gemma4/src/program.rs`; the head is
`[RmsnormAutomatic, Gemv, Softcap]`. Executors live in the new model-neutral
`hipfire-dispatch/src/pipeline/sandwich.rs` with `rows == 1` (decode) and
`rows > 1` (batched prefill / EAGLE verify) bodies moved verbatim from
forward.rs, so every route issues the same launches. RoPE is an explicit typed
`Rope { kind: RotateHalf | PartialHalved { rot_pairs }, theta }`. KV-shared
layers and the K=V full layers are op fields, not hand branches.
`AttentionFamily::run_attend_only` serves read-only (shared-KV) attention.
The fused-route developer knobs (HIPFIRE_GEMMA4_FUSED_*, EAGLE strict) are
read once in dispatch with unchanged names and defaults.

Deleted from forward.rs: sliding/full layer arms, attn_q8_swa,
finish_attn_and_ffn, both PLE branches, batch_attn_block, batch_ffn_block,
proj_gemm_batched and the fused-key helpers (-1.6k lines). Dispatch-bypass
ledger: gemma4 17 -> 6.

Gate (gfx1151, pre-port daemon 8a6c4710 vs this build, greedy, 5 prompts x
128 tokens, HIPFIRE_GEMMA4_LOGIT_TRACE_DIR top-16 logits compared with cmp):
- gemma4-12b.mq4 decode (graph): tokens + logit trace byte-identical
- gemma4-12b.mq4 HIPFIRE_GEMMA4_EAGLE=1 (unfused route): byte-identical
- gemma4-12b.mq4 HIPFIRE_GEMMA4_PREFILL_BATCH=64 (rows>1 prefill): identical
- gemma4-e2b / e4b q8 decode and PREFILL_BATCH=64 (PLE, KV sharing): identical
- gemma4-12b.mq4 + gemma-4-12B-it-assistant drafter, EAGLE spec 3: identical
  text, tau per prompt unchanged (3.368/3.459/2.667/3.447/3.514)
gfx1100/gfx1201-only batched fusions (Q8 fused prefill, wide PLE GEMM, fused
PLE activation) moved verbatim; not exercised on this host.
…; drop super-op path

arch-22 EAGLE/MTP draft head now run the same typed step program as the
eager stack:

- `ParallelMoeMlp` (dispatch `sandwich.rs`): dense GeGLU MLP in parallel with
  the routed experts (router norm/scale, softmax top-8 renorm, indexed
  gate/up + scaled indexed down, per-expert scale, triple post-norms), moved
  verbatim from lowered.rs `apply_moe_branch`. Expert formats without an
  indexed kernel pair refuse at load instead of taking a host D2H loop.
- `program.rs` binds either weight stack (eager, lowered, drafter) through
  one `Geometry` / `LayerRefs` / `ProgramBinding`; K/V projection is a typed
  `Option<KvProjection>` (query-only layers read a cache they do not write).
- Draft step = `[Gemv(pre_projection), query-only [SandwichAttention,
  SandwichMlp, Scale?] x n, RmsnormAutomatic, Gemv(lm_head),
  Gemv(post_projection)]` over the target's last slot of each type.
- Deleted: lowered.rs super-op facade (`Gemma4Variant`, `lower_variant`,
  `Gemma4Bindings`, `HIPFIRE_FORWARD_LOWERED` gate), the lowered decode hand
  arms, dead v1 prefill, `drafter_layer`, and the Gemma-only dead super-op
  vocabulary in superop.rs (`AttnFlavor`, `ActFlavor`, `RopeFlavor`,
  `EscapeKind::GemmaLogitSoftcap`). lowered.rs -2.5k lines.
- Residual / K=V copies are `copy_f32_buffer` launches (visible to hipGraph
  capture and the Redline recorder); a D2D memcpy made PM4 replay inexact.

Gate (gfx1151, pre-port daemon 8a6c4710, greedy 5 prompts x 128 tokens):
- 12B, 12B EAGLE-strict, 12B PREFILL_BATCH=64, E2B/E4B decode and
  PREFILL_BATCH=64: tokens + top-16 logit traces byte-identical.
- 12B + gemma-4-12B-it-assistant (q8f16) EAGLE spec 3: text identical, tau
  per prompt unchanged 3.368/3.459/2.667/3.447/3.514.
- gemma4-26b-a4b.q8-experts-reference: with HIPFIRE_GEMMA4_FUSED_QK_ROPE=0
  HIPFIRE_GEMMA4_FUSED_POSTNORM=0 byte-identical (faithful move); default
  takes the eager fused routes: 2/5 identical, 3/5 diverge late, coherent.
- 12B HIPFIRE_BATCHED_PREFILL=1 (lowered opt-in): byte-identical with every
  HIPFIRE_GEMMA4_FUSED_* at 0.
- 26B redline_daemon_harness --pm4 --skip-prefill: shadow exact, 1083
  launches (was 1263), tape hash 249f24e5f39d425d (was 0909f3792962c37d).
- gemma4-26b-a4b.mq4 (MQ6G256 gate_up / HFQ4G128 down) produced garbage
  pre-port via the host expert loop; it now refuses at load.

Not ported: lowered::forward_prefill_batch (tool-only batched prefill with
calibration taps); see docs/design/gemma4-declarative-program.md.

Rebased onto beta: beta's lowered hipGraph capture/replay
(lowered_graph_binding, gfx1201 default) now captures the program body;
the fingerprint drops the deleted HIPFIRE_FORWARD_LOWERED switch.
The fused Gemma 4 routes (attn-norm+FWHT, Q8 q+k, qk-norm+RoPE,
post-norm+residual, gate+up) are always on; only greedy EAGLE
(HIPFIRE_GEMMA4_EAGLE=1) still turns off the two non-bitwise fusions.
Default behaviour unchanged: 12B decode tokens and top-16 logits bitwise
vs the pre-port daemon, EAGLE text and tau (3.368/3.459/2.667/3.447/3.514)
identical. Daemon md5 d02826078f1f343392fed7517defbfea.
through lowered::forward_prefill_batch, now a thin driver over the shared
layer program (rows <= 64 per chunk, Q8 KV). The hand-written batched
prefill (v1/v2 layer bodies, run_prefill_gemm, batched MoE, pb_* scratch)
is deleted.

- Calibration taps live in dispatch: the sandwich projection helpers (gemv,
  gemm_rows) and the fused-QKV family call maybe_capture_activation with the
  unrotated input, so any architecture on these steps calibrates without
  per-architecture taps. While a collector is armed, rows == 1 fusions that
  never materialize the unrotated input step aside.
- gemm_rows stages BF16 teacher inputs for the MFMA GEMM and runs formats
  without a batched kernel as one GEMV per row.
- ParallelMoeMlp gains a rows > 1 executor (batched dense half, routed
  experts row by row); rows == 1 issues the same launches as before.
- HIPFIRE_BATCHED_PREFILL / HIPFIRE_WMMA_PREFILL are removed: dense Gemma 4
  always loads eager, MoE on the lowered stack.
- eval_hipfire's Gemma arm sat inside the llama branch, so Gemma models
  scored 0 tokens; it is its own branch again.

Evidence (gfx1151, daemon 21384cc30cb4df9580cb6a057fd91649):
- prefill_parity_gemma4, 725-token prompt: batched vs per-token same argmax
  and identical 24-token continuation; 12B mq4 35.2s -> 4.4s, 26B-A4B
  q8-experts 19.1s -> 7.9s.
- calib_sweep 12B mq4 (2048 tokens, seq 512): coverage 328/328, consistency
  0, q/k/(v) and gate/up identity PASS.
- eval_hipfire 12B mq4, synthetic ref: prefill KLD 21.170 vs per-token
  21.122.
- Serve gates vs pre-port daemon 8a6c4710: 12B, 12B pb64, E2B pb64, E4B
  tokens and top-16 logits bitwise; EAGLE text and tau
  3.368/3.459/2.667/3.447/3.514 unchanged; 26B text identical to the first
  port; 26B Redline PM4 shadow exact, 1083 launches, tape 249f24e5f39d425d.

Rebased onto beta: lowered.rs keeps beta's gemma4_decode_flash_partials_len
(and its test), the loader files keep beta's formatting with only the
gemma4_source_uses_lowered / opt-in removal applied, and the carrier log
line reports beta's full-tier KV mode.
…s in the hybrid SwiGLU op

swiglu_down_residual replaced llama::weight_gemv_swiglu_residual for Qwen3.5
dense decode but kept only the MQ, HFQ and unrotated arms: MQ8G256, MQ4G128,
ParoQ4G128, MFP4/MFP3/MFP2 and the other rotated formats refused with
UnsupportedVariant on the first decoded token. They now run a plain GEMV with
the weight's own rotation and AWQ sidecar plus an in-place add, as before.

Also carries beta's memory.offload_exec=cpu arm: the fused GPU arms hide the
down GEMV from the execute_steps cpu_exec seam, so host-mapped layers split
here (SiLU on the GPU, down projection and residual on the CPU).
…beta

- CHANGELOG: the port's entries move from the released v0.4.0 section into
  v0.4.1 (check-changelog), and list every removed Gemma 4 switch.
- debt-dispatch-bypass: qwen35 184 -> 90 and gemma4 17 -> 2 (registry 0,
  now debt) as measured; bypass_slack is 0 again.
- env-vars inventory and crate maps regenerated.
…edicates

The declarative port moved mq_f16_projection_fast_route, the batched attend
step and the decode attend into hipfire-dispatch (batched.rs,
batched_attention.rs, hybrid.rs) and deleted the Gemma 4 lowered hand arms,
so their SITES rows move with them.

The sandwich executor's calibration-collector check was named
`capturing`, which the inventory scanner reads as the graph-slot capture
predicate; it is renamed `calibrating`.
…e ports; correct stale design-doc vocabulary
@fivetide
fivetide force-pushed the feat/gemma-declarative branch from f7ad7a0 to ae9a02f Compare October 3, 2026 17:44
@fivetide fivetide changed the title refactor(gemma4): run Gemma 4 layers as declarative Steps; calibration tools on the step program (stacked on #795) refactor(gemma4): run Gemma 4 layers as declarative Steps; calibration tools on the step program (carries #795) Oct 3, 2026
…HFQ4-G128

The local gemma4-26b-a4b.mq4 produced garbage on beta for two reasons:

- Its routed-expert down_proj is HFQ4-G128 with K = 704 and was quantized
  before b484628, which packed groups across rows. Every per-row kernel
  reads such a tensor as garbage. The lowered loader now refuses HFQ4-G128
  tensors whose byte size is the cross-row layout, naming the cause.
- 4 of its 30 layers carry MQ6G256 expert gate/up, which had no indexed
  route. MQ6G256 is the HFQ6G256 container over a FWHT-rotated input.

The current quantizer emits MQ4G256V2 (qt 44) for the remaining layers,
which the lowered loader did not map at all. Routed experts now accept
MQ4G256V2 and MQ6G256 gate/up (rotate, then the shared indexed kernel), the
loader maps qt 44, and the hipGraph eligibility check uses
RoutedExperts::supports instead of a second format list.

Verified on gfx1151 with a requantization of google/gemma-4-26B-A4B-it:
coherent greedy text, serve battery clean, Redline --pm4 shadow exact (1202
launches), hipGraph decode equal to direct decode, prefill_parity same
continuation. Q8-expert 26B and 12B token streams are unchanged.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant