Skip to content

MiniMax-H3: sage takes the fused RoPE op, faster video and audio VAE decode (B200 max render 22.6 to 15.4 s, video bit-identical) - #19

Open
danielhanchen wants to merge 7 commits into
perf/h3-gguf-speed-v2from
perf/h3-gguf-speed-v3
Open

danielhanchen wants to merge 7 commits into
perf/h3-gguf-speed-v2from
perf/h3-gguf-speed-v3

Conversation

@danielhanchen

Copy link
Copy Markdown
Member

Stacked on #18 (base branch perf/h3-gguf-speed-v2). Review that first.

Problem

An Nsight Systems profile of a warm MiniMax-H3 render on a B200 (960x544, 124 frames, 4 steps, UD-Q3_K_XL DiT) put this fork's max mode at 21.4 s against 14.6 to 15.2 s for ComfyUI's int8 graph. The gap had three parts:

  • Denoiser, about 3.0 s. With sage on, forward_head_major() returns nullptr, so the fused RoPE / permute op is skipped. About 765 ms per step comes back as f32 RoPE, permute copies, casts and elementwise kernels.
  • Video VAE decode, about 3.6 s:
    • 2 s of GPU idle between tiles, from CPU-side split, blend and graph build;
    • 416 cudaMemGetInfo calls per decode;
    • per-head RMS norm and RoPE as separate kernels.
  • Audio VAE decode, about 1.1 s. The anti-aliasing depthwise convs run through mul_mat_vec_f plus im2col (1.0 s of GPU time, against 37 ms in ComfyUI).

Change

Each lever has an environment switch; =0 turns it off.

lever switch numerics
Sage takes the fused RoPE / head-major op, writing the layout sage reads (F32 Q / K, F16 V with the kv scale) SD_H3_FAST_SAGE_QKV bit-identical
Tile blend runs on a worker thread while the next tile computes (every tiled VAE decode) SD_TILE_ASYNC_MERGE bit-identical
Each H3 temporal chunk is assembled straight into the output frames on a worker thread SD_H3_VAE_ASYNC_ASSEMBLY bit-identical
One free-memory reading per decode for checks that allocate nothing (420 to 3 cudaMemGetInfo calls) SD_H3_VAE_REUSE_MEMQUERY bit-identical
Decoder q / k / v through the fused head-major RoPE op SD_H3_VAE_FUSED_QKV bit-identical
Per-head q / k RMS norm folded into that op (new ggml patch 0005, CUDA only; HIP and Vulkan keep the separate norm) SD_H3_VAE_FUSED_QK_NORM bit-identical
H3 audio VAE anti-aliasing depthwise convs use CONV_2D_DW in F32 (LTX unchanged) SD_H3_AUDIO_DIRECT_DW closer to F32, see Accuracy
Latent tile size for the H3 video VAE SD_H3_VAE_TILE=N opt-in only
  • RoPE tables are built once per tile shape.
  • scripts/unsloth/ggml-patches/0005-ggml-rope-pe-permute-rms.patch adds the fused norm. Patches 0001 to 0005 apply cleanly to the pinned ggml and reproduce the patched tree; the ggml gitlink is unchanged.

Results

Warm renders (renders 1 and 2), seconds:

GPU / mode s/step audio decode video decode render
B200 max, before / after 3.12 / 2.38 to 2.46 1.21 / 0.19 to 0.21 7.22 to 7.58 / 3.95 to 4.20 22.4 to 22.9 / 15.4
B200 default, before / after 7.18 / 7.17 1.21 / 0.20 to 0.23 7.27 to 7.38 / 3.83 to 3.89 38.8 / 34.2
G4 (sm120) max, before / after 4.33 / 3.54 1.17 / 0.14 7.08 / 5.19 26.8 / 20.7
A100 40 GB (sm80) max, before / after 7.16 / 6.12 1.36 / 0.31 13.10 / 9.07 45.4 / 36.0
  • B200 max mode is now 2.39 s per step under the profiler, level with ComfyUI's 2.37 s.
  • B200 video decode varied 6.8 to 7.6 s before and 3.8 to 4.5 s after across sessions.
  • Video decode share per lever on the B200 (each turned off alone): async assembly 1.2 s, fused q / k / v 1.0 s, async tile merge 0.45 s, fused norm 0.35 s, memory reading 0.25 s.
  • G4 and A100 default mode: sampling unchanged, video decode 1.9 s and 4.0 s faster, video bit-identical.
  • Tile 20 decodes in 3.07 s against 3.77 s at tile 16, but scores 36.2 dB PSNR / 0.09 LPIPS against tile 16, so it stays opt-in.

Accuracy

  • Video: the output stream is bit-identical before and after on B200, G4 and A100, in max and default modes.

  • All switches off: video and audio reproduce the base build exactly.

  • Audio: the direct depthwise path changes the audio on purpose. SNR against the base build: 41.5 dB on B200, 41.7 / 42.7 dB on G4, 35.0 / 41.6 dB on A100 (max / default). Against the same graph run fully in F32 it is closer than before:

    vs F32 (anti-alias convs only) vs F32 (all convs)
    before 41.3 dB 40.8 dB
    after 49.8 dB 44.3 dB

Tests

  • test-backend-ops on CUDA (B200): ROPE_PE_PERMUTE (80 cases, including the new norm variants), RMS_NORM and CONV_2D_DW pass.
  • Builds at the final commit: CUDA sm75 / 80 / 86 / 89 / 90 / 100 / 120, HIP gfx1100 / gfx1151 and Vulkan.

Not covered

  • The A100 max-mode audio SNR (35.0 dB) is lower than the other cards. The audio latents match there, so it comes from the depthwise change; it is likely content-dependent but not proven.
  • HIP and Vulkan were built, not run.
  • Colab G4 and A100 timings are one session per arm, without repeated base runs. The Colab runs used the commit before the last one, which only hardens the reuse paths.
  • Not done here, since they change numerics or need larger ggml work: bias fused into the cuBLASLt GEMM, F16 VAE activations and GPU-side tile blending.
  • The BF16 cuBLAS route for the DiT (GGML_CUDA_QUANT_CUBLAS_MIN_BATCH) was measured on the B200 without sage: 7.20 to 3.12 s per step, 0.4 GB more VRAM, and closer to an F32-dequant reference than MMQ (44.4 vs 37.9 dB PSNR, 0.017 vs 0.044 LPIPS, 23.2 vs 14.7 dB audio SNR). Making it the default is a Studio-side change, proposed separately.

With --sage-attn, forward_head_major returned nullptr, so every block fell back to the
unfused chunk / slice / rope / concat / scale / cast chain (about 0.75 s per 960x544x124
step on B200). ggml_rope_pe_permute already writes [head_dim, tokens, heads, batch], the
layout ggml_sage_attn reads: emit F32 Q, F32 K * kv_scale and F16 V * kv_scale and call
sage directly. Same arithmetic and rounding points as the unfused sage branch; the
rendered video and audio are bit-identical. SD_H3_FAST_SAGE_QKV=0 restores the old chain.
The 960x544x124 decode left the B200 idle for about 2 s between its 105 tile graphs:
the host blended each tile, zero-filled and assembled the frames, rebuilt the rotary
tables and queried free device memory four times per tile while the GPU waited.

- tiling: the blend of a batch into the output runs on a worker thread while the next
  batch is split and computed; merges stay serial and in tile order with the same
  arithmetic (SD_TILE_ASYNC_MERGE=0 merges inline).
- H3 decode: each temporal chunk is trimmed, cross-faded and copied straight into the
  trimmed output on a worker thread while the next chunk decodes, instead of
  concatenating and slicing at the end (SD_H3_VAE_ASYNC_ASSEMBLY=0 restores it).
- rotary tables are rebuilt only when the tile shape changes.
- DeviceMemoryRequest::reuse_device_query: while a runner repeats identical computes,
  a capacity check that allocates nothing new (no pending bytes, all parameters
  resident) reuses the owner's earlier free-memory reading; runner_end() drops it.
  The H3 decode turns it on for its duration (SD_H3_VAE_REUSE_MEMQUERY=0 disables).

Warm decode on B200 6.8-7.0 s -> 5.06 s; frames and audio bit-identical.
…onvs

The up/down-sampling filters of every Activation1D went through ggml_conv_1d /
ggml_conv_1d_dw: an F16 im2col plus a single-column matmul per call. Route them
through CONV_2D_DW (H=1, F32 per-channel kernel) instead, for the MiniMax-H3
audio VAE only. Falls back to the im2col graph when the backend lacks the op.
SD_H3_AUDIO_DIRECT_DW=0 restores the previous graph.
…tion

The decoder attention chunk-copied q/k/v out of the fused projection, applied the
partial RoPE as table mul/add/concat chains and then permuted and cast K/V for flash
attention. With one tile per graph, normalise q/k in place on the projection and let
ggml_rope_pe_permute (already used by the DiT) write the F32 Q and F16 K/V head-major
tensors ggml_ext_attention_prepared reads. Same products, sums and casts as the table
path, so the decoded frames are bit-identical. SD_H3_VAE_FUSED_QKV=0 restores the old
graph; batched tiles, sage and non-flash decodes keep it.
Carry ggml patch 0005 (ggml_rope_pe_permute_rms) and use it for the decoder's q/k:
the per-head norm was 7560 launches of a 64-thread kernel, about 0.45 s per
960x544x124 decode on B200. The fused op normalises each head row with the same block
reduction as the standalone norm, so the frames stay bit-identical. Backends without
it (HIP, Vulkan, GGML_CUDA_NORM_SMALL_ROWS=0) keep the separate norm;
SD_H3_VAE_FUSED_QK_NORM=0 forces that.
Fewer, larger tiles cut the per-tile fixed cost (20x20 decodes 960x544x124 in 3.07 s
vs 3.77 s at 16x16 on B200) but move the seams: 36.2 dB PSNR / 0.09 LPIPS vs the
16x16 decode, slightly closer to an untiled decode than 16x16 is. Default stays 16.
- The saved device free-memory reading is now taken only by a check that allocates
  nothing, so it is never the reading from before this owner's compute buffer was
  allocated, and any check that does not fit drops the owner's readings so the
  reclaim / evict retries see fresh ones.
- If a worker thread cannot be started, the tile merge and the chunk assembly run
  inline instead of throwing.
@chatgpt-codex-connector

chatgpt-codex-connector Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codex Review Summary

This comment shows the latest Codex review activity on this pull request.

Review Status Commit Review trigger
📝 Code Review ✅ Completed 2026-10-02T09:51:07.208594Z ee09dfa PR opened
🔒 Security Review ✅ Completed 2026-10-02T09:51:41.992028Z ee09dfa PR opened
ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review" or "@codex security review".

Codex reacts with 👀 while any review is running, comments if it has suggestions, and reacts with 👍 once all reviews finish with no findings.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant