中文 · Install · How it works · Benchmarks · Workflows
A ComfyUI plugin with source-built INT8 kernels for RTX 5090 D v2. Keep Larry v4, eight steps, your seed and your sampler. Add a model node for fused DiT execution and optional video nodes for INT8 VAE, RGB8 transfer and MP4 output.
Kitchen 0.2.36: 264.44 → 233.63 seconds per measured request — 11.65% less time. Same INT8 VAE on both sides: 6.72% less time. These are measured averages on two private representative clips, two formal repeats per arm. Timing starts from prepared conditioning and ends with a local MP4; it excludes reference/text encoding, cold checkpoint reads and network. The large comparison includes FP16 → INT8 VAE. Full protocol and raw records →
| What changes | Measured result |
|---|---|
| DiT operators, same INT8 precision and 8 steps | 6.49% less DiT time |
| Full measured cycle, native FP16 VAE → optimized INT8 VAE | 11.65% less time |
| Inverse-latency capacity equivalent | +13.19%; not a sustained serving benchmark |
| Same INT8 path, video/audio latents + RGB + PCM | Exact signatures on all four formal .36 requests |
| FP16 → INT8 VAE, historical 15 identical latents | 58.67 dB raw PSNR, MP4 SSIM 0.98907, LPIPS 0.00759 |
INT8 VAE has a very small measured difference from FP16; it is not pixel-identical. The DiT fusion preserves the selected INT8 baseline on tested inputs, not BF16 dense arithmetic. Audio stays FP32 on GPU.
Public-input check: 768×512, 124 frames, public merged Larry: 19.31 → 16.76 s (−13.19%), all four output signatures equal. First-use verification excluded; separate records.
Initial source release: Linux, RTX 5090 D v2 (SM120), Torch 2.12.0+cu130, Triton 3.7.0, Kitchen 0.2.36 and the pinned ComfyUI H3 revision. See the exact compatibility matrix. Use a separate environment; this package does not upgrade Torch automatically.
cd ComfyUI/custom_nodes
git clone https://github.com/Tokha233/ComfyUI-H3-SpeedKit.git
cd ComfyUI-H3-SpeedKit
python -m pip install comfy-kitchen==0.2.36 av
git clone --branch v4.5.0 --depth 1 https://github.com/NVIDIA/cutlass.git /tmp/h3-cutlass
python tools/build_kernels.py --cutlass /tmp/h3-cutlass
python -m h3_speedkit --probe-cudaRestart ComfyUI with --disable-cuda-malloc. Load an INT8 ConvRot H3 model, then connect:
flowchart LR
L[Load H3 INT8 / merged Larry] --> S[H3 Sigma Shift]
S --> O[H3 SpeedKit · Optimize DiT]
O --> K[KSampler: 8 · Euler · beta · CFG 1]
K --> V[H3 SpeedKit · Decode to RGB8]
K --> A[FP32 Audio VAE Decode]
V --> M[H3 SpeedKit · Save Video]
A --> M
Use H3 SpeedKit · Video VAE Loader for the upstream INT8 video VAE. It retains the old reference encoder. For Larry, follow the one-time merge instructions; do not apply a LoRA twice. Unknown compatible layouts run a slower first-use exact comparison before fast reuse. Conflicting patches and unsupported inputs are reported explicitly.
Need regular IMAGE nodes for grading or upscaling? Keep standard VAE Decode and output nodes; the DiT node can be used independently. Detailed setup, model links and troubleshooting →
- SM120 dense attention: three-slot TMA pipeline, resident Q, direct BSHD output and an exact last-block query window.
- Fused DiT preparation: native norm + modulation + ConvRot; QKV + RMS/RoPE + quantization; current-input K anchors; fused V preparation.
- INT8 GEMM: FC1 raster scheduling and outproj/FC2 gate-residual epilogues with the original BF16 rounding barriers.
- Video output: upstream INT8 decoder, old encoder, final-blend RGB8 conversion and two-slot asynchronous D2H.
- Service integration: bounded CPU MP4 export queue, with explicit completion and error propagation. Normal ComfyUI nodes wait for the saved file.
- Reproducibility: complete CUDA source/build manifests, per-request checks, public-input benchmark CLI and 30 historical optimization records.
Every optimization and its numerical contract → · Historical experiments, including rejected approaches →
Upstream work: ComfyUI #16677 merged by kijai. 1 merged ComfyUI PR and 11 open Kitchen/ComfyUI PRs: 9 ready, 2 awaiting API releases. The October 1 audit also covers new GEMM scheduling, VAE kernel opportunities and recent H3 adapters; these research candidates have no new SpeedKit GPU benchmark yet. October 1 integration test: 15.737 → 14.201 s (−9.76% sampler time) on a newer fixed Kitchen/ComfyUI baseline, with identical video/audio latent, RGB8 and PCM hashes. This experimental integration is separate from the released plugin benchmarks above. An optional Indexed Gate node isolates PR #219 in a real H3 workflow. It requires a source build of the unmerged PR and is independent of the complete Optimize DiT node. Full upstream scope and the 6.49% DiT comparison (中文) distinguish submitted kernels, remaining model integration and newer ComfyUI changes.
python -m unittest discover -s tests -v
python tools/verify_kitchen036.py
python tools/verify_metrics.py
python tools/check_repository.py
python benchmarks/run.py --helpCPU checks verify tooling and evidence, not GPU performance. benchmarks/run.py takes your own reference, public models and prompt, then compares stock/optimized outputs and complete local-file timing. First-use shape verification is intentionally slower. Other GPUs, Windows, newer Torch/ComfyUI and arbitrary model patches require separate qualification. No precompiled wheel or model weights are bundled.
Hugging Face overview · Kitchen upstream contribution analysis (中文)
Built on ComfyUI, Comfy Kitchen, SageAttention, CUTLASS, PyTorch and SGLang. Larry v4 and the INT8 VAE are upstream work. This project contributes hardware-specific adaptations, integration, output engineering and measured validation.
Combined plugin: GPL-3.0-or-later. Kernel and other file-level licenses are preserved in NOTICE and the provenance index. MiniMax H3 weights have a separate community license with geographic and commercial conditions; obtain them under upstream terms. Private prompts, business media and merged weights are not distributed.
Weight provenance: Current INT8 Larry merge and the BF16-first INT8 candidate. Completed 15-clip evaluation: small mean SSIM/audio gains, mixed per-clip results, and no meaningful speedup. The deployed weight baseline remains unchanged.
October 1 follow-up: PDMD quality and D64 attention experiments. Two PDMD INT8 weight paths were evaluated on all 15 business segments against BF16 50-step and frozen deployment media; neither met the current Larry8 quality target. Kitchen #227 adds a narrow exact SM120 D64 tile: 4.6–5.8% direct attention time reduction, and 0.96% full-decoder reduction on one representative latent in independent-process testing. These are separate from the released plugin numbers.
Kitchen follow-up: V quantization measurements and PR status. New #231 reduces the isolated 8-step sampler time by 0.96% with identical latents. Both ComfyUI drafts are now ready for review; their API release dependencies remain. This is not an additive gain over fully fused R85.