DiffusionGemma (google/diffusiongemma-26B-A4B-it) is a text generation model
that uses block-diffusion rather than left-to-right autoregressive decoding. This
page explains the generation mechanism, the CLI flags that control it, measured
throughput, and current limitations.
Autoregressive models predict one token at a time, each conditioned on all previous tokens. DiffusionGemma decodes a canvas of tokens at once through iterative denoising: the canvas starts as uniformly random token ids, and the model refines the whole block each step until tokens reach stable, high-confidence predictions and the step loop stops early or exhausts the step budget.
The generation loop works as follows. First, the prompt is encoded into a read-only prefix stored in dense FP16 KV caches. Then, for each output block:
- A canvas is initialized from random token ids (seeded by
--seedfor reproducibility) with lengthmin(max_canvas, max(remaining, min_canvas)). - Up to
--max-denoising-stepsdenoising steps run. Each step embeds the current canvas with a learned self-conditioning signal, attends bidirectionally within the canvas while attending causally to the prefix, and uses the sampler to accept new token predictions. - After each step the stable-and-confident early stop checks whether the canvas
contents have been unchanged for
stability_thresholdconsecutive steps. If so, the block is committed and the loop moves on. - The committed canvas is appended to the prefix and, if any output remains, the next canvas block begins.
Two per-step acceptance samplers are available via --diffusion-sampler:
-
entropy-bound(default): For each canvas position, compute a per-token entropy. Sort positions by ascending entropy. Accept positions in that order until the running prefix sum of entropies exceedsentropy_bound(read from the checkpoint'sgeneration_config). At least one position is always accepted. This reproduces the mlx-vlm reference(cumsum - cummax) <= boundrule. -
confidence-threshold: Accept all unrevealed positions where the top-1 token probability exceeds--diffusion-threshold(default 0.9). If no position qualifies, the highest-confidence unrevealed position is accepted unconditionally so the loop always makes progress.
All diffusion flags appear in the mlxcel generate --help output under the
Diffusion Options heading. Autoregressive models ignore them entirely.
| Flag | Default | Notes |
|---|---|---|
--max-denoising-steps N |
checkpoint generation_config (typically 48) |
Maximum denoising iterations per canvas block. Lower values trade quality for speed. |
--diffusion-sampler entropy-bound|confidence-threshold |
entropy-bound |
Per-step token acceptance strategy. See above. |
--diffusion-threshold FLOAT |
0.9 |
Confidence threshold for confidence-threshold sampler. |
--diffusion-min-canvas-length N |
64 |
Shortest canvas for the generation tail. Prevents tiny final blocks. |
--diffusion-max-canvas-length N |
model canvas_length (typically 256) |
Cap on the per-block canvas length. |
--diffusion-full-canvas |
off | Always allocate the model's full canvas_length per block. |
--seed N |
unset (nondeterministic) | Seeds MLX's global RNG, which also controls canvas initialization. Two runs with the same seed and temperature 0 produce identical output. |
Per-step cost dominates throughput. The stable-and-confident early stop fires earlier when the model converges quickly, so tokens/sec varies with prompt and output difficulty.
Measured on M1 Ultra with diffusiongemma-26B-A4B-it-4bit:
- ~20 tok/s for short answers (64-token canvas, converging in about 12 of 48 steps)
- ~8 tok/s for longer outputs requiring most of the 48-step budget on a 256-token canvas
Per-step time on the same hardware: 254 ms/step for a 64-token canvas (103% of the mlx-vlm Python reference at 261 ms/step) and 651 ms/step for a 256-token canvas (96% of the reference at 626 ms/step).
Model load: 13.5 GiB resident (4-bit quantized checkpoint, M1 Ultra).
Canvas noise is the main source of nondeterminism. Pass --seed N to seed MLX's
global RNG before generation begins. Two runs with the same seed, same prompt, and
--max-tokens produce the same commit stream at temperature 0. The seed also
affects the autoregressive sampling path, so it is useful for both model families.
The environment variable MLXCEL_DIFFUSION_DEBUG_CANVAS=1 replaces all random
canvas initialization with a fixed deterministic pattern. This is intended for
cross-implementation parity testing against the Python reference; output is not
useful for normal generation. See Environment variables
for details.
Pass one or more --image <path> flags to include images in the prompt. The flag is repeatable for multi-image prompts.
mlxcel generate -m google/diffusiongemma-26B-A4B-it \
--image photo.jpg \
-p "Describe what you see." \
--max-tokens 256 \
--seed 42
Preprocessing reuses the Gemma 4 vision pipeline: images are resized and padded to 768x768, then projected into 256 soft tokens per image. Each image block receives bidirectional attention during prefill, matching the same overlay used by Gemma 4 Unified. The image tokens are inserted into the prompt sequence between the boi and eoi sentinel tokens.
Video and audio inputs are not supported and are rejected with a clear error message.
mlxcel-server --model <diffusion-model> (or mlxcel serve -m <diffusion-model>) loads the checkpoint and serves /v1/chat/completions and /v1/completions. Both streaming (SSE) and non-streaming responses are supported. Requests are served serially: DiffusionGemma cannot join the batch scheduler, so the server queues them in arrival order and processes one at a time. The pending request queue still honors --max-queue-depth and returns the same overloaded response as batched serving when full. That is the intended design, not a temporary limitation.
Serve-level flags that affect diffusion models:
| Flag | Default | Notes |
|---|---|---|
--diffusion-sampler entropy-bound|confidence-threshold |
entropy-bound |
Per-step token acceptance strategy. |
--diffusion-threshold FLOAT |
0.9 |
Confidence threshold; validated to [0, 1] at startup. |
--max-denoising-steps N |
checkpoint default (typically 48) | Maximum denoising iterations per canvas block. |
Image input follows the standard OpenAI image_url content part format. Requests that include video or audio are rejected with a clear error; the connection stays open and the next request is served normally.
The checkpoint is a thinking-channel model. The assistant fills a reasoning_content channel before writing the visible reply. A few practical consequences:
- Give requests enough
max_tokens. With a budget of 64 tokens a simple question may produce only reasoning, leavingcontentempty withfinish_reason: "length". A budget of 256 produces the full answer for most short prompts. - Reasoning is surfaced on both paths. Streaming responses carry it as
delta.reasoning_contentand an identicaldelta.reasoningalias; non-streaming/v1/chat/completionsresponses include both fields on the assistant message. They are omitted when the model produces no reasoning, and--reasoning-alias-field nonesuppresses only thereasoningduplicate. /v1/completionsworks but raw prompts degrade. The endpoint accepts un-templated text mechanically, but the model depends on its chat structure, so output quality degrades significantly. Prefer/v1/chat/completions.- Cancellation resolution. A cancelled request is aborted within at most one denoising step, not instantaneously.
- Tensor parallelism: The TP planner has a placeholder arm for this family.
mlxcel generate -m google/diffusiongemma-26B-A4B-it \
-p "Why is the sky blue?" \
--max-tokens 256 \
--seed 42
For a 4-bit quantized snapshot in the model store:
mlxcel generate -m diffusiongemma-26B-A4B-it-4bit \
-p "Explain TCP vs UDP in one paragraph." \
--max-tokens 512 \
--diffusion-sampler entropy-bound \
--max-denoising-steps 32