Repository navigation
test: Apple Silicon verification pass for 2026-10-06 CUDA-only changes #2158
Description
Activity
- addedstatus:readyReady to be worked onReady to be worked ontype:testTest related changesTest related changespriority:highHigh priorityHigh priorityarea:inferenceGeneration, sampling, decoding (incl. speculative, DRY)Generation, sampling, decoding (incl. speculative, DRY)platform:macosmacOS (Apple Silicon) specificmacOS (Apple Silicon) specific
on Oct 6, 2026 Adding one item from PR #2163 (#1611, merged): Youtu-VL now resizes images up to the AutoProcessor cap of 36864 patches (was 4096).
- On M5 Max, run
mlxcel generate -m models/mlx/youtu-vl-4b-instruct --image <a 3000x4000 image> -p "Describe this image." -n 32 --temp 0and record peak memory. On GB10 (CUDA) the 36520-patch case peaked at about 30.1 GB with the query-chunked SDPA. If MLX's fused Metal SDPA rejects head_dim 72, the four full-attention layers would materialize about 43 GB of scores. Pass: the run completes with a correct description and peak memory is reported; if it OOMs, open a fix issue rather than lowering the cap silently.
- On M5 Max, run
- added a commit that references this issue
on Oct 6, 2026 Adding items from PR #2182 (#2159) and PR #2185 (#2160), both merged. Neither was run on Metal.
- PR perf(scheduler): restore Gemma 3 decode lookahead via ring rewind #2182: Gemma 3 decode lookahead now runs again with a depth-2 ring undo log on rotating caches. On M5 Max, gemma-3-4b-it-4bit via
mlxcel-server: greedy output with default flags must be byte-identical toMLXCEL_FORCE_SYNC=1for a short prompt, a prompt past the 1024 sliding window, and 4 concurrent requests, at--parallel 4and--parallel 1. Report decode tok/s for both arms (GB10: 87.3 vs 76.2 short, 72.7 vs 67.3 past the window). - PR fix(gemma4): make MTP verify byte-identical to decode on CUDA #2185: shared-code changes that reach Metal. The
Gemma4Attention::attendgate refactor (31B behaviour intended unchanged, plus a new per-row fallback for a sliding layer with no ring cursor), and the history-boundary prefill split in the row-wise MTP prefill. On M5 Max, run the 31B MTP startup probe (--draft-kind mtp --draft-block-size 4 --parallel 1, no inexact override) andcargo test --profile test-fast --features metal,accelerate --test speculative_parity -- --ignored greedy_parity_mtp_gemma4 --test-threads=1. Pass: the probe passes as before fix(gemma4): make MTP verify byte-identical to decode on CUDA #2185 and the parity tests pass. Noteb1_batched_baseline_probealready failed on main before fix(gemma4): make MTP verify byte-identical to decode on CUDA #2185 and is not part of this item.
- PR perf(scheduler): restore Gemma 3 decode lookahead via ring rewind #2182: Gemma 3 decode lookahead now runs again with a depth-2 ring undo log on rotating caches. On M5 Max, gemma-3-4b-it-4bit via
Adding the Metal-reaching changes from PR #2203 (#2190). Metal 31B is a row-wise geometry (
mtp_row_verifyon,mtp_row_ropeoff), so it takes these paths too. None runs in Metal serving: the new batched-window gate declines every row-wise window off CUDA.- PR fix(gemma4): make batched MTP verify decode-exact on row-wise geometries #2203: the batched MTP adapter now marks verify forwards (
mtp_verify), so at B=1 it takes the row-wise linear verify; on row-wise targets it prefills each row on its own cache over the classic partition and stacks the caches; at B>1 it verifies each row as its own[1, K]call (projections, uniform RoPE chain on Metal,attend_verify_rows,o_proj, MLP, LM head) with only the cache write batched; andDecoderLayer::finish_layeris extracted fromforward_with_profileunchanged. On M5 Max, runcargo test --profile test-fast --features metal,accelerate --test speculative_parity -- --ignored --test-threads=1 b1_batched_baseline_probe gemma4_batched(the new test skips the 12B on Metal, where it is not row-wise) and the lib testsmodels::gemma4_mtp_target::rowsandmodels::gemma4::verify_rows. Pass: the 31B probe rows are byte-identical and the new test's 31B cases match classic decode. If they pass, Metal 31B batched throughput against classic batched decode (thedocs/benchmark_results/data/gemma4-mtp-batched-gb10-2026-10-07/harnessmethod) decides whetherRowWiseBatchedWindowshould admit Metal.
- PR fix(gemma4): make batched MTP verify decode-exact on row-wise geometries #2203: the batched MTP adapter now marks verify forwards (
M1 Ultra verification — 2026-10-07
Host: Apple M1 Ultra, 128 GB, macOS 27.0 (26A428).
xcrun -f metal:/Applications/Xcode-27.0.0-Release.Candidate.app/Contents/Developer/Toolchains/XcodeDefault.xctoolchain/usr/bin/metal. MLX pin:81ba1c6a0e50a9268b931579c2d4f1158b9aab5a; Metal kernels were compiled, not taken from a release metallib. Main baseline:b34d517bc422dd0465a6c404cbb2b84d4181124d; PR #2138 rebased/tested head:25202198c3ea7444cf9aa6b14c059f79683b75ab.Item Result 1, 2, 3 NOT VERIFIED: M5 required. 4 PASS on M1: main reproduces the known #1599 failures (3 passed, 2 failed); PR #2138 passes all 5 mxfp tests. Qwen3-30B-A3B-4bit greedy 64-token main/main-repeat/PR outputs are byte-identical. 5 PASS on M1: the four numeric tests plus sentinel pass (5/5); explicit MLX_ENABLE_TF32=1sentinel also passes (1/1, expected on M1; error is below 1e-4, exact value is not printed on success). Decode comparisons below all pass. M5 and cross-host comparison remain unverified.6, 7, 8, 9, 10, 11, 12 NOT VERIFIED: M5 work deferred at the user's request. No throughput, MTP diagnostic, or large-image memory claims. Item 5 decode
generate --temp 0 -n 256 --show-reasoning, each prompt under default (variable unset),MLX_ENABLE_TF32=0, and default again. After removing loader banners and timing lines, all six default/disabled pairs and all six default/repeat controls are byte-identical; no divergence. Plamo has 218 F32 tensors; the locally namedllama-3.1-8b-instruct-bf16checkpoint actually has 291 F16 tensors (SafeTensors headers checked).Prompt Plamo tokens Llama tokens Both comparisons Explain why the sky appears blue during the day in simple terms. 256 231 (EOS) Identical Write a short story about a traveler who discovers a hidden garden. 256 256 Identical 日本の四季について、それぞれの季節の特徴を説明してください。 256 256 Identical Item 4 prompt:
Explain in simple terms how a rainbow forms after rain and why different colors appear in a fixed order. /no_think; output SHA-256 across all three runs:d2f40b2d5831c3d7790d467fb51123f0d1f615fb7300aa2f92d86a9d8831a0de. Used the existingscripts/ab_output_equality.sh, including its distinct-binary guard and baseline-repeat control.Both release builds succeeded. The PR build initially hit sccache
Too many open files; rerunning withRUSTC_WRAPPER=succeeded without code changes. This was build infrastructure, not an inference failure. No new product failure was found; the red baseline is already tracked by #1599. Local raw logs/outputs:/tmp/mlxcel-2158/. This issue remains open for M5 validation; #1599 remains open pending a post-merge Metal nightly.PR #2138 is now ready for review, but not merged: CI run
37637901820has four jobs queued for the GB10 runner (CPU-only link, CUDA sm_70 compile, CUDA block-float converter, OpenXLA feature link). Merge remains gated on those results; no failing CI job was observed.
Problem / Background
Seven changes merged or prepared on 2026-10-06 were verified only on CUDA (GB10). Each PR body lists the Apple Silicon checks it could not run under "Not verified on this host", and several of them touch paths that only run on Metal: the tile-padded prefill gated by
should_align_prefill()(M5 or newer with Neural Accelerators,src/server/batch/scheduler/pad_trim.rs), the Metalgather_qmm_rhsoverlay, and the NAX f32 GEMM path thatMLX_ENABLE_TF32reaches on Apple GPU generation 17. This issue collects those checks into one pass.Run all eight items in a single session on one build, then post one comment here with a result per item. Do not fix anything inside this issue: for any failure, open a separate
fix:issue with the reproduction and link it from the result comment. Item 4 is the one exception that changes repository state (marking a draft PR ready and merging it when green).Setup (once, for all items)
xcrun -f metaloutput, and commit SHA in the result comment.KVCache::trimfor pool-backed caches #2097 found that a Command Line Tools-only host links a stale releasemlx.metallibinstead of compiling the kernels atmain's MLX pin, which makes three mxfp/fp8 tests fail for reasons unrelated to the change under test.mainat1e561f1eor later. Checkpoints are referenced asmodels/mlx/<name>(the./modelssymlink to the local model store); adjust the prefix if the host lays them out differently.git checkout main && git pull cargo build --release --features metal,accelerate --bins1. Gemma 3 tile-padded prefill on M5 (PR #2140, #1755)
PR #2140 routes the scheduler's pad trims through
LanguageModel::trim_sequence_stateand addsRotatingKVCache::rewind_padded_prefill, so Gemma 3 (model-owned layout) takes tile-padded prefill again by default on M5. GB10 could only exercise it throughMLXCEL_FORCE_PADDED_PREFILL. Reference:TECHNICAL_REPORTS/pr-2140-model-owned-pad-trim.en.md.Server,
models/mlx/gemma-3-4b-it-4bit,--prefill-chunk-size 500, greedy (temperature 0),max_tokens 48,logprobswithtop_logprobs 2, three prompts of about 176, 730 and 1852 tokens (none a multiple of 32; the last is past the 1024 sliding window). Arms: default (padded on M5) andMLXCEL_NO_PADDED_PREFILL=1, each with--parallel 1(sequential) and--parallel 4(all three prompts concurrent). Fresh server per arm.CLI:
mlxcel generate -m models/mlx/gemma-3-4b-it-4bit --temp 0 -n 40 -p <the 176-token prompt>, default vsMLXCEL_NO_PADDED_PREFILL=1.Padded and unpadded outputs are identical, or the first differing token has its top-2 logprobs equal (report the gap at every divergence). A divergence at token 1, or degenerate repetition like the
texIdloop seen with the rewind disabled, is a failure.CLI output identical across the two arms.
2. Paged vs dense decode storage on M5 (PR #2097, #2096)
PR #2097 makes
KVCache::trimrewind the pool for pool-backed caches. The contributor's M5 Pro run is the only M5 evidence. Reference:TECHNICAL_REPORTS/pr-2097-pool-backed-kv-trim.en.md.mlxcel-server -m <model>(default, resolves to paged) vs--decode-storage-backend dense, fresh server per run,temperature 0,max_tokens 24, five chat prompts none of whose token counts is a multiple of 32, onLlama-3.2-1B-Instruct-4bit,SmolLM2-135M-InstructandQwen3-0.6B-4bit. Run the five sequentially, then all five concurrently (twice).5 of 5 identical to dense, sequential and concurrent, on all three checkpoints. "What is the capital of France?" on SmolLM2 answers
The capital of France is Paris.3. Whole-prompt replay on M5 (PR #2136 / #2139, #1760 / #1754)
#2096 reported that on M5 replaying a 32-token prompt (both backends) or a 48-token prompt (
--parallel 1) aborted the server with[broadcast_shapes] Shapes (32,63) and (1,9,32,64) cannot be broadcast. PR #2136 caps a whole-prompt hit atlen - 1; PR #2139 keys completion snapshots by the tokens their state holds.mlxcel-server -m models/mlx/qwen3-0.6b-4bit(prompt cache on, default): send a 32-token prompt cold, then replay it twice in the same process; repeat with a 48-token prompt under--parallel 1. Repeat both onmodels/mlx/gemma-3-4b-it-4bit.temperature 0,seed 0, 200 tokens.cargo test --test prompt_cache_e2e --release --features metal,accelerate -- --ignored --nocapture --test-threads=1 identical_prompt_replay_restores_all_but_the_last_tokenNo abort on either length or checkpoint; warm replies byte-identical to cold, or the first divergence is a top-2 tie (gap 0.0), reported.
The e2e test passes and its output does not contain
Skipping: model or mlxcel-server binary not present.4. Metal
gather_qmm_rhsrow count (draft PR #2138, #1599)PR #2138 passes
B * Minstead ofx.size() / Ktogather_qmm_rhsin the overlaidquantized.cpp. It was never compiled on Metal. Run on the PR branch, preferably on the M1 Ultra nightly runner.mxfp_testspass, includingmxfp4_gather_qmm_matches_host_referenceandmxfp8_gather_qmm_matches_host_reference.mainbuild (sorted prefill path unchanged).gh pr ready 2138, then merge. [nightly-verify] main is red #1599 stays open until a Metal nightly (make verify-test) is green after the merge.5. TF32 pin and sentinel (PR #2132, #1065)
Reference:
TECHNICAL_REPORTS/2132-tf32-test-pin-sentinel-20261006.en.md. On M5 Max and M1 Ultra:mlxcel generate --temp 0 -n 256onplamo-2-1b(f32) and one f16/bf16 checkpoint (for examplellama-3.1-8b-bf16), three prompts, default vsMLX_ENABLE_TF32=0, on both hosts. Report where outputs diverge between arms and between M5 and M1.6. Granite Vision descriptive prompts (PR #2131, #1683)
PR #2131 registers
added_tokens_decoderentries missing fromtokenizer.json, so<image>(id 49155) is one token again. Reference:TECHNICAL_REPORTS/1683-granite-vision-config-added-tokens-20261006.en.md.CLI:
mlxcel generate -m models/mlx/granite-vision-3.2-2b-4bit --image tests/fixtures/test_image.png --temp 0 -n 64 -p "<prompt>"for "What is in this image? Describe it briefly.", "Describe this image.", "What do you see?" and a colour question.Server: the same four prompts through
/v1/chat/completionswith the image as a data URI,temperature 0.No refusal ("I can't see the image") on any prompt, CLI or server; answers describe the fixture image (a solid colour field).
Server
prompt_tokensare 1545 / 1538 / 1539 / 1545 (the GB10 and mlx-vlm values), and the preparation summary does not print theno <image> placeholder in the prompt; splicedfallback.7. SwitchGLU shared gather indices throughput (PR #2137, #1713)
Reference:
TECHNICAL_REPORTS/2137-switchglu-shared-gather-indices-20261006.en.md. Arms: base =7c3bece5(parent of the merge), new =main, null = a byte-identical copy of the base binary. Models:granite-4.0-h-350m-4bit,granite-4.0-h-tiny-4bit, controlsqwen3-30b-a3b-4bit,qwen3.5-35b-a3b-4bit. Three interleaved rounds per arm, host idle.benchmarks/pylm_m5max_2026-09-06.csv(mlx-lm: 572.86, 235.44, 143.81, 152.72 tok/s respectively), plus[PIPELINE_DETAIL]forwardms/token per arm. A base/new difference inside the null arm's spread is reported as no change, not as a gain or regression.8. Gemma 4 12B MTP diagnostic (#1986, reference only)
Run the Mac procedure in #1986 (comment) on M5 Max: apply its diagnostic patch,
MLXCEL_QMV_WIDE=0, probe lengths8,8,8,1056then1056alone, then the sliding gate widening to 16/8/256 (and the full-attention gate to 16/1/512 if the divergence moves). Do not commit the patch.[DIAG1986]lines and localized first divergence for each run are posted on fix: investigate Gemma 4 12B long-context MTP parity on Metal #1986, with a link from the result comment here. The fix stays in fix: investigate Gemma 4 12B long-context MTP parity on Metal #1986.Acceptance Criteria
Related
PRs #2140, #2097, #2136, #2139, #2138, #2132, #2131, #2137. Issues #1755, #2096, #1760, #1754, #1599, #1065, #1683, #1713, #1986.