Skip to content

test: Apple Silicon verification pass for 2026-10-06 CUDA-only changes #2158

Description

@inureyes

Problem / Background

Seven changes merged or prepared on 2026-10-06 were verified only on CUDA (GB10). Each PR body lists the Apple Silicon checks it could not run under "Not verified on this host", and several of them touch paths that only run on Metal: the tile-padded prefill gated by should_align_prefill() (M5 or newer with Neural Accelerators, src/server/batch/scheduler/pad_trim.rs), the Metal gather_qmm_rhs overlay, and the NAX f32 GEMM path that MLX_ENABLE_TF32 reaches on Apple GPU generation 17. This issue collects those checks into one pass.

Run all eight items in a single session on one build, then post one comment here with a result per item. Do not fix anything inside this issue: for any failure, open a separate fix: issue with the reproduction and link it from the result comment. Item 4 is the one exception that changes repository state (marking a draft PR ready and merging it when green).

Setup (once, for all items)

  • Hosts: M5 Max (primary, every item) and M1 Ultra (items 4 and 5). Record chip, macOS version, xcrun -f metal output, and commit SHA in the result comment.
  • Full Xcode with the Metal shader compiler is required. PR fix(core): rewind the pool in KVCache::trim for pool-backed caches #2097 found that a Command Line Tools-only host links a stale release mlx.metallib instead of compiling the kernels at main's MLX pin, which makes three mxfp/fp8 tests fail for reasons unrelated to the change under test.
  • Build main at 1e561f1e or later. Checkpoints are referenced as models/mlx/<name> (the ./models symlink to the local model store); adjust the prefix if the host lays them out differently.
git checkout main && git pull
cargo build --release --features metal,accelerate --bins

1. Gemma 3 tile-padded prefill on M5 (PR #2140, #1755)

PR #2140 routes the scheduler's pad trims through LanguageModel::trim_sequence_state and adds RotatingKVCache::rewind_padded_prefill, so Gemma 3 (model-owned layout) takes tile-padded prefill again by default on M5. GB10 could only exercise it through MLXCEL_FORCE_PADDED_PREFILL. Reference: TECHNICAL_REPORTS/pr-2140-model-owned-pad-trim.en.md.

  • Server, models/mlx/gemma-3-4b-it-4bit, --prefill-chunk-size 500, greedy (temperature 0), max_tokens 48, logprobs with top_logprobs 2, three prompts of about 176, 730 and 1852 tokens (none a multiple of 32; the last is past the 1024 sliding window). Arms: default (padded on M5) and MLXCEL_NO_PADDED_PREFILL=1, each with --parallel 1 (sequential) and --parallel 4 (all three prompts concurrent). Fresh server per arm.

  • CLI: mlxcel generate -m models/mlx/gemma-3-4b-it-4bit --temp 0 -n 40 -p <the 176-token prompt>, default vs MLXCEL_NO_PADDED_PREFILL=1.

  • Padded and unpadded outputs are identical, or the first differing token has its top-2 logprobs equal (report the gap at every divergence). A divergence at token 1, or degenerate repetition like the texId loop seen with the rewind disabled, is a failure.

  • CLI output identical across the two arms.

2. Paged vs dense decode storage on M5 (PR #2097, #2096)

PR #2097 makes KVCache::trim rewind the pool for pool-backed caches. The contributor's M5 Pro run is the only M5 evidence. Reference: TECHNICAL_REPORTS/pr-2097-pool-backed-kv-trim.en.md.

  • mlxcel-server -m <model> (default, resolves to paged) vs --decode-storage-backend dense, fresh server per run, temperature 0, max_tokens 24, five chat prompts none of whose token counts is a multiple of 32, on Llama-3.2-1B-Instruct-4bit, SmolLM2-135M-Instruct and Qwen3-0.6B-4bit. Run the five sequentially, then all five concurrently (twice).

  • 5 of 5 identical to dense, sequential and concurrent, on all three checkpoints. "What is the capital of France?" on SmolLM2 answers The capital of France is Paris.

3. Whole-prompt replay on M5 (PR #2136 / #2139, #1760 / #1754)

#2096 reported that on M5 replaying a 32-token prompt (both backends) or a 48-token prompt (--parallel 1) aborted the server with [broadcast_shapes] Shapes (32,63) and (1,9,32,64) cannot be broadcast. PR #2136 caps a whole-prompt hit at len - 1; PR #2139 keys completion snapshots by the tokens their state holds.

  • mlxcel-server -m models/mlx/qwen3-0.6b-4bit (prompt cache on, default): send a 32-token prompt cold, then replay it twice in the same process; repeat with a 48-token prompt under --parallel 1. Repeat both on models/mlx/gemma-3-4b-it-4bit. temperature 0, seed 0, 200 tokens.

  • cargo test --test prompt_cache_e2e --release --features metal,accelerate -- --ignored --nocapture --test-threads=1 identical_prompt_replay_restores_all_but_the_last_token

  • No abort on either length or checkpoint; warm replies byte-identical to cold, or the first divergence is a top-2 tie (gap 0.0), reported.

  • The e2e test passes and its output does not contain Skipping: model or mlxcel-server binary not present.

4. Metal gather_qmm_rhs row count (draft PR #2138, #1599)

PR #2138 passes B * M instead of x.size() / K to gather_qmm_rhs in the overlaid quantized.cpp. It was never compiled on Metal. Run on the PR branch, preferably on the M1 Ultra nightly runner.

gh pr checkout 2138
cargo test --profile test-fast --features metal,accelerate -p mlxcel --lib models::switch_layers::mxfp_tests -- --test-threads=1
cargo build --release --features metal,accelerate --bins
target/release/mlxcel generate -m models/mlx/qwen3-30b-a3b-4bit --temp 0 -n 64 -p "<prompt of at least 16 tokens, so tokens x top_k 8 exceeds 64 slots>"
  • All mxfp_tests pass, including mxfp4_gather_qmm_matches_host_reference and mxfp8_gather_qmm_matches_host_reference.
  • The MoE smoke output is byte-identical to the same command on the main build (sorted prefill path unchanged).
  • If both are green: gh pr ready 2138, then merge. [nightly-verify] main is red #1599 stays open until a Metal nightly (make verify-test) is green after the merge.

5. TF32 pin and sentinel (PR #2132, #1065)

Reference: TECHNICAL_REPORTS/2132-tf32-test-pin-sentinel-20261006.en.md. On M5 Max and M1 Ultra:

cargo test --profile test-fast --features metal,accelerate -p mlxcel-core --lib -- --test-threads=1 layers::tests::chunked_causal_attention_matches_fast_causal layers::tests::chunked_query_attention_matches_unchunked_with_mask mla::decode::decode_tests::absorbed_attention_respects_an_additive_causal_mask mla::decode::decode_tests::expand_latent_reproduces_the_up_projection tf32_pin_tests
MLX_ENABLE_TF32=1 cargo test --profile test-fast --features metal,accelerate -p mlxcel-core --lib -- --test-threads=1 tf32_pin_tests
  • The first command passes on both hosts.
  • On M5 Max the second command fails with the measured error above the 1e-4 bound (report it). On M1 Ultra report the result; a pass is expected because it has no TF32-class path.
  • Decode quality: mlxcel generate --temp 0 -n 256 on plamo-2-1b (f32) and one f16/bf16 checkpoint (for example llama-3.1-8b-bf16), three prompts, default vs MLX_ENABLE_TF32=0, on both hosts. Report where outputs diverge between arms and between M5 and M1.

6. Granite Vision descriptive prompts (PR #2131, #1683)

PR #2131 registers added_tokens_decoder entries missing from tokenizer.json, so <image> (id 49155) is one token again. Reference: TECHNICAL_REPORTS/1683-granite-vision-config-added-tokens-20261006.en.md.

  • CLI: mlxcel generate -m models/mlx/granite-vision-3.2-2b-4bit --image tests/fixtures/test_image.png --temp 0 -n 64 -p "<prompt>" for "What is in this image? Describe it briefly.", "Describe this image.", "What do you see?" and a colour question.

  • Server: the same four prompts through /v1/chat/completions with the image as a data URI, temperature 0.

  • No refusal ("I can't see the image") on any prompt, CLI or server; answers describe the fixture image (a solid colour field).

  • Server prompt_tokens are 1545 / 1538 / 1539 / 1545 (the GB10 and mlx-vlm values), and the preparation summary does not print the no <image> placeholder in the prompt; spliced fallback.

7. SwitchGLU shared gather indices throughput (PR #2137, #1713)

Reference: TECHNICAL_REPORTS/2137-switchglu-shared-gather-indices-20261006.en.md. Arms: base = 7c3bece5 (parent of the merge), new = main, null = a byte-identical copy of the base binary. Models: granite-4.0-h-350m-4bit, granite-4.0-h-tiny-4bit, controls qwen3-30b-a3b-4bit, qwen3.5-35b-a3b-4bit. Three interleaved rounds per arm, host idle.

MLXCEL_PROFILE_PIPELINE_DETAIL=1 target/release/mlxcel-bench-decode -m models/mlx/<model> --prompt-tokens 512 -n 128 --ignore-eos
./scripts/ab_output_equality.sh --baseline target/release/mlxcel.before --arm target/release/mlxcel --model models/mlx/<model>
  • Greedy output byte-identical base vs new on all four models.
  • Report decode tok/s medians (base / new / null) and the ratio against benchmarks/pylm_m5max_2026-09-06.csv (mlx-lm: 572.86, 235.44, 143.81, 152.72 tok/s respectively), plus [PIPELINE_DETAIL] forward ms/token per arm. A base/new difference inside the null arm's spread is reported as no change, not as a gain or regression.

8. Gemma 4 12B MTP diagnostic (#1986, reference only)

Run the Mac procedure in #1986 (comment) on M5 Max: apply its diagnostic patch, MLXCEL_QMV_WIDE=0, probe lengths 8,8,8,1056 then 1056 alone, then the sliding gate widening to 16/8/256 (and the full-attention gate to 16/1/512 if the divergence moves). Do not commit the patch.

Acceptance Criteria

  • One result comment on this issue with host details, commit SHA, and a pass or fail per item with the numbers each item asks for.
  • Every failure has its own linked fix issue; this issue closes when every item is either green or linked to one.

Related

PRs #2140, #2097, #2136, #2139, #2138, #2132, #2131, #2137. Issues #1755, #2096, #1760, #1754, #1599, #1065, #1683, #1713, #1986.

Activity

  1. added
    type:testTest related changes
    area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)
    platform:macosmacOS (Apple Silicon) specific
    on Oct 6, 2026
  2. inureyes commented on Oct 6, 2026

    @inureyes
    MemberAuthor

    Adding one item from PR #2163 (#1611, merged): Youtu-VL now resizes images up to the AutoProcessor cap of 36864 patches (was 4096).

    1. On M5 Max, run mlxcel generate -m models/mlx/youtu-vl-4b-instruct --image <a 3000x4000 image> -p "Describe this image." -n 32 --temp 0 and record peak memory. On GB10 (CUDA) the 36520-patch case peaked at about 30.1 GB with the query-chunked SDPA. If MLX's fused Metal SDPA rejects head_dim 72, the four full-attention layers would materialize about 43 GB of scores. Pass: the run completes with a correct description and peak memory is reported; if it OOMs, open a fix issue rather than lowering the cap silently.
  3. inureyes commented on Oct 6, 2026

    @inureyes
    MemberAuthor

    Adding items from PR #2182 (#2159) and PR #2185 (#2160), both merged. Neither was run on Metal.

    1. PR perf(scheduler): restore Gemma 3 decode lookahead via ring rewind #2182: Gemma 3 decode lookahead now runs again with a depth-2 ring undo log on rotating caches. On M5 Max, gemma-3-4b-it-4bit via mlxcel-server: greedy output with default flags must be byte-identical to MLXCEL_FORCE_SYNC=1 for a short prompt, a prompt past the 1024 sliding window, and 4 concurrent requests, at --parallel 4 and --parallel 1. Report decode tok/s for both arms (GB10: 87.3 vs 76.2 short, 72.7 vs 67.3 past the window).
    2. PR fix(gemma4): make MTP verify byte-identical to decode on CUDA #2185: shared-code changes that reach Metal. The Gemma4Attention::attend gate refactor (31B behaviour intended unchanged, plus a new per-row fallback for a sliding layer with no ring cursor), and the history-boundary prefill split in the row-wise MTP prefill. On M5 Max, run the 31B MTP startup probe (--draft-kind mtp --draft-block-size 4 --parallel 1, no inexact override) and cargo test --profile test-fast --features metal,accelerate --test speculative_parity -- --ignored greedy_parity_mtp_gemma4 --test-threads=1. Pass: the probe passes as before fix(gemma4): make MTP verify byte-identical to decode on CUDA #2185 and the parity tests pass. Note b1_batched_baseline_probe already failed on main before fix(gemma4): make MTP verify byte-identical to decode on CUDA #2185 and is not part of this item.
  4. inureyes commented on Oct 7, 2026

    @inureyes
    MemberAuthor

    Adding the Metal-reaching changes from PR #2203 (#2190). Metal 31B is a row-wise geometry (mtp_row_verify on, mtp_row_rope off), so it takes these paths too. None runs in Metal serving: the new batched-window gate declines every row-wise window off CUDA.

    1. PR fix(gemma4): make batched MTP verify decode-exact on row-wise geometries #2203: the batched MTP adapter now marks verify forwards (mtp_verify), so at B=1 it takes the row-wise linear verify; on row-wise targets it prefills each row on its own cache over the classic partition and stacks the caches; at B>1 it verifies each row as its own [1, K] call (projections, uniform RoPE chain on Metal, attend_verify_rows, o_proj, MLP, LM head) with only the cache write batched; and DecoderLayer::finish_layer is extracted from forward_with_profile unchanged. On M5 Max, run cargo test --profile test-fast --features metal,accelerate --test speculative_parity -- --ignored --test-threads=1 b1_batched_baseline_probe gemma4_batched (the new test skips the 12B on Metal, where it is not row-wise) and the lib tests models::gemma4_mtp_target::rows and models::gemma4::verify_rows. Pass: the 31B probe rows are byte-identical and the new test's 31B cases match classic decode. If they pass, Metal 31B batched throughput against classic batched decode (the docs/benchmark_results/data/gemma4-mtp-batched-gb10-2026-10-07/harness method) decides whether RowWiseBatchedWindow should admit Metal.
  5. inureyes commented on Oct 7, 2026

    @inureyes
    MemberAuthor

    M1 Ultra verification — 2026-10-07

    Host: Apple M1 Ultra, 128 GB, macOS 27.0 (26A428). xcrun -f metal: /Applications/Xcode-27.0.0-Release.Candidate.app/Contents/Developer/Toolchains/XcodeDefault.xctoolchain/usr/bin/metal. MLX pin: 81ba1c6a0e50a9268b931579c2d4f1158b9aab5a; Metal kernels were compiled, not taken from a release metallib. Main baseline: b34d517bc422dd0465a6c404cbb2b84d4181124d; PR #2138 rebased/tested head: 25202198c3ea7444cf9aa6b14c059f79683b75ab.

    Item Result
    1, 2, 3 NOT VERIFIED: M5 required.
    4 PASS on M1: main reproduces the known #1599 failures (3 passed, 2 failed); PR #2138 passes all 5 mxfp tests. Qwen3-30B-A3B-4bit greedy 64-token main/main-repeat/PR outputs are byte-identical.
    5 PASS on M1: the four numeric tests plus sentinel pass (5/5); explicit MLX_ENABLE_TF32=1 sentinel also passes (1/1, expected on M1; error is below 1e-4, exact value is not printed on success). Decode comparisons below all pass. M5 and cross-host comparison remain unverified.
    6, 7, 8, 9, 10, 11, 12 NOT VERIFIED: M5 work deferred at the user's request. No throughput, MTP diagnostic, or large-image memory claims.

    Item 5 decode

    generate --temp 0 -n 256 --show-reasoning, each prompt under default (variable unset), MLX_ENABLE_TF32=0, and default again. After removing loader banners and timing lines, all six default/disabled pairs and all six default/repeat controls are byte-identical; no divergence. Plamo has 218 F32 tensors; the locally named llama-3.1-8b-instruct-bf16 checkpoint actually has 291 F16 tensors (SafeTensors headers checked).

    Prompt Plamo tokens Llama tokens Both comparisons
    Explain why the sky appears blue during the day in simple terms. 256 231 (EOS) Identical
    Write a short story about a traveler who discovers a hidden garden. 256 256 Identical
    日本の四季について、それぞれの季節の特徴を説明してください。 256 256 Identical

    Item 4 prompt: Explain in simple terms how a rainbow forms after rain and why different colors appear in a fixed order. /no_think; output SHA-256 across all three runs: d2f40b2d5831c3d7790d467fb51123f0d1f615fb7300aa2f92d86a9d8831a0de. Used the existing scripts/ab_output_equality.sh, including its distinct-binary guard and baseline-repeat control.

    Both release builds succeeded. The PR build initially hit sccache Too many open files; rerunning with RUSTC_WRAPPER= succeeded without code changes. This was build infrastructure, not an inference failure. No new product failure was found; the red baseline is already tracked by #1599. Local raw logs/outputs: /tmp/mlxcel-2158/. This issue remains open for M5 validation; #1599 remains open pending a post-merge Metal nightly.

    PR #2138 is now ready for review, but not merged: CI run 37637901820 has four jobs queued for the GB10 runner (CPU-only link, CUDA sm_70 compile, CUDA block-float converter, OpenXLA feature link). Merge remains gated on those results; no failing CI job was observed.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    area:inferenceGeneration, sampling, decoding (incl. speculative, DRY)platform:macosmacOS (Apple Silicon) specificpriority:highHigh prioritystatus:readyReady to be worked ontype:testTest related changes

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions