Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
29 changes: 17 additions & 12 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -557,10 +557,13 @@ jobs:
contents: read
timeout-minutes: 90
env:
# What releases build for this runner's architecture. Pinned rather than
# left to auto-detection, which yields `121a` here and would compile every
# kernel architecture-specific (issue #1934), so the artifact this job
# verifies carries the same device code the published one does.
# What releases build for this runner's architecture, pinned rather than
# inferred so the artifact this job verifies carries the same device code
# the published one does whatever GPU the runner turns out to have. Since
# issue #1943 auto-detection produces this same `121` on a GB10, so the pin
# no longer corrects a shape mismatch: it states the list instead of
# inferring it, which is what makes the job reviewable and what keeps it
# right if this ever runs on something other than a GB10.
MLX_CUDA_ARCHITECTURES: "121"
steps:
- uses: actions/checkout@v7
Expand Down Expand Up @@ -929,10 +932,11 @@ jobs:
contents: read
timeout-minutes: 120
env:
# The value shipped for this runner's architecture. Auto-detection yields
# `121a`, which is not what releases build and would compile every kernel
# architecture-specific (issue #1934), so the gate compiles what ships. The
# hardware NVFP4/MXFP4 converters come from the per-source injection in
# The value shipped for this runner's architecture, so the gate compiles
# what ships rather than whatever the runner's GPU implies. Since issue
# #1943 auto-detection produces this same `121` on a GB10; the pin remains
# because a gate should name the list it compiles instead of inferring it.
# The hardware NVFP4/MXFP4 converters come from the per-source injection in
# `src/lib/mlx-cpp/CMakeLists.txt`, not from this list.
MLX_CUDA_ARCHITECTURES: "121"
# Every rustc lint, not just the `unused_imports` this job started with.
Expand Down Expand Up @@ -1113,10 +1117,11 @@ jobs:
cancel-in-progress: ${{ github.ref != 'refs/heads/main' }}
timeout-minutes: 120
env:
# The value shipped for this runner's architecture. Auto-detection yields
# `121a`, which is not what releases build and would compile every kernel
# architecture-specific (issue #1934), so the gate links what ships. The
# hardware NVFP4/MXFP4 converters come from the per-source injection in
# The value shipped for this runner's architecture, so the gate links what
# ships rather than whatever the runner's GPU implies. Since issue #1943
# auto-detection produces this same `121` on a GB10; the pin remains
# because a gate should name the list it compiles instead of inferring it.
# The hardware NVFP4/MXFP4 converters come from the per-source injection in
# `src/lib/mlx-cpp/CMakeLists.txt`, not from this list.
MLX_CUDA_ARCHITECTURES: "121"
steps:
Expand Down
23 changes: 17 additions & 6 deletions .github/workflows/release.yml
Original file line number Diff line number Diff line change
Expand Up @@ -534,9 +534,15 @@ jobs:
# One aarch64 fat binary covering the NVIDIA aarch64 targets in a single
# build: GH200 (Grace Hopper, sm_90a), GB200 (Grace Blackwell, sm_100), and
# GB10 / DGX Spark (Blackwell, sm_121), cross-compiled on the GB10 runner.
# 90a is load-bearing for MLX's Hopper quantized kernel gate, and build.rs
# forwards the list verbatim. The CLI and server (~347 MB each) ship as
# separate archives, both with the CCCL headers.
# build.rs forwards the list verbatim, so this is exactly what nvcc is asked
# for. Hopper is spelled 90a because that is the spelling this project has
# shipped and the one auto-detection mirrors, not because it gates a kernel
# at the current MLX pin: upstream 44540d12 moved the Hopper quantized kernel
# to runtime NVRTC compilation and deleted MLX_CUDA_SM90A_ENABLED, which this
# comment used to cite. CUTLASS_ARCH_MMA_SM90A_ENABLED still keys on
# __CUDA_ARCH_FEAT_SM90_ALL, so a later pin can make it matter again
# (issue #1943). The CLI and server (~347 MB each) ship as separate archives,
# both with the CCCL headers.
#
# Blackwell stays plain here even though MLX's hardware NVFP4/MXFP4
# converters are gated on `__CUDA_ARCH_SPECIFIC__` and a plain target
Expand Down Expand Up @@ -766,9 +772,14 @@ jobs:
permissions:
contents: write

# One fat binary covering Ampere and later. 90a is load-bearing: MLX gates
# its Hopper quantized kernel on "90a" being in the list (MLX_CUDA_SM90A_ENABLED),
# and build.rs forwards MLX_CUDA_ARCHITECTURES verbatim (no auto-suffix).
# One fat binary covering Ampere and later. Hopper is spelled 90a, and
# build.rs forwards MLX_CUDA_ARCHITECTURES verbatim, so this list is exactly
# what nvcc is asked for. At the current MLX pin that suffix gates nothing
# by itself: upstream 44540d12 moved the Hopper quantized kernel to runtime
# NVRTC compilation and deleted MLX_CUDA_SM90A_ENABLED, which the rationale
# here used to cite. It stays because CUTLASS_ARCH_MMA_SM90A_ENABLED still
# keys on __CUDA_ARCH_FEAT_SM90_ALL, so a later pin can make it matter, and
# because auto-detection now mirrors this list (issue #1943).
# Blackwell stays plain for the reason spelled out in the aarch64 job
# above: the hardware NVFP4/MXFP4 converters are added to `fp_quantize.cu`
# alone by `src/lib/mlx-cpp/CMakeLists.txt` rather than to the whole target.
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -21,7 +21,7 @@ docs/benchmark_results/data/draft-block-width-gb10-2026-09-20/harness/run_sessio

**The binary is renamed before it runs.** `run_session.sh` copies it to `mlxcel1797-server`. The imported host gate blocks on a foreign inference process by `/proc/<pid>/comm`, matching `mlxcel` and `mlxcel-server`, which is what keeps a peer session's model off the GPU during a timed run. A server started under the stock name matches that list and gates itself forever.

**`MLX_CUDA_ARCHITECTURES=121` is pinned for the build and for every server.** `build.rs` auto-detects `121a` on this host, and every earlier GB10 record here pins plain `121`. The suffix does not change the decode kernels: in the pinned MLX tree only the NVFP4 weight-quantization converter keys on `__CUDA_ARCH_SPECIFIC__`. The pin is for comparability with those records, and it happens to match the release configuration as of 2026-09-20.
**`MLX_CUDA_ARCHITECTURES=121` is pinned for the build and for every server.** Every earlier GB10 record here uses plain `121`, and the pin states what was compiled rather than leaving a reader to infer it. Since issue #1943 `build.rs` auto-detects that same plain `121` on this host, so a run of this harness without the pin is comparable with those records; before #1943 auto-detection yielded `121a`, which is why rows recorded between #1934 and #1943 on an unpinned build are the ones that do not line up. The suffix does not change the decode kernels either way: in the pinned MLX tree only the NVFP4 weight-quantization converter keys on `__CUDA_ARCH_SPECIFIC__`.

**The host gate is imported from `../../sdpa-plan-bucket-gb10-2026-09-12/harness/hostgate.py`, not copied.** It encodes three predicates that were each paid for with a lost measurement: contention means processes that can consume CPU now (match `/proc/<pid>/comm` exactly, drop state `T`, exclude the CI runner's permanent `RunnerService.js` and `Runner.Listener` daemons, never a `%cpu` threshold because `ps` reports a lifetime average); a foreign model process is a reason to wait; and the cumulative `NV_ERR_NO_MEMORY` count is the freeze precursor on this host and does not decay. A second copy would drift from the original.

Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -28,11 +28,14 @@ printf 'nvidia_driver: %s\n' \
"$(nvidia-smi --query-gpu=driver_version --format=csv,noheader 2>/dev/null | head -1)"
printf 'gpu: %s\n' \
"$(nvidia-smi --query-gpu=name,compute_cap --format=csv,noheader 2>/dev/null | head -1)"
# The architecture list the build was PINNED to, which is not what a default
# build produces here: build.rs auto-detects and yields `121a`, while every
# prior GB10 record pins plain `121`. The suffix does not change the decode
# kernels (only the NVFP4 weight-quantization converter keys on
# __CUDA_ARCH_SPECIFIC__); it is recorded so two sessions can be told apart.
# The architecture list the build was compiled for, recorded so two sessions can
# be told apart rather than assumed comparable. Since issue #1943 a default
# build here auto-detects plain `121`, the same list every GB10 record pins, so
# an auto-detected row and a pinned row agree. Before #1943 auto-detection
# yielded `121a`, so rows recorded between #1934 and #1943 on an unpinned build
# are the ones to check this field on. The suffix does not change the decode
# kernels either way (only the NVFP4 weight-quantization converter keys on
# __CUDA_ARCH_SPECIFIC__).
printf 'MLX_CUDA_ARCHITECTURES: %s\n' "${MLX_CUDA_ARCHITECTURES:-unset (auto-detected)}"
# nvcc is not on the default PATH on this host; the CUDA install root is.
NVCC=$(command -v nvcc || echo /usr/local/cuda/bin/nvcc)
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -15,11 +15,15 @@
# unambiguously not the name a peer's server carries.
#
# 2. `MLX_CUDA_ARCHITECTURES=121` is exported for the build AND for every
# server process. build.rs auto-detects `121a` on this host, which is not
# what any earlier GB10 record used. The suffix does not change the decode
# kernels this harness measures (only the NVFP4 weight-quantization
# converter keys on __CUDA_ARCH_SPECIFIC__); it is pinned so these numbers
# sit alongside every earlier GB10 record's.
# server process, so these numbers sit alongside every earlier GB10 record's
# and the list is stated rather than inferred. Since issue #1943 a default
# build here auto-detects this same plain `121`, so the pin no longer
# corrects a shape; it is kept because a harness should name what it
# compiled. Before #1943 auto-detection yielded `121a`, so rows recorded
# between #1934 and #1943 on an unpinned build are the ones that are not
# directly comparable. The suffix does not change the decode kernels this
# harness measures either way (only the NVFP4 weight-quantization converter
# keys on __CUDA_ARCH_SPECIFIC__).
set -uo pipefail
BIN=${1:?usage: run_session.sh <mlxcel-server binary> [outdir]}
D="$(cd "$(dirname "$0")" && pwd)"
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -190,9 +190,12 @@ def run_arm(a, width, prompt, tag):
# left unset, so the record states it instead of the reader assuming it.
env.setdefault("MLX_ENABLE_TF32", "1")
env.setdefault("RUST_LOG", "info")
# Pinned, not auto-detected: build.rs yields `121a` here while every prior
# GB10 record uses plain `121`. The suffix does not change the decode
# kernels; the pin is for comparability with those records.
# Stated rather than inferred, for comparability with every prior GB10
# record, which uses plain `121`. Since issue #1943 auto-detection here
# produces that same `121`, so the pin no longer corrects a shape; before
# #1943 it yielded `121a`, which is why rows recorded between #1934 and
# #1943 on an unpinned build are the ones that do not line up. The suffix
# does not change the decode kernels either way.
env.setdefault("MLX_CUDA_ARCHITECTURES", "121")
env.update(arm_env(width))
t0 = time.time()
Expand Down
54 changes: 37 additions & 17 deletions docs/installation.md
Original file line number Diff line number Diff line change
Expand Up @@ -164,18 +164,23 @@ CUDA_HOME=/opt/cuda cargo build --release --features cuda


`src/lib/mlxcel-core/build.rs` reads `MLX_CUDA_ARCHITECTURES`. If it is unset,
the build script tries to detect the compute capability with `nvidia-smi` and
falls back to `90a` when detection fails. For SM 90 and above it appends CUDA's
architecture-specific `a` suffix (so `90` becomes `90a`), because the dedicated
Hopper quantized kernel (`qmm_sm90`) is only compiled when `90a` is in the arch
list. An explicitly set `MLX_CUDA_ARCHITECTURES` is used verbatim, so include the
suffix yourself for Hopper (`90a`).

On Blackwell (`sm_100`, `sm_120`, `sm_121`) the `a` suffix decides something
else, and you should not set it. MLX compiles the hardware block-float
converters in `mlx/backend/cuda/quantized/nvfp4_quantize.cuh`, the ones that
issue `cvt.rn.satfinite.e2m1x2.f32`, only when nvcc is compiling for an
architecture-specific target, which it signals with `__CUDA_ARCH_SPECIFIC__`.
the build script detects the compute capability with `nvidia-smi` and spells it
the way the release workflow spells it: Hopper gets CUDA's architecture-specific
`a` suffix (`90` becomes `90a`), and Blackwell (`sm_100`, `sm_120`, `sm_121`)
stays plain. Detection failure falls back to `90a`. An auto-detected build and a
published one therefore differ in which architectures they cover and never in
what machine code those architectures get, which is the point: before issue
#1943 the rule suffixed everything from SM 90 up, so a default build on a GB10
produced `121a` while the release archives for the same card carried `121`.

An explicitly set `MLX_CUDA_ARCHITECTURES` is used verbatim, so spell it the same
way by hand: `90a` for Hopper, plain for Blackwell. The rest of this section is
why Blackwell is plain.

On Blackwell the `a` suffix decides something else. MLX compiles the hardware
block-float converters in `mlx/backend/cuda/quantized/nvfp4_quantize.cuh`, the
ones that issue `cvt.rn.satfinite.e2m1x2.f32`, only when nvcc is compiling for
an architecture-specific target, which it signals with `__CUDA_ARCH_SPECIFIC__`.
A plain `121` build therefore compiles them out, and both NVFP4 and MXFP4
quantization fall back to a scalar CUTLASS conversion sequence with nothing in
the build output saying so.
Expand All @@ -192,13 +197,28 @@ So the list stays plain and `src/lib/mlx-cpp/CMakeLists.txt` adds the
architecture-specific image to `fp_quantize.cu` alone, which is the only
translation unit that can reach those converters. You get the hardware path
without asking for it, and without it reaching anything else. Building with
`121a` is a step backwards, not a step forwards; `121f` is worse still, because
it satisfies the converters' dispatcher gate but not their own gate and fails
to compile outright.
`121a` is a step backwards, not a step forwards, and it is also self-defeating:
the injection skips a capability the list already names architecture-specific,
because asking nvcc for the same `--generate-code` twice is an error rather than
a no-op. `121f` is worse still, because it satisfies the converters' dispatcher
gate but not their own gate and fails to compile outright.

Hopper's `90a` is a different case: it is what both release lists ship, so the
auto-detected spelling matches it. At the current MLX pin the suffix no longer
gates anything on its own. Upstream commit `44540d12` moved `qmm_sm80`,
`qmm_sm90` and `gather_gemm` to runtime NVRTC compilation and removed the
`MLX_CUDA_SM90A_ENABLED` definition an earlier version of this section cited,
and `jit_module.cpp` now derives the NVRTC `--gpu-architecture` from the running
device, appending `a` itself from compute capability 9 up. Cross-compiling the
pinned tree at `90` and at `90a` agrees: `qmm_sm90.cu`, `qmm.cu` and
`qmm_sm80.cu` emit no device function at either spelling, and `qmv.cu` and
`fp_qmv.cu` emit identical SASS apart from the `EF_CUDA_ACCELERATORS` header
flag that marks a cubin architecture-specific.
`CUTLASS_ARCH_MMA_SM90A_ENABLED` still keys on `__CUDA_ARCH_FEAT_SM90_ALL`, so
a later pin can make it matter again.

```bash
# Hopper / GH200-style target. The `a` suffix is required for the Hopper
# quantized kernel; plain `90` builds without it.
# Hopper / GH200-style target, spelled the way the release workflow spells it.
MLX_CUDA_ARCHITECTURES=90a cargo build --release --features cuda

# GB10 / DGX Spark-style target used by the release workflow. Plain: the
Expand Down
11 changes: 8 additions & 3 deletions scripts/bench_block_width.sh
Original file line number Diff line number Diff line change
Expand Up @@ -111,9 +111,14 @@ MEM=$(( $(sysctl -n hw.memsize 2>/dev/null || echo 0) / 1073741824 ))
# 2026-09-20 and the resulting driver change was invisible to every harness
# here, so a cross-session comparison had nothing to key on (issue #1797).
# These three travel together as the host identity. The architecture belongs
# with them because build.rs auto-detects `121a` on GB10 while every earlier
# GB10 record pins plain `121`, so a row measured on an auto-detected build is
# not directly comparable with those records.
# with them because it decides what machine code the kernels are, so two rows
# built for different lists are not comparable whatever else matches. Since
# issue #1943 an auto-detected GB10 build resolves to plain `121`, the same list
# every earlier GB10 record pins, so an auto-detected row IS comparable with
# those records. That was not true before #1943, when auto-detection produced
# `121a` and gave every kernel a second, architecture-specific image; rows
# recorded between #1934 and #1943 on an unpinned build are the ones to treat
# with suspicion. The field is recorded rather than assumed either way.
KERNEL=$(uname -r)
DRIVER=$(nvidia-smi --query-gpu=driver_version --format=csv,noheader 2>/dev/null | head -1)
ARCHES=${MLX_CUDA_ARCHITECTURES:-auto-detected}
Expand Down
Loading
Loading