Problem
While gating PR #2077 on GB10 (2026-09-30), cargo test --release --features cuda --lib -- --test-threads=1 gave 8741 passed / 6 failed. The same six fail identically on the pre-#2077 base, so they are not regressions from it. The cuda test build compiles on current main (the sanitize_tests cfg fix landed in 929c80a).
Items (one PR, one verification pass at the end, minimal pushes)
1. f32 op bf16 returns bf16 on CUDA (5 tests)
dtype::FLOAT32 == 10, dtype::BFLOAT16 == 12 (src/lib/mlxcel-core/src/dtype.rs:27-29), so left: 12, right: 10 means a mixed f32/bf16 op came back bf16. MLX promotion says f32. Affected: precast_matmul_and_addmm_match_the_promoted_graph (src/audio/f32_weights.rs:117), precast_conv_and_layer_norm_match_the_promoted_graph (f32_weights.rs:132, byte mismatch, likely the same cause), perception_f32_weights_match_the_promoted_bf16_path (src/audio/fastconformer/tests.rs:492), and depthsum_ignores_mask_index / mog_infer_shapes_are_finite_with_guidance (src/models/nemotron_voicechat/tts/tts_tests.rs:107,349; rvq.rs:113-116 adds bf16 take rows to an f32 zeros). This is a dtype defect, not TF32: TF32 changes mantissa, not output dtype.
First step: a probe test asserting array_dtype(add(f32, bf16)) and array_dtype(matmul(f32, bf16)) are FLOAT32 on CPU and CUDA. If it fails in plain MLX at the current pin, fix or overlay upstream (see #1932) or make the production code cast explicitly; if an mlxcel-core wrapper demotes, fix the wrapper.
2. gelu_approx 1-ulp difference
gelu_approx_matches_mlx_nn_bit_for_bit (src/models/gemma3_backbone_tests.rs:181): -0.12607098 vs -0.12607102, 4.0999565 vs 4.099957. The reference's doc comment already names compile/fusion rounding as a 1-ulp source on Metal. Decide whether CUDA fuses or uses FMA here; either make the port match on CUDA or assert a documented backend-specific contract (exact on CPU/Metal, at most 1 f32 ulp on CUDA, with the reason in the test).
No #[ignore] as a fix.
Verification
Build without the GPU lock, run under it:
cargo test --release --features cuda --lib --no-run
gpu-lock run --tag cuda-lib-tests -- ./target/release/deps/mlxcel-<hash> --test-threads=1
cargo test --release --lib
Pass: zero failures in both runs.
Notes (2026-10-02): never wrap cargo in gpu-lock (the sccache daemon cargo spawns inherits the lock); run the built test binary under it instead. cargo test --release --lib does not link on Linux main without a GPU feature: src/lib/mlx-cpp/turbo/kv_inplace_write.cpp (#1961) references mlx::core::copy_gpu_inplace, which a no-GPU MLX build does not define.
Problem
While gating PR #2077 on GB10 (2026-09-30),
cargo test --release --features cuda --lib -- --test-threads=1gave 8741 passed / 6 failed. The same six fail identically on the pre-#2077 base, so they are not regressions from it. The cuda test build compiles on current main (the sanitize_tests cfg fix landed in 929c80a).Items (one PR, one verification pass at the end, minimal pushes)
1. f32 op bf16 returns bf16 on CUDA (5 tests)
dtype::FLOAT32 == 10,dtype::BFLOAT16 == 12(src/lib/mlxcel-core/src/dtype.rs:27-29), soleft: 12, right: 10means a mixed f32/bf16 op came back bf16. MLX promotion says f32. Affected:precast_matmul_and_addmm_match_the_promoted_graph(src/audio/f32_weights.rs:117),precast_conv_and_layer_norm_match_the_promoted_graph(f32_weights.rs:132, byte mismatch, likely the same cause),perception_f32_weights_match_the_promoted_bf16_path(src/audio/fastconformer/tests.rs:492), anddepthsum_ignores_mask_index/mog_infer_shapes_are_finite_with_guidance(src/models/nemotron_voicechat/tts/tts_tests.rs:107,349;rvq.rs:113-116adds bf16takerows to an f32zeros). This is a dtype defect, not TF32: TF32 changes mantissa, not output dtype.First step: a probe test asserting
array_dtype(add(f32, bf16))andarray_dtype(matmul(f32, bf16))areFLOAT32on CPU and CUDA. If it fails in plain MLX at the current pin, fix or overlay upstream (see #1932) or make the production code cast explicitly; if an mlxcel-core wrapper demotes, fix the wrapper.cudaand bf16 with it, because the cause is mlxcel's ownsrc/lib/mlx-cpp/patches-cuda/dtype.cppoverlay (perf(cuda): single-dtype decode graph: eliminate per-token AsType conversions #636), which is kept; the production code that needed f32 (RVQ depth-sum, MoG gathered slabs) now casts explicitly.--features cudaand without it (PR fix(tts): cast bf16 tables explicitly where the CUDA overlay demotes #2093).2. gelu_approx 1-ulp difference
gelu_approx_matches_mlx_nn_bit_for_bit(src/models/gemma3_backbone_tests.rs:181): -0.12607098 vs -0.12607102, 4.0999565 vs 4.099957. The reference's doc comment already names compile/fusion rounding as a 1-ulp source on Metal. Decide whether CUDA fuses or uses FMA here; either make the port match on CUDA or assert a documented backend-specific contract (exact on CPU/Metal, at most 1 f32 ulp on CUDA, with the reason in the test).--features cudaand without it, with the reason documented if the contract changes. Already fixed on main by fix(test): repair gate failures left by #2034 and #2037 #2080 (backendtanh/powerrounding vs hardcoded Metal values; no fusion, FMA or TF32 involved); verified on CUDA and CPU in PR fix(tts): cast bf16 tables explicitly where the CUDA overlay demotes #2093.No
#[ignore]as a fix.Verification
Build without the GPU lock, run under it:
Pass: zero failures in both runs.
Notes (2026-10-02): never wrap
cargoingpu-lock(the sccache daemon cargo spawns inherits the lock); run the built test binary under it instead.cargo test --release --libdoes not link on Linux main without a GPU feature:src/lib/mlx-cpp/turbo/kv_inplace_write.cpp(#1961) referencesmlx::core::copy_gpu_inplace, which a no-GPU MLX build does not define.