Skip to content

fix: six lib tests fail under --features cuda on main #2087

Description

@inureyes

Problem

While gating PR #2077 on GB10 (2026-09-30), cargo test --release --features cuda --lib -- --test-threads=1 gave 8741 passed / 6 failed. The same six fail identically on the pre-#2077 base, so they are not regressions from it. The cuda test build compiles on current main (the sanitize_tests cfg fix landed in 929c80a).

Items (one PR, one verification pass at the end, minimal pushes)

1. f32 op bf16 returns bf16 on CUDA (5 tests)

dtype::FLOAT32 == 10, dtype::BFLOAT16 == 12 (src/lib/mlxcel-core/src/dtype.rs:27-29), so left: 12, right: 10 means a mixed f32/bf16 op came back bf16. MLX promotion says f32. Affected: precast_matmul_and_addmm_match_the_promoted_graph (src/audio/f32_weights.rs:117), precast_conv_and_layer_norm_match_the_promoted_graph (f32_weights.rs:132, byte mismatch, likely the same cause), perception_f32_weights_match_the_promoted_bf16_path (src/audio/fastconformer/tests.rs:492), and depthsum_ignores_mask_index / mog_infer_shapes_are_finite_with_guidance (src/models/nemotron_voicechat/tts/tts_tests.rs:107,349; rvq.rs:113-116 adds bf16 take rows to an f32 zeros). This is a dtype defect, not TF32: TF32 changes mantissa, not output dtype.

First step: a probe test asserting array_dtype(add(f32, bf16)) and array_dtype(matmul(f32, bf16)) are FLOAT32 on CPU and CUDA. If it fails in plain MLX at the current pin, fix or overlay upstream (see #1932) or make the production code cast explicitly; if an mlxcel-core wrapper demotes, fix the wrapper.

2. gelu_approx 1-ulp difference

gelu_approx_matches_mlx_nn_bit_for_bit (src/models/gemma3_backbone_tests.rs:181): -0.12607098 vs -0.12607102, 4.0999565 vs 4.099957. The reference's doc comment already names compile/fusion rounding as a 1-ulp source on Metal. Decide whether CUDA fuses or uses FMA here; either make the port match on CUDA or assert a documented backend-specific contract (exact on CPU/Metal, at most 1 f32 ulp on CUDA, with the reason in the test).

No #[ignore] as a fix.

Verification

Build without the GPU lock, run under it:

cargo test --release --features cuda --lib --no-run
gpu-lock run --tag cuda-lib-tests -- ./target/release/deps/mlxcel-<hash> --test-threads=1
cargo test --release --lib

Pass: zero failures in both runs.

Notes (2026-10-02): never wrap cargo in gpu-lock (the sccache daemon cargo spawns inherits the lock); run the built test binary under it instead. cargo test --release --lib does not link on Linux main without a GPU feature: src/lib/mlx-cpp/turbo/kv_inplace_write.cpp (#1961) references mlx::core::copy_gpu_inplace, which a no-GPU MLX build does not define.

Activity

  1. added
    type:bugBug fixes, error corrections, or issue resolutions
    and removed on Oct 1, 2026
  2. added a commit that references this issue on Oct 2, 2026
  3. added 2 commits that reference this issue on Oct 2, 2026
    df686a7
    b77d006
  4. added a commit that references this issue on Oct 5, 2026
    324a83d
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    priority:mediumMedium prioritystatus:doneCompletedtype:bugBug fixes, error corrections, or issue resolutions

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions