You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
left is max_abs_diff(&gate, &ref_gate), so the gate_proj leg produced by sanitize_gemma4_unified_weights dequantizes to something structurally different from the reference split(dequantize(gate_up_proj)). The up_proj assertion that follows is never reached, so its status is unknown.
The rest of the workspace is green: make verify-test reports 7439 passed and this 1 failed, and both make verify-fmt and make verify-clippy (--workspace --all-targets --features metal,accelerate -- -D warnings) pass.
Already ruled out (do not re-run these)
Not the Rust 1.97.1 toolchain bump (chore(ci): bump the pinned Rust toolchain to 1.97.1 #1066). Holding the source constant at e9f1f191 and compiling with the previously pinned toolchain reproduces the failure with byte-identical values (3.7108002 against 0.0):
cargo +1.93.1-aarch64-apple-darwin test --profile test-fast --features metal,accelerate --lib unified_sanitize_quantized_split_dequant_equivalence
The compiler version is not the variable.
Not an f32 reassociation tolerance problem. Reassociation differences of this kind land around 1e-6. A left-hand value of 3.71 against an expected 0.0 is a structural disagreement, not a rounding one. fix(tests): tolerate f32 reassociation in the mllama tile-selection parity test #1067 did fix a genuine reassociation case in the mllama tile-selection parity test immediately after the toolchain bump, which makes this failure easy to misfile as the same class. It is not.
The root cargo test gate historically covered only one of the workspace members (#1007), and the nightly run has not completed within its timeout (#1000). This test may have been red for some time and is only now visible because the workspace-wide make verify-test gate exercises it.
Determine the root cause: whether sanitize_gemma4_unified_weights splits quantized fused MoE experts in a way that no longer commutes with dequantization for the gate projection (a real sanitize defect), or whether the test expectation is stale relative to a deliberate later change (a stale expectation). The commuting property is what fix: split quantized fused MoE experts in gemma4_unified sanitize #156 originally implemented, so the split must partition weight, scales, and biases at the same output-axis half boundary with no quantization group straddling.
Land the fix, or the corrected expectation with a justification for why the previous invariant no longer holds.
Confirm the up_proj leg as well, since the current failure masks it.
Acceptance Criteria
Regression window identified: the last commit at which unified_sanitize_quantized_split_dequant_equivalence passed is named.
Root cause stated explicitly as either a real sanitize defect or a stale test expectation, with the evidence for that classification.
Fix (or corrected expectation) is merged, covering both the gate_proj and up_proj legs.
make verify-test is green on main.
make verify-fmt and make verify-clippy remain green.
Technical Considerations
Affected files: src/loading/vlm_gemma_unified.rs (the sanitize_gemma4_unified_weights split) and src/loading/vlm_gemma_unified_tests.rs (the parity test).
The test builds a synthetic quantized gate_up_proj via make_quantized_gate_up(2, 8, 64) and does not need a real quantized-MoE gemma4_unified checkpoint, so reproduction is cheap and deterministic.
If the defect is real, quantized gemma4_unified MoE checkpoints load with silently wrong expert weights, which argues for validating against a real quantized checkpoint before closing.
Problem / Background
loading::vlm::gemma_unified::tests::unified_sanitize_quantized_split_dequant_equivalencefails onmainat commite9f1f191.The failing assertion is at
src/loading/vlm_gemma_unified_tests.rs:354:leftismax_abs_diff(&gate, &ref_gate), so thegate_projleg produced bysanitize_gemma4_unified_weightsdequantizes to something structurally different from the referencesplit(dequantize(gate_up_proj)). Theup_projassertion that follows is never reached, so its status is unknown.The rest of the workspace is green:
make verify-testreports 7439 passed and this 1 failed, and bothmake verify-fmtandmake verify-clippy(--workspace --all-targets --features metal,accelerate -- -D warnings) pass.Already ruled out (do not re-run these)
Not the Rust 1.97.1 toolchain bump (chore(ci): bump the pinned Rust toolchain to 1.97.1 #1066). Holding the source constant at
e9f1f191and compiling with the previously pinned toolchain reproduces the failure with byte-identical values (3.7108002 against 0.0):The compiler version is not the variable.
Not an f32 reassociation tolerance problem. Reassociation differences of this kind land around 1e-6. A left-hand value of 3.71 against an expected 0.0 is a structural disagreement, not a rounding one. fix(tests): tolerate f32 reassociation in the mllama tile-selection parity test #1067 did fix a genuine reassociation case in the mllama tile-selection parity test immediately after the toolchain bump, which makes this failure easy to misfile as the same class. It is not.
Not epic epic: add Florence-2 (florence2) VLM support #850 (Florence-2). Its five merged PRs (feat(models): Florence-2 BART seq2seq engine and text core #1060, feat(models): Florence-2 DaViT vision backbone #1063, feat(models): Florence-2 vision-language fusion and full weight loading #1064, feat(models): Florence-2 processor, task prompts, and location tokens #1069, feat(models): Florence-2 end-to-end integration and real-checkpoint validation #1071) touch no gemma and no quantization code.
src/loading/vlm_gemma_unified.rsandsrc/loading/vlm_gemma_unified_tests.rswere last modified by feat: add video input support for Gemma 4 Unified (gemma4_unified) #400 (d35ef06f), well before that epic.Why this may have gone unnoticed
The root
cargo testgate historically covered only one of the workspace members (#1007), and the nightly run has not completed within its timeout (#1000). This test may have been red for some time and is only now visible because the workspace-widemake verify-testgate exercises it.Proposed Solution
d35ef06f) and fix: split quantized fused MoE experts in gemma4_unified sanitize #156 (6c1861e1) to find when it last passed. This is the cheapest first step and it decides everything downstream.sanitize_gemma4_unified_weightssplits quantized fused MoE experts in a way that no longer commutes with dequantization for thegateprojection (a real sanitize defect), or whether the test expectation is stale relative to a deliberate later change (a stale expectation). The commuting property is what fix: split quantized fused MoE experts in gemma4_unified sanitize #156 originally implemented, so the split must partitionweight,scales, andbiasesat the same output-axis half boundary with no quantization group straddling.up_projleg as well, since the current failure masks it.Acceptance Criteria
unified_sanitize_quantized_split_dequant_equivalencepassed is named.sanitizedefect or a stale test expectation, with the evidence for that classification.gate_projandup_projlegs.make verify-testis green onmain.make verify-fmtandmake verify-clippyremain green.Technical Considerations
src/loading/vlm_gemma_unified.rs(thesanitize_gemma4_unified_weightssplit) andsrc/loading/vlm_gemma_unified_tests.rs(the parity test).gate_up_projviamake_quantized_gate_up(2, 8, 64)and does not need a real quantized-MoE gemma4_unified checkpoint, so reproduction is cheap and deterministic.