test_kmeans_device_buffer_samples_host_path asserts that the host-streaming and device KMeans fit paths agree on inertia_ to within rtol=1e-3 with no atol (python/cuml/tests/test_kmeans.py). For the init_method="default" parametrisation it fails intermittently in CI.
Most recent occurrence: nightly run (on main) https://github.com/NVIDIA/cuml/actions/runs/35706866779 conda-python-tests-singlegpu / 13.3.0, 3.14, amd64, ubuntu26.04, rtxpro6000.
This is a flake-y test because the two things being compared (host streaming and device) start from a different sampled centre. I think because there is subsampling/streaming in one case it can not/very unlikely to sample the same init as the "no streaming" case. The test some times passes and sometimes fails.
I'm not sure what exactly we are trying to check with this test. Maybe we can adjust the test based on what we are trying to check. For example, is there a way to figure out what the subsample is for the streaming approach and use that subsample for the other case?
I asked AI and after a bit of back and forth it proposed the following idea for a fix. The fix sounds reasonable to me, but it does not try to address the question of "what is it we are trying to check?".
Make the default-init assertion one-sided, and keep the tight two-sided check where it is
justified. This mirrors the pattern already used twice in this same file (:119-120 and :161-165, introduced by #8099 for the same reason), and would have passed the observed failure.
if init_method == "explicit":
# Both paths start from the same centers, so this is one optimization
# run two ways and the results should agree closely.
np.testing.assert_allclose(
float(host_model.inertia_), float(dev_model.inertia_), rtol=1e-3
)
else:
# The two paths seed differently and neither fit is reproducible
# run-to-run, so they can converge to different local optima. Only
# flag the host path being materially worse than the device path.
if float(host_model.inertia_) > float(dev_model.inertia_):
np.testing.assert_allclose(
float(host_model.inertia_), float(dev_model.inertia_), rtol=1e-2
)
The cluster_centers_ comparison and the adjusted_rand_score >= 0.97 check stay as they are;
the latter already covers the "did the two paths find the same clustering" question for the
default branch.
test_kmeans_device_buffer_samples_host_pathasserts that the host-streaming and device KMeans fit paths agree oninertia_to withinrtol=1e-3with noatol(python/cuml/tests/test_kmeans.py). For theinit_method="default"parametrisation it fails intermittently in CI.Most recent occurrence: nightly run (on
main) https://github.com/NVIDIA/cuml/actions/runs/35706866779conda-python-tests-singlegpu / 13.3.0, 3.14, amd64, ubuntu26.04, rtxpro6000.This is a flake-y test because the two things being compared (host streaming and device) start from a different sampled centre. I think because there is subsampling/streaming in one case it can not/very unlikely to sample the same
initas the "no streaming" case. The test some times passes and sometimes fails.I'm not sure what exactly we are trying to check with this test. Maybe we can adjust the test based on what we are trying to check. For example, is there a way to figure out what the subsample is for the streaming approach and use that subsample for the other case?
I asked AI and after a bit of back and forth it proposed the following idea for a fix. The fix sounds reasonable to me, but it does not try to address the question of "what is it we are trying to check?".