Describe the bug
cuml.preprocessing.TargetEncoder does not preserve or match a None category learned during fit when transforming new data.
In the reproducer below, None occurs in the training feature with target value 1.5. With smooth=0, transforming the same category should therefore produce its learned category mean, 1.5.
Instead, cuML returns 2.25, which is the overall mean of the complete target array:
(1.0 + 2.0 + 1.5 + 3.0 + 2.5 + 3.5) / 6 = 2.25
This indicates that cuML treats the previously fitted None value as an unseen category and incorrectly applies the global fallback. The equivalent scikit-learn estimator correctly returns 1.5.
Steps/Code to reproduce bug
cuML reproducer:
import numpy as np
from cuml.preprocessing import TargetEncoder
X = np.array([
["a"],
["b"],
[None],
["c"],
["b"],
["c"],
], dtype=object)
y = np.array([1.0, 2.0, 1.5, 3.0, 2.5, 3.5])
X_test = np.array([
["a"],
[None],
], dtype=object)
k = TargetEncoder(
smooth=0,
).fit(X, y)
print(k.transform(X_test))
Output:
For comparison, the equivalent scikit-learn code:
import numpy as np
from sklearn.preprocessing import TargetEncoder
X = np.array([
["a"],
["b"],
[None],
["c"],
["b"],
["c"],
], dtype=object)
y = np.array([1.0, 2.0, 1.5, 3.0, 2.5, 3.5])
X_test = np.array([
["a"],
[None],
], dtype=object)
k = TargetEncoder(
smooth=0,
).fit(X, y)
print(k.transform(X_test))
Output:
Expected behavior
A missing-value category that occurs during fit should remain a known category during transform.
With smooth=0, each known category should be encoded using its observed target mean. For this input:
- category
"a" should be encoded as 1.0;
- category
None should be encoded as 1.5.
The expected result, consistent with scikit-learn, is:
The global target statistic 2.25 should be used only for genuinely unseen categories, not for a None category present in the fitted data.
Environment details (please complete the following information):
- Environment location: Docker
- Linux Distro/Architecture: Ubuntu 24.04 / x86_64
- GPU Model/Driver: NVIDIA GeForce RTX 4090 / 595.71.05
- CUDA: 13.2
- Method of cuDF & cuML install: conda
conda list:
conda list
# packages in environment at /opt/conda/envs/rapids-26.08:
#
# Name Version Build Channel
# Name Version Build Channel
python 3.14.6 h242f9ac_102_cp314 conda-forge
numpy 2.4.6 py314h2b28147_0 conda-forge
scipy 1.16.3 py314hf07bd8e_2 conda-forge
scikit-learn 1.9.0 np2py314hf09ca88_0 conda-forge
rapids 26.08.00 cuda13_260806_c2656556 rapidsai
cuml 26.08.00 cuda13_cp311_abi3_260805_265b9da6 rapidsai
libcuml 26.08.00 cuda13_260805_265b9da6 rapidsai
cudf 26.08.00 cuda13_cp311_abi3_260805_ff5b362d rapidsai
libraft 26.08.00 cuda13_260805_ebf92684 rapidsai
libraft-headers 26.08.00 cuda13_260805_ebf92684 rapidsai
pylibraft 26.08.00 cuda13_cp311_abi3_260805_ebf92684 rapidsai
cuvs 26.08.01 cuda13_cp311_abi3_260806_25b1be43 rapidsai
libcuvs 26.08.01 cuda13_260806_25b1be43 rapidsai
cupy 14.1.1 py314hdea9c46_0 conda-forge
cupy-core 14.1.1 py314hcd3b49b_0 conda-forge
numba 0.64.0 py314h8169c2f_0 conda-forge
numba-cuda 0.30.4 py314h42812f9_0 conda-forge
rmm 26.08.00 cuda13_cp311_abi3_260805_42d059f1 rapidsai
librmm 26.08.00 cuda13_260805_42d059f1 rapidsai
cuda-version 13.3 hcbadf70_3 conda-forge
cuda-bindings 13.3.1 py314h42812f9_1 conda-forge
cuda-cudart 13.3.29 hecca717_0 conda-forge
cuda-nvrtc 13.3.33 hecca717_0 conda-forge
libcublas 13.6.0.2 h676940d_0 conda-forge
libcusolver 12.2.6.9 h676940d_0 conda-forge
libcusparse 12.8.2.51 hecca717_0 conda-forge
libcurand 10.4.3.29 h676940d_0 conda-forge
Additional context
- The training input is a NumPy object array with shape
(6, 1).
None appears once during fitting and is also present in X_test.
smooth=0 removes smoothing as a possible explanation for the discrepancy.
- The value returned by cuML for
None, 2.25, is exactly the global mean of y, strongly indicating that the category lookup did not match the fitted missing-value key.
- This may result from null values being excluded during the groupby aggregation, or from inconsistent null-key handling between the fitted encoding table and the transform-time join.
- The ordinary string category
"a" is encoded correctly as 1.0.
Describe the bug
cuml.preprocessing.TargetEncoderdoes not preserve or match aNonecategory learned duringfitwhen transforming new data.In the reproducer below,
Noneoccurs in the training feature with target value1.5. Withsmooth=0, transforming the same category should therefore produce its learned category mean,1.5.Instead, cuML returns
2.25, which is the overall mean of the complete target array:This indicates that cuML treats the previously fitted
Nonevalue as an unseen category and incorrectly applies the global fallback. The equivalent scikit-learn estimator correctly returns1.5.Steps/Code to reproduce bug
cuML reproducer:
Output:
For comparison, the equivalent scikit-learn code:
Output:
Expected behavior
A missing-value category that occurs during
fitshould remain a known category duringtransform.With
smooth=0, each known category should be encoded using its observed target mean. For this input:"a"should be encoded as1.0;Noneshould be encoded as1.5.The expected result, consistent with scikit-learn, is:
The global target statistic
2.25should be used only for genuinely unseen categories, not for aNonecategory present in the fitted data.Environment details (please complete the following information):
conda list:Additional context
(6, 1).Noneappears once during fitting and is also present inX_test.smooth=0removes smoothing as a possible explanation for the discrepancy.None,2.25, is exactly the global mean ofy, strongly indicating that the category lookup did not match the fitted missing-value key."a"is encoded correctly as1.0.