Skip to content

[BUG] TargetEncoder treats a fitted None category as unseen during transform #8697

Description

@poiuyt009

Describe the bug

cuml.preprocessing.TargetEncoder does not preserve or match a None category learned during fit when transforming new data.

In the reproducer below, None occurs in the training feature with target value 1.5. With smooth=0, transforming the same category should therefore produce its learned category mean, 1.5.

Instead, cuML returns 2.25, which is the overall mean of the complete target array:

(1.0 + 2.0 + 1.5 + 3.0 + 2.5 + 3.5) / 6 = 2.25

This indicates that cuML treats the previously fitted None value as an unseen category and incorrectly applies the global fallback. The equivalent scikit-learn estimator correctly returns 1.5.

Steps/Code to reproduce bug

cuML reproducer:

import numpy as np
from cuml.preprocessing import TargetEncoder 

X = np.array([
    ["a"],
    ["b"],
    [None],
    ["c"],
    ["b"],
    ["c"],
], dtype=object)

y = np.array([1.0, 2.0, 1.5, 3.0, 2.5, 3.5])

X_test = np.array([
    ["a"],
    [None],
], dtype=object)

k = TargetEncoder(
    smooth=0,
).fit(X, y)

print(k.transform(X_test))

Output:

[[1.  ]
 [2.25]]

For comparison, the equivalent scikit-learn code:

import numpy as np
from sklearn.preprocessing import TargetEncoder 

X = np.array([
    ["a"],
    ["b"],
    [None],
    ["c"],
    ["b"],
    ["c"],
], dtype=object)

y = np.array([1.0, 2.0, 1.5, 3.0, 2.5, 3.5])

X_test = np.array([
    ["a"],
    [None],
], dtype=object)

k = TargetEncoder(
    smooth=0,
).fit(X, y)

print(k.transform(X_test))

Output:

[[1. ]
 [1.5]]

Expected behavior

A missing-value category that occurs during fit should remain a known category during transform.

With smooth=0, each known category should be encoded using its observed target mean. For this input:

  • category "a" should be encoded as 1.0;
  • category None should be encoded as 1.5.

The expected result, consistent with scikit-learn, is:

[[1. ]
 [1.5]]

The global target statistic 2.25 should be used only for genuinely unseen categories, not for a None category present in the fitted data.

Environment details (please complete the following information):

  • Environment location: Docker
  • Linux Distro/Architecture: Ubuntu 24.04 / x86_64
  • GPU Model/Driver: NVIDIA GeForce RTX 4090 / 595.71.05
  • CUDA: 13.2
  • Method of cuDF & cuML install: conda

conda list:

conda list
# packages in environment at /opt/conda/envs/rapids-26.08:
#
# Name                              Version          Build                                         Channel
# Name              Version       Build                                      Channel
python              3.14.6        h242f9ac_102_cp314                         conda-forge
numpy               2.4.6         py314h2b28147_0                            conda-forge
scipy               1.16.3        py314hf07bd8e_2                            conda-forge
scikit-learn        1.9.0         np2py314hf09ca88_0                         conda-forge
rapids              26.08.00      cuda13_260806_c2656556                     rapidsai
cuml                26.08.00      cuda13_cp311_abi3_260805_265b9da6          rapidsai
libcuml             26.08.00      cuda13_260805_265b9da6                     rapidsai
cudf                26.08.00      cuda13_cp311_abi3_260805_ff5b362d          rapidsai
libraft             26.08.00      cuda13_260805_ebf92684                     rapidsai
libraft-headers     26.08.00      cuda13_260805_ebf92684                     rapidsai
pylibraft           26.08.00      cuda13_cp311_abi3_260805_ebf92684          rapidsai
cuvs                26.08.01      cuda13_cp311_abi3_260806_25b1be43          rapidsai
libcuvs             26.08.01      cuda13_260806_25b1be43                     rapidsai
cupy                14.1.1        py314hdea9c46_0                            conda-forge
cupy-core           14.1.1        py314hcd3b49b_0                            conda-forge
numba               0.64.0        py314h8169c2f_0                            conda-forge
numba-cuda          0.30.4        py314h42812f9_0                            conda-forge
rmm                 26.08.00      cuda13_cp311_abi3_260805_42d059f1          rapidsai
librmm              26.08.00      cuda13_260805_42d059f1                     rapidsai
cuda-version        13.3           hcbadf70_3                                 conda-forge
cuda-bindings       13.3.1        py314h42812f9_1                            conda-forge
cuda-cudart         13.3.29       hecca717_0                                 conda-forge
cuda-nvrtc          13.3.33       hecca717_0                                 conda-forge
libcublas           13.6.0.2      h676940d_0                                 conda-forge
libcusolver         12.2.6.9      h676940d_0                                 conda-forge
libcusparse         12.8.2.51     hecca717_0                                 conda-forge
libcurand           10.4.3.29     h676940d_0                                 conda-forge

Additional context

  • The training input is a NumPy object array with shape (6, 1).
  • None appears once during fitting and is also present in X_test.
  • smooth=0 removes smoothing as a possible explanation for the discrepancy.
  • The value returned by cuML for None, 2.25, is exactly the global mean of y, strongly indicating that the category lookup did not match the fitted missing-value key.
  • This may result from null values being excluded during the groupby aggregation, or from inconsistent null-key handling between the fitted encoding table and the transform-time join.
  • The ordinary string category "a" is encoded correctly as 1.0.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

bugSomething isn't working

Type

No type

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions