Skip to content

normalizers: implement the Lowercase normalizer - #72

Merged
AlonKejzman merged 2 commits into
crusoecloud:mainfrom
GuyStone:guys/lowercase-normalizer
Oct 8, 2026
Merged

AlonKejzman merged 2 commits into
crusoecloud:mainfrom
GuyStone:guys/lowercase-normalizer

Conversation

@GuyStone

Copy link
Copy Markdown
Contributor

Summary

Adds the Lowercase normalizer. Tokenizers declaring {"type": "Lowercase"}
previously failed to load with unsupported normalizer type: Lowercase — the
config was parsed but had no runtime implementation.

Notes

Parity. Lowercasing is per character via char::to_lowercase, mirroring HF's
NormalizedString::lowercase. Deliberately not str::to_lowercase, which also
applies the Greek final-sigma rule — ὈΔΥΣΣΕΎΣ → ὀδυσσεύς where HF gives
ὀδυσσεύσ. U+03A3 is the only character the two disagree on.

Testing

9 unit tests covering the Cow::Borrowed guarantee, titlecase letters, the 1→2
expansion of İ U+0130, final sigma, and a capital at every one of 28 byte
offsets to pin the SWAR lane indexing. Plus
local_tests::lowercase_normalizer_matches_huggingface, which diffs directly
against the tokenizers crate. Full suite passes; fmt and clippy -D warnings
clean.

Also re-exports Prepend at the crate root, missed when it landed in edccd31.

GuyStone and others added 2 commits October 8, 2026 12:07
`NormalizerConfig::Lowercase` is already parsed but has no runtime
implementation, so tokenizers declaring `{"type": "Lowercase"}` fail to load
with "unsupported normalizer type: Lowercase". Add the Lowercase normalizer
and wire it into `Normalizer`. This unlocks standalone `Lowercase` and the
CLIP / sentence-transformers shape `Sequence[NFC, Replace, Lowercase]`;
uncased BERT still needs `NFD` + `StripAccents`.

Lowercasing is done per character with `char::to_lowercase`, mirroring HF's
`NormalizedString::lowercase`. This is deliberately not `str::to_lowercase`:
the std method additionally applies the Greek final-sigma rule, rendering
`ὈΔΥΣΣΕΎΣ` as `ὀδυσσεύς` where HF produces `ὀδυσσεύσ`. U+03A3 is the only
character the two disagree on. The "would this character change?" predicate
compares against the mapping itself rather than `char::is_uppercase()`, which
is false for titlecase letters such as `Dž` U+01C5 and would produce wrong
output.

Normalization runs single-threaded over the whole document before rayon
parallelism starts at BPE, so both the scan and the transform take a SWAR fast
path over ASCII, mirroring `ascii_lower_run_end` in `pre_tokenizers::scan`. The
transform is branch-free: `0x80 >> 2 == 0x20` is exactly the ASCII case bit.
Measured against a straightforward `chars().flat_map(char::to_lowercase)` on
~100 KiB inputs: 68x faster on already-lowercase ASCII, 45x on mixed case and
on ALL CAPS, and 0.95x on CJK (both dominated by std's LOWERCASE_TABLE binary
search, which HF pays too).

Tests cover the Cow::Borrowed guarantee, a capital at every one of 28 byte
offsets to pin the SWAR lane indexing and window boundary, titlecase letters,
the 1->2 expansion of `İ` U+0130, and final sigma — the last asserting
inequality with `str::to_lowercase` so a refactor towards it fails loudly.
`local_tests::lowercase_normalizer_matches_huggingface` diffs against the
`tokenizers` crate directly.

Also re-exports `Prepend` at the crate root, which was missed when it landed
in edccd31 (`Nfc` and `Replace` are both there).

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Rebasing onto main brings in `Normalizer::is_identity_on` (crusoecloud#74), whose match
did not cover the new `Lowercase` variant, so the crate no longer compiled.
`Lowercase` is the identity exactly when `first_change` finds nothing to
lower, which also holds for every substring, so expose that scan as
`Lowercase::is_normalized` (as `Nfc::is_normalized`) and use it: already
lowercase text keeps the in-place encode path instead of being copied
through normalization.

The HF parity test now also asserts that `is_identity_on` agrees with
whether HF's Lowercase changed the text, since the encode path skips
normalization when it says so.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
@AlonKejzman
AlonKejzman force-pushed the guys/lowercase-normalizer branch from 025d9c2 to ba21355 Compare October 8, 2026 09:17
@AlonKejzman
AlonKejzman merged commit 9d66ab2 into crusoecloud:main Oct 8, 2026
22 checks passed
@AlonKejzman

Copy link
Copy Markdown
Collaborator

Thank you for your contribution!

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants