Repository navigation
fix: treat literal tiktoken special tokens as ordinary text in TiktokenCounter - #12871
HarshRajSinghania wants to merge 1 commit into
Conversation
TiktokenCounter.count() used Encoding.encode(), which raises ValueError when user text contains markers such as <|endoftext|>. Switch to encode_ordinary() so those strings are counted as regular text, matching DocumentSplitter after deepset-ai#12853. Fixes deepset-ai#12869
|
Someone is attempting to deploy a commit to the deepset Team on Vercel. A member of the Team first needs to authorize it. |
|
Harsh Raj Singhania seems not to be a GitHub user. You need a GitHub account to be able to sign the CLA. If you have already a GitHub account, please add the email address used for this commit to your account. You have signed the CLA already but the status is still pending? Let us recheck it. |
|
Hi @HarshRajSinghania, thanks a lot for your contribution! 🙏 We noticed that the Contributor License Agreement (CLA) check ( To get your PR reviewed, please sign the CLA via the link in the |
|
Hi @HarshRajSinghania, just a friendly reminder: this PR is still in draft because the Contributor License Agreement (CLA) hasn't been signed yet. We'd love to review your contribution! Please sign the CLA via the link in the |
|
Thank you for your efforts! We're closing this PR because it duplicates #13009, which makes the same one-line fix to |
Summary
TiktokenCounter.count()encoded rendered conversations withEncoding.encode(), which defaults todisallowed_special="all"and raisesValueErrorwhen the text contains a literal special-token marker such as<|endoftext|>.This switches the call to
encode_ordinary(), which counts those markers as ordinary text. That matches the approach already used byDocumentSplitter/RecursiveDocumentSplitterafter #12853.TiktokenCounterwas the remaininghaystack/caller that used a bare.encode()on user-supplied text.Motivation
Fixes #12869.
TiktokenCounteris used by Agent conversation-compaction hooks. The counted text can include tool results and retrieved documents the developer does not control. A page that quotes<|endoftext|>currently crashes compaction instead of returning a count.Implementation
haystack/token_counters/tiktoken_counter.py: useself._encoder.encode_ordinary(...)instead of.encode(...).test/token_counters/test_tiktoken_counter.py: addencode_ordinaryon the fake encoder so unit tests still record the rendered text; add two integration regression tests covering a message body and a tool description that contain the marker.No behavior change for inputs that already worked, because
encode()never emitted special tokens with the previous defaults.Testing
Environment limits prevented installing Haystack's full test extras here (PyPI 502s on this runner). The change is the one-line API swap requested in #12869 plus tests that follow the existing file's patterns. Please rely on CI for
pytest test/token_counters/test_tiktoken_counter.py.