Repository navigation
Conversation
…ry text TiktokenCounter.count() encoded the rendered conversation and tool schemas with Encoding.encode(), which defaults to disallowed_special="all" and so raised ValueError on any source text containing a string such as <|endoftext|>. Because allowed_special also defaults to set(), the call could never emit a special token: the strict check could only turn ordinary text into a crash. Encode with encode_ordinary() instead, so such strings are measured as ordinary text. This matches the fix already applied to DocumentSplitter and RecursiveDocumentSplitter in deepset-ai#12853. Refs deepset-ai#12869 Co-Authored-By: Claude Code <noreply@anthropic.com>
|
@SEVEN-us is attempting to deploy a commit to the deepset Team on Vercel. A member of the Team first needs to authorize it. |
|
Hi @SEVEN-us, thanks for your interest in contributing to Haystack! 🙏 This is an automated message to help us keep the review queue healthy. |
|
|
|
Hi @SEVEN-us, thanks a lot for your contribution! 🙏 We noticed that the Contributor License Agreement (CLA) check ( To get your PR reviewed, please sign the CLA via the link in the |
Related Issues
<|endoftext|>#12869Proposed Changes:
TiktokenCounter.count()encoded the rendered conversation and tool schemas withEncoding.encode(), which defaults todisallowed_special="all"and therefore raisesValueErroron any source text containing a string such as<|endoftext|>. Becauseallowed_specialalso defaults toset(), the call can never emit a special token — the strict check can only turn ordinary text into a crash.The counter now encodes with
encode_ordinary(), so such strings are measured as ordinary text.TiktokenCounterwas the last place inhaystack/still using the strict call; #12853 applied the same fix toDocumentSplitterandRecursiveDocumentSplitter.haystack/token_counters/tiktoken_counter.py: useencode_ordinary()and document the behaviour in the class docstring.test/token_counters/test_tiktoken_counter.py:_FakeEncodernow recordsencode_ordinary(), plus an integration test asserting a message containing<|endoftext|>is measured rather than rejected.docs-website/docs/token-counters/tiktokencounter.mdx: note that special-token strings are counted as ordinary text.How did you test it?
Locally with Hatch:
I also verified the new test catches the bug: with
encode()restored, it fails withNotes for the reviewer
Renaming
_FakeEncoder.encodetoencode_ordinarymirrors the mock change in #12853 and is what keeps the unit tests meaningful — if the counter went back toencode(), the stub no longer has that attribute.This PR was fully generated with an AI assistant. I have reviewed the changes and run the relevant tests.
Checklist
fix:,feat:,build:,chore:,ci:,docs:,style:,refactor:,perf:,test:and added!in case the PR includes breaking changes.