fix: support BOM-prefixed encodings in KB TextParser - #9919
Draft
lxfight wants to merge 1 commit into
Draft
Conversation
Windows Notepad's "UTF-8 with BOM" and "Unicode" (UTF-16 LE) save options produce files that TextParser could not decode, failing KB uploads with an unclear error. Detect UTF-8/UTF-16 BOMs before falling back to the plain utf-8/gbk sequence. BOM detection is intentional: blindly trying utf-16 before GBK would silently mis-decode BOM-less GBK text as garbage.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #9904
TextParseronly tried["utf-8", "gbk", "gb2312", "gb18030"], so txt files saved from Windows Notepad as "UTF-8 with BOM" or "Unicode" (UTF-16 LE) failed to decode and KB upload reported an unclear "无法读取或解析上传文件" error.Modifications / 改动点
astrbot/core/knowledge_base/parsers/text_parser.py: detectUTF-8/UTF-16 LE|BEBOMs first and decode accordingly; BOM-less files keep the originalutf-8 → gbk → gb2312 → gb18030sequence.tests/test_text_parser.py: add unit tests covering plain UTF-8, UTF-8 BOM, UTF-16 LE/BE BOM, GBK, and undecodable input.Note: BOM detection (instead of blindly adding
utf-16to the try sequence) is intentional —utf-16decodes almost any byte stream, so trying it before GBK would silently mis-decode BOM-less GBK text into garbage instead of raisingUnicodeDecodeError.Screenshots or Test Results / 运行截图或测试结果
uv run ruff format/uv run ruff checkpass.Checklist / 检查清单
😊 If there are new features added in the PR, I have discussed it with the authors through issues/emails, etc.
/ 如果 PR 中有新加入的功能,已经通过 Issue / 邮件等方式和作者讨论过。
👀 My changes have been well-tested, and "Verification Steps" and "Screenshots" have been provided above.
/ 我的更改经过了良好的测试,并已在上方提供了“验证步骤”和“运行截图”。
🤓 I have ensured that no new dependencies are introduced, OR if new dependencies are introduced, they have been added to the appropriate locations in
requirements.txtandpyproject.toml./ 我确保没有引入新依赖库,或者引入了新依赖库的同时将其添加到
requirements.txt和pyproject.toml文件相应位置。😮 My changes do not introduce malicious code.
/ 我的更改没有引入恶意代码。
Summary by Sourcery
Support BOM-prefixed text encodings in the knowledge-base text parser while retaining reliable fallback decoding for BOM-less files.
Bug Fixes:
Tests: