Skip to content

fix: support BOM-prefixed encodings in KB TextParser - #9919

Draft
lxfight wants to merge 1 commit into
AstrBotDevs:masterfrom
lxfight:fix/kb-text-parser-encodings
Draft

fix: support BOM-prefixed encodings in KB TextParser#9919
lxfight wants to merge 1 commit into
AstrBotDevs:masterfrom
lxfight:fix/kb-text-parser-encodings

Conversation

@lxfight

@lxfight lxfight commented Sep 2, 2026

Copy link
Copy Markdown
Member

Fixes #9904

TextParser only tried ["utf-8", "gbk", "gb2312", "gb18030"], so txt files saved from Windows Notepad as "UTF-8 with BOM" or "Unicode" (UTF-16 LE) failed to decode and KB upload reported an unclear "无法读取或解析上传文件" error.

Modifications / 改动点

  • astrbot/core/knowledge_base/parsers/text_parser.py: detect UTF-8 / UTF-16 LE|BE BOMs first and decode accordingly; BOM-less files keep the original utf-8 → gbk → gb2312 → gb18030 sequence.
  • tests/test_text_parser.py: add unit tests covering plain UTF-8, UTF-8 BOM, UTF-16 LE/BE BOM, GBK, and undecodable input.

Note: BOM detection (instead of blindly adding utf-16 to the try sequence) is intentional — utf-16 decodes almost any byte stream, so trying it before GBK would silently mis-decode BOM-less GBK text into garbage instead of raising UnicodeDecodeError.

  • This is NOT a breaking change. / 这不是一个破坏性变更。

Screenshots or Test Results / 运行截图或测试结果

$ uv run pytest tests/test_text_parser.py -q
......                                                                   [100%]
6 passed in 0.41s
  • uv run ruff format / uv run ruff check pass.
  • Verification steps: in Windows Notepad, save a txt via "另存为 → 编码: Unicode" (UTF-16 LE with BOM) and upload it to a knowledge base; it now ingests correctly instead of failing.

Checklist / 检查清单

  • 😊 If there are new features added in the PR, I have discussed it with the authors through issues/emails, etc.
    / 如果 PR 中有新加入的功能,已经通过 Issue / 邮件等方式和作者讨论过。

  • 👀 My changes have been well-tested, and "Verification Steps" and "Screenshots" have been provided above.
    / 我的更改经过了良好的测试,并已在上方提供了“验证步骤”和“运行截图”

  • 🤓 I have ensured that no new dependencies are introduced, OR if new dependencies are introduced, they have been added to the appropriate locations in requirements.txt and pyproject.toml.
    / 我确保没有引入新依赖库,或者引入了新依赖库的同时将其添加到 requirements.txtpyproject.toml 文件相应位置。

  • 😮 My changes do not introduce malicious code.
    / 我的更改没有引入恶意代码。

Summary by Sourcery

Support BOM-prefixed text encodings in the knowledge-base text parser while retaining reliable fallback decoding for BOM-less files.

Bug Fixes:

  • Enable knowledge-base text uploads to correctly decode UTF-8, UTF-16 LE, and UTF-16 BE files with byte-order marks.
  • Preserve the existing encoding fallback behavior for BOM-less files and report undecodable content consistently.

Tests:

  • Add coverage for UTF-8, BOM-prefixed UTF-8 and UTF-16, GBK, and undecodable text parsing.

Windows Notepad's "UTF-8 with BOM" and "Unicode" (UTF-16 LE) save
options produce files that TextParser could not decode, failing KB
uploads with an unclear error.

Detect UTF-8/UTF-16 BOMs before falling back to the plain utf-8/gbk
sequence. BOM detection is intentional: blindly trying utf-16 before
GBK would silently mis-decode BOM-less GBK text as garbage.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug] 知识库解析:TextParser 不支持 UTF-16/BOM 编码,Windows 记事本另存的 txt 无法入库

1 participant