Repository navigation
EPUB / PDF / 检索网页 / 离线版:把超长的收益、备注、来源拆成子条目 - #96
Merged
Merged
Conversation
电子书里不少「收益」一条就是一整段一千多字。比如第 31 节第 1 条,把十二条路的 门槛(兵役法、公务员录用规定、消防员招录办法……)全写在一行里,阅读器里就是一整块, 找不到哪句讲的是哪条路。 新增 tools/lib/split-items.mjs,只在生成 EPUB 时处理,正文 md 一个字不动: - 收益、备注:超过 300 字的,在句号处、且下一句开始引用新的依据时断开 (「X看《…》」「《某条例》」「刑法第…条」「另一项试验」「案例二」「芬兰…研究」等)。 没有明确换话题点的长段不硬拆。 - 来源:一篇文献一行,按「<网址>;」后面接下一条来切。 - 计括号层级时跳过「」《》()和 <http…> 里面的标点,「P<0.0001」这种不当成括号。 效果:742 条超过 300 字的收益/备注/来源里拆了 413 条。44 个源文件拆前拆后 去掉列表符号逐字比对一致(来源栏文献之间的「;」换成了换行)。epubcheck 5.4.0 无错误无警告。 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X6s254TjQjvF9YgR8uK1LE
上一个提交只改了 EPUB,网页、离线单文件、PDF 上同一条收益仍是一整块。这次四个出口都拆, 正文 md 仍不动(index.html 按行解析、sync-stats 按「- 来源:」行数链接都依赖一个字段一行)。 - tools/lib/split-items.mjs:导出 splitGain / splitSrc。来源栏改用 index.html 原有 splitSrc 的规则(括号外的「;」一条一行),四个出口的来源拆法从此一致。 - index.html:在原有 splitSrc 旁边加同一份 splitGain,收益、备注拆开时渲染成一块一行的列表, 两份规则用 // <split-rules> 标出来。离线单文件用的是同一个 index.html,自动跟上。 - tools/pdf/build.mjs:read 之后同样套 splitLongItems。 - tools/check-split.mjs:从 index.html 抽出 <split-rules> 那段,和 split-items.mjs 在全书 2013 个收益/备注/来源字段上逐条比对;也查拆完能否逐字拼回原文、** 有没有被拆开。 CI 新增「拆分规则一致性检查」job,和交叉引用检查一样不挡发布。 - 换话题规则排除「该 / 本 / 此 / 这 / 同 / 上述」开头:「该办法第五条」是接着上文讲的,不另起。 - CLAUDE.md 电子版那段记一笔:规则两份要同步。 验证:check-split 0 处问题(拆开 554 个字段);check-plain、check-refs --check、 sync-stats --check 全过;epubcheck 5.4.0 0 errors / 0 warnings;pandoc 3.11 + typst 0.15.1 出 PDF 正常;本地起服务器和 file:// 打开离线单文件,第 31 节第 1 条的收益都拆成 11 行,控制台无报错。 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01X6s254TjQjvF9YgR8uK1LE
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
问题
不少「收益」「备注」一条就是一整段,有的上千字。比如第 31 节第 1 条,把十二条路的门槛(兵役法、公务员录用规定、消防员招录办法……)全写在一行里。EPUB、PDF、检索网页、离线单文件看到的都是一整块,很难找到哪句讲的是哪条路。
改法
四个出口用同一套规则拆,正文 md 一个字不动(index.html 按行解析、sync-stats 按行数链接,都依赖一个字段一行)。
splitSrc的规则,括号外的「;」一条一行。改动的文件:
tools/lib/split-items.mjs:新增,EPUB、PDF 构建时调用。index.html:加同一份splitGain,收益、备注渲染成一块一行的列表。离线单文件自动跟上。tools/check-split.mjs+ CI job「拆分规则一致性检查」:在全书逐条比对两份规则的结果,也查拆完能否逐字拼回原文。不挡发布。CLAUDE.md:记一笔,规则两份要同步改。验证