Repository navigation
docs: 在「衍生工具」中新增结构化数据集 htlb-dataset - #99
Merged
Merged
Conversation
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
是什么
htlb-dataset 把全书正文解析成机器可读的结构化数据(JSON / CSV / SQLite),每天自动同步上游。
为什么提这个 PR
上游只发布 Markdown,不发布数据;于是每个衍生工具(小程序、App、翻译、本地化)都在各自重写解析器,条数口径互相打架(528 / 612 / 672 同时在流传)。本仓库只做一件事:当所有衍生工具的公共数据源。
口径对齐
解析规则逐字对齐上游
tools/sync-stats.mjs与index.html的COST_W/e.ratio:脚本启动时比对上游
index.html的权重行,上游一旦改规则就报错退出,不会静默算出对不上的数。合规
如果上游更愿意把数据入口放在别处(比如独立小节或 docs/),我也可以改。谢谢维护!