Skip to content

docs(docs): Lead with what mache is for and what it can measure - #24

Merged
aywrite merged 2 commits into
mainfrom
claude/stoic-ramanujan-xqxekr
Sep 28, 2026
Merged

aywrite merged 2 commits into
mainfrom
claude/stoic-ramanujan-xqxekr

Conversation

@aywrite

@aywrite aywrite commented Sep 27, 2026

Copy link
Copy Markdown
Owner

What changed

The README opened on the name and a line about benchmarking, and the case for using mache came several paragraphs later. It now starts with what mache does and which engines it suits, and links to two new sections.

What hosted runners can measure. How many pairs a default run plays, the margin an estimate reaches at 250, 1,000 and 4,000 pairs, and how long a normalized sequential test ran in simulation at [0, 10] and [0, 5]. The margins are computed from the code. The test lengths come from 300 simulated tests a row using mache's own code, judged every 250 pairs as with the workflow's default batch. It also says that strength.yml takes logistic bounds, that batch time depends on the time control and the engines, and what hosted runners cost.

Running a match without CI. How to split a match across machines with book-slice, play each shard with fastchess, and pool the shards with match-estimate. These commands were run while writing it, with fastchess built from the pinned tag and a test engine: two shards of 20 pairs pooled into 40 pairs from 80 games.

Discoverability. Two badges in the README (the Tests workflow and the PyPI version), and a summary, keywords and classifiers in pyproject.toml. The classifiers are checked against the published list and name only the Python versions the tests run on. PyPI shows these from the next release.

Corrections to existing text. The README said six composite actions where there are seven, and the workflow comments now say "the actions" with no number. A claim about fishtest that I could not check is gone. A few phrasings are narrowed to what the code does, and two three-part lists are rewritten.

How it was checked

An independent read checked each claim against the code, the simulation output and the walkthrough runs, and looked for phrasing that reads as generated. Its findings are the second commit. Among them: the clocks table reads as zeros without nodes=true (with a dash for time left) rather than as all zeros, fastchess appends to an existing pgn so each shard needs an empty directory, and one table cell was rounded up past one of its two results.

The two sentences about GitHub billing were checked against GitHub's published pricing, not against its documentation page itself, which this environment could not reach.

Not in this pull request

Repository topics and the repository description are settings, not files, so they are for later.

🤖 Generated with Claude Code

https://claude.ai/code/session_018ccJFsmZ4jAKYmFNgNBh9u


Generated by Claude Code

The README opened on a name and a line about benchmarking, and the case
for using it came several paragraphs later. It now starts with what it
does and who it suits, and links to two new sections.

What hosted runners can measure gives the margins an estimate reaches at
a given number of pairs, and how long a normalized sequential test ran
in simulation at two sets of bounds. Running a match without CI walks
through playing shards with fastchess and pooling them, as it was run
while writing it.

pyproject.toml gains a summary, keywords and classifiers that say this
is for chess engines, so PyPI's search can find it. The README says
seven actions where it said six, and a few claims are narrowed to what
the code does.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ccJFsmZ4jAKYmFNgNBh9u
An independent read against the code and the test runs found a few
things to fix. The clocks table reads as zeros without nodes=true, with
a dash for time left, rather than as all zeros. fastchess appends to a
pgn that is already there, which the walkthrough now warns about. The
[0, 5] row said 14,000 pairs where one of its two results was 13,087.
The simulation now says how many tests it ran and what it drew from,
and the note on logistic bounds names strength.yml rather than every
reusable workflow.

The PyPI summary now describes the package, which is the tools and not
the actions. The two new sections are linked by full address, since
PyPI's copy of the README may not keep heading anchors, and some
phrasing is plainer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ccJFsmZ4jAKYmFNgNBh9u
@aywrite
aywrite merged commit 05bb4e9 into main Sep 28, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants