Repository navigation
docs(docs): Lead with what mache is for and what it can measure - #24
Merged
Merged
Conversation
The README opened on a name and a line about benchmarking, and the case for using it came several paragraphs later. It now starts with what it does and who it suits, and links to two new sections. What hosted runners can measure gives the margins an estimate reaches at a given number of pairs, and how long a normalized sequential test ran in simulation at two sets of bounds. Running a match without CI walks through playing shards with fastchess and pooling them, as it was run while writing it. pyproject.toml gains a summary, keywords and classifiers that say this is for chess engines, so PyPI's search can find it. The README says seven actions where it said six, and a few claims are narrowed to what the code does. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018ccJFsmZ4jAKYmFNgNBh9u
An independent read against the code and the test runs found a few things to fix. The clocks table reads as zeros without nodes=true, with a dash for time left, rather than as all zeros. fastchess appends to a pgn that is already there, which the walkthrough now warns about. The [0, 5] row said 14,000 pairs where one of its two results was 13,087. The simulation now says how many tests it ran and what it drew from, and the note on logistic bounds names strength.yml rather than every reusable workflow. The PyPI summary now describes the package, which is the tools and not the actions. The two new sections are linked by full address, since PyPI's copy of the README may not keep heading anchors, and some phrasing is plainer. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018ccJFsmZ4jAKYmFNgNBh9u
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
The README opened on the name and a line about benchmarking, and the case for using mache came several paragraphs later. It now starts with what mache does and which engines it suits, and links to two new sections.
What hosted runners can measure. How many pairs a default run plays, the margin an estimate reaches at 250, 1,000 and 4,000 pairs, and how long a normalized sequential test ran in simulation at [0, 10] and [0, 5]. The margins are computed from the code. The test lengths come from 300 simulated tests a row using mache's own code, judged every 250 pairs as with the workflow's default batch. It also says that
strength.ymltakes logistic bounds, that batch time depends on the time control and the engines, and what hosted runners cost.Running a match without CI. How to split a match across machines with
book-slice, play each shard with fastchess, and pool the shards withmatch-estimate. These commands were run while writing it, with fastchess built from the pinned tag and a test engine: two shards of 20 pairs pooled into 40 pairs from 80 games.Discoverability. Two badges in the README (the Tests workflow and the PyPI version), and a summary, keywords and classifiers in
pyproject.toml. The classifiers are checked against the published list and name only the Python versions the tests run on. PyPI shows these from the next release.Corrections to existing text. The README said six composite actions where there are seven, and the workflow comments now say "the actions" with no number. A claim about fishtest that I could not check is gone. A few phrasings are narrowed to what the code does, and two three-part lists are rewritten.
How it was checked
An independent read checked each claim against the code, the simulation output and the walkthrough runs, and looked for phrasing that reads as generated. Its findings are the second commit. Among them: the clocks table reads as zeros without
nodes=true(with a dash for time left) rather than as all zeros, fastchess appends to an existing pgn so each shard needs an empty directory, and one table cell was rounded up past one of its two results.The two sentences about GitHub billing were checked against GitHub's published pricing, not against its documentation page itself, which this environment could not reach.
Not in this pull request
Repository topics and the repository description are settings, not files, so they are for later.
🤖 Generated with Claude Code
https://claude.ai/code/session_018ccJFsmZ4jAKYmFNgNBh9u
Generated by Claude Code