Repository navigation
feat(package): Let the sequential test take its error rates - #28
Merged
Merged
Conversation
match-estimate takes --alpha and --beta, the chances of a wrong pass and a wrong fail the test accepts. Both are five percent unless given, as they were before, and must lie between nought and a half. They set Wald's bounds and nothing else: the ratio is the same whatever they are. A verdict at the default rates reads as it did. One reached at other rates names them in the report, the line and the trailer (`SPRT [0, 10] alpha=0.01 beta=0.01 passed`), since a pass at one rate is not a pass at another. An inconclusive one asks for the same rates again. The json adds alpha and beta beside the bounds they gave, which the format allows. The rates were fixed until now, on the grounds that they are what a verdict means. Naming them wherever a verdict is quoted keeps that meaning while letting a caller choose them. actions/summarise-match takes them as alpha and beta. The reusable workflows pin the actions at a release, so strength.yml can pass them once a release holds this. Checked against fastchess 1.8.2 on one match at alpha 0.2 and beta 0.01. Both printed bounds of (-4.38, 1.60) and a final ratio of 1.60, and both passed it. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018ccJFsmZ4jAKYmFNgNBh9u
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What changed
match-estimatetakes--alphaand--beta, the chances of a wrong pass and a wrong fail the test accepts.--elo0is refused too.SPRT [0, 10] alpha=0.01 beta=0.01 passed. A pass at one rate is not a pass at another, so a quoted verdict has to say which.alphaandbetabesidelowerandupper. Adding fields is allowed by format 1.actions/summarise-matchtakes them asalphaandbeta, both defaulting to0.05.One wording change at the defaults: the report now says "a 5 percent error rate" where it said "a five percent error rate".
A decision this reverses
The comment beside
ALPHA = BETA = 0.05said the rates were fixed on purpose, because they are what a verdict means. This keeps that meaning by making every verdict at other rates name them. If the rates should stay fixed, this PR is the one to close.Not in this PR
strength.ymlandbatch.ymlpin their actions atv0.6.0, so they cannot pass the new inputs until a release holds them. That is the same routesprt_modeltook. Wiringalphaandbetathrough the reusable workflows, and into the manifest line that now hardcodesalpha=0.05 beta=0.05, is a follow-up after the next release.How it was checked
python3 -m pytest testspasses (420 tests), andpre-commit run --all-filesis clean.--alpha 0.05reading the same as leaving it out-sprt elo0=0 elo1=50 alpha=0.2 beta=0.01: fastchess printedLLR: 1.60 (-4.38, 1.60)and accepted H1.match-estimate --alpha 0.2 --beta 0.01on the same games printedLLR 1.60 (-4.38, 1.60)and passed.🤖 Generated with Claude Code
https://claude.ai/code/session_018ccJFsmZ4jAKYmFNgNBh9u
Generated by Claude Code