Skip to content

feat(package): Let the sequential test take its error rates - #28

Merged
aywrite merged 1 commit into
mainfrom
claude/stoic-ramanujan-xqxekr
Sep 30, 2026
Merged

aywrite merged 1 commit into
mainfrom
claude/stoic-ramanujan-xqxekr

Conversation

@aywrite

@aywrite aywrite commented Sep 30, 2026

Copy link
Copy Markdown
Owner

What changed

  • match-estimate takes --alpha and --beta, the chances of a wrong pass and a wrong fail the test accepts.
    • Both are 0.05 unless given, as before.
    • Each must be above 0 and below 0.5. Anything else is refused with a plain error, and a rate given without --elo0 is refused too.
    • They move Wald's bounds and nothing else. The ratio is the same whatever they are.
  • A verdict at the default rates reads as it did.
  • A verdict at other rates names them in the report, the line and the trailer, e.g. SPRT [0, 10] alpha=0.01 beta=0.01 passed. A pass at one rate is not a pass at another, so a quoted verdict has to say which.
    • The report says which rate goes with which verdict ("at error rates of 20 percent for a wrong pass and 1 percent for a wrong fail").
    • An inconclusive verdict asks for the same rates again, as it already does for the normalized model.
  • The json adds alpha and beta beside lower and upper. Adding fields is allowed by format 1.
  • actions/summarise-match takes them as alpha and beta, both defaulting to 0.05.
  • The README describes the flags. It now says the simulated test lengths are at the default rates.

One wording change at the defaults: the report now says "a 5 percent error rate" where it said "a five percent error rate".

A decision this reverses

The comment beside ALPHA = BETA = 0.05 said the rates were fixed on purpose, because they are what a verdict means. This keeps that meaning by making every verdict at other rates name them. If the rates should stay fixed, this PR is the one to close.

Not in this PR

strength.yml and batch.yml pin their actions at v0.6.0, so they cannot pass the new inputs until a release holds them. That is the same route sprt_model took. Wiring alpha and beta through the reusable workflows, and into the manifest line that now hardcodes alpha=0.05 beta=0.05, is a follow-up after the next release.

How it was checked

  • python3 -m pytest tests passes (420 tests), and pre-commit run --all-files is clean.
  • New tests cover:
    • the bounds against Wald's formula at four pairs of rates, including uneven ones
    • one set of pairs (ratio 3.40) that passes at the defaults and stays inconclusive at 1 percent each
    • each output naming non-default rates and leaving the defaults unnamed
    • --alpha 0.05 reading the same as leaving it out
    • the refusals
    • the json fields
    • the action passing the rates through, and passing the defaults when they are unset
  • Against fastchess 1.8.2, on one match played with -sprt elo0=0 elo1=50 alpha=0.2 beta=0.01: fastchess printed LLR: 1.60 (-4.38, 1.60) and accepted H1. match-estimate --alpha 0.2 --beta 0.01 on the same games printed LLR 1.60 (-4.38, 1.60) and passed.

🤖 Generated with Claude Code

https://claude.ai/code/session_018ccJFsmZ4jAKYmFNgNBh9u


Generated by Claude Code

match-estimate takes --alpha and --beta, the chances of a wrong pass and
a wrong fail the test accepts. Both are five percent unless given, as
they were before, and must lie between nought and a half. They set
Wald's bounds and nothing else: the ratio is the same whatever they are.

A verdict at the default rates reads as it did. One reached at other
rates names them in the report, the line and the trailer
(`SPRT [0, 10] alpha=0.01 beta=0.01 passed`), since a pass at one rate
is not a pass at another. An inconclusive one asks for the same rates
again. The json adds alpha and beta beside the bounds they gave, which
the format allows.

The rates were fixed until now, on the grounds that they are what a
verdict means. Naming them wherever a verdict is quoted keeps that
meaning while letting a caller choose them.

actions/summarise-match takes them as alpha and beta. The reusable
workflows pin the actions at a release, so strength.yml can pass them
once a release holds this.

Checked against fastchess 1.8.2 on one match at alpha 0.2 and beta
0.01. Both printed bounds of (-4.38, 1.60) and a final ratio of 1.60,
and both passed it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ccJFsmZ4jAKYmFNgNBh9u
@aywrite
aywrite merged commit 89f935d into main Sep 30, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants