Skip to content

feat(ci): Let strength.yml take its sequential test in normalized elo - #26

Merged
aywrite merged 1 commit into
mainfrom
claude/stoic-ramanujan-xqxekr
Sep 29, 2026
Merged

aywrite merged 1 commit into
mainfrom
claude/stoic-ramanujan-xqxekr

Conversation

@aywrite

@aywrite aywrite commented Sep 28, 2026

Copy link
Copy Markdown
Owner

What changed

v0.6.0 gave actions/summarise-match an sprt_model input. The reusable workflows pin their actions at a release, so they could not pass it until a release held it. Now one does.

  • strength.yml takes sprt_model, logistic by default, and hands it to each stage of the batch ladder.
  • batch.yml takes it and passes it to summarise-match.
  • Before anything is built, strength.yml checks that the model is logistic or normalized and that normalized bounds lie within 100 either side of nought. The summary refuses both as well, but only after the batch has been played.
  • batch.yml's manifest recorded model=logistic whatever the test was. It now records the model the summary used.
  • The main README and .github/workflows/README.md say that strength.yml takes the model, and that a test carried on in a second run should be given the same one. Nothing checks that.

calibrate.yml runs no sequential test and is unchanged.

How it was checked

  • The pin tests read each action's inputs at the pinned v0.6.0 tag and pass, so summarise-match at that tag does declare sprt_model.
  • The new check was run as a script on six inputs. It accepted logistic 0 10, normalized 0 5, normalized -3 100 and logistic 0 500, and refused normalized 0 150 and a misspelled normalised.
  • 380 tests and the lints pass.

🤖 Generated with Claude Code

https://claude.ai/code/session_018ccJFsmZ4jAKYmFNgNBh9u


Generated by Claude Code

v0.6.0 gave summarise-match an sprt_model input, and the reusable
workflows could not pass it until a release held it. strength.yml now
takes sprt_model and hands it through batch.yml to the summary, so a
caller of the workflow can state its bounds in normalized elo.

The model is checked before anything is built: it is logistic or
normalized, and normalized bounds lie within 100 either side of nought.
The summary refuses both too, but only after the batch has been played.
batch.yml's manifest recorded model=logistic whatever the test was, and
now records the model the summary used.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018ccJFsmZ4jAKYmFNgNBh9u
@aywrite
aywrite merged commit 6a98504 into main Sep 29, 2026
9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants