Skip to content

R13: position-effect pairing, one character measure, un-softened zh README - #33

Merged
LeoLin990405 merged 1 commit into
mainfrom
r13/cleanup
Aug 3, 2026
Merged

LeoLin990405 merged 1 commit into
mainfrom
r13/cleanup

Conversation

@LeoLin990405

Copy link
Copy Markdown
Owner

The three non-blocking items from the PR #32 pre-merge review.

Position effect dropped every judge after the first (P2). forward/reversed
were keyed by regime alone and kept the first entry, so with --judges 2 only
one provider's swap movement survived. Reproduced: codex scores both arms 7 in
both passes while a second judge swings 8 points on the same civ, and
meanAbsDelta came back 0.

That figure is the noise floor a paired score gap has to clear in the E1/E2
analysis. Under-reporting it marks order-sensitive results as resolved, which
is the opposite of what the position check exists for. Passes are now paired
within a provider; perProvider is reported; maxAbsDelta is added because the
floor a gap must clear is the worst judge's movement, not the average across
judges. Revert-verified — collapsing providers into one bucket reddens it.

72,828 vs 72,830 (P3). Both correct for the same transcript: JavaScript
String.length counts UTF-16 code units, Python len() counts code points, and
the transcript contains two astral characters (🔄). The docs now use the figure
the code reports, and E1-results.md states which measure that is. Mixing the
two in one document was the defect, not either number.

The Chinese README softened §9.3 (P3). It dropped the findings table and kept
only "found 5 real errors" — four of which are still unfixed, and a reader of the
Chinese page had no way to know which. The table is restored with the same four
pending rows the English page carries.

CI: 520 backend + 25 frontend + 57 regimes + end-to-end smoke.

…ten zh README

Three items left open by the pre-merge review.

P2 — the position effect dropped every judge after the first. forward/reversed
were keyed by regime alone and kept the first entry, so with two judges only
one contributed. Reproduced: codex scores both arms 7 in both passes while a
second judge swings 8 points on the same civ, and meanAbsDelta came back 0.
That number is the noise floor a paired score gap must clear in the E1/E2
analysis, so under-reporting it marks order-sensitive results as resolved —
exactly backwards. Passes are now paired within a provider, perProvider
breakdown is reported, and maxAbsDelta is added because the floor a gap must
clear is the worst judge's movement, not the average across judges.
Revert-verified.

P3 — 72,828 and 72,830 were both in the docs for the same transcript. Both are
correct: JavaScript String.length counts UTF-16 code units, Python len() counts
code points, and the transcript contains two astral characters (🔄). The docs
now use the figure the code reports throughout, and E1-results states which
measure that is, because mixing the two in one document is the actual defect.

P3 — the Chinese README dropped the §9.3 table entirely and kept only "found 5
real errors". Four of those five are still unfixed, and a reader of the Chinese
page could not know which. The table is restored with the same four ⚠️ pending
rows the English page carries. The rule this violated is one I set: the Chinese
page is a translation of the same claims, not a softened summary.

CI: 519 backend + 25 frontend + 57 regimes + smoke.
@LeoLin990405
LeoLin990405 merged commit 3369ce1 into main Aug 3, 2026
3 checks passed
@LeoLin990405
LeoLin990405 deleted the r13/cleanup branch August 3, 2026 08:09
LeoLin990405 pushed a commit that referenced this pull request Aug 6, 2026
feat: install.sh 跨平台支持 (macOS/CentOS/RHEL/Alpine)
LeoLin990405 added a commit that referenced this pull request Aug 6, 2026
…ten zh README (#33)

Three items left open by the pre-merge review.

P2 — the position effect dropped every judge after the first. forward/reversed
were keyed by regime alone and kept the first entry, so with two judges only
one contributed. Reproduced: codex scores both arms 7 in both passes while a
second judge swings 8 points on the same civ, and meanAbsDelta came back 0.
That number is the noise floor a paired score gap must clear in the E1/E2
analysis, so under-reporting it marks order-sensitive results as resolved —
exactly backwards. Passes are now paired within a provider, perProvider
breakdown is reported, and maxAbsDelta is added because the floor a gap must
clear is the worst judge's movement, not the average across judges.
Revert-verified.

P3 — 72,828 and 72,830 were both in the docs for the same transcript. Both are
correct: JavaScript String.length counts UTF-16 code units, Python len() counts
code points, and the transcript contains two astral characters (🔄). The docs
now use the figure the code reports throughout, and E1-results states which
measure that is, because mixing the two in one document is the actual defect.

P3 — the Chinese README dropped the §9.3 table entirely and kept only "found 5
real errors". Four of those five are still unfixed, and a reader of the Chinese
page could not know which. The table is restored with the same four ⚠️ pending
rows the English page carries. The rule this violated is one I set: the Chinese
page is a translation of the same claims, not a softened summary.

CI: 519 backend + 25 frontend + 57 regimes + smoke.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant