Area is a dimension, so a per-m2 rate is not a per-nothing rate - #36
Merged
Merged
Conversation
`kWh/yr` and `kWh per m2 per year` are different quantities and the scorer was reconciling them. Ten answers giving a whole-building annual total were scored against a per-square-metre truth, the worst off by a factor of 200, and counted as scoreable-but-wrong instead of unscoreable. Fixes #35. The cause was a missing dimension table, not a missing rule. Area had none, so a per-m2 denominator fell through to the catch-all `other:m2` and was compared as a STRING -- which put the pair in front of the containment rule, where `{year}` is a subset of `{m2, year}` and the two matched. `reconcile`'s own docstring already said this must not happen: "per dwelling per year vs per m2 per year is two different quantities that merely share the word year". So the fix is one table and one line in the denominator scan. Area is a dimension; saying so lets the existing den_dim comparison refuse the pair, and it makes `per hectare` and `per m2` convert properly instead of being two unrelated `other:` strings. Two tempting fixes are deliberately NOT here, and verify/test_denominator_basis.py pins both so they cannot be introduced quietly: - A guard keyed on a hand-written list of "dimension words". It reproduces the same ten-answer change, and it is a second list to maintain that fails silently and open -- the same shape as the duplicated currency table in #34. - A guard comparing every recognised unit in the denominator. That one refuses `kgCO2e/km T&D loss` against `kg CO2e per km`, because `T&D` cleans to a bare `t` and `t` is tonnes. It killed two correct answers in a draft of this fix. Corpus effect, measured against main: 10 answers refused, 0 newly scored, 0 scored differently. scoreable <=10% <=50% >50% off right src/wrong no claude-opus-5 315 -> 310 45.7 -> 46.5 77.5 -> 78.7 22.5 -> 21.3 54.1 -> 53.1 gemini-3.6-flash 312 -> 307 42.0 -> 42.7 77.9 -> 78.8 22.1 -> 21.2 58.8 -> 58.0 Unscoreable rises, deliberately: claude 110 -> 115, gemini 88 -> 93. Paired, Claude unaided: within 10% 37.1% -> 37.7%, off by >50% 25.8% -> 24.6% (n 62 -> 61). gpt-5.5, grok-4.6 and gemini-3.1-pro-preview do not move, and results/absolute.json does not move at all -- all ten were wrong either way, so no `correct` count changes. The corpus-size correlation moves with the scored set: r -0.2756 -> -0.2905, reported as -0.29 rather than -0.28, p 0.19 -> 0.17. The sensitivity sweep still holds across all 48 specifications, negative throughout, -0.11 to -0.43. Published figures updated in FINDINGS.md, README.md and paper/main.tex, including the abstract's accuracy range, its right-source/wrong-number range and its paired sentence. The bug log is now eight, and says plainly that this correction and the last one both RAISED the headline -- which is the direction that deserves the most scrutiny, and not the reason to make them. The coverage paragraph now records a step that made the unscoreable pile bigger rather than smaller. Three figures the bug log narrates are allowlisted as the history they now are, beside the existing 46.7% and 41.9%. Still open, documented in #35 and pinned by a test: a bare `per km` reconciles with a `per tonne-km` truth, because the denominator scan stops at `km` before reaching `tonne` so both report den_dim "length". No answer in the corpus is affected. The `T&D` case above is why it is not a one-line fix. NOT DONE, needs the merge first: the live guides still show the old figures and `r = -0.28`. The online gate fails on them, which is the correct signal. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NRueWxopDXHoWY2dPvsmLG
This was referenced Sep 20, 2026
jeremiahsay
added a commit
to Navneshwar/greencalculus-benchmark
that referenced
this pull request
Sep 20, 2026
Rebased onto main after greencalculus#36 (the area/denominator fix), so these are the figures the alias produces against final scorer behaviour rather than the ones from before that landed. The alias does what greencalculus#18 asked: four freight answers recover, all None -> a value, no already-scored answer changes. claude-opus-5 scoreable 310 -> 314 within 10% 46.5% -> 46.2% unscoreable 115 -> 111 absolute correct 144 -> 145 of all 467 30.8% -> 31.0% paired, unaided within 10% 37.7% -> 38.1% (n 61 -> 63), >50% off 24.6% -> 23.8% within 50% and >50%-off are unchanged at this rounding, and no other model moves. The headline goes DOWN, which is the honest direction: four answers that were being thrown away are now scored and most of them are wrong. Figures updated in FINDINGS.md, README.md and paper/main.tex, including the abstract's paired sentence and the results table. The coverage paragraph records this as one more step in the sequence -- 115 to 111, 46.5% to 46.2% -- and 46.5% joins the allowlist as the history it now is, beside 45.7%, 42.0% and 41.9%. No current figure is allowlisted. verify/test_tkm_alias.py covers tkm in both directions, the existing spellings, scale carry-through, and that a unit merely containing those letters is untouched. One test documents an open hole rather than asserting a fix: a bare `per km` still reconciles with a tonne-km truth, because the denominator scan stops at km before reaching tonne. That predates this alias and greencalculus#36 did not close it -- see issue greencalculus#35 -- and no answer in the corpus is affected, though the alias does widen it latently since `per tkm` used to be refused as other:tkm. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01NRueWxopDXHoWY2dPvsmLG
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #35.
1 · The failure
kWh/yrandkWh per m2 per yearare different quantities, and the scorer was reconciling them.Ten answers giving a whole-building annual total were scored against a per-square-metre truth — the worst 20,500 against 102.1, a factor of 200 — and counted as scoreable-but-wrong rather than unscoreable. Five
claude-opus-5, fivegemini-3.6-flash, allbuilding_energy.gbr.floorband.*.The route is the
both_othercontainment rule:{year}is a subset of{m2, year}, so it passed. That rule exists for a real case — "per year vs per person per year is the same basis stated at different length" — and the same comment states the case it was supposed to exclude:That held when both sides were qualified and failed when one side was bare, because a subset test cannot distinguish a qualifier (
person,dwelling,room) from a measured dimension (m2).The reason to fix this is that the scorer was comparing quantities that are not comparable. It also raises two headline figures. That is documented in full in §5 and is not the motive.
2 · Why area gets a table, rather than a list of dimension words
Area had no dimension table. So a per-m² denominator fell through to the catch-all
other:m2and was compared as a string, which is what put the pair in front of the containment rule at all.Area is a dimension. Saying so is the whole fix: the existing
den_dimcomparison then refuses the pair with no new rule, and as a side effectper hectareandper m2convert properly instead of being two unrelatedother:strings. The table is grounded in the corpus — 19 truth units are per-m², 5 per-hectare, 1 per-acre.The alternative was a guard that refuses when the differing denominator tokens intersect a hand-written set of dimension words. It reproduces exactly the same ten-answer change, so it is not wrong on this data. It was rejected because it is a second list to maintain that fails silently and open: the next area or volume unit nobody thought to list reconciles as though it matched. That is the same shape as the duplicated currency table in #34, and this module already has an authority for what counts as a measured dimension — the tables at the top of the file.
3 · Why comparing every recognised denominator unit was rejected
The stronger rule — collect every recognised unit in the denominator and require the sets of dimensions to match — is more principled on paper. It catches this bug and the related
per kmvsper tonne-kmcase in one stroke.It is unsafe with the current tokenisation, because the denominator text includes trailing prose:
T&Dloses its ampersand and leaves a baret.tis tonnes. So the denominator reads as mass and length and the answer is refused against akg CO2e per kmtruth.4 · The regression that caught it — three answers, not two
I built that version first. The corpus diff came back 13 refusals instead of 10, which is the only reason I found it:
ev_charging.car.segment.mpv.phev.td_loss_ukkgCO2e/km T&D loss for an MPV PHEVper length+mass vs lengthev_charging.car.size.large.phev.td_loss_ukkgCO2e/km T&D loss for a large PHEV carper length+mass vs lengthtransmission_distribution.gbr.electricitykgCO2e/kWh for UK electricity T&D lossesper energy+mass vs energyAll three were being scored correctly — as wrong answers, 39%, 126% and 46% off. Refusing them would have moved three wrong answers out of the denominator, flattering the result further on top of the improvement this fix already produces. That is the worst direction for a scorer bug to fail in, and it is why "compare every recognised unit" is not in this PR.
Correction to the commit message: it says "two correct answers". It is three, and they were correctly-scored wrong answers. The table above is right.
verify/test_denominator_basis.pypins both rejected approaches so neither can return quietly —test_prose_after_the_unit_does_not_invent_a_dimensionis the one that would fail.5 · Exact benchmark impact
10 answers refused, 0 newly scored, 0 scored differently.
gpt-5.5,grok-4.6andgemini-3.1-pro-previewdo not move. Paired, Claude unaided: within 10% 37.1% → 37.7%, off by >50% 25.8% → 24.6% (n 62 → 61).results/absolute.jsondoes not move at all — all ten were wrong either way, so nocorrectcount changes.The corpus-size correlation moves with the scored set:
r−0.2756 → −0.2905, reported as −0.29 rather than −0.28,p0.19 → 0.17. The sensitivity sweep still holds across all 48 specifications, negative throughout, −0.11 to −0.43 — so the claim is unchanged, but the figure in the abstract paragraph and on the live by-category page is not.6 · Unscoreable goes up, and that is the point
Every previous scorer fix in the log made this number smaller by reconciling units the scorer had failed to understand. This one makes it bigger, by refusing ten pairs that were never comparable in the first place. The benchmark's stated rule is that an unreconcilable pair leaves the denominator rather than counting as a wrong answer, and these ten were being counted as wrong answers to a question the model had not been asked.
FINDINGS.md's coverage paragraph now records that step explicitly, so the sequence reads 137 → 110 → 115 rather than stopping at the last improvement.7 · Which figures are historical, and which are current
Every number in the tables above is current and reproduces from
results/*.json. Nothing current is allowlisted.Three figures appear in
FINDINGS.mdonly as narration inside the bug log and the coverage sequence, and no committed script produces them any more. Those are allowlisted underfindings, beside the entries that were already there:45.7%42.0%41.9%46.7%,46.9%,47%,35.6%Each reason names the figure it has been superseded by.
20,500was reworded out of the bug-log entry rather than allowlisted, because a raw answer value that no script produces should not be in published prose at all.The bug log is now eight, and says plainly that bugs 7 and 8 both raised the headline, that this is the direction deserving the most scrutiny, and that the reason to make the corrections is incomparable quantities rather than better numbers.
Verification
verify/figures.py --check-freshverify/check_claims.py --offlinepaper/corpus_size.py --sensitivityverify/dataset.py/verify/licences.pyverify/test_denominator_basis.pyverify/test_claims_gate.py/test_currency_mismatch.pyThe new test goes red without the fix:
AssertionError: kWh/yr reconciled with 'kWh per m2 per year' as 20500.0.claims-are-backedfails only on its Published-pages step. The live guides still show the old figures andr = −0.28; that is the correct signal, and the live update follows the merge.Still open
A bare
per kmreconciles with aper tonne-kmtruth — the denominator scan stops atkmbefore reachingtonne, so both reportden_dim"length". No answer in the corpus is affected. Tracked in #35 and pinned bytest_a_bare_per_km_against_a_tonne_km_truth_is_a_KNOWN_OPEN_HOLE, with §3 above explaining why it is not a one-liner.Sequencing
Not stacked on #29 or #30. Both rebase onto this once it merges and regenerate from actual scorer output — #29 moves Claude again (four freight answers recovered), #30 moves Claude, GPT and Grok.
🤖 Generated with Claude Code
https://claude.ai/code/session_01NRueWxopDXHoWY2dPvsmLG