Repository navigation
Default the judge to the skill's backend; warn in compare on different judges - #225
Merged
Merged
Conversation
…are on different judges The judge keeps its fixed claude-code default (ADR 0004). A spec with expect: now stops before the first attempt when the judge's CLI isn't installed, naming --judge-model as the way out, instead of paying for every attempt and recording judge_error on each. compare warns (judge_mismatch) when two runs were graded by a different judge backend or model. Closes #213
With no --judge-model, the judge now runs on the --model backend, on that CLI's default model, instead of always on claude-code. A bare --judge-model <model> still means a claude-code model, like a bare --model. The missing-judge-CLI refusal and the compare judge_mismatch warning stay, the latter now the guard for cross-engine comparisons where each engine grades itself. ADR 0034 records the decision and supersedes ADR 0004's judge default.
judge_model is None for a codex, pi or hermes judge on its CLI default, and set for an assert-only run given --judge-model, so it was the wrong signal for whether a judge ran. Read it from judge_seconds instead, and label an unnamed model as '<backend> (default model)'.
- The missing-judge refusal now keys its advice on whether --judge-model was passed, so an explicit judge on the same missing CLI as --model is told to point --judge-model elsewhere, not to change --model. - RunMeta.judge_model prefers the model the judge reported over the one requested, so compare no longer reads an alias (opus) and the id it resolved to as two different judges.
…13-0b03f2 # Conflicts: # caliper/runner.py
…dge model - The missing-judge-CLI refusal now runs from inside run(), after the environment resolved and before any attempt is scheduled, so a bad skill source keeps its own exit 1 instead of being masked by exit 2. - RunMeta.judge_model is recorded like the skill model: the most common model the autorater reported (ties alphabetical), with a warning when attempts disagree, rather than whichever attempt finished first.
This was referenced Oct 1, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #213. The judge now defaults to the backend of the model being evaluated (
--model), instead of alwaysclaude-code. The fairness concern the issue raised is handled by a warning incompare, and a missing judge CLI now stops the run before any attempt is paid for.1. The judge follows
--modelWith no
--judge-model, the judge runs on the--modelbackend, on that CLI's default model. It follows the backend only:--model codex:<cheap model>doesn't also get a cheap judge.--judge-modelstill picks any backend or model. A bare--judge-model <model>still means aclaude-codemodel, the same way a bare--modeldoes. So a Codex-, pi- or hermes-only user needs one CLI and one flag.caliper run spec.eval.yaml--model codex--model codex:gpt-5-codex--model codex --judge-model pi:mmm--model codex --judge-model opusopusopusNew ADR 0034 records the decision and supersedes ADR 0004's
claude-codejudge default. The--judge-modelhelp text now states the default.2.
comparewarns when two runs had different judgesNow that each engine grades itself by default, a cross-engine delta can partly come from a stricter or looser judge.
RunComparison.judge_mismatchis set, and a header warning is shown, when an LLM judge graded both runs and the(judge_backend, judge_model)pair differs.judge_seconds, which is only set when the judge's model call ran.judge_modelisn't a reliable signal for this: it's empty for a codex, pi or hermes judge on its default model, and it's set for anassert:-only run that was given--judge-model. Soassert:-only runs and runs saved beforejudge_secondsnever trigger the warning.<backend> (default model).Example: a claude-code run and a codex run, each graded by its own engine:
Passing the same
--judge-modelto both runs removes the warning.3. The recorded judge model is the one that actually ran
RunMeta.judge_modelis now recorded the same way as the agent's model (_recorded_model):--judge-model claude-code:opusrecords the full idopusresolved to, andcomparedoesn't read an alias and its id as two different judges.Judge:warning, rather than keeping whichever finished first.4. A missing judge CLI stops the run before any attempt
A spec with
expect:now exits2when the judge's CLI isn't installed, instead of paying for every attempt and recordingjudge_erroron each.HarnessBackend.prompt_cli_missing().CliHarnesscheckscli_path(), andClaudeCodeHarnesslooks upclaudeon the PATH its prompt call will use (including the nvm prefix).run()through a newbefore_attemptshook, after skill sources andmcp:servers are resolved and before any attempt is scheduled. So a bad spec keeps its own exit1instead of being hidden behind the missing judge.assert:-only specs never call the judge, so they aren't checked.The message depends on how the judge was chosen:
--judge-model(judge follows--model)--model--judge-modelon a different backend than--model--judge-modeland the--modelbackend grades too--judge-modelon the same missing backend as--model--judge-modelat an installed backendExample: an
expect:spec where--judge-modelnames a CLI that isn't installed:Docs
docs/backends.md,docs/spec-reference.md,docs/CONTEXT.md, both skillREFERENCE.mdfiles andevaluate-skill/SKILL.md, calling--model"the model being evaluated".docs/results.md, with the example above.docs/backends.mdanddocs/CONTEXT.mdnow say it's the model the judge reports.Tests
--modelbackend and not its model;--judge-modeloverrides; a bare--judge-modelmodel means claude-code.before_attempts; is skipped forassert:-only specs; the message covers all three cases.1when the judge CLI is also missing (test_run_refuses_an_unfetchable_git_source).ruff formatandruff checkare clean. On this Windows machine the full suite has the same 72 failures as untouchedmain(sandbox and filesystem harness tests) and no new ones.