Skip to content

Default the judge to the skill's backend; warn in compare on different judges - #225

Merged
edonadei merged 10 commits into
mainfrom
claude/github-issue-213-0b03f2
Sep 28, 2026
Merged

edonadei merged 10 commits into
mainfrom
claude/github-issue-213-0b03f2

Conversation

@edonadei

@edonadei edonadei commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Closes #213. The judge now defaults to the backend of the model being evaluated (--model), instead of always claude-code. The fairness concern the issue raised is handled by a warning in compare, and a missing judge CLI now stops the run before any attempt is paid for.

1. The judge follows --model

With no --judge-model, the judge runs on the --model backend, on that CLI's default model. It follows the backend only: --model codex:<cheap model> doesn't also get a cheap judge. --judge-model still picks any backend or model. A bare --judge-model <model> still means a claude-code model, the same way a bare --model does. So a Codex-, pi- or hermes-only user needs one CLI and one flag.

Command Agent runs on Judge runs on (before) Judge runs on (now)
caliper run spec.eval.yaml claude-code claude-code claude-code
--model codex codex claude-code codex, default model
--model codex:gpt-5-codex codex, gpt-5-codex claude-code codex, default model
--model codex --judge-model pi:m codex pi, m pi, m
--model codex --judge-model opus codex claude-code, opus claude-code, opus

New ADR 0034 records the decision and supersedes ADR 0004's claude-code judge default. The --judge-model help text now states the default.

2. compare warns when two runs had different judges

Now that each engine grades itself by default, a cross-engine delta can partly come from a stricter or looser judge. RunComparison.judge_mismatch is set, and a header warning is shown, when an LLM judge graded both runs and the (judge_backend, judge_model) pair differs.

  • When it applies: "An LLM judge graded the run" is read from attempts' judge_seconds, which is only set when the judge's model call ran. judge_model isn't a reliable signal for this: it's empty for a codex, pi or hermes judge on its default model, and it's set for an assert:-only run that was given --judge-model. So assert:-only runs and runs saved before judge_seconds never trigger the warning.
  • Unnamed models: a judge on its CLI's default with no reported model is labelled <backend> (default model).

Example: a claude-code run and a codex run, each graded by its own engine:

$ caliper compare <claude-code run> <codex run>
──────────────────────── CALIPER  —  compare  —  hello ────────────────────────
    2026-09-27T10-00-00Z (claude-code) → 2026-09-27T11-00-00Z (codex)   ·   k=3
 ⚠ different judges: claude-code:claude-opus-5-5 vs codex:gpt-5-codex — part of
the delta may be a stricter or looser grader rather than the agent; re-run with
the same --judge-model for a like-for-like comparison

Passing the same --judge-model to both runs removes the warning.

3. The recorded judge model is the one that actually ran

RunMeta.judge_model is now recorded the same way as the agent's model (_recorded_model):

  • Reported over requested: it records the model the judge reported, falling back to the requested one. So --judge-model claude-code:opus records the full id opus resolved to, and compare doesn't read an alias and its id as two different judges.
  • Most common, not first: attempts finish in timing order, so if they report different judge models, the run records the most common one (ties alphabetical) and prints a Judge: warning, rather than keeping whichever finished first.

4. A missing judge CLI stops the run before any attempt

A spec with expect: now exits 2 when the judge's CLI isn't installed, instead of paying for every attempt and recording judge_error on each.

  • The check: new HarnessBackend.prompt_cli_missing(). CliHarness checks cli_path(), and ClaudeCodeHarness looks up claude on the PATH its prompt call will use (including the nvm prefix).
  • When it runs: from inside run() through a new before_attempts hook, after skill sources and mcp: servers are resolved and before any attempt is scheduled. So a bad spec keeps its own exit 1 instead of being hidden behind the missing judge.
  • Which specs: assert:-only specs never call the judge, so they aren't checked.

The message depends on how the judge was chosen:

Judge chosen by Way out offered
No --judge-model (judge follows --model) install the CLI, or pick an installed one with --model
--judge-model on a different backend than --model install the CLI, or remove --judge-model and the --model backend grades too
--judge-model on the same missing backend as --model install the CLI, or point --judge-model at an installed backend

Example: an expect: spec where --judge-model names a CLI that isn't installed:

$ caliper run hello.eval.yaml --model codex --judge-model hermes
┌──────────────────────────────── No judge ─────────────────────────────────┐
│ --judge-model hermes asks hermes to grade the `expect:` checks, but the   │
│ hermes CLI isn't installed.                                               │
│                                                                           │
│ Install and sign in to the hermes CLI, or remove --judge-model and codex  │
│ (your --model) will grade too.                                            │
└───────────────────────────────────────────────────────────────────────────┘

Docs

  • Judge default: described in README (options table, Choosing an engine, Troubleshooting), docs/backends.md, docs/spec-reference.md, docs/CONTEXT.md, both skill REFERENCE.md files and evaluate-skill/SKILL.md, calling --model "the model being evaluated".
  • Compare warning: documented in docs/results.md, with the example above.
  • Missing-CLI refusal: a README troubleshooting entry with the example above.
  • Recorded judge model: docs/backends.md and docs/CONTEXT.md now say it's the model the judge reports.
  • ADRs: new ADR 0034, and a pointer to it from ADR 0004.

Tests

  • Judge default: follows the --model backend and not its model; --judge-model overrides; a bare --judge-model model means claude-code.
  • Refusal: fires from before_attempts; is skipped for assert:-only specs; the message covers all three cases.
  • Error order: a bad skill source still exits 1 when the judge CLI is also missing (test_run_refuses_an_unfetchable_git_source).
  • Claude PATH lookup.
  • Compare warning: different models; different backends on default models; same judge; a judge that never graded.
  • Recorded judge model: the resolved model beats a requested alias; mixed judge models record the most common one and warn.

ruff format and ruff check are clean. On this Windows machine the full suite has the same 72 failures as untouched main (sandbox and filesystem harness tests) and no new ones.

…are on different judges

The judge keeps its fixed claude-code default (ADR 0004). A spec with
expect: now stops before the first attempt when the judge's CLI isn't
installed, naming --judge-model as the way out, instead of paying for
every attempt and recording judge_error on each. compare warns
(judge_mismatch) when two runs were graded by a different judge backend
or model.

Closes #213
devin-ai-integration[bot]

This comment was marked as resolved.

With no --judge-model, the judge now runs on the --model backend, on that
CLI's default model, instead of always on claude-code. A bare
--judge-model <model> still means a claude-code model, like a bare --model.
The missing-judge-CLI refusal and the compare judge_mismatch warning stay,
the latter now the guard for cross-engine comparisons where each engine
grades itself. ADR 0034 records the decision and supersedes ADR 0004's
judge default.
@edonadei edonadei changed the title Refuse an expect: run when the judge CLI is missing; warn in compare on different judges Default the judge to the skill's backend; warn in compare on different judges Sep 28, 2026
judge_model is None for a codex, pi or hermes judge on its CLI default,
and set for an assert-only run given --judge-model, so it was the wrong
signal for whether a judge ran. Read it from judge_seconds instead, and
label an unnamed model as '<backend> (default model)'.
- The missing-judge refusal now keys its advice on whether --judge-model
  was passed, so an explicit judge on the same missing CLI as --model is
  told to point --judge-model elsewhere, not to change --model.
- RunMeta.judge_model prefers the model the judge reported over the one
  requested, so compare no longer reads an alias (opus) and the id it
  resolved to as two different judges.
devin-ai-integration[bot]

This comment was marked as resolved.

…dge model

- The missing-judge-CLI refusal now runs from inside run(), after the
  environment resolved and before any attempt is scheduled, so a bad skill
  source keeps its own exit 1 instead of being masked by exit 2.
- RunMeta.judge_model is recorded like the skill model: the most common
  model the autorater reported (ties alphabetical), with a warning when
  attempts disagree, rather than whichever attempt finished first.
@edonadei
edonadei merged commit 40029d0 into main Sep 28, 2026
8 checks passed
@edonadei
edonadei deleted the claude/github-issue-213-0b03f2 branch September 28, 2026 04:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Judge engine: default to the skill's engine instead of claude-code?

1 participant