Skip to content

Record the judge's input and script code on the attempt - #224

Merged
edonadei merged 5 commits into
mainfrom
claude/caliper-issue-216-0236ed
Sep 28, 2026
Merged

edonadei merged 5 commits into
mainfrom
claude/caliper-issue-216-0236ed

Conversation

@edonadei

@edonadei edonadei commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Closes #216.

A saved run kept the judge's reasoning but not what it saw or what it ran. The transcript is already saved on every attempt, so this doesn't store the prompt a second time. Instead it saves the pieces that were missing, so the exact prompt can be rebuilt.

What's saved

Field What it is
TaskResult.expect The task's expect: as it read at run time, so a later spec edit doesn't lose it.
RunMeta.judge_prompt_version A 12-char hash of what render_judge_prompt produces for a fixed sample transcript, so it changes whenever the templates or the transcript formatting change, with no manual bump. null when no task has an expect:, like judge_model.
AttemptRecord.autorater_script The Python the judge wrote in script mode, kept whether it passed, failed or timed out. The workdir it asserted on is deleted with the attempt, so nothing else can rebuild it.

render_judge_prompt(expect, transcript) in caliper/judge/eval_judge.py is now public, and the judge builds its own prompt with it. It accepts saved TranscriptTurns, so a saved run's prompt can be rebuilt exactly, as long as its judge_prompt_version matches JUDGE_PROMPT_VERSION.

The default terminal report doesn't show these fields. --verbose shows the expect and the judge script. All three are optional, so older runs still load.

The Judge protocol gains prompt_version, which is None on the test doubles.

Example output

Real output from this branch: a made-up task run through the actual runner and judge. The agent claims it wrote notes.md, the judge chooses to check with a script, and the script fails because this example never creates the file.

Saved JSON (.caliper/results/<spec>/<timestamp>.json, trimmed to the relevant fields). The new fields are judge_prompt_version, expect and autorater_script:

{
  "run": {
    "judge_backend": "claude-code",
    "judge_model": "claude-sonnet-5",
    "judge_prompt_version": "bbecb87f7798"
  },
  "task_results": [
    {
      "task_id": "write-notes",
      "expect": "The agent writes a notes.md file summarizing the meeting",
      "attempts": [
        {
          "attempt": 1,
          "outcome": "task_fail",
          "autorater_passed": false,
          "autorater_reasoning": "The file's existence is checkable directly. | script: Traceback ... AssertionError: notes.md missing",
          "autorater_script": "import os\nassert os.path.exists('notes.md'), 'notes.md missing'"
        }
      ]
    }
  ]
}

For a direct verdict, autorater_script is null.

Rebuilt from that JSON with render_judge_prompt(expect, transcript). This is the prompt the judge was sent, minus the fixed instructions at the top:

<expectation>
The agent writes a notes.md file summarizing the meeting
</expectation>

<transcript>
[assistant] I'll write the notes.
[tool_use: Write] {"file_path": "notes.md", "content": "# Notes"}
[tool_result] File written
[assistant] Done: notes.md is ready.
</transcript>

Evaluate the transcript. Respond with JSON.

Terminal report. The default report is unchanged: failure panels show the judge's reasoning only. --verbose adds the task's expect at the top of its panel and a judge script row per attempt:

┌─ Writes meeting notes  ✗ FAIL ────────────────────────────────────────┐
│ expect            The agent writes a notes.md file summarizing the    │
│                   meeting                                             │
│ ✗ Attempt 1       0.1s                                                │
│     output        Done: notes.md is ready.                            │
│     judge         The file's existence is checkable directly. |       │
│                   script: Traceback (most recent call last): ...      │
│                   AssertionError: notes.md missing                    │
│     judge script  import os                                           │
│                   assert os.path.exists('notes.md'), 'notes.md        │
│                   missing'                                            │
└───────────────────────────────────────────────────────────────────────┘

Tests

  • A prompt rebuilt from a saved transcript matches the one sent to the judge, byte for byte (including a tool result past the 2000-char cut).
  • Changing the transcript formatting changes JUDGE_PROMPT_VERSION.
  • The script is kept on both pass and fail, and is None for a direct verdict.
  • The runner records expect and the prompt version, and records no version on an assert-only run.
  • The default report leaves the expect and script out, --verbose shows them, and the JSON round-trips them.

ruff format --check and ruff check are clean. The touched test files pass (119 tests). Locally on Windows the full suite has 73 failures, and the same 73 fail on unchanged main (git sources, the user-settings layer, harnesses), so CI is the real check.

Docs

  • docs/results.md: new Judge input fields section.
  • docs/spec-reference.md: a pointer to it from Judging.
  • skills/evaluate-skill/REFERENCE.md: saved-run contents and the --verbose comment.
  • README.md: the --verbose flag row and its description.

The flip-rate / voting half of the original issue moved to #222.

A saved run kept the judge's reasoning but not what it saw or ran. The
transcript was already saved, so rather than storing the prompt a second
time, the run now stores the pieces that were missing:

- TaskResult.expect: the task's expect: as it read at run time.
- RunMeta.judge_prompt_version: a hash of the prompt templates.
- AttemptRecord.autorater_script: the Python the judge wrote in script
  mode, whether it passed, failed or hung. The workdir it asserted on
  is deleted with the attempt, so nothing else can rebuild it.

render_judge_prompt(expect, transcript) is now public, and the judge
builds its own prompt with it, so a saved run's prompt can be rebuilt
exactly while the versions match. The fields are JSON-only: the
terminal report does not show them, even with --verbose.

Closes #216

@devin-ai-integration devin-ai-integration Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Devin Review found 1 potential issue.

1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)

Devin Review

Comment thread caliper/judge/eval_judge.py Outdated
JUDGE_PROMPT_VERSION hashed only _SYSTEM and _USER_TMPL, but
_format_transcript decides the judge's input too: a new tool-result cut
would change every prompt while leaving the version unchanged, so a saved
run would pass the check and rebuild a different prompt. The version is
now a hash of render_judge_prompt's output for a fixed probe transcript
that exercises every formatting branch.
The default report still shows only the judge's reasoning. --verbose,
which already opens a panel for every task, now adds the task's expect:
at the top of its panel and, per attempt, any script the judge wrote.
README, the evaluate-skill reference and docs/results.md describe the
new --verbose rows.
@edonadei
edonadei merged commit b452a8b into main Sep 28, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Record the judge's input and script code on the attempt

1 participant