Repository navigation
Record the judge's input and script code on the attempt - #224
Merged
Merged
Conversation
A saved run kept the judge's reasoning but not what it saw or ran. The transcript was already saved, so rather than storing the prompt a second time, the run now stores the pieces that were missing: - TaskResult.expect: the task's expect: as it read at run time. - RunMeta.judge_prompt_version: a hash of the prompt templates. - AttemptRecord.autorater_script: the Python the judge wrote in script mode, whether it passed, failed or hung. The workdir it asserted on is deleted with the attempt, so nothing else can rebuild it. render_judge_prompt(expect, transcript) is now public, and the judge builds its own prompt with it, so a saved run's prompt can be rebuilt exactly while the versions match. The fields are JSON-only: the terminal report does not show them, even with --verbose. Closes #216
There was a problem hiding this comment.
Devin Review found 1 potential issue.
1 flag not posted on this PR by your GitHub settings — view it in Devin Review. (Configure)
JUDGE_PROMPT_VERSION hashed only _SYSTEM and _USER_TMPL, but _format_transcript decides the judge's input too: a new tool-result cut would change every prompt while leaving the version unchanged, so a saved run would pass the check and rebuild a different prompt. The version is now a hash of render_judge_prompt's output for a fixed probe transcript that exercises every formatting branch.
The default report still shows only the judge's reasoning. --verbose, which already opens a panel for every task, now adds the task's expect: at the top of its panel and, per attempt, any script the judge wrote. README, the evaluate-skill reference and docs/results.md describe the new --verbose rows.
This was referenced Oct 1, 2026
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Closes #216.
A saved run kept the judge's reasoning but not what it saw or what it ran. The transcript is already saved on every attempt, so this doesn't store the prompt a second time. Instead it saves the pieces that were missing, so the exact prompt can be rebuilt.
What's saved
TaskResult.expectexpect:as it read at run time, so a later spec edit doesn't lose it.RunMeta.judge_prompt_versionrender_judge_promptproduces for a fixed sample transcript, so it changes whenever the templates or the transcript formatting change, with no manual bump.nullwhen no task has anexpect:, likejudge_model.AttemptRecord.autorater_scriptrender_judge_prompt(expect, transcript)incaliper/judge/eval_judge.pyis now public, and the judge builds its own prompt with it. It accepts savedTranscriptTurns, so a saved run's prompt can be rebuilt exactly, as long as itsjudge_prompt_versionmatchesJUDGE_PROMPT_VERSION.The default terminal report doesn't show these fields.
--verboseshows theexpectand the judge script. All three are optional, so older runs still load.The
Judgeprotocol gainsprompt_version, which isNoneon the test doubles.Example output
Real output from this branch: a made-up task run through the actual runner and judge. The agent claims it wrote
notes.md, the judge chooses to check with a script, and the script fails because this example never creates the file.Saved JSON (
.caliper/results/<spec>/<timestamp>.json, trimmed to the relevant fields). The new fields arejudge_prompt_version,expectandautorater_script:{ "run": { "judge_backend": "claude-code", "judge_model": "claude-sonnet-5", "judge_prompt_version": "bbecb87f7798" }, "task_results": [ { "task_id": "write-notes", "expect": "The agent writes a notes.md file summarizing the meeting", "attempts": [ { "attempt": 1, "outcome": "task_fail", "autorater_passed": false, "autorater_reasoning": "The file's existence is checkable directly. | script: Traceback ... AssertionError: notes.md missing", "autorater_script": "import os\nassert os.path.exists('notes.md'), 'notes.md missing'" } ] } ] }For a direct verdict,
autorater_scriptisnull.Rebuilt from that JSON with
render_judge_prompt(expect, transcript). This is the prompt the judge was sent, minus the fixed instructions at the top:Terminal report. The default report is unchanged: failure panels show the judge's reasoning only.
--verboseadds the task'sexpectat the top of its panel and ajudge scriptrow per attempt:Tests
JUDGE_PROMPT_VERSION.Nonefor a direct verdict.expectand the prompt version, and records no version on an assert-only run.expectand script out,--verboseshows them, and the JSON round-trips them.ruff format --checkandruff checkare clean. The touched test files pass (119 tests). Locally on Windows the full suite has 73 failures, and the same 73 fail on unchangedmain(git sources, the user-settings layer, harnesses), so CI is the real check.Docs
docs/results.md: new Judge input fields section.docs/spec-reference.md: a pointer to it from Judging.skills/evaluate-skill/REFERENCE.md: saved-run contents and the--verbosecomment.README.md: the--verboseflag row and its description.The flip-rate / voting half of the original issue moved to #222.