chore(tests): move the claude-code evals to Opus 5.5 - #3514
Conversation
Picks up the command_executed fix that searches the whole command instead of only its first 2000 characters, plus the Opus 5.5 rate card. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Measured each skill's full positive set on claude-opus-5-5 at max_turns: 1, the gate's own method, for the move of the claude-code evals from Sonnet 5 to Opus 5.5. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
|
Claude finished @bai-uipath's task in 1m 57s —— View job SummaryMoves the claude-code eval suite from Sonnet 5 to Opus 5.5 by bumping coder-eval to 0.12.7 and re-baselining the activation gate's per-skill recall thresholds on the new model. No skill content changes. Change-by-Change Review1.
|
| Skill | Sonnet 5 | Opus 5.5 | Delta |
|---|---|---|---|
| uipath-rpa | 100 | 90 | -10 |
| uipath-review | 95 | 90 | -5 |
| uipath-maestro-bpmn | 95 | 100 | +5 |
| uipath-admin | 95 | 100 | +5 |
| uipath-planner | 60 | 95 | +35 |
| uipath-solution | 90 | 100 | +10 |
| uipath-agents | 90 | 100 | +10 |
| uipath-maestro-case | 90 | 95 | +5 |
| uipath-api-workflow | 90 | 100 | +10 |
| uipath-functions | 95 | 100 | +5 |
Net positive: most skills improved, two dropped within the DROP_PP = 10 tolerance band. The large uipath-planner jump (60→95) is consistent with the PR description's note that the old 60% baseline reflected Sonnet 5's poor performance on that skill's ambiguous positives. The long explanatory comment block for uipath-planner is correctly removed — it was model-specific context that no longer applies.
DROP_PP stays at 10, which remains reasonable as the safety margin.
What's Missing
Nothing identified. The PR description includes clear merge-time instructions for updating CLAUDE_CODE_MODEL in both the repo variable and the ADO variable group — those are operational steps that correctly live outside the code diff.
Area Ratings
| Area | Status | Notes |
|---|---|---|
| Frontmatter | N/A | No skill frontmatter changed |
| E2E Tests | OK | Test infrastructure change only |
| Skill Body | N/A | No skill content changed |
| References & Assets | N/A | No references changed |
| Repo Hygiene | OK | No secrets, no stray changes, well-scoped |
Issues for Manual Review
- Merge ordering: The PR correctly flags that coder-eval 0.12.7 must exist before merge. Verify that fix(criteria): search the whole command in command_executed and bump harness SDKs for Opus 5.5 coder_eval#196 has been merged and the 0.12.7 release is published before merging this PR.
- Repo variable flip: At merge time,
CLAUDE_CODE_MODELmust be changed fromclaude-sonnet-5toclaude-opus-5-5(repo variable) andeu.anthropic.claude-opus-5-5(ADO). If done before merge, the gate will compare Sonnet runs against Opus baselines.
Conclusion
Clean, well-scoped infrastructure PR. No code issues found. The baseline re-measurements are internally consistent, the comment updates match the new model, and the merge-order dependency is clearly documented. Approve.
|
Moves the claude-code evals (the nightly and this repo's PR gates) from Sonnet 5 to Opus 5.5.
Why
Opus 5.5 full-suite run (ADO build 13505870, on coder-eval 0.12.4), compared with each task's latest Sonnet 5 nightly result from 2026-09-17 to 2026-09-23:
The only skill that regresses is uipath-agents: 95% → 76% as graded by 0.12.4, 99% with the grader fix. Most of its failures were the long-script grading bug.
What changes
1. coder-eval 0.12.4 → 0.12.7
command_executedsearches the whole command, not the first 2,000 charactersuip agent refresh && uip agent validatewas graded as never running them. 35 Opus 5.5 tasks flip fail → pass on re-grade; the Sonnet 5 nightly doesn't change.hello_dateon Opus 5.5 reports $0.1275934, exactly its tokens at $4/$20. The old CLI billed Opus 5.5 at Opus 5 rates (1.65x high) and Sonnet 5 at Sonnet 4 rates.max_turnsin the harness--max-turnson some routes (221 turns against a cap of 75). The turn now ends at the cap.system_one_judge, post-failure grading forrun_commandThe results above ran on 0.12.4, so the first Opus 5.5 nightly changes the model and the Claude Code version at once. A move in pass rate can come from either.
2. Activation gate re-baselined on Opus 5.5
Baselines are per model. These were measured on 2026-09-23 with the gate's own method: every skill's full positive set at
max_turns: 1. The gate fails a skill whose recall drops more than 10 points below its baseline.Every other skill keeps its baseline.
Validation
coder-eval 0.12.7 is published on PyPI and GHCR (
coder-eval-agent:0.12.7).PR CI at this head (
5a466b60e): all checks pass. smoke-skills installed 0.12.7, pulled the 0.12.7 image and passed 39/40 (97.5%, gate 95%). It runs on codex (gpt-5.6-luna), so it checks the install and image, not Opus 5.5.Nightly pipeline dry runs at this head, with
coderEvalVersion=pinned(so 0.12.7),dryRun=trueandrunWindows=true. That covers the Linux and Windows slices, the merge and the upload to the isolatedruns-dryruncontainer. Each dry run takes one smoke task per skill and one activation case per skill, so it checks the pipeline, not the pass rate at scale.The gate needs zero errors. Every error traces to something outside this PR:
troubleshoot-smoke-manifest-commandsuip tools installgets "No compatible version … for the CLI's 1.204.x line". It has errored on every run that included it since 2026-09-23, on antigravity, delegate-sdk, codex and claude-code, including the Opus 5.5 sweep above.migrator-migrate-with-fixes-e2e(Windows)llm_judgegoes through the experiments'route: litellm, and coder_eval_uipath'sdaily-windows.ps1installscoder-eval[dev,uipath]without thelitellmextra. The task landed on main today (#3501), so tonight's nightly hits it on any pin.migrator-migrate-classic-windows(codex only)The dry-run summary also reports "Activation: 0% avg recall". That is the rollup's 20-prompt minimum per skill against one sampled case, and the 0.12.4 and 0.12.1 dry runs show the same 0%.
At merge time
Flip these at merge, not before. Otherwise the gate compares Sonnet runs against Opus baselines, and uipath-planner fails (Sonnet recalls about 60 against a floor of 85).
CLAUDE_CODE_MODELclaude-opus-5-5(fromclaude-sonnet-5). The us-east-2 Bedrock account already servesus.anthropic.claude-opus-5-5.coder-eval-athena-config(nightly)CLAUDE_CODE_MODELeu.anthropic.claude-opus-5-5BEDROCK_MODELMaturity is tracked per (harness, model), so the first five Opus 5.5 nightlies run the full suite before skipping resumes.
🤖 Generated with Claude Code