Skip to content

chore(tests): move the claude-code evals to Opus 5.5 - #3514

Merged
bai-uipath merged 3 commits into
mainfrom
bai/nightly-opus-5-5
Sep 25, 2026
Merged

bai-uipath merged 3 commits into
mainfrom
bai/nightly-opus-5-5

Conversation

@bai-uipath

@bai-uipath bai-uipath commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Moves the claude-code evals (the nightly and this repo's PR gates) from Sonnet 5 to Opus 5.5.

Why

Opus 5.5 full-suite run (ADO build 13505870, on coder-eval 0.12.4), compared with each task's latest Sonnet 5 nightly result from 2026-09-17 to 2026-09-23:

Sonnet 5 Opus 5.5
Pass rate, with the coder_eval#196 grader fix 95.2% 96.5%
Pass rate, as graded by 0.12.4 95.2% 94.0%
Cost, full suite at list $1,196 $761 (0.64x)
Mean wall time per task 290s 121s (0.42x)
Mean turns per task 26.2 14.5 (0.55x)

The only skill that regresses is uipath-agents: 95% → 76% as graded by 0.12.4, 99% with the grader fix. Most of its failures were the long-script grading bug.

What changes

1. coder-eval 0.12.4 → 0.12.7

Change Example Source
command_executed searches the whole command, not the first 2,000 characters A heredoc script ending in uip agent refresh && uip agent validate was graded as never running them. 35 Opus 5.5 tasks flip fail → pass on re-grade; the Sonnet 5 nightly doesn't change. coder_eval#196
Claude Code 2.1.216 → 2.1.281, other harness SDKs to current releases claude-code costs now come out at list. hello_date on Opus 5.5 reports $0.1275934, exactly its tokens at $4/$20. The old CLI billed Opus 5.5 at Opus 5 rates (1.65x high) and Sonnet 5 at Sonnet 4 rates. coder_eval#196
claude-code enforces max_turns in the harness The CLI skipped --max-turns on some routes (221 turns against a cap of 75). The turn now ends at the cap. coder_eval#194
Rate card adds Opus 5.5 and GPT-6 coder_eval#195
New opt-in criteria: system_one_judge, post-failure grading for run_command No effect unless a task uses them. coder_eval#192, #197

The results above ran on 0.12.4, so the first Opus 5.5 nightly changes the model and the Claude Code version at once. A move in pass rate can come from either.

2. Activation gate re-baselined on Opus 5.5

Baselines are per model. These were measured on 2026-09-23 with the gate's own method: every skill's full positive set at max_turns: 1. The gate fails a skill whose recall drops more than 10 points below its baseline.

Skill Was Now Gate fails below
uipath-planner 60 95 85
uipath-rpa 100 90 80
uipath-review 95 90 80
uipath-maestro-case 90 95 85
uipath-maestro-bpmn 95 100 90
uipath-admin 95 100 90
uipath-functions 95 100 90
uipath-solution 90 100 90
uipath-agents 90 100 90
uipath-api-workflow 90 100 90

Every other skill keeps its baseline.

Validation

coder-eval 0.12.7 is published on PyPI and GHCR (coder-eval-agent:0.12.7).

PR CI at this head (5a466b60e): all checks pass. smoke-skills installed 0.12.7, pulled the 0.12.7 image and passed 39/40 (97.5%, gate 95%). It runs on codex (gpt-5.6-luna), so it checks the install and image, not Opus 5.5.

Nightly pipeline dry runs at this head, with coderEvalVersion=pinned (so 0.12.7), dryRun=true and runWindows=true. That covers the Linux and Windows slices, the merge and the upload to the isolated runs-dryrun container. Each dry run takes one smoke task per skill and one activation case per skill, so it checks the pipeline, not the pass rate at scale.

claude-code, Opus 5.5 codex, nightly default model
ADO build 13537467 13537468
Versions coder-eval 0.12.7, Claude Code 2.1.281, uip 1.204.0-dev.8852 coder-eval 0.12.7, gpt-5.6-terra, uip 1.204.0-dev.8852
Linux skills slice 24/25 (96%, floor 90%) 24/25 (96%)
Activation cases 27/27 25/27
Windows slice (2-task smoke) 1/2 0/2
Other harnesses on the same install codex and antigravity OK on Linux and Windows antigravity OK on Linux, codex and antigravity OK on Windows
Merge + upload OK OK
Cost $7.62 $4.24
Dry-run gate FAIL on 1 error (pass rate OK) FAIL on 1 error (pass rate OK)

The gate needs zero errors. Every error traces to something outside this PR:

Task Result Cause
troubleshoot-smoke-manifest-commands ERROR, both builds Its pre-run uip tools install gets "No compatible version … for the CLI's 1.204.x line". It has errored on every run that included it since 2026-09-23, on antigravity, delegate-sdk, codex and claude-code, including the Opus 5.5 sweep above.
migrator-migrate-with-fixes-e2e (Windows) ERROR, both builds Its llm_judge goes through the experiments' route: litellm, and coder_eval_uipath's daily-windows.ps1 installs coder-eval[dev,uipath] without the litellm extra. The task landed on main today (#3501), so tonight's nightly hits it on any pin.
migrator-migrate-classic-windows (codex only) FAILURE, score 0.86 Content, not infrastructure.

The dry-run summary also reports "Activation: 0% avg recall". That is the rollup's 20-prompt minimum per skill against one sampled case, and the 0.12.4 and 0.12.1 dry runs show the same 0%.

At merge time

Flip these at merge, not before. Otherwise the gate compares Sonnet runs against Opus baselines, and uipath-planner fails (Sonnet recalls about 60 against a floor of 85).

Where Variable Set to
This repo's Actions variables CLAUDE_CODE_MODEL claude-opus-5-5 (from claude-sonnet-5). The us-east-2 Bedrock account already serves us.anthropic.claude-opus-5-5.
ADO variable group coder-eval-athena-config (nightly) CLAUDE_CODE_MODEL eu.anthropic.claude-opus-5-5
Both BEDROCK_MODEL Unchanged, stays Sonnet 5. It is the judge and simulated-user model, not the agent under test.

Maturity is tracked per (harness, model), so the first five Opus 5.5 nightlies run the full suite before skipping resumes.

🤖 Generated with Claude Code

bai-uipath and others added 2 commits September 23, 2026 14:54
Picks up the command_executed fix that searches the whole command instead of only its first 2000 characters, plus the Opus 5.5 rate card.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Measured each skill's full positive set on claude-opus-5-5 at max_turns: 1, the gate's own method, for the move of the claude-code evals from Sonnet 5 to Opus 5.5.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@github-actions

github-actions Bot commented Sep 24, 2026 •

Copy link
Copy Markdown
Contributor

Claude finished @bai-uipath's task in 1m 57s —— View job


Summary

Moves the claude-code eval suite from Sonnet 5 to Opus 5.5 by bumping coder-eval to 0.12.7 and re-baselining the activation gate's per-skill recall thresholds on the new model. No skill content changes.

Change-by-Change Review

1. tests/.coder-eval-version

Severity: OK

Version bump from 0.12.4 to 0.12.7. Single-line change, straightforward.

The PR description notes 0.12.7 depends on UiPath/coder_eval#196 being merged and released first. The IMPORTANT banner at the top of the PR description correctly flags that installs will fail until that release lands — good flagging of the merge-order dependency.

2. tests/scripts/activation_gate.py

Severity: OK

Two changes:

  1. Comment update (lines 30–31): Date and model updated from 2026-06-17 / claude-sonnet-4-6 to 2026-09-23 / claude-opus-5-5. Matches the re-baseline described in the PR.

  2. Baseline re-measurement (lines 39–62): 22 skills baselined. All 22 skill names in BASELINES_PCT correspond to existing folders under skills/. Five newer skills on disk (uipath-aops, uipath-automationhub, uipath-connector-builder, uipath-insights, uipath-process-mining) are absent from the dict — the script handles these gracefully with SKIP: no baseline for {skill!r} at line 122.

    Baseline moves across models:

    Skill Sonnet 5 Opus 5.5 Delta
    uipath-rpa 100 90 -10
    uipath-review 95 90 -5
    uipath-maestro-bpmn 95 100 +5
    uipath-admin 95 100 +5
    uipath-planner 60 95 +35
    uipath-solution 90 100 +10
    uipath-agents 90 100 +10
    uipath-maestro-case 90 95 +5
    uipath-api-workflow 90 100 +10
    uipath-functions 95 100 +5

    Net positive: most skills improved, two dropped within the DROP_PP = 10 tolerance band. The large uipath-planner jump (60→95) is consistent with the PR description's note that the old 60% baseline reflected Sonnet 5's poor performance on that skill's ambiguous positives. The long explanatory comment block for uipath-planner is correctly removed — it was model-specific context that no longer applies.

    DROP_PP stays at 10, which remains reasonable as the safety margin.

What's Missing

Nothing identified. The PR description includes clear merge-time instructions for updating CLAUDE_CODE_MODEL in both the repo variable and the ADO variable group — those are operational steps that correctly live outside the code diff.

Area Ratings

Area Status Notes
Frontmatter N/A No skill frontmatter changed
E2E Tests OK Test infrastructure change only
Skill Body N/A No skill content changed
References & Assets N/A No references changed
Repo Hygiene OK No secrets, no stray changes, well-scoped

Issues for Manual Review

Conclusion

Clean, well-scoped infrastructure PR. No code issues found. The baseline re-measurements are internally consistent, the comment updates match the new model, and the merge-order dependency is clearly documented. Approve.
|

@uipreliga uipreliga left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🚢

@bai-uipath
bai-uipath merged commit 68dd36e into main Sep 25, 2026
39 checks passed
@bai-uipath
bai-uipath deleted the bai/nightly-opus-5-5 branch September 25, 2026 21:41
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants