Skip to content

Latest commit

 

History

History
3968 lines (2819 loc) · 201 KB

File metadata and controls

3968 lines (2819 loc) · 201 KB

CHANGELOG

v0.12.3 (2026-09-17)

Bug Fixes

  • Code review fixes for container-contract-and-command-surface (#178, 2f01c9d)

  • Code review fixes for tests-slim-prose (#180, a249877)

  • docs: Wrap an over-long line in the CE063 docstring (#180, a249877)

  • harbor: Allow touch's absolute-path form in trajectory_criteria (#187, 3addf46)

  • harbor: Reject --resume with --workspace-dir, sync design notes (#186, c9b96a8)

  • harbor: Remove the Write/content ambiguity from trajectory_criteria (#187, 3addf46)

  • harbor: Require Harbor E2E on PRs, de-flake trajectory_criteria fixture (#187, 3addf46)

  • harbor: Resolve workdir dynamically at run time instead of guessing at export (#182, 7af839b)

  • harbor: Trim packager docstrings under the prose budget (#182, 7af839b)

  • lint: Fail the prose gate cleanly on a missing root, and name the blank-line blind spot (#180, a249877)

  • lint: Include untracked files in the prose-only proof (#180, a249877)

  • lint: Make the prose-only proof see added directives and reject typos (#180, a249877)

  • lint: Point main's exempt-pair test at the repo root (#180, a249877)

  • sandbox: Close the adopt() installer-disclosure gap and fix the review's other findings (#187, 3addf46)

  • sandbox: Reprovision env_packages that a workspace capture stripped (#187, 3addf46)

  • sandbox: Trim adopt()'s re-provisioning comment under the prose budget cap (#187, 3addf46)

Continuous Integration

  • harbor: Zip and upload a failing scenario's full export/+jobs/ tree (#187, 3addf46)

Documentation

  • Apply main's prose rules to the container-contract branch (#178, 2f01c9d)

  • Record two prose-budget guard gaps the final review found (#180, a249877)

  • harbor: Trim two docstrings over the prose budget (#186, c9b96a8)

  • lint: Answer review — cut runs off the cap, drop lint-rules.md overlap (#180, a249877)

  • lint: Slim the five essays main's reports split brought in (#180, a249877)

  • tests: 2/8 — move the CE catalogue from notes/README.md to lint-rules.md (#180, a249877)

  • tests: 3/8 — move lint-rule defect stories into notes/lint-rules.md (#180, a249877)

  • tests: 4/8 — move doc-surface rule rationale into notes/lint-rules.md (#180, a249877)

  • tests: 5/8 — move golden-sensor and bracket-clock rationale into notes (#180, a249877)

  • tests: 6/8 — move plain test-module rationale into the subsystem notes (#180, a249877)

  • tests: 7/8 — delete HISTORY prose from tests, fix dangling docstring citations (#180, a249877)

Features

  • Give the host→container boundary a contract and fold aggregate into report --rebuild (#178, 2f01c9d)

  • cli: 4/6 — replace aggregate with report --rebuild (#178, 2f01c9d)

  • evaluate: 5/6 — refresh the run-level run.json after a detached grade (#178, 2f01c9d)

  • harbor: Unblock trajectory-dependent criteria and always allow credentials at export (#186, c9b96a8)

  • harbor: Unblock trajectory-dependent criteria, always allow credentials at export (#186, c9b96a8)

  • isolation: 1/6 — parse context.json as a ContainerContext contract (#178, 2f01c9d)

  • isolation: 2/6 — refuse image skew and assert the container echoed its contract (#178, 2f01c9d)

  • lint: 1/8 — scan a _ROOTS tuple in the prose budget (#180, a249877)

  • lint: 8/8 — turn the prose budget gate on for tests/ (#180, a249877)

  • lint: Cap the comment RUN, replacing the per-file comment budget (#180, a249877)

  • lint: Gate tests/ on the prose rules, and cap the comment run as well as the file total (#180, a249877)

  • lint: Keep the file-total comment budget as a backstop under the run cap (#180, a249877)

Refactoring

  • isolation: 3/6 — move the docker→tempdir driver rewrite host-side (#178, 2f01c9d)

  • lint: Delete CE023, which guards a package that no longer exists (#180, a249877)

Testing

  • Guard the contract echo over a maximal sandbox and the Typer-command exemption list (#178, 2f01c9d)

  • cli: 6/6 — pin every shared run/execute flag as declared identically (#178, 2f01c9d)

  • harbor: Dump agent-phase recorded commands on scenario failure (#187, 3addf46)

v0.12.2 (2026-09-16)

Bug Fixes

  • Code review fixes for antigravity-message-id (#165, 6166b1e)

  • Code review fixes for the reports consolidation (#176, 396c22c)

  • Code review fixes for the reports-consolidation review fixes (#176, 396c22c)

  • Code review fixes for timing-architecture-standardization (#165, 6166b1e)

  • Code review fixes for turn head/tail timing (#165, 6166b1e)

  • Code review fixes for turn-timing-consolidation (#165, 6166b1e)

  • Code review fixes for turn-timing-p0-p3 (#165, 6166b1e)

  • Reconcile the reports split with main's docs restructure (#176, 396c22c)

  • Review pass B — the tool bucket's dash survives the language boundary (#165, 6166b1e)

  • antigravity: 1/3 — give every generation a message_id (#165, 6166b1e)

  • claude-code: Subtract tool execution from the generation windows (#165, 6166b1e)

  • docs: Restore contracts the prose refactor lost or misstated (#177, 30f78f4)

  • docs: Restore the qualifier that made each absolute true (#177, 30f78f4)

  • docs: Stop asserting a tmpfs mask that no longer exists (#177, 30f78f4)

  • evalboard: 3/8 — an unbounded tool call contributes to no bucket (#165, 6166b1e)

  • lint: 1/4 — is_core_path covers the whole core layer (#176, 396c22c)

  • lint: 2/4 — declare CE044 and CE065 to ruff, and pin the id space (#176, 396c22c)

  • lint: CE004 checks reports/ — stop borrowing CE066's exemption set (#176, 396c22c)

  • lint: CE004 never fired on the relative import spelling (#176, 396c22c)

  • lint: Let CE045 globs reach .claude/*.md surfaces (#174, 9af3db3)

  • lint: Resolve a relative import against the importing file (#176, 396c22c)

  • lint: Stop the prose budget penalising usage examples (#177, 30f78f4)

  • pricing: Carry the exemption set into the generated mirror (#176, 396c22c)

  • timing: 2/8 — a published window must match its own bounds (#165, 6166b1e)

  • timing: 3/6 — a tool that closes between two windows is not model time (#165, 6166b1e)

  • timing: 3/7 — a naive/aware mix names the pair that disagreed (#165, 6166b1e)

  • timing: 5/6 — one clock basis per turn on antigravity and pi (#165, 6166b1e)

  • timing: A TurnClock for claude-code, and pi's two captured defects (#165, 6166b1e)

  • timing: Bracket the head and tail on the main thread only (#165, 6166b1e)

  • timing: Claude-code's windows tile across a tool result (#165, 6166b1e)

  • timing: Setup_ms marks from the task's start, not from _setup() (#165, 6166b1e)

  • timing: Stamp the turn bracket off the turn clock (CE064) (#165, 6166b1e)

  • timing: The bracket fixture's anchor cannot expire, and the tail is checked (#165, 6166b1e)

Chores

  • Record two deferred harness candidates from the reports consolidation (#176, 396c22c)

  • timing: Decompose_run's dead datetime import (#165, 6166b1e)

Documentation

  • 2/7 — move timing and permissions rationale into .claude/notes (#177, 30f78f4)

  • 3/7 — move agent-adapter rationale into .claude/notes (#177, 30f78f4)

  • 4/4 — retarget prose that names deleted report modules (#176, 396c22c)

  • 4/7 — move orchestration rationale into .claude/notes (#177, 30f78f4)

  • 5/5 — retarget every reports_* reference and record the rationale (#176, 396c22c)

  • 5/7 — move isolation, sandbox and CLI rationale into .claude/notes (#177, 30f78f4)

  • 6/7 — move criteria, routing and judging rationale into .claude/notes (#177, 30f78f4)

  • 7/7 — move reporting, harbor and telemetry rationale into .claude/notes (#177, 30f78f4)

  • Move design rationale out of src/ into .claude/notes, and gate it (#177, 30f78f4)

  • Record the four harness gaps the prose run did not close (#177, 30f78f4)

  • Restore the self-containment clause on reports_html (#177, 30f78f4)

  • Slim CLAUDE.md and split long-form rationale into .claude (#174, 9af3db3)

  • Slim CLAUDE.md, and drop a directory that never existed (#177, 30f78f4)

  • Update CLAUDE.md with communication style and development command clarifications (#177, 30f78f4)

  • harbor: Apply the branch's prose rules to the bind-mount rewrite (#177, 30f78f4)

  • harness: 3/3 — message_id is what splits the timeline (#165, 6166b1e)

  • harness: 4/8 — say what each harness's clock basis actually is (#165, 6166b1e)

  • harness: 6/6 — the timing architecture as it now stands (#165, 6166b1e)

  • harness: Claude-code's generation windows are not tool-subtracted (#165, 6166b1e)

  • harness: Record three guards the head/tail review could not close (#165, 6166b1e)

  • harness: Register that the golden corpus cannot see a timing value move (#165, 6166b1e)

  • harness: Register the message_id gaps the final review surfaced (#165, 6166b1e)

  • harness: Register what the turn-timing run could not guard (#165, 6166b1e)

  • harness: Widen the measured head/tail figures to six turns per harness (#165, 6166b1e)

  • notes: Cut the duplicated catalogue and the unbuilt design (#177, 30f78f4)

  • timing: Every harness subtracts tool time now, not two (#165, 6166b1e)

Features

  • evalboard: 3/4 — name the harness head and tail in the timeline strip (#165, 6166b1e)

  • harbor: Bind-mount task.yaml/plugins/templates/extra_mounts instead of COPY, skip Dockerfile when unneeded, translate pre_run (#175, 7bf8966)

  • lint: 1/7 — prose budget ratchet and one home for rationale (#177, 30f78f4)

  • lint: CE067 — CLAUDE.md's tree must name every top-level package member (#176, 396c22c)

  • lint: Replace the prose baseline with two self-adjusting rules (#177, 30f78f4)

  • pricing: 1/5 — generate the evalboard rate table from pricing.py (#176, 396c22c)

  • reports: 6-7/7 — the offline report carries the buckets; a TS None-vs-0 guard (#165, 6166b1e)

  • reports: 7/8 — the buckets reach every surface through one function (#165, 6166b1e)

  • timing: 1/6 — a two-sided residual gate for the four-bucket identity (#165, 6166b1e)

  • timing: 1/7 — a committed, ms-exact magnitude sensor (#165, 6166b1e)

  • timing: 4/4 — assert the buckets in replay, and record what they contain (#165, 6166b1e)

  • timing: 4/7 — one meaning for harness_startup_ms, on all five (#165, 6166b1e)

  • timing: Book each turn's head and tail as their own buckets (#165, 6166b1e)

  • timing: Name the setup and grading phases; union a row's tool time (#165, 6166b1e)

Refactoring

  • opencode: 8/8 — one duplicate parse (#165, 6166b1e)

  • reports: 2/5 — two DRY fixes and hoist 18 function-local imports (#176, 396c22c)

  • reports: 3/5 — split reports_stats.py into stats, result_metrics and helpers (#176, 396c22c)

  • reports: 4/5 — reports/ package, durations.py, and CE066 (#176, 396c22c)

  • reports: Generate the pricing mirror and split the reports layer (#176, 396c22c)

  • timing: 2/6 — one close_window() for the tiling reducers (#165, 6166b1e)

  • timing: 5/7 — one tool-subtraction, at the collector seam (#165, 6166b1e)

  • timing: 5/8 — one home for the rules, and a stored tool union (#165, 6166b1e)

  • timing: 6/8 — the two sensors share one selection rule (#165, 6166b1e)

Testing

  • harness: 2/7 — OpenCode and Pi get a corpus worth replaying (#165, 6166b1e)

  • harness: No test may read the pinned timing corpus (#165, 6166b1e)

  • harness: Pin why claude-code's zero head is left as a clamp (#165, 6166b1e)

  • harness: Stamp the two codex fixtures that timed themselves with now() (#165, 6166b1e)

  • lint: 2/3 — CE060, an AssistantMessage must declare its message_id (#165, 6166b1e)

  • lint: 2/4 — widen CE058 to the turn head/tail buckets (#165, 6166b1e)

  • lint: 4/6 — CE061, a window must come from the shared helper (#165, 6166b1e)

  • lint: CE058 form 6 — the zero its guard does not vouch for (#165, 6166b1e)

  • lint: Drop a personal path and de-duplicate the isolation pin (#176, 396c22c)

  • lint: Drop two function-local json imports that shadow the module one (#176, 396c22c)

  • lint: Guard the Rationale pointer's placement, not just its target (#177, 30f78f4)

  • reports: 3/4 — cover the HTML slowest-commands truncation branch (#176, 396c22c)

  • timing: 1/8 — the bracket's SOURCE, not just its presence (#165, 6166b1e)

  • timing: Unify the fixture clocks and assert the four-bucket identity (#165, 6166b1e)

v0.12.1 (2026-09-12)

Bug Fixes

  • Code review fixes for record-cli-review-fixes (#150, 3aca8e8)

  • Resolve py/import-and-import-from CodeQL alerts in test_regrade.py + test_detached_grading_boundaries.py (#161, 6ab6b87)

  • criteria: 2/4 — a shim rule fault is an eval-config error, not an agent failure (#150, 3aca8e8)

  • evaluate: Close the PR review findings on container-based detached grading (#161, 6ab6b87)

  • evaluate: Close the review blockers on container-based detached grading (#161, 6ab6b87)

  • harbor: Address code review findings from full 8-axis review (#166, 9caffcb)

  • harbor: Address PR review blockers (score-integrity, security, drops) (#166, 9caffcb)

  • harbor: Derive a prebuilt image's real WORKDIR instead of guessing /app (#166, 9caffcb)

  • harbor: Suppress pyright reportMissingImports for the harbor package (#166, 9caffcb)

  • models: Reference FlagPredicate at runtime so the cast is a real use (#150, 3aca8e8)

  • record_cli: Close three real holes the Copilot review found (#150, 3aca8e8)

  • record_cli: Reject an unusable response rule at load, and trace a shim fault (#150, 3aca8e8)

  • sandbox: Give the sandbox venv system site packages (#162, e7d8326)

  • test: Strip ANSI color before asserting --workspace-dir in CLI output (#166, 9caffcb)

  • tests: Harden harbor CLI-error assertions and Windows chmod checks (#166, 9caffcb)

  • timing: Account for generation and tool time on every harness (#164, 87102a3)

Build System

  • docker: Bake the harbor extra into the coder-eval-agent image (#166, 9caffcb)

Continuous Integration

  • harbor: Add Harbor E2E workflow (export + CoderEvalAgent round trip) (#166, 9caffcb)

Features

  • evaluate: Grade a driver: docker row inside a container of its own image (#161, 6ab6b87)

  • harbor: Coder-eval as a Harbor agent (C1.2) (#166, 9caffcb)

  • harbor: Export coder-eval tasks to Harbor, with coder-eval as the grader (#166, 9caffcb)

  • harbor: Export experiment.yaml variants to Harbor task directories (#166, 9caffcb)

  • harbor: Fix agent workspace alignment, env passthrough, and template sources (#166, 9caffcb)

  • harbor: Task.yaml/experiment.yaml → Harbor export + coder-eval as a Harbor agent (#166, 9caffcb)

  • record_cli: Serve a different canned response per invocation (#150, 3aca8e8)

Refactoring

  • record_cli: 1/4 — copy argv_match beside the shim instead of splicing it (#150, 3aca8e8)

Testing

  • Close three harness gaps this review surfaced (#150, 3aca8e8)

  • lint: 4/4 — put the record_cli authoring surface under CE030 doc parity (#150, 3aca8e8)

  • tasks: 3/4 — add the record_cli per-invocation-response probe (#150, 3aca8e8)

v0.12.0 (2026-09-09)

Bug Fixes

  • Code review fixes for pi-harness (#159, 57c9e33)

  • container: Arm the heartbeat watchdog only inside the container (#154, 0fe8cb0)

  • criteria: Refuse an out-of-sandbox criterion path instead of scoring it 0.0 (#154, 0fe8cb0)

  • eval: Address the medium and low findings from the branch review (#154, 0fe8cb0)

  • eval: Close the verdict-correctness gaps in detached grading (#154, 0fe8cb0)

  • evalboard: Show per-row cost for open-weight harnesses (Pi) (#159, 57c9e33)

  • execute: Close the verdict-changing and trust-boundary defects in detached grading (#154, 0fe8cb0)

  • execute: Close the verdict-divergence, trust-gate and fabricated-rate defects (#154, 0fe8cb0)

  • execute: Move post_run to the grading phase and stop the record lying about its driver (#154, 0fe8cb0)

  • pi: Address bai-uipath review — apportionment caveat + token/telemetry minors (#159, 57c9e33)

  • pi: Do not forward allowed_tools/disallowed_tools to Pi (#159, 57c9e33)

  • pi: Multi-model review — gate error-crash on intentional cuts + dead dup + doc/guard nits (#159, 57c9e33)

  • pi: Review blockers — session-id sanitize, error crash, telemetry warn, tool map, CE047 (#159, 57c9e33)

  • security: Close the three CodeQL findings on the detached-grading diff (#154, 0fe8cb0)

  • tasks: Make the two remaining absolute criterion paths reachable, and add CE055 (#154, 0fe8cb0)

  • tests: Kill the CodeQL taint source and the Windows mode assertion (#154, 0fe8cb0)

Documentation

  • agents: 4/4 — add Pi harness docs + enumeration-surface parity (#159, 57c9e33)

  • docker: Fix env_passthrough model name + list Pi in baked-toolchain docs (#159, 57c9e33)

  • pi: 3/4 — document docker support for Pi (#159, 57c9e33)

  • pi: Correct plugins->--skill support + OPENROUTER passthrough claims (#159, 57c9e33)

  • pi: Fix tool-enforcement claims + cost docstring after review (#159, 57c9e33)

Features

  • agents: 2/4 — add PiAgent harness (pi --mode json) + registration (#159, 57c9e33)

  • cli: coder-eval execute + detached grading via evaluate <run_dir> (#154, 0fe8cb0)

  • cli: Add coder-eval execute — run tasks without grading them (#154, 0fe8cb0)

  • cli: Grade an executed run afterwards — evaluate <run_dir> + Sandbox.adopt (#154, 0fe8cb0)

  • cli: Make --resume distinguish "executed" from "graded" (#154, 0fe8cb0)

  • docker: 1/4 — bake the pinned Pi CLI into docker/Dockerfile (#159, 57c9e33)

  • evalboard: Show the Pi logo + "Pi" in the harness view (#159, 57c9e33)

  • models: 1/4 — add AgentKind.PI + PiAgentConfig model (#159, 57c9e33)

  • pi: Add Pi harness (--type pi) with docker + skill injection (#159, 57c9e33)

  • pi: Load agent.plugins skills via --skill + harden turn/token handling (#159, 57c9e33)

  • sandbox: 2/4 — forward OPENROUTER_API_KEY into docker containers (#159, 57c9e33)

  • tasks: 3/4 — add local pi_smoke_test task (#159, 57c9e33)

Refactoring

  • agents: Hoist shared plugins->skills resolver into agents/_skills.py (#159, 57c9e33)

Testing

  • Fix two CI-only failures in the detached-grading tests (#154, 0fe8cb0)

  • docker: 4/4 — guard that the Dockerfile bakes a pinned Pi CLI (#159, 57c9e33)

  • pi: Assign the awaited cancel result to silence CodeQL 'no effect' (#159, 57c9e33)

  • pi: Install shutil.which patch so the env-info test doesn't need the real CLI (#159, 57c9e33)

  • pi: Port OpenCode teardown + cost-fallback matrices (blockers 3, 4) (#159, 57c9e33)

  • pi: Scrub personal scratchpad path from happy-stream fixture (#159, 57c9e33)

v0.11.7 (2026-09-08)

Bug Fixes

  • deps: Bump google-antigravity 0.1.7 -> 0.1.8 (Defender FP on the harness) (#158, 48a4d53)

  • deps: Bump google-antigravity to 0.1.8 to clear a Defender false positive on the bundled harness [PILOT-7463] (#158, 48a4d53)

  • docs: Update Antigravity version in docs (#158, 48a4d53)

  • evalboard: Always show every known harness in the filter (#152, 5e7d2b6)

  • evalboard: Stop counting one model as two in the run header (#152, 5e7d2b6)

  • pricing: Refresh the rate card and add gemini 3.7/3.8 Flash (#155, be98f9d)

Documentation

  • Add intro video and make the README agent-agnostic (#157, d715f85)

  • Widen the framing past skills-only and guard the agent roster with CE047 (#157, d715f85)

  • deps: Drop the Defender rationale from the antigravity pin comment (#158, 48a4d53)

  • evalboard: Trim the comments added by this branch (#152, 5e7d2b6)

  • pricing: Trim the rate-card comments to what affects an edit (#155, be98f9d)

  • readme: Add intro video and make the README agent-agnostic (#157, d715f85)

  • readme: Play the intro video inline, with YouTube as the fallback (#157, d715f85)

  • readme: Tell viewers to unmute the inline video (#157, d715f85)

  • stub: Make the Pages stub and package metadata agent-agnostic (#157, d715f85)

  • tutorial: Stop sending first-timers through the contributor toolchain (#157, d715f85)

Features

  • evalboard: Add loading skeletons to the slow routes (#152, 5e7d2b6)

Performance Improvements

  • evalboard: Cut page load time and show loading state while pages load (#152, 5e7d2b6)

  • evalboard: Move repeated per-row utility classes into the stylesheet (#152, 5e7d2b6)

  • evalboard: Stop re-reading the whole run store on every render (#152, 5e7d2b6)

v0.11.6 (2026-09-01)

Bug Fixes

  • action: Drop the literal expression syntax from a run-step comment (#147, 7b81456)

  • action: Stop the env passthrough mutating the action's own shell (#147, 7b81456)

  • docs: Drop the last references to inputs that no longer exist (#147, 7b81456)

  • evalboard: Even out the variant tiles and keep a task's arms adjacent (#145, 6780227)

  • evalboard: Resolve a variant-less task link to the task's first arm (#145, 6780227)

Documentation

  • action: Trim the comment bulk on the eight-input surface (#147, 7b81456)

  • evalboard: Trim variant commentary and redundant tests (#145, 6780227)

Features

  • action: Add working-directory, extras, extra-packages, prerelease and args inputs (#147, 7b81456)

  • action: Refactor github action input surface around args; harden env passthrough (#147, 7b81456)

  • evalboard: Register an unlisted gha source for ad-hoc dispatch runs (#147, 7b81456)

  • evalboard: Render multi-variant runs, one row per (task, arm) (#145, 6780227)

Refactoring

  • action: Drop every forwarding input; flags go through args (#147, 7b81456)

  • evalboard: Drop the variant colour palette and the spread helper (#145, 6780227)

  • evalboard: Report pass rate per arm instead of alongside a blended one (#145, 6780227)

Testing

  • Drop the skip-guard unit test (#146, d679ff2)

  • Skip the enforcement live tests when the agent declines to try (#146, d679ff2)

  • action: Record stub argv from bash so Windows argv is byte-exact (#147, 7b81456)

  • action: Stop Git Bash rewriting the absolute path in the install-spec tests (#147, 7b81456)

v0.11.5 (2026-08-28)

Bug Fixes

  • opencode: Bound the post-EOF reap, reap the whole process group, test every failure path (#115, 5532b87)

  • opencode: Close the remaining review findings on the harness (#115, 5532b87)

  • opencode: Gate on captured tokens, canonicalize tool args, pin max_turns (#115, 5532b87)

  • opencode: Inject plugin skills so skill suites measure the skills (#115, 5532b87)

  • opencode: Keep the smoke task out of the CI smoke-pass bucket (#115, 5532b87)

  • opencode: Map apply_patch to Write so GPT-family edits are seen by criteria (#115, 5532b87)

  • opencode: Reap the CLI on every turn exit, test the sandbox env contract (#115, 5532b87)

  • opencode: Spread super() in get_environment_info; guard with CE046 (#115, 5532b87)

  • opencode: Typecheck on Windows, satisfy both CodeQL findings (#115, 5532b87)

  • plugin: Close the review's verified gaps — regex blind spot, parallel surfaces, runtime guard (#143, c565ebb)

  • plugin: Correct the skill-reachability path — every generated activation suite reports recall 0.0 (#143, c565ebb)

  • plugin: Correct the skill-reachability path, and reuse PR #109's measured descriptions (#143, c565ebb)

  • routing: Decouple simulator route from checker_context.api_route (#144, 88ff0f0)

  • routing: Reinstate litellm+agent_judge rejection guard (#144, 88ff0f0)

  • routing: Restore LiteLLM->Claude pin on simulator_route (#144, 88ff0f0)

  • test: Add CE045, document the plugin-path divergence, unpin a test from ordering (#143, c565ebb)

  • utils: Use explicit concatenation in the plugin-root warning (#143, c565ebb)

Chores

  • opencode: Standardize on deepseek-v4-pro, drop the flash-0731 rate entry (#115, 5532b87)

  • plugin: Answer "skills only?", fold tags into keywords, add CE044 (#141, 2ae9b7b)

  • plugin: Give both author objects the same contact address (#141, 2ae9b7b)

  • plugin: Give the marketplace owner a contact address (#141, 2ae9b7b)

  • plugin: Lead both manifests with skill evaluation, add discovery metadata (#141, 2ae9b7b)

  • plugin: Pin both manifests to their published JSON schemas (#141, 2ae9b7b)

Documentation

  • opencode: Note _TERM_GRACE_SECONDS's second role as the post-EOF exit grace (#115, 5532b87)

Features

  • agents: Add OpenCode harness with opt-in [opencode] extra (#115, 5532b87)

  • opencode: Add require_token_telemetry, an escape hatch for the zero-token guard (#115, 5532b87)

v0.11.4 (2026-08-27)

Chores

  • Add @CarlesUIPath as a code owner (#140, 8fff865)

  • Sync PR-review comment allowlist with CODEOWNERS (#140, 8fff865)

Features

  • docker: Bake litellm extra into the docker image (#142, 3f4ae36)

v0.11.3 (2026-08-27)

Bug Fixes

  • checker-context: Typed model, reject litellm+agent_judge/sim, live test (#137, 727bb7b)

  • eval-routing: Restore DEFAULT_JUDGE_MODEL floor + add litellm judge transport (#137, 727bb7b)

  • evalboard: Exclude carried-forward passes from the wall-clock aggregates (#125, dd918e6)

  • release: Pin python-semantic-release + GitPython to unbreak version bump (#139, ef74014)

  • routing: Make _resolve_backend_route's match exhaustive (#137, 727bb7b)

Continuous Integration

  • Retrigger checks (GH Actions appeared stalled repo-wide) (#137, 727bb7b)

  • Retrigger checks (previous push did not trigger CI) (#137, 727bb7b)

  • fix: Install litellm extra for pyright, address CodeQL findings (#137, 727bb7b)

Documentation

  • timing: Trim justification prose from the wall-clock comments (#125, dd918e6)

Features

  • eval-routing: Decouple judge/agent_judge backend+model from the agent's own route (#137, 727bb7b)

  • eval-routing: Decouple judge/agent_judge backend+model from the agent's route (#137, 727bb7b)

  • evalboard: Chart seconds per passed task on the overview (#125, dd918e6)

  • evalboard: Read the time ratio without hovering (#125, dd918e6)

  • evalboard: Replace the turn-budget signal with time per passed task (#125, dd918e6)

  • evalboard: Run the wall-clock signal beside the turn budget, behind tabs (#125, dd918e6)

  • litellm-judge: Support arbitrary litellm kwargs via params/auth (#137, 727bb7b)

Refactoring

  • litellm-judge: Drop settings coupling, rename auth to env_params (#137, 727bb7b)

v0.11.2 (2026-08-24)

Bug Fixes

  • docker: Expand ~ and $VAR in extra_mounts destinations (#128, 02dbf69)

  • docker: Reject expansions that inject a ':' into a mount path (#128, 02dbf69)

Documentation

  • docker: Trim the extra_mounts destination comments to scope (#128, 02dbf69)

v0.11.1 (2026-08-21)

Bug Fixes

  • ci: Pin pip floor to 26.2 to close PYSEC-2026-3721 (#133, beecedd)

  • codex: Store command output whole in result_summary (+ code-review fixes) (#127, 3d09f0f)

  • deps: Address PR #133 review — anthropic 1.0.0 dropped temperature kwarg (#133, beecedd)

  • deps: Bump anthropic to 1.0.0 and migrate Bedrock judge path to httpx2 (#133, beecedd)

  • deps: Bump anthropic to 1.0.0, migrate Bedrock judge path to httpx2 (#133, beecedd)

  • deps: Re-lock pip to 26.2.1 to actually close PYSEC-2026-3721 (#133, beecedd)

v0.11.0 (2026-08-19)

Bug Fixes

  • docker: Restore container access after the DAC cap drop, shield the task dir (#106, 56bffad)

  • reference: Address code review and CodeQL findings (#106, 56bffad)

  • reference: Address PR review — scoring correctness, fail-closed anti-cheat (#106, 56bffad)

  • reference: Clear remaining CodeQL alerts (#106, 56bffad)

Documentation

  • reference: Record why READ_ONLY_MODE exists, and what is left to wire (#106, 56bffad)

Features

  • reference: Directory-only references + anti-cheat permission window (#106, 56bffad)

Testing

  • early-stop: CE036 enforces the live_verdict determinism + monotonicity contract (#126, d854004)

  • reference: Skip host-side chmod assertions on Windows (#106, 56bffad)

v0.10.2 (2026-08-18)

Continuous Integration

  • Exclude attestation sidecars from the PyPI artifact-identity assert (#124, 7d8a771)

v0.10.1 (2026-08-18)

Continuous Integration

  • Bump gh-action-pypi-publish to v1.14.2 to accept Metadata-Version 2.5 (#123, a65ee69)

v0.10.0 (2026-08-18)

Bug Fixes

  • Code review fixes for the Claude Code plugin marketplace (#82, 02e5151)

  • Code review fixes for the plugin generic-adopter plan (#82, 02e5151)

  • Code review fixes for the plugin-audit P0/P1 plan (#82, 02e5151)

  • Derive the poll loop's exit bound from the turn's actual timeout (#111, d3f1432)

  • Make lint-tasks' read-only rule outlive the frontmatter deny (#82, 02e5151)

  • Reconcile the plugin branch with main after rebase (#82, 02e5151)

  • Remove redundant asyncio re-import flagged by CodeQL (#111, d3f1432)

  • agent: Address PR review — simulator replace mode, preset-aware reports, replace validator (#92, fddd3c5)

  • agent: Address PR review — unconditional preset, judge replace seam (#92, fddd3c5)

  • agent: Align Codex system_prompt with the append-only contract (#92, fddd3c5)

  • agent: Append system_prompt to the Claude Code preset instead of replacing (#92, fddd3c5)

  • agent: Reject a blank system_prompt_file under replace mode (#92, fddd3c5)

  • agent: Resolve system_prompt_file atomically and reject blank prompts (#92, fddd3c5)

  • antigravity: Poll for backgrounded work instead of grading it incomplete (#111, d3f1432)

  • ci: Address PR #81 review — undefined step output, dead job gate, CE035 (#81, 9cc45da)

  • ci: Close the five gate-correctness findings from the code review (#81, 9cc45da)

  • ci: Code review fixes for the published-action verification (#81, 9cc45da)

  • ci: Fail the published-action gate on a run-limit breach (#81, 9cc45da)

  • ci: Harden promote ordering and stop preflight misdiagnosing healthy lag (#81, 9cc45da)

  • claude: Make Claude speak English, not Claudish (#118, 5006c91)

  • cli-called: Move alternation to verb_any_of, close review findings (#103, f7e9fda)

  • codex: Fold sub-agent tokens on a turn-cap stop (#110, 12c5031)

  • codex: Report per-turn tokens instead of the thread-cumulative total (#113, 80f3523)

  • deps: Bump sqlparse 0.5.5 -> 0.6.0 to clear the pip-audit gate (#120, ea5a3fc)

  • deps: Bump sqlparse to 0.6.0 and close pre/post-run subprocess transports (#120, ea5a3fc)

  • evalboard: Address review findings on the Scribe source layer (#116, d6f8d7b)

  • evaluation: Harden post-failure evidence (#119, 636e87d)

  • orchestrator: Close pre/post-run subprocess transports so Windows CI stops leaking (#120, ea5a3fc)

  • plugin: Address PR #82 review — reachable activation suites, least-privilege CI (#82, 02e5151)

  • plugin: Address the three PR #82 findings left open (#82, 02e5151)

  • plugin: Make analyze compute its numbers, and weight smoke criteria honestly (#82, 02e5151)

  • plugin: Name the tautological-criterion trap in init and task (#82, 02e5151)

  • plugin: Wire the skill source into the ci skill's scheduled drift run (#82, 02e5151)

  • tasks: Armed positives must require success, guarded by CE034 (#82, 02e5151)

  • test: Resolve bash by absolute path so Windows CI stops hitting WSL (#81, 9cc45da)

Chores

  • Reconcile the published-action verification with main (#81, 9cc45da)

  • Renumber a rebase-collided lint-rule candidate (#81, 9cc45da)

  • agents: Drop config_support and the Antigravity tool mapping (#110, 12c5031)

  • agents: Drop the system_prompt changes from this PR (#110, 12c5031)

  • plugin: Pin plugin.json to 0.9.6 after rebasing onto the release (#82, 02e5151)

  • plugin: Regenerate the bundled criteria reference (#103, f7e9fda)

Code Style

  • Sort claude_agent_sdk.types import (#92, fddd3c5)

Continuous Integration

  • release: Promote v0 only after PyPI publish, verify the published action (#81, 9cc45da)

Documentation

  • Add Tutorial 07 for the plugin, and fix two gaps in PLUGIN.md (#82, 02e5151)

  • Defer one harness candidate from the plugin-audit run (#82, 02e5151)

  • Rework Tutorial 07 after review — accuracy fixes and far less narration (#82, 02e5151)

  • Stop teaching the recursive task glob the ci skill forbids (#82, 02e5151)

  • Use current-generation models in examples (#110, 12c5031)

  • claude: Drop the config_support contract from the repo guide (#110, 12c5031)

  • cli-called: Drop the hardcoded subcommand count (#103, f7e9fda)

  • cli-called: Fix the guide's orphaned exact_positional prerequisites (#103, f7e9fda)

  • cli-called: Fix two unclear field descriptions (#103, f7e9fda)

  • cli-called: Name the bug the detail renderer's source avoids (#103, f7e9fda)

  • cli-called: Put the offset comment's two reasons on their own lines (#103, f7e9fda)

  • cli-called: Trim comments to the whys (#103, f7e9fda)

  • parity: State the real final_status of a capped run (#110, 12c5031)

  • run-limits: Keep the contract on the page, the measurements in the PR (#110, 12c5031)

  • run-limits: Re-measure the antigravity timeout case after the poll loop (#110, 12c5031)

  • run-limits: Record the measured cross-harness parity results (#110, 12c5031)

Features

  • agent: Emit system_prompt_semantics from the Agent base (#92, fddd3c5)

  • agent: Record system_prompt_semantics marker in environment_info (#92, fddd3c5)

  • agents: Honor run_limits.max_turns on codex and antigravity (#110, 12c5031)

  • agents: Make a base-config field mean the same thing on every harness (#110, 12c5031)

  • cli-called: Accept a list of verb spellings (#103, f7e9fda)

  • cli-called: Accept alternative verbs via verb_any_of (#103, f7e9fda)

  • cli-called: Add exact_positional to pin the argument tail (#103, f7e9fda)

  • evalboard: Add a Scribe tab, reading the Autopilot suite's own blob container (#116, d6f8d7b)

  • evaluation: Preserve criteria after agent failures (#119, 636e87d)

  • lint: CE026 clause 4 — snippet with: keys must be real action inputs (#82, 02e5151)

  • plugin: 1/5 — the bundled criteria reference explains optional fields (#82, 02e5151)

  • plugin: 1/6 — discover the eval tree instead of assuming tasks/ and runs/latest (#82, 02e5151)

  • plugin: 1/6 — marketplace + coder-eval plugin skeleton (#82, 02e5151)

  • plugin: 2/5 — shared adversarial task rubric, applied by task (#82, 02e5151)

  • plugin: 2/6 — analyze reads the run's actual schema, not one generation's (#82, 02e5151)

  • plugin: 2/6 — generate the bundled criteria reference, guard it with CE032 (#82, 02e5151)

  • plugin: 3/5 — a real run becomes part of done in task (#82, 02e5151)

  • plugin: 3/6 — resolve the version a project pins before validating anything (#82, 02e5151)

  • plugin: 3/6 — skill-check skill and the canonical activation suite (#82, 02e5151)

  • plugin: 4/5 — /coder-eval:lint-tasks, a read-only reviewer of existing tasks (#82, 02e5151)

  • plugin: 4/6 — init and task skills (#82, 02e5151)

  • plugin: 4/6 — look before you write: skill-check, init, lint-tasks (#82, 02e5151)

  • plugin: 5/5 — activation budgets in skill-check, layer routing in analyze (#82, 02e5151)

  • plugin: 5/6 — analyze and ci skills (#82, 02e5151)

  • plugin: 5/6 — repo convention wins, and the CI gate stops measuring the wrong set (#82, 02e5151)

  • plugin: 6/6 — document the criterion aliases the loader accepts, from the models (#82, 02e5151)

  • plugin: 6/6 — validate the plugin in CI, extend CE026, document it (#82, 02e5151)

  • plugin: CLI-driving skills offer to install coder-eval, asking first (#82, 02e5151)

  • plugin: Ship coder_eval as a Claude Code plugin + marketplace (#82, 02e5151)

Refactoring

  • Retire the repo-local twins of the plugin's authoring skills (#82, 02e5151)

  • agent: Resolve the system prompt to a value, not a mode string (#92, fddd3c5)

  • cli-called: Collapse the pairwise verb check to itertools.combinations (#103, f7e9fda)

  • plugin: Rename skill-check to check-skill, and write the naming rule down (#82, 02e5151)

  • reports: Read the recorded prompt regime instead of sniffing its shape (#92, fddd3c5)

Testing

  • Defer the plugin-skill repo-file containment guard (#82, 02e5151)

  • antigravity: Clear the CodeQL findings on the fake SDK helper (#110, 12c5031)

  • antigravity: Pin the SDK half of the env seam contract (#110, 12c5031)

  • plugin: Make a new skill declare whether it needs eval-root discovery (#82, 02e5151)

  • run-limits: Add cross-harness max_turns / turn_timeout fixtures (#110, 12c5031)

  • run-limits: Make the max_turns fixture assert the cap bound (#110, 12c5031)

  • run-limits: Tag the parity fixtures (#110, 12c5031)

v0.9.6 (2026-08-11)

Bug Fixes

  • Code review fixes for evalboard path-to-ga stale tags and mature passes (#94, ce006c1)

  • Review round 2 — gate evalboard in CI, correct GPT-5.6 rates (#94, ce006c1)

  • ci: Correct defects in the uipath runner migration (#86, 57556af)

  • criteria: Make glob path resolution literal-first and ignore-filtered (#65, b3bba2b)

  • criteria: Split clustered short flags and keep negative numbers positional (#73, a7ec3ea)

  • evalboard: 1/3 — drop de-tagged tasks and score only executed runs (#94, ce006c1)

  • evalboard: Path-to-GA shows only still-tagged tasks, scored on runs that executed (#94, ce006c1)

  • evalboard: Resync lib/pricing.ts with the authoritative pricing.py table (#94, ce006c1)

  • sandbox: Close the remaining record_cli findings from #73 (#73, a7ec3ea)

  • sandbox: Repair record_cli defects found reviewing #73 (#73, a7ec3ea)

  • sandbox: Stop the recorder dir defeating the PLUGIN_TOOLS_DIR pin (#73, a7ec3ea)

Chores

  • Change to centralized managed GitHub pool (#86, 57556af)

Documentation

  • harness: Defer four TS-side evalboard invariants from the path-to-ga fix (#94, ce006c1)

Features

  • criteria: Accept glob patterns in criterion path fields (#65, b3bba2b)

  • evalboard: 2/3 — surface last-seen and maturity on the Path-to-GA table (#94, ce006c1)

  • sandbox: Generate CLI recording shims via record_cli (#73, a7ec3ea)

v0.9.5 (2026-08-05)

Bug Fixes

  • command-executed: Keep whole argv-joined payload in shell unwrap (#77, 7abd080)

  • command-executed: Match patterns against shell-normalized commands (#77, 7abd080)

  • command-executed: Narrow command param to str before shell-normalizing (#77, 7abd080)

  • command-executed: Recognize shell wrappers by predicate, not allowlist (#77, 7abd080)

  • criteria: Add present predicate so asserting a switch cannot weaken a guard (#72, 8574ded)

  • criteria: Make cli_called guards fail loud instead of vacuously passing (#72, 8574ded)

  • criteria: Stop ignore_flags re-opening the guard false-PASS (#72, 8574ded)

  • early-stop: Address PR review — trajectory parity, reason determinism, doc restore (#78, 4cf8092)

  • lint: Derive CE030 criteria from the source union literal, not runtime (#77, 7abd080)

  • lint: Enumerate in-tree criteria by module attribute, not the union (#77, 7abd080)

  • lint: Scope CE030 criterion parity to in-tree criteria only (#77, 7abd080)

  • reports: Explicit return on every early_stop_gate_note path (CodeQL py/mixed-returns) (#78, 4cf8092)

  • reports: Pre-initialize the gate note so CodeQL sees it bound on every path (#78, 4cf8092)

Chores

  • deps-dev: Bump postcss from 8.5.18 to 8.5.23 in /evalboard (#75, cdced15)

Documentation

  • Surface the Marketplace listing and make the Action quickstarts self-sufficient (#80, 401245a)

  • command-executed: Document shell-normalization contract + gate it (CE030) (#77, 7abd080)

Features

  • criteria: Add cli_called for structured invocation matching (#72, 8574ded)

  • criteria: Match a flag across spellings with FlagMatch.aliases (#72, 8574ded)

  • early-stop: Per-criterion arming via stop_early blocks on live criteria (#78, 4cf8092)

Refactoring

  • command-executed: Total _match_haystacks, shared window, memoized (#77, 7abd080)

v0.9.4 (2026-08-04)

Bug Fixes

  • litellm: Pin litellm[proxy]==1.95.0 + fastapi==0.140.0 for proxy startup (#76, f2f8580)

  • litellm: Pin proxy deps (litellm 1.95.0 + fastapi 0.140.0) to fix startup crash (#76, f2f8580)

Chores

  • action: Rename Marketplace listing to coder_eval, add author (a9c274d)

Documentation

  • litellm: Surface the proxy dep-pin override vars in start script (#76, f2f8580)

Refactoring

  • litellm: Address PR review — pin SSOT guard, rename, doc ripple (#76, f2f8580)

v0.9.3 (2026-08-04)

Bug Fixes

  • early-stop: Address PR review — polarity-blind budget, pass_threshold displacement, gate-semantic split (#74, 800ac77)

Chores

  • deps: Bump aiohttp 3.14.1→3.14.3, cryptography 49.0.0→50.0.0 (#74, 800ac77)

Features

  • early-stop: Weighted ceiling/floor bounds + decision-step budget (#74, 800ac77)

v0.9.2 (2026-07-31)

Bug Fixes

  • cost: A task timeout with no preserved turn is unrecorded spend, not free (#63, 93c7fc0)

  • cost: Book spend on the error and timeout paths, flag what is unpriced (#63, 93c7fc0)

  • cost: Flag every hard-killed task as a cost floor, not just the empty ones (#63, 93c7fc0)

  • evalboard: Honest scoped counts, and one definition of a run's scope (#69, 0bdac0b)

  • litellm: Gate cost_log_tags on agent capability, not route (fixes non-Claude crash) (#66, 4131a2a)

  • litellm: Make the orphaned-spend warning actually fire (#66, 4131a2a)

  • litellm: Per-attempt cost-log scoping + single run-id accessor + no-match warning (#66, 4131a2a)

  • litellm: Pin each open-weight model to a vetted provider set (no silent fallback) (#66, 4131a2a)

  • litellm: Proxy-authoritative token buckets + all-priced gate + transactional join (#66, 4131a2a)

  • litellm: Sanitize cost headers, reject non-finite cost, drop debug scaffolding (#66, 4131a2a)

  • orchestrator: Recover the in-flight turn's spend on a hard kill (#63, 93c7fc0)

  • pricing: Add the claude-opus-5 rate so killed turns stop booking zero (#63, 93c7fc0)

  • pricing: Add the five unpriced codex tiers still on OpenAI's rate card (#63, 93c7fc0)

  • pricing: Correct every wrong rate-card entry and close the alias gaps (#63, 93c7fc0)

  • pricing: Refresh the rate card and correct gemini-3-flash-preview (#63, 93c7fc0)

  • reports: Count errors as misses and stop losing cost on error paths (#63, 93c7fc0)

  • reports: Count errors as misses in one canonical pass rate (#63, 93c7fc0)

Code Style

  • evalboard: Drop the swatch dots and the scope caption from the header (#69, 0bdac0b)

Documentation

  • cost: Describe the per-turn backfill as the net it is (#63, 93c7fc0)

  • cost: Describe the unpriced-crash mechanism accurately and keep comments framework-general (#63, 93c7fc0)

  • litellm: Correct the cost contract after cutting per-message distribution (#66, 4131a2a)

  • litellm: Document LITELLM_COST_LOG wiring + correct the reconciliation-cost contract (#66, 4131a2a)

Features

  • cost: Publish one accurate total on every reporting surface (#63, 93c7fc0)

  • docker: Bind-mount the LiteLLM cost log so --driver docker joins actual cost (#66, 4131a2a)

  • evalboard: Compare every harness on the overview, and scope the whole page to one (#69, 0bdac0b)

  • evalboard: Compare harnesses on the overview, and identify each run (#69, 0bdac0b)

  • evalboard: Lift the harness scope to the page header, in vendor colors (#69, 0bdac0b)

  • evalboard: Make each turn's provider-call table a collapsed dropdown (#66, 4131a2a)

  • evalboard: Mark a partly-priced run total as a floor, not the bill (#63, 93c7fc0)

  • evalboard: One set of pass-rate cutoffs, and a run table that pages through all history (#69, 0bdac0b)

  • evalboard: Per-call cost/cache table from provider_call_costs (replaces inline) (#66, 4131a2a)

  • evalboard: Read the canonical pass rate and surface incomplete cost (#63, 93c7fc0)

  • evalboard: Say which harness, model, and framework version a run used (#69, 0bdac0b)

  • litellm: Actual per-call cost + cache accounting for the open-weight backend (#66, 4131a2a)

Refactoring

  • cost: Correct the simulator-cost bound and drop the unread variant error share (#63, 93c7fc0)

  • cost: Cut the commentary and drop unreachable rate-card keys (#63, 93c7fc0)

  • cost: Define the unpriced-row test once, and only for new runs (#63, 93c7fc0)

  • cost: Total_cost_usd means the whole bill everywhere (#63, 93c7fc0)

  • evalboard: Call the UiPath harness Delegate (#69, 0bdac0b)

  • litellm: Cut per-message distribution; turn-level join + per-call audit record (#66, 4131a2a)

  • litellm: Drop the provider field/column — unavailable on the streaming path (#66, 4131a2a)

  • litellm: Stream the cost log + de-duplicate the OpenRouter config comment (#66, 4131a2a)

Testing

  • litellm: Cover config shape, join ordering, and defensive cost branches (#66, 4131a2a)

v0.9.1 (2026-07-29)

Features

  • agents: Extend cooperative early stop to codex and antigravity (b849421)

v0.9.0 (2026-07-28)

Features

  • criteria: Async-primary BaseCriterion contract with CheckerMisuseError escalation (#60)

v0.8.10 (2026-07-24)

Bug Fixes

  • Remove dead SimulationConfig.parallel_trials; add CE031 to guard the class (947acd3)

  • early-stop: Add stop_when 'auto' — per-instance arming + pass-armed-subset stop rule (#51, 08e21e6)

  • early-stop: Defer fail-stop while a pass-armed criterion is undecided (#51, 08e21e6)

  • early-stop: Return assert_never explicitly to satisfy CodeQL (#51, 08e21e6)

  • reports: Code review fixes for the JUnit CI gate (#37, 74db6fa)

  • reports: JUnit CI-gate review fixes + CE027 env-var lint (#37, 74db6fa)

  • reports: Make skipped-task JUnit names platform-independent (#37, 74db6fa)

Chores

  • Re-trigger CI (GitHub dropped the force-push event) (#37, 74db6fa)

  • deps: Lock defusedxml (dev-only, test-side XML parsing) (#37, 74db6fa)

Continuous Integration

  • Disable Docs gh-pages auto-publish on push (Pages not enabled yet) (c289d46)

Documentation

  • 1/8 — add DATASETS.md and a task-schema dataset: section (e3d37ac)

  • 2/8 — retire BYOD.md into DOCKER_ISOLATION.md (1a4a2a5)

  • 3/8 — one complete run_limits reference; document skip (bcb7e24)

  • 4/8 — add DIALOG_MODE.md and correct four stale simulation claims (18dcbe9)

  • 5/8 — fix prompt_mutations example; add CE029 (b524009)

  • 7/8 — generate flat indexes from the mkdocs nav; add CE028 (1514bcb)

  • Add CI Gate reference (GitHub Action + JUnit) and wire into indexes (b8c6301)

  • Agent guides, extending & report-schema references, and fixes (ce74824)

  • Fold nav long tail into one Advanced group; align index ordering (e2ff053)

  • Point docs links to coder-eval.com/docs; drop Ruff badge (60430e5)

  • Point pyproject Documentation URL to coder-eval.com/docs (be39df2)

  • Reword CODER_EVAL_RAW_SDK_LOG to prose form (satisfy CE027) (d7d5b59)

  • Use the brand name "Coder Eval" in prose and titles (821f11b)

Features

  • Packaged CI gate — JUnit XML output + composite GitHub Action (#37, 74db6fa)

  • action: Generic env passthrough + minimum-task-score gate (#37, 74db6fa)

  • ci: 3/3 — publish composite action, release automation, PR dogfood (#37, 74db6fa)

  • cli: 2/3 — wire run --junit-xml and report -f junit (#37, 74db6fa)

  • reports: 1/3 — add reports_junit.py disk-driven JUnit XML writer (#37, 74db6fa)

Testing

  • 6/8 — CE030 documents-or-exempts model fields (c9a3b16)

v0.8.9 (2026-07-23)

Bug Fixes

  • Code review fixes for welch-t-test-exact (#38, 6df6e9b)

  • Render weight:0 criteria as informational on every display surface (#34, 9a34e90)

  • Weight:0 un-gates criteria (informational criteria) (#34, 9a34e90)

  • Weight:0 un-gates criteria and renders as informational (#34, 9a34e90)

  • early-stop: Decide skill activation on the tool call, not its result (#43, d34aa97)

  • early-stop: Latch skill activation on any engagement, not first (#43, d34aa97)

  • evalboard: Match watchlist skeleton header to avoid layout shift (#45, acc1c86)

  • reports: 1/2 — exact Student-t p-values in welch_t_test (#38, 6df6e9b)

  • reports: Exact Student-t p-values and a paired comparison section (#38, 6df6e9b)

  • reports: Fail loud on t* overflow; surface excluded paired tasks (#38, 6df6e9b)

  • reports: One source of truth for variant series and paired stats (#38, 6df6e9b)

  • reports: Validate confidence and n_resamples in bootstrap_mean_ci (#38, 6df6e9b)

Chores

  • harness: Defer two guards from the welch-t-test-exact run (#38, 6df6e9b)

Documentation

  • Add adopter issue template and ADOPTERS.md (#40, bfbed4e)

  • Switch multi-model review from codex to gpt-5 alias (#36, 0a3f2a7)

Features

  • evalboard: Make all pages harness-aware and stream tables (#45, acc1c86)

  • evalboard: Make analytics surfaces harness-aware and stream tables (#45, acc1c86)

  • evalboard: Scope task trends to one harness (#45, acc1c86)

  • reports: 2/2 — add a Paired Comparison section to experiment reports (#38, 6df6e9b)

Refactoring

  • evalboard: Address review nits on harness plumbing (#45, acc1c86)

Testing

  • early-stop: Cover second-review items (two-AgentStart, golden corpus, parity) (#43, d34aa97)

v0.8.8 (2026-07-22)

Bug Fixes

  • codex: Create CODEX_HOME before pinning it in the app-server env (#39, 8d98b91)

Documentation

  • Add Website badge and point PyPI Homepage at coder-eval.com (#41, 0551534)

v0.8.7 (2026-07-22)

Bug Fixes

  • Detect Windows skill paths in telemetry (#24, c57a6b0)

  • agents: Run codex + antigravity harnesses on the tempdir/host path (#33, 04498a9)

  • antigravity: Pin google-antigravity 0.1.7 to load on glibc 2.35 (#33, 04498a9)

  • codex: Always run full-access; drop the in-process OS sandbox (#33, 04498a9)

  • codex: Cover zsh login shells (macOS default) in the mock-PATH home shim (#26, 73f0db2)

  • codex: Keep mock CLIs shadowed in bash login shells (#26, 73f0db2)

  • codex: Make login-shell temp-home lifecycle exception-safe incl. kill_sync (#26, 73f0db2)

  • codex: Restore mock-CLI PATH prepend in login shells via per-task HOME profile (#26, 73f0db2)

  • codex: Restore original HOME inside generated login-shell profile (#26, 73f0db2)

  • codex: Use full-access sandbox on coder_eval-managed tempdir path (#33, 04498a9)

  • deps: Bump pyasn1 to 0.6.4 for GHSA-8ppf-4f7h-5ppj / GHSA-hm4w-wwcw-mr6r (#33, 04498a9)

Chores

  • deps: Bump pyasn1 0.6.3 -> 0.6.4 (GHSA-8ppf-4f7h-5ppj, GHSA-hm4w-wwcw-mr6r) (#32, f1f4f9f)

  • deps: Upgrade agent SDKs to latest (claude 0.2.124, codex 0.144.4) (#33, 04498a9)

Continuous Integration

  • release: Publish a prerelease from a non-main branch (#33, 04498a9)

Documentation

  • Add MkDocs docs site, comparison page, and SEO metadata (#32, f1f4f9f)

  • Address PR review — Coder Eval naming, cleaner table, gh-deploy workflow (#32, f1f4f9f)

  • Address review — git identity for gh-pages, single strict deploy, table glyphs, uv tool install in Tutorial 01 (#32, f1f4f9f)

Refactoring

  • codex: Drop dead sandbox branch + honest full-access messaging (#33, 04498a9)

Testing

  • Accept sandbox_managed kwarg in agent test doubles (#33, 04498a9)

  • codex: Pin start-to-CodexConfig env composition and tidy test imports (#26, 73f0db2)

v0.8.6 (2026-07-20)

Features

  • Add Sonnet 5 + GPT-5.6 pricing; default antigravity to gemini-3.5-flash (#31, 4c40acd)

v0.8.5 (2026-07-20)

Bug Fixes

  • deps: Bump mcp to >=1.28.1 for CVE-2026-52869/52870/59950 (#25, fb1ad4c)

  • orchestration: Isolate per-task config-resolution failures from the suite (#25, fb1ad4c)

  • orchestration: Normalize the all-fail config-resolution abort to ValueError (#25, fb1ad4c)

  • orchestration: Re-raise ValueError verbatim in the all-fail abort (#25, fb1ad4c)

  • orchestrator: Interrupt-proof teardown so a timeout can't drop task.json (#29, 89ec0d0)

  • sandbox: Prune capture-ignored entries on every preservation path; interrupt-proof teardown (#29, 89ec0d0)

Chores

  • evalboard: Address npm Dependabot alerts (#18, 3d5e7b7)

  • evalboard: Batch github-actions bumps into one grouped PR (#18, 3d5e7b7)

Documentation

  • orchestrator: Correct teardown comment to scope of the fix (#29, 89ec0d0)

Features

  • evalboard: Show conversation transcript for simulation tasks (#23, 252722a)

Refactoring

  • orchestrator: Trim to interrupt-proof teardown; drop preservation-prune (#29, 89ec0d0)

v0.8.4 (2026-07-13)

Bug Fixes

  • deps: Upgrade click to 8.4.2 to resolve PYSEC-2026-2132 (#19, a240d4e)

  • deps: Upgrade click to 8.4.2 to resolve PYSEC-2026-2132 (#20, bba8645)

  • evalboard: Address search-box code review comments (#13, 382d80d)

  • evalboard: Fix search bar clear issue (#13, 382d80d)

  • evalboard: Prevent search bar from resetting mid-type (#13, 382d80d)

  • sandbox: Exclude home-dir dotfiles from capture_to artifacts (#19, a240d4e)

  • sandbox: Extend capture_to denylist with credential stores (#19, a240d4e)

Code Style

  • Fix ruff formatting in _WORKSPACE_CAPTURE_IGNORE (#19, a240d4e)

Documentation

  • readme: Clarify framing before hero gif (#16, b7dee1c)

  • readme: Reframe title toward agents & their skills (#16, b7dee1c)

Features

  • early-stop: Opt-in early stop once armed criteria are decided (#14, b0c1ade)

Testing

  • sandbox: Cover home-dir dotfile exclusion in capture_to (#19, a240d4e)

v0.8.3 (2026-07-09)

Bug Fixes

  • evalboard: Fall back to agent_config type when run_config omits harness (#10, 59a240c)

Continuous Integration

  • Remove Azure Artifacts publishing (#11, 2a5124f)

  • Restore auto-bump release (semantic-release + app token) (#12, 0b9d378)

Features

  • evalboard: Add a Harness (RunConfig) column to the runs tables (#10, 59a240c)

  • evalboard: Add Gemini rates to the frontend pricing table (#10, 59a240c)

  • evalboard: Show harness as a vendor logo (internal, main table only) (#10, 59a240c)

  • evalboard: Show the harness (RunConfig) column on the runs tables (#10, 59a240c)

v0.8.2 (2026-07-07)

Bug Fixes

  • antigravity: Apply env_path_prepend so mock CLIs shadow real ones (#487, 412d28f)

  • antigravity: Take harness-spawn lock unconditionally so no-prepend spawns wait out mutated-PATH windows (#487, 412d28f)

  • ci: Restore claude-pr-review git-fetch auth broken by persist-credentials (#485, 155ecb2)

Continuous Integration

  • Address PR #485 review — helper host-scoping, test linkage, doc version drift (#485, 155ecb2)

Documentation

  • Add Docker isolation tutorial (04) + venv activation note (#482, 47b3e0b)

  • Add docker isolation tutorials (#482, 47b3e0b)

  • Optimize README + pyproject for discoverability (#485, 155ecb2)

  • OSS discoverability + packaging metadata, and a claude-pr-review CI fix (#485, 155ecb2)

  • OSS-readiness follow-ups — version drift, positioning, PyPI install (#485, 155ecb2)

  • antigravity: Describe the PATH-mutation window as the full harness context-entry, not just the Popen (#487, 412d28f)

  • tutorials: Add a section about docker instalation (#482, 47b3e0b)

  • tutorials: Changed some texts to make more clear (#482, 47b3e0b)

  • tutorials: Fix links (#482, 47b3e0b)

  • tutorials: Fix stale uipath-credentials claim + review nits (#482, 47b3e0b)

  • tutorials: Rename docker tutorials (#482, 47b3e0b)

Refactoring

  • antigravity: Type the spawn-lock loop global as AbstractEventLoop | None instead of Any (#487, 412d28f)

Testing

  • antigravity: Cover PATH restore when the spawn guard body raises (#487, 412d28f)

v0.8.1 (2026-07-06)

Bug Fixes

  • Address PR #468 review — OSS-prep docs/CI/packaging papercuts (#468, 7e75cd8)

  • Code review fixes for open-source-docs-cleanup (#481, 1810741)

  • Results of Fable code review: discriminated unions, judge retry/ERROR escalation, DEFAULT_* removal, resilience fixes (#483, 6d2564f)

  • agents: Address Antigravity review — loud turn-status conversion + test/lint hardening (#461, 74e4774)

  • agents: Let Antigravity read skill files inside workspace_only sandbox (#461, 74e4774)

  • agents: Normalize Antigravity tool-call params to canonical keys (#461, 74e4774)

  • agents: Silence pyright on optional google-antigravity import (#461, 74e4774)

  • ci: Close token-exfil + comment-injection gaps in claude-pr-review (#472, 9e374de)

  • ci: Harden claude-pr-review with tool + comment allowlists (#472, 9e374de)

  • codex: Apply env_path_prepend to app-server PATH so mock CLIs shadow real ones (#480, 8a53bbe)

  • docker: Real multi-stage runtime kit + drop unused label (review) (#466, 15712a6)

  • errors: 6/6 — categorization fall-through + sandbox-cleanup guard (#483, 6d2564f)

  • evalboard: Collapsed replicate row shows Passed if any replicate passed (#474, caa9cee)

  • evalboard: Colour all failure statuses red, not grey (#474, caa9cee)

  • evalboard: Reconcile window-summary Runs denominator and share passClass (#477, 936731b)

Chores

  • docs: 1/3 — delete internal doc trees, CLA, and feature-doc process (#481, 1810741)

Code Style

  • samples: Restore upstream curly quotes in pddl prompt (#473, 6bc706f)

Documentation

  • 2/3 — purge all dangling references to the deleted docs/features tree (#481, 1810741)

  • Add Contributor License Agreement file and reference it in CONTRIBUTING (#468, 7e75cd8)

  • Address PR #468 review round 2 — remove workflow residue from public tree (#468, 7e75cd8)

  • Address PR #481 review — scrub residues, purge dangling refs (#481, 1810741)

  • Defer type-Literal-default lint candidate from top5-review-fixes run (#483, 6d2564f)

  • Open-source prep — community-health files, docs, and review fixes (#468, 7e75cd8)

  • Open-source prep — delete internal doc trees, purge references, add tutorials 04/05 (#481, 1810741)

  • Sweep stale .env DEFAULT_* references after layer-5 removal (#483, 6d2564f)

  • samples: Add step-by-step run guide to SkillsBench README (#473, 6bc706f)

  • tasks: Lead the tasks README with samples, sentinels below (#473, 6bc706f)

  • tutorials: 3/3 — add 04 Writing a task and 05 Comparing two models (#481, 1810741)

Features

  • agents: Add Antigravity (Gemini) backend via google-antigravity SDK (#461, 74e4774)

  • agents: Wire Antigravity skill discovery via native skills_paths (#461, 74e4774)

  • config: 5/6 — remove the .env DEFAULT_* layer-5 knobs (#483, 6d2564f)

  • docker: Make docker-images — build agent + runtime kit in one command (#466, 15712a6)

  • docker: Relocatable runtime kit for inject-mode tasks (#466, 15712a6)

  • evalboard: K/N ✓ pass-count badge on replicated task rows (#474, caa9cee)

  • evalboard: Per-task pass rate + k/N ✓ badge for replicated runs (#474, caa9cee)

  • evalboard: Per-task pass rate across replicates (#474, caa9cee)

  • evalboard: Window cost + run summary on the front page (#477, 936731b)

  • evaluation: 4/6 — Bedrock judge retry + JudgeInfrastructureError escalation (#483, 6d2564f)

  • lint: 2/6 — CE024 discriminated-union rule + TemplateSource wrap (#483, 6d2564f)

  • models: 1/6 — discriminated SuccessCriterion union + fail-loud validate_registry (#483, 6d2564f)

  • models: 3/6 — extra=forbid on mutation models + CE009 scope extension (#483, 6d2564f)

  • samples: Add vendored SkillsBench sample tasks (#473, 6bc706f)

  • samples: Swap pddl-tpp-planning for court-form-filling (#473, 6bc706f)

Refactoring

  • evalboard: Address PR review — consistent replicate rollup (#474, caa9cee)

  • tasks: Group agent feature-tests under tasks/agents/ + add index (#473, 6bc706f)

Testing

  • agents: Drop redundant module import in Antigravity timeout test (#461, 74e4774)

  • codex: Cover env_path_prepend PATH shadowing in _build_codex_env and start() (#480, 8a53bbe)

  • config: Isolate defaults test from ambient TELEMETRY_ENABLED (#483, 6d2564f)

  • docker: Drift guards + harden runtime kit (review feedback) (#466, 15712a6)

v0.8.0 (2026-07-01)

Bug Fixes

  • ci: Disable telemetry workflow-wide in pr-checks (baked-in default would emit) (#456, fe27c5f)

  • codex: Fall back to full-access sandbox under the docker driver (#459, 371ab28)

  • codex: Run end-to-end under the docker driver (#459, 371ab28)

  • deps: Bump python-socketio/engineio to clear pip-audit CVEs (#456, fe27c5f)

  • docker: Copy ~/.claude symlinks verbatim so a plugin-cache loop can't abort setup (#460, b192ca2)

  • enums: Make FinalStatus.category exhaustive (no silent "failed" fall-through) (#456, fe27c5f)

  • evalboard: Align run-view count + trends with replicate collapse (#458, a9f93e5)

  • pricing: Inline proxy-shim deprecation message to satisfy pyright (#463, 75a07f9)

  • pricing: Keep coder_eval.proxy.pricing as a deprecated alias (#463, 75a07f9)

  • sandbox: Root Windows tempdir sandboxes off the user temp tree (#465, f857c64)

  • telemetry: First-run disclosure notice, README, caller-settable Source, SchemaVersion (#456, fe27c5f)

  • telemetry: Keep docker container silent to avoid double-counted Task.End (#456, fe27c5f)

  • telemetry: Never emit real telemetry from the test suite (#456, fe27c5f)

  • telemetry: One CoderEval.Task.End event per task + Category dim (drop divergent buckets) (#456, fe27c5f)

  • telemetry: Single Task.End event + Category dim, test isolation, exhaustiveness, baked-in default connection string (#456, fe27c5f)

Build System

  • docker: Bake the codex agent into the default image (#459, 371ab28)

Chores

  • Bump socketio/engineio (CVE fixes) + address review feedback (#462, 80991c8)

  • Delete coder_eval.proxy.pricing shim + finish gateway-identifier cleanup (#467, 32e4e63)

  • Finish LLMGW residual cleanup left by the proxy removal (#467, 32e4e63)

  • Finish LLMGW residual cleanup left by the proxy removal (#463) (#467, 32e4e63)

Code Style

  • Fix ruff E501 and invalid noqa warning (#465, f857c64)

  • sandbox: Wrap long mkdtemp dir= line to satisfy formatter (#465, f857c64)

Documentation

  • Keep proxy design docs as historical records (#463, 75a07f9)

  • telemetry: Drop remaining stale Task.End/.Failed references (#456, fe27c5f)

Features

  • BUILD_FAILED status + capture build log for failed image builds (#462, 80991c8)

  • evalboard: Per-replicate task detail with a sticky run selector (#458, a9f93e5)

  • telemetry: Bake in a default App Insights connection string (env overrides) (#456, fe27c5f)

Refactoring

  • Remove the LLM Gateway proxy backend and command (#463, 75a07f9)

v0.7.2 (2026-06-25)

Bug Fixes

  • Address PR #450 review — CE020→CE021 rename, else-clause guard, round-trip + ctor-leak tests (#450, faa91bd)

  • Address PR #451 review findings (CodeQL + type/diagnostic nits) (#451, 5c285d6)

  • Address PR #451 round-2 review findings (CodeQL + Docker test gap) (#451, 5c285d6)

  • Code review fixes for decouple-base-agent-config-sdk-types (#452, 6283346)

  • Code review fixes for malformed-task-json/sim-leak plan (#450, faa91bd)

  • Degrade-not-crash on malformed task.json + simulator scratch-dir leak (#450, faa91bd)

  • Full code review fixes for harness-lint-improvements (#440, 6dd7fae)

  • Harden CE019 scoping per final code review (#451, 5c285d6)

  • High-priority code-review fixes (review 260622-1009) (#439, 3e2d668)

  • PR-gate + review fixes (merge main, finalize hardening) (#440, 6dd7fae)

  • cli: Drop codex+backend rejection — --backend routes the judges too (#440, 6dd7fae)

  • docker: 1/3 — degrade-not-crash on malformed task.json (#450, faa91bd)

  • evalboard: Don't let mature-source scan crash the run page (#455, b5d838d)

  • lint: Full review fix — CE018 also flags list/set membership denylists (#440, 6dd7fae)

  • models: Correct misleading type:none validation suggestion (#440, 6dd7fae)

  • orchestrator: Code review fixes for phase 1 (#439, 3e2d668)

  • reports: Avoid implicit string concat in denominator render line (#453, c2412a7)

  • reports: Re-evaluate thresholds on missing-aggregator path so completion_rate is consistent (#453, c2412a7)

  • reports-html: Phase 3 — _status_badge dispatches on FinalStatus.category (#439, 3e2d668)

  • results: Phase 1 — calculate_weighted_score fails loud on length mismatch (#439, 3e2d668)

  • sampling: Keep --sample-per-stratum nondeterministic by default (#439, 3e2d668)

  • scoring: Phase 7 — fail-loud weighted score + single-sourced all_criteria_passed gate (#440, 6dd7fae)

  • simulation: 2/3 — self-cleaning UserSimulator.start() (no scratch leak) (#450, faa91bd)

  • task-loader: Phase 2 — CLI --sample-per-stratum reproducible-by-default (#439, 3e2d668)

  • telemetry: Code review fixes for phase 1 (#441, 2f84789)

  • telemetry: Code review fixes for phase 4 (#441, 2f84789)

  • telemetry: Full code review fixes — docker per-task events (#441, 2f84789)

  • telemetry: Review fixes + reshape to generic product telemetry (#441, 2f84789)

  • tests: Reset deadline-break scenario state between replays (#451, 5c285d6)

  • types: Phase 1 — promote pyright string-concat + missing-type-arg to error (#440, 6dd7fae)

Chores

  • Harness & lint improvements (code-review 2026-06-22) (#440, 6dd7fae)

  • Remove suite-aggregate-error-row-denominator scratch docs from repo root (#453, c2412a7)

  • claude: Shared axis catalog, cr-workflow scoring fix + 3-way change class (#445, 21b7386)

  • evalboard: Gitignore the coverage/ output dir (#439, 3e2d668)

  • evalboard: Remove deploy plumbing (moving to coder_eval_uipath) (#447, 50f570a)

  • lint: Phase 5 — enable PLR0915/PLR0912 ceiling with tracked debt markers (#440, 6dd7fae)

  • lint: Renumber dialog-loop statement-cap rule CE019 -> CE020 (#451, 5c285d6)

Code Style

  • test: Drop extra blank line left by finalize-test removal (#440, 6dd7fae)

Continuous Integration

  • claude-review: Drop dead cross-repo skills-YAML check from prompt (#444, 6c4bc35)

  • claude-review: Harden PR-review workflow for public open-sourcing (#444, 6c4bc35)

  • claude-review: Harden PR-review workflow for public repo (#444, 6c4bc35)

Documentation

  • Mark suite-aggregate-error-row-denominator plan complete (#453, c2412a7)

  • evalboard: Make public README local-only, drop deploy/Azure details (#447, 50f570a)

  • evalboard: Repoint README deploy section after plumbing move (#447, 50f570a)

  • orchestrator: Full code review fix — clarify finalize score-wrap asymmetry (#439, 3e2d668)

Features

  • cli: Phase 12 — reject codex+bedrock/proxy, document sampling reproducibility (#440, 6dd7fae)

  • cli: Phase 6 — constrain proxy --vendor with click.Choice (#440, 6dd7fae)

  • docker: Mount a lean RW copy of ~/.claude instead of host dir read-only (#446, dfcbf5b)

  • evalboard: Add EVALBOARD_EDITION edition flag (#447, 50f570a)

  • evalboard: Add OSS edition and move deploy plumbing to coder_eval_uipath (#447, 50f570a)

  • evalboard: Clickable mature test links + trends maturity (#455, b5d838d)

  • evalboard: Gate more internal-only surfaces in OSS edition (#447, 50f570a)

  • evalboard: Hide internal-only nav links in OSS edition (#447, 50f570a)

  • evalboard: Mature-link popover + simpler trends maturity (#455, b5d838d)

  • lint: Phase 3 — CE018 no-final-status-name-denylist + dispatch _status_badge on category (#440, 6dd7fae)

  • reports: 1/2 — record ERROR-row exclusion + gateable completion_rate in suite aggregates (#453, c2412a7)

  • reports: Record ERROR-row exclusion & expose gateable completion_rate in suite aggregates (#453, c2412a7)

  • telemetry: Opt-out usage telemetry via OpenTelemetry → Azure App Insights customEvents (#441, 2f84789)

  • telemetry: Phase 1 — telemetry module + config + core deps (#441, 2f84789)

  • telemetry: Phase 2 — emit run-start + task-end/failed events (#441, 2f84789)

  • telemetry: Phase 3 — per-command CoderEval.Cli. events (#441, 2f84789)

  • telemetry: Phase 4 — CE018 lint rule + feature doc (#441, 2f84789)

Refactoring

  • Decompose agent turn-loops + the five remaining god-functions (#451, 5c285d6)

  • agents: Phase 4 — shared finalize/raise/partial-record kernels on Agent (#451, 5c285d6)

  • claude-agent: Phase 2 — extract _ClaudeTurnState + _build_claude_query (#451, 5c285d6)

  • cli: Phase 9 — extract pure heartbeat_is_fresh watchdog predicate (#440, 6dd7fae)

  • codex-agent: Drop dead prev_prompt_tokens + fix post-watchdog timeout state (#451, 5c285d6)

  • codex-agent: Phase 3 — extract _CodexTurnState (#451, 5c285d6)

  • docker: 4/5 — decompose DockerRunner.run into staged helpers (#451, 5c285d6)

  • errors: 1/5 — decompose categorize_error into group classifiers (#451, 5c285d6)

  • errors: Phase 8 — delete dead ErrorCategory members + add liveness contract test (#440, 6dd7fae)

  • models: 1/2 — decouple BaseAgentConfig from claude-code-sdk types (#452, 6283346)

  • models: Decouple BaseAgentConfig from claude-code-sdk types (+ CE020 guard) (#452, 6283346)

  • models: Phase 4 — declare type on BaseSuccessCriterion, drop getattr type-holes (#440, 6dd7fae)

  • orchestrator: 5/5 — decompose _simulation_dialog_loop + add CE019 ratchet (#451, 5c285d6)

  • orchestrator: Phase 2 — drop run_batch wrapper, promote reportImportCycles to error (#440, 6dd7fae)

  • orchestrator: Phase 6 — DRY simulation telemetry via one builder (#439, 3e2d668)

  • reports: 2/5 — decompose generate_markdown into section helpers (#451, 5c285d6)

  • reports: 3/5 — decompose generate_experiment_report into section helpers (#451, 5c285d6)

Testing

  • Address review — guard SettingSource mirror + freeze codex task.json shape (#452, 6283346)

  • Move setting_sources merge-strategy assertion to Claude subclass (#452, 6283346)

  • Resolve CodeQL mixed-import alert on agent_config (#452, 6283346)

  • Stop codex CLI test leaking settings.api_backend into later tests (#440, 6dd7fae)

  • agents: Phase 1 review fix — keep Claude goldens collectable without codex (#451, 5c285d6)

  • agents: Phase 1 — golden-master characterization harness (#451, 5c285d6)

  • agents: Phase 4 review fix — lock _state-before-finalize ordering (#451, 5c285d6)

  • cli: Code review fixes for phase 4 (#439, 3e2d668)

  • cli: Phase 11 — CliRunner coverage for the report command (#440, 6dd7fae)

  • cli: Phase 4 — report CLI tests + extract heartbeat_is_alive helper (#439, 3e2d668)

  • codex: Split SDK-independent unit tests into an ungated file (#451, 5c285d6)

  • evalboard: Cover EVALBOARD_EDITION gate (#447, 50f570a)

  • lint: 2/2 — add CE020 banning SDK-typed BaseAgentConfig fields (#452, 6283346)

  • lint: 3/3 — add CE020 guarding EvaluationResult.model_validate_json (#450, faa91bd)

  • lint: Code review fix for phase 3 — pin CE018 denylist to FinalStatus (#440, 6dd7fae)

  • models: Code review fix for phase 4 — accurate union test name + bogus-tag assertion (#440, 6dd7fae)

  • proxy: Phase 10 — TokenManager._acquire_token HTTP unit test via httpx.MockTransport (#440, 6dd7fae)

  • reports: 2/2 — ERROR-row denominator + completion_rate gate tests (#453, c2412a7)

  • telemetry: Strip ANSI before asserting on --help output (#441, 2f84789)

v0.7.1 (2026-06-22)

Bug Fixes

  • cli: Harden aggregate recovery path and disclose rebuild scope (#438, 5f32a13)

  • deps: Bump pydantic-settings and msgpack to patch CVEs (#435, 24a0e06)

  • evalboard: Date and sort ad-hoc runs by run start_time (#432, f058ea8)

  • evalboard: Exclude mature-skipped tasks from run-page cost/duration metrics (#427, 18f303c)

Build System

  • Sync uv.lock after dropping pylint/radon scoring deps (#430, 1abf449)

Chores

  • Relicense from MIT to Apache 2.0 with UiPath copyright notice (#435, 24a0e06)

Features

  • Restore HTML report generation and the report command (#430, 1abf449)

  • cli: Add aggregate to rebuild run.json/run.md from task.json (#438, 5f32a13)

  • cli: Add summarize to rebuild run.json/run.md from task.json (#438, 5f32a13)

  • evalboard: Cap ad-hoc runs with a show-all toggle (#432, f058ea8)

  • evalboard: Date, sort, and paginate the ad-hoc runs section (#432, f058ea8)

  • evalboard: Mark skipped-mature tasks with a non-clickable badge (#427, 18f303c)

  • evalboard: Show mature tasks as green passes, keep trend averages honest (#427, 18f303c)

Refactoring

  • Remove dormant features ahead of open-sourcing (#430, 1abf449)

  • Remove dormant features ahead of open-sourcing (#422) (#430, 1abf449)

  • Scrub stale references to removed features (#430, 1abf449)

  • cli: Rename summarize command to aggregate (#438, 5f32a13)

Testing

  • Isolate version-info tests from the host UiPath CLI (#430, 1abf449)

v0.7.0 (2026-06-16)

Bug Fixes

  • Address pr:424 review nits (pricing seam + error-category contract) (#424, b24010e)

  • Code review findings for pricing seam + error category (#424, b24010e)

  • Full code review fixes for delegate-sdk base primitives (phases 1-3) (#424, b24010e)

  • errors: Code review fixes for phase 1 (#424, b24010e)

Documentation

  • Correct the is-not-None lookup rationale (review follow-up) (#424, b24010e)

  • Surface register_pricing extension point + BYOA worked example (#424, b24010e)

  • pricing: Add feature spec for the pricing registration seam (#424, b24010e)

Features

  • Delegate-sdk base primitives — AgentConfigError + pricing registration seam (#424, b24010e)

  • errors: Add generic AgentConfigError + AGENT_CONFIG_ERROR category (#424, b24010e)

  • pricing: Add register_pricing seam for plugin-contributed model rates (#424, b24010e)

v0.6.2 (2026-06-16)

Build System

  • deps: Bump aiohttp, cryptography, python-multipart, starlette for CVEs (#426, 633628c)

Continuous Integration

  • release: Drop dry-run input; dispatch always cuts a release (#426, 633628c)

  • release: Make releases manual via dispatch with dry-run preview (#426, 633628c)

  • release: Manual dispatch-triggered release + versioned agent image (#426, 633628c)

  • release: Pick bump level at dispatch, not from commit messages (#426, 633628c)

  • release: Publish versioned agent image from the release job (#426, 633628c)

v0.6.1 (2026-06-15)

Bug Fixes

  • docker: Pin Claude Code CLI version (was @latest) (#425, ae47e87)

Continuous Integration

  • Remove validate-skills-yamls job from PR checks (#413, c205207)

Documentation

  • docker: Trim historical narrative from Claude Code pin comment (#425, ae47e87)

v0.6.0 (2026-06-15)

Chores

  • Drop leftover eval-runner config refs from coder_eval (#419, b862794)

Continuous Integration

Features

  • review: Add workflow-orchestrated code-review skill + harden the full review command (#423, ac7b2e0)

Refactoring

  • Move dashboard CI to coder_eval_uipath (as eval-runner) (#419, b862794)

  • Remove dashboard CI (moved to coder_eval_uipath as eval-runner) (#419, b862794)

v0.5.0 (2026-06-12)

Bug Fixes

  • agents: Address PR #416 review feedback (#416, df9ac29)

  • agents: Break plugins<->registry import cycle (CodeQL) (#416, df9ac29)

  • agents: Code review fixes for phase 1 (#416, df9ac29)

  • agents: Code review fixes for phase 2 — uniform ResolvedAgentConfig (#416, df9ac29)

  • agents: Code review fixes for phase 3 (#416, df9ac29)

  • agents: Full code review fixes for BYOA SPI (#416, df9ac29)

  • agents: Guard duplicate plugin-kind registration + review polish (#416, df9ac29)

  • ci: Byoa fixture plugin uses hatchling only-include, not setuptools py-modules (#416, df9ac29)

Features

  • agents: Bring-your-own-agent (BYOA) plugin SPI (#416, df9ac29)

  • agents: Phase 1 — entry-point plugin discovery + string-keyed registry (#416, df9ac29)

  • agents: Phase 2 — registry-driven agent config dispatch (#416, df9ac29)

  • agents: Phase 3 — worked BYOA plugin, live test + CI, docs (#416, df9ac29)

v0.4.0 (2026-06-12)

Bug Fixes

  • docker: Pin framework entrypoint via docker run --entrypoint (#418, a86ad2b)

  • docker: Restore actionable runtime-image guard; reconcile docs (#418, a86ad2b)

Features

  • docker: Pin framework entrypoint via docker run --entrypoint (#418, a86ad2b)

Refactoring

  • docker: Coder-eval-specific entrypoint, drop baked ENTRYPOINT (#418, a86ad2b)

v0.3.0 (2026-06-12)

Chores

Continuous Integration

  • Add --strict so no-op release exits non-zero (#417, 4f6876b)

  • Add conventional commits checker for PR titles and commits (#406, ff73e9a)

  • Configure git identity before amending release commit (#420, 37923e8)

  • Fix false-positive release on non-releasable commits (#417, 4f6876b)

  • Fix orphaned tag after uv.lock amend (#415, fdce101)

  • Regenerate uv.lock in the release commit (#415, fdce101)

Features

  • Skill-activation nightly — stratified sampling, per-skill recall, evalboard view (#391, ecc0403)

  • activation: Compute activation score + add Slack line (#391, ecc0403)

  • activation: Enrich case rows + rework the evalboard activation views (#391, ecc0403)

  • activation: Keep activation rows out of run-level metrics (#391, ecc0403)

  • activation: Merge skills + activation into one nightly run (#391, ecc0403)

  • activation: Per-skill recall aggregation + evalboard view (#391, ecc0403)

  • activation: Run 20/skill nightly via --sample-per-stratum (#391, ecc0403)

  • activation: Run as a nested sub-run instead of merging into run.json (#391, ecc0403)

  • activation: Surface activation in the evalboard (front page, run card, dedicated page) (#391, ecc0403)

  • dataset: Stratified random sampling; make --sample random (#391, ecc0403)

  • evalboard: Hide activation rows from run view by default (#391, ecc0403)

v0.2.1 (2026-06-11)

Bug Fixes

  • Pin BYOD smoke template to coder-eval-agent:latest (#401, 6405391)

  • Sync uv.lock with the 0.2.0 version bump (#401, 6405391)

Refactoring

  • Remove UiPath eval content moved to coder-eval-uipath (#401, 6405391)

  • Sweep remaining references to moved eval content (#401, 6405391)

v0.2.0 (2026-06-11)

Bug Fixes

  • sandbox: Address PR review — restore -p alias, close test gaps, doc deltas (#398, cbcfb85)

  • sandbox: Clear stale artifacts on resume re-run; drop CLI aliases; review polish (#398, cbcfb85)

Features

  • sandbox: Explicit --preservation-mode (NONE/MOVE_ON_WRITE/DIRECT_WRITE) (#398, cbcfb85)

Testing

  • Replace tautological evaluate-mapping test (CodeQL constant-in-conditional) (#398, cbcfb85)

v0.1.0 (2026-06-10)

  • Initial Release