-
Code review fixes for container-contract-and-command-surface (#178,
2f01c9d) -
docs: Wrap an over-long line in the CE063 docstring (#180,
a249877) -
harbor: Allow touch's absolute-path form in trajectory_criteria (#187,
3addf46) -
harbor: Reject --resume with --workspace-dir, sync design notes (#186,
c9b96a8) -
harbor: Remove the Write/content ambiguity from trajectory_criteria (#187,
3addf46) -
harbor: Require Harbor E2E on PRs, de-flake trajectory_criteria fixture (#187,
3addf46) -
harbor: Resolve workdir dynamically at run time instead of guessing at export (#182,
7af839b) -
harbor: Trim packager docstrings under the prose budget (#182,
7af839b) -
lint: Fail the prose gate cleanly on a missing root, and name the blank-line blind spot (#180,
a249877) -
lint: Include untracked files in the prose-only proof (#180,
a249877) -
lint: Make the prose-only proof see added directives and reject typos (#180,
a249877) -
lint: Point main's exempt-pair test at the repo root (#180,
a249877) -
sandbox: Close the adopt() installer-disclosure gap and fix the review's other findings (#187,
3addf46) -
sandbox: Reprovision env_packages that a workspace capture stripped (#187,
3addf46) -
sandbox: Trim adopt()'s re-provisioning comment under the prose budget cap (#187,
3addf46)
-
Apply main's prose rules to the container-contract branch (#178,
2f01c9d) -
Record two prose-budget guard gaps the final review found (#180,
a249877) -
harbor: Trim two docstrings over the prose budget (#186,
c9b96a8) -
lint: Answer review — cut runs off the cap, drop lint-rules.md overlap (#180,
a249877) -
lint: Slim the five essays main's reports split brought in (#180,
a249877) -
tests: 2/8 — move the CE catalogue from notes/README.md to lint-rules.md (#180,
a249877) -
tests: 3/8 — move lint-rule defect stories into notes/lint-rules.md (#180,
a249877) -
tests: 4/8 — move doc-surface rule rationale into notes/lint-rules.md (#180,
a249877) -
tests: 5/8 — move golden-sensor and bracket-clock rationale into notes (#180,
a249877) -
tests: 6/8 — move plain test-module rationale into the subsystem notes (#180,
a249877) -
tests: 7/8 — delete HISTORY prose from tests, fix dangling docstring citations (#180,
a249877)
-
Give the host→container boundary a contract and fold aggregate into report --rebuild (#178,
2f01c9d) -
cli: 4/6 — replace
aggregatewithreport --rebuild(#178,2f01c9d) -
evaluate: 5/6 — refresh the run-level run.json after a detached grade (#178,
2f01c9d) -
harbor: Unblock trajectory-dependent criteria and always allow credentials at export (#186,
c9b96a8) -
harbor: Unblock trajectory-dependent criteria, always allow credentials at export (#186,
c9b96a8) -
isolation: 1/6 — parse context.json as a ContainerContext contract (#178,
2f01c9d) -
isolation: 2/6 — refuse image skew and assert the container echoed its contract (#178,
2f01c9d) -
lint: 1/8 — scan a _ROOTS tuple in the prose budget (#180,
a249877) -
lint: 8/8 — turn the prose budget gate on for tests/ (#180,
a249877) -
lint: Cap the comment RUN, replacing the per-file comment budget (#180,
a249877) -
lint: Gate tests/ on the prose rules, and cap the comment run as well as the file total (#180,
a249877) -
lint: Keep the file-total comment budget as a backstop under the run cap (#180,
a249877)
-
isolation: 3/6 — move the docker→tempdir driver rewrite host-side (#178,
2f01c9d) -
lint: Delete CE023, which guards a package that no longer exists (#180,
a249877)
-
Guard the contract echo over a maximal sandbox and the Typer-command exemption list (#178,
2f01c9d) -
cli: 6/6 — pin every shared run/execute flag as declared identically (#178,
2f01c9d) -
harbor: Dump agent-phase recorded commands on scenario failure (#187,
3addf46)
-
Code review fixes for antigravity-message-id (#165,
6166b1e) -
Code review fixes for the reports consolidation (#176,
396c22c) -
Code review fixes for the reports-consolidation review fixes (#176,
396c22c) -
Code review fixes for timing-architecture-standardization (#165,
6166b1e) -
Code review fixes for turn-timing-consolidation (#165,
6166b1e) -
Reconcile the reports split with main's docs restructure (#176,
396c22c) -
Review pass B — the tool bucket's dash survives the language boundary (#165,
6166b1e) -
antigravity: 1/3 — give every generation a message_id (#165,
6166b1e) -
claude-code: Subtract tool execution from the generation windows (#165,
6166b1e) -
docs: Restore contracts the prose refactor lost or misstated (#177,
30f78f4) -
docs: Restore the qualifier that made each absolute true (#177,
30f78f4) -
docs: Stop asserting a tmpfs mask that no longer exists (#177,
30f78f4) -
evalboard: 3/8 — an unbounded tool call contributes to no bucket (#165,
6166b1e) -
lint: 1/4 — is_core_path covers the whole core layer (#176,
396c22c) -
lint: 2/4 — declare CE044 and CE065 to ruff, and pin the id space (#176,
396c22c) -
lint: CE004 checks reports/ — stop borrowing CE066's exemption set (#176,
396c22c) -
lint: CE004 never fired on the relative import spelling (#176,
396c22c) -
lint: Let CE045 globs reach .claude/*.md surfaces (#174,
9af3db3) -
lint: Resolve a relative import against the importing file (#176,
396c22c) -
lint: Stop the prose budget penalising usage examples (#177,
30f78f4) -
pricing: Carry the exemption set into the generated mirror (#176,
396c22c) -
timing: 2/8 — a published window must match its own bounds (#165,
6166b1e) -
timing: 3/6 — a tool that closes between two windows is not model time (#165,
6166b1e) -
timing: 3/7 — a naive/aware mix names the pair that disagreed (#165,
6166b1e) -
timing: 5/6 — one clock basis per turn on antigravity and pi (#165,
6166b1e) -
timing: A TurnClock for claude-code, and pi's two captured defects (#165,
6166b1e) -
timing: Bracket the head and tail on the main thread only (#165,
6166b1e) -
timing: Claude-code's windows tile across a tool result (#165,
6166b1e) -
timing: Setup_ms marks from the task's start, not from _setup() (#165,
6166b1e) -
timing: Stamp the turn bracket off the turn clock (CE064) (#165,
6166b1e) -
timing: The bracket fixture's anchor cannot expire, and the tail is checked (#165,
6166b1e)
-
Record two deferred harness candidates from the reports consolidation (#176,
396c22c) -
timing: Decompose_run's dead datetime import (#165,
6166b1e)
-
2/7 — move timing and permissions rationale into .claude/notes (#177,
30f78f4) -
3/7 — move agent-adapter rationale into .claude/notes (#177,
30f78f4) -
4/4 — retarget prose that names deleted report modules (#176,
396c22c) -
4/7 — move orchestration rationale into .claude/notes (#177,
30f78f4) -
5/5 — retarget every reports_* reference and record the rationale (#176,
396c22c) -
5/7 — move isolation, sandbox and CLI rationale into .claude/notes (#177,
30f78f4) -
6/7 — move criteria, routing and judging rationale into .claude/notes (#177,
30f78f4) -
7/7 — move reporting, harbor and telemetry rationale into .claude/notes (#177,
30f78f4) -
Move design rationale out of src/ into .claude/notes, and gate it (#177,
30f78f4) -
Record the four harness gaps the prose run did not close (#177,
30f78f4) -
Restore the self-containment clause on reports_html (#177,
30f78f4) -
Slim CLAUDE.md and split long-form rationale into .claude (#174,
9af3db3) -
Slim CLAUDE.md, and drop a directory that never existed (#177,
30f78f4) -
Update CLAUDE.md with communication style and development command clarifications (#177,
30f78f4) -
harbor: Apply the branch's prose rules to the bind-mount rewrite (#177,
30f78f4) -
harness: 3/3 — message_id is what splits the timeline (#165,
6166b1e) -
harness: 4/8 — say what each harness's clock basis actually is (#165,
6166b1e) -
harness: 6/6 — the timing architecture as it now stands (#165,
6166b1e) -
harness: Claude-code's generation windows are not tool-subtracted (#165,
6166b1e) -
harness: Record three guards the head/tail review could not close (#165,
6166b1e) -
harness: Register that the golden corpus cannot see a timing value move (#165,
6166b1e) -
harness: Register the message_id gaps the final review surfaced (#165,
6166b1e) -
harness: Register what the turn-timing run could not guard (#165,
6166b1e) -
harness: Widen the measured head/tail figures to six turns per harness (#165,
6166b1e) -
notes: Cut the duplicated catalogue and the unbuilt design (#177,
30f78f4) -
timing: Every harness subtracts tool time now, not two (#165,
6166b1e)
-
evalboard: 3/4 — name the harness head and tail in the timeline strip (#165,
6166b1e) -
harbor: Bind-mount task.yaml/plugins/templates/extra_mounts instead of COPY, skip Dockerfile when unneeded, translate pre_run (#175,
7bf8966) -
lint: 1/7 — prose budget ratchet and one home for rationale (#177,
30f78f4) -
lint: CE067 — CLAUDE.md's tree must name every top-level package member (#176,
396c22c) -
lint: Replace the prose baseline with two self-adjusting rules (#177,
30f78f4) -
pricing: 1/5 — generate the evalboard rate table from pricing.py (#176,
396c22c) -
reports: 6-7/7 — the offline report carries the buckets; a TS None-vs-0 guard (#165,
6166b1e) -
reports: 7/8 — the buckets reach every surface through one function (#165,
6166b1e) -
timing: 1/6 — a two-sided residual gate for the four-bucket identity (#165,
6166b1e) -
timing: 1/7 — a committed, ms-exact magnitude sensor (#165,
6166b1e) -
timing: 4/4 — assert the buckets in replay, and record what they contain (#165,
6166b1e) -
timing: 4/7 — one meaning for harness_startup_ms, on all five (#165,
6166b1e) -
timing: Book each turn's head and tail as their own buckets (#165,
6166b1e) -
timing: Name the setup and grading phases; union a row's tool time (#165,
6166b1e)
-
reports: 2/5 — two DRY fixes and hoist 18 function-local imports (#176,
396c22c) -
reports: 3/5 — split reports_stats.py into stats, result_metrics and helpers (#176,
396c22c) -
reports: 4/5 — reports/ package, durations.py, and CE066 (#176,
396c22c) -
reports: Generate the pricing mirror and split the reports layer (#176,
396c22c) -
timing: 2/6 — one close_window() for the tiling reducers (#165,
6166b1e) -
timing: 5/7 — one tool-subtraction, at the collector seam (#165,
6166b1e) -
timing: 5/8 — one home for the rules, and a stored tool union (#165,
6166b1e) -
timing: 6/8 — the two sensors share one selection rule (#165,
6166b1e)
-
harness: 2/7 — OpenCode and Pi get a corpus worth replaying (#165,
6166b1e) -
harness: No test may read the pinned timing corpus (#165,
6166b1e) -
harness: Pin why claude-code's zero head is left as a clamp (#165,
6166b1e) -
harness: Stamp the two codex fixtures that timed themselves with now() (#165,
6166b1e) -
lint: 2/3 — CE060, an AssistantMessage must declare its message_id (#165,
6166b1e) -
lint: 2/4 — widen CE058 to the turn head/tail buckets (#165,
6166b1e) -
lint: 4/6 — CE061, a window must come from the shared helper (#165,
6166b1e) -
lint: CE058 form 6 — the zero its guard does not vouch for (#165,
6166b1e) -
lint: Drop a personal path and de-duplicate the isolation pin (#176,
396c22c) -
lint: Drop two function-local json imports that shadow the module one (#176,
396c22c) -
lint: Guard the Rationale pointer's placement, not just its target (#177,
30f78f4) -
reports: 3/4 — cover the HTML slowest-commands truncation branch (#176,
396c22c) -
timing: 1/8 — the bracket's SOURCE, not just its presence (#165,
6166b1e) -
timing: Unify the fixture clocks and assert the four-bucket identity (#165,
6166b1e)
-
Code review fixes for record-cli-review-fixes (#150,
3aca8e8) -
Resolve py/import-and-import-from CodeQL alerts in test_regrade.py + test_detached_grading_boundaries.py (#161,
6ab6b87) -
criteria: 2/4 — a shim rule fault is an eval-config error, not an agent failure (#150,
3aca8e8) -
evaluate: Close the PR review findings on container-based detached grading (#161,
6ab6b87) -
evaluate: Close the review blockers on container-based detached grading (#161,
6ab6b87) -
harbor: Address code review findings from full 8-axis review (#166,
9caffcb) -
harbor: Address PR review blockers (score-integrity, security, drops) (#166,
9caffcb) -
harbor: Derive a prebuilt image's real WORKDIR instead of guessing /app (#166,
9caffcb) -
harbor: Suppress pyright reportMissingImports for the harbor package (#166,
9caffcb) -
models: Reference FlagPredicate at runtime so the cast is a real use (#150,
3aca8e8) -
record_cli: Close three real holes the Copilot review found (#150,
3aca8e8) -
record_cli: Reject an unusable response rule at load, and trace a shim fault (#150,
3aca8e8) -
sandbox: Give the sandbox venv system site packages (#162,
e7d8326) -
test: Strip ANSI color before asserting --workspace-dir in CLI output (#166,
9caffcb) -
tests: Harden harbor CLI-error assertions and Windows chmod checks (#166,
9caffcb) -
timing: Account for generation and tool time on every harness (#164,
87102a3)
-
evaluate: Grade a
driver: dockerrow inside a container of its own image (#161,6ab6b87) -
harbor: Export coder-eval tasks to Harbor, with coder-eval as the grader (#166,
9caffcb) -
harbor: Export experiment.yaml variants to Harbor task directories (#166,
9caffcb) -
harbor: Fix agent workspace alignment, env passthrough, and template sources (#166,
9caffcb) -
harbor: Task.yaml/experiment.yaml → Harbor export + coder-eval as a Harbor agent (#166,
9caffcb) -
record_cli: Serve a different canned response per invocation (#150,
3aca8e8)
-
Close three harness gaps this review surfaced (#150,
3aca8e8) -
lint: 4/4 — put the record_cli authoring surface under CE030 doc parity (#150,
3aca8e8) -
tasks: 3/4 — add the record_cli per-invocation-response probe (#150,
3aca8e8)
-
container: Arm the heartbeat watchdog only inside the container (#154,
0fe8cb0) -
criteria: Refuse an out-of-sandbox criterion path instead of scoring it 0.0 (#154,
0fe8cb0) -
eval: Address the medium and low findings from the branch review (#154,
0fe8cb0) -
eval: Close the verdict-correctness gaps in detached grading (#154,
0fe8cb0) -
evalboard: Show per-row cost for open-weight harnesses (Pi) (#159,
57c9e33) -
execute: Close the verdict-changing and trust-boundary defects in detached grading (#154,
0fe8cb0) -
execute: Close the verdict-divergence, trust-gate and fabricated-rate defects (#154,
0fe8cb0) -
execute: Move post_run to the grading phase and stop the record lying about its driver (#154,
0fe8cb0) -
pi: Address bai-uipath review — apportionment caveat + token/telemetry minors (#159,
57c9e33) -
pi: Do not forward allowed_tools/disallowed_tools to Pi (#159,
57c9e33) -
pi: Multi-model review — gate error-crash on intentional cuts + dead dup + doc/guard nits (#159,
57c9e33) -
pi: Review blockers — session-id sanitize, error crash, telemetry warn, tool map, CE047 (#159,
57c9e33) -
security: Close the three CodeQL findings on the detached-grading diff (#154,
0fe8cb0) -
tasks: Make the two remaining absolute criterion paths reachable, and add CE055 (#154,
0fe8cb0) -
tests: Kill the CodeQL taint source and the Windows mode assertion (#154,
0fe8cb0)
-
agents: 4/4 — add Pi harness docs + enumeration-surface parity (#159,
57c9e33) -
docker: Fix env_passthrough model name + list Pi in baked-toolchain docs (#159,
57c9e33) -
pi: Correct plugins->--skill support + OPENROUTER passthrough claims (#159,
57c9e33) -
pi: Fix tool-enforcement claims + cost docstring after review (#159,
57c9e33)
-
agents: 2/4 — add PiAgent harness (pi --mode json) + registration (#159,
57c9e33) -
cli:
coder-eval execute+ detached grading viaevaluate <run_dir>(#154,0fe8cb0) -
cli: Add
coder-eval execute— run tasks without grading them (#154,0fe8cb0) -
cli: Grade an executed run afterwards —
evaluate <run_dir>+Sandbox.adopt(#154,0fe8cb0) -
cli: Make --resume distinguish "executed" from "graded" (#154,
0fe8cb0) -
docker: 1/4 — bake the pinned Pi CLI into docker/Dockerfile (#159,
57c9e33) -
evalboard: Show the Pi logo + "Pi" in the harness view (#159,
57c9e33) -
models: 1/4 — add AgentKind.PI + PiAgentConfig model (#159,
57c9e33) -
pi: Add Pi harness (
--type pi) with docker + skill injection (#159,57c9e33) -
pi: Load agent.plugins skills via --skill + harden turn/token handling (#159,
57c9e33) -
sandbox: 2/4 — forward OPENROUTER_API_KEY into docker containers (#159,
57c9e33)
-
Fix two CI-only failures in the detached-grading tests (#154,
0fe8cb0) -
docker: 4/4 — guard that the Dockerfile bakes a pinned Pi CLI (#159,
57c9e33) -
pi: Assign the awaited cancel result to silence CodeQL 'no effect' (#159,
57c9e33) -
pi: Install shutil.which patch so the env-info test doesn't need the real CLI (#159,
57c9e33) -
pi: Port OpenCode teardown + cost-fallback matrices (blockers 3, 4) (#159,
57c9e33) -
pi: Scrub personal scratchpad path from happy-stream fixture (#159,
57c9e33)
-
deps: Bump google-antigravity 0.1.7 -> 0.1.8 (Defender FP on the harness) (#158,
48a4d53) -
deps: Bump google-antigravity to 0.1.8 to clear a Defender false positive on the bundled harness [PILOT-7463] (#158,
48a4d53) -
evalboard: Always show every known harness in the filter (#152,
5e7d2b6) -
evalboard: Stop counting one model as two in the run header (#152,
5e7d2b6) -
pricing: Refresh the rate card and add gemini 3.7/3.8 Flash (#155,
be98f9d)
-
Add intro video and make the README agent-agnostic (#157,
d715f85) -
Widen the framing past skills-only and guard the agent roster with CE047 (#157,
d715f85) -
deps: Drop the Defender rationale from the antigravity pin comment (#158,
48a4d53) -
evalboard: Trim the comments added by this branch (#152,
5e7d2b6) -
pricing: Trim the rate-card comments to what affects an edit (#155,
be98f9d) -
readme: Add intro video and make the README agent-agnostic (#157,
d715f85) -
readme: Play the intro video inline, with YouTube as the fallback (#157,
d715f85) -
readme: Tell viewers to unmute the inline video (#157,
d715f85) -
stub: Make the Pages stub and package metadata agent-agnostic (#157,
d715f85) -
tutorial: Stop sending first-timers through the contributor toolchain (#157,
d715f85)
-
evalboard: Cut page load time and show loading state while pages load (#152,
5e7d2b6) -
evalboard: Move repeated per-row utility classes into the stylesheet (#152,
5e7d2b6) -
evalboard: Stop re-reading the whole run store on every render (#152,
5e7d2b6)
-
action: Drop the literal expression syntax from a run-step comment (#147,
7b81456) -
action: Stop the env passthrough mutating the action's own shell (#147,
7b81456) -
docs: Drop the last references to inputs that no longer exist (#147,
7b81456) -
evalboard: Even out the variant tiles and keep a task's arms adjacent (#145,
6780227) -
evalboard: Resolve a variant-less task link to the task's first arm (#145,
6780227)
-
action: Trim the comment bulk on the eight-input surface (#147,
7b81456) -
evalboard: Trim variant commentary and redundant tests (#145,
6780227)
-
action: Add working-directory, extras, extra-packages, prerelease and args inputs (#147,
7b81456) -
action: Refactor github action input surface around args; harden env passthrough (#147,
7b81456) -
evalboard: Register an unlisted gha source for ad-hoc dispatch runs (#147,
7b81456) -
evalboard: Render multi-variant runs, one row per (task, arm) (#145,
6780227)
-
action: Drop every forwarding input; flags go through args (#147,
7b81456) -
evalboard: Drop the variant colour palette and the spread helper (#145,
6780227) -
evalboard: Report pass rate per arm instead of alongside a blended one (#145,
6780227)
-
Skip the enforcement live tests when the agent declines to try (#146,
d679ff2) -
action: Record stub argv from bash so Windows argv is byte-exact (#147,
7b81456) -
action: Stop Git Bash rewriting the absolute path in the install-spec tests (#147,
7b81456)
-
opencode: Bound the post-EOF reap, reap the whole process group, test every failure path (#115,
5532b87) -
opencode: Close the remaining review findings on the harness (#115,
5532b87) -
opencode: Gate on captured tokens, canonicalize tool args, pin max_turns (#115,
5532b87) -
opencode: Inject plugin skills so skill suites measure the skills (#115,
5532b87) -
opencode: Keep the smoke task out of the CI smoke-pass bucket (#115,
5532b87) -
opencode: Map apply_patch to Write so GPT-family edits are seen by criteria (#115,
5532b87) -
opencode: Reap the CLI on every turn exit, test the sandbox env contract (#115,
5532b87) -
opencode: Spread super() in get_environment_info; guard with CE046 (#115,
5532b87) -
opencode: Typecheck on Windows, satisfy both CodeQL findings (#115,
5532b87) -
plugin: Close the review's verified gaps — regex blind spot, parallel surfaces, runtime guard (#143,
c565ebb) -
plugin: Correct the skill-reachability path — every generated activation suite reports recall 0.0 (#143,
c565ebb) -
plugin: Correct the skill-reachability path, and reuse PR #109's measured descriptions (#143,
c565ebb) -
routing: Decouple simulator route from checker_context.api_route (#144,
88ff0f0) -
routing: Reinstate litellm+agent_judge rejection guard (#144,
88ff0f0) -
routing: Restore LiteLLM->Claude pin on simulator_route (#144,
88ff0f0) -
test: Add CE045, document the plugin-path divergence, unpin a test from ordering (#143,
c565ebb) -
utils: Use explicit concatenation in the plugin-root warning (#143,
c565ebb)
-
opencode: Standardize on deepseek-v4-pro, drop the flash-0731 rate entry (#115,
5532b87) -
plugin: Answer "skills only?", fold tags into keywords, add CE044 (#141,
2ae9b7b) -
plugin: Give both author objects the same contact address (#141,
2ae9b7b) -
plugin: Give the marketplace owner a contact address (#141,
2ae9b7b) -
plugin: Lead both manifests with skill evaluation, add discovery metadata (#141,
2ae9b7b) -
plugin: Pin both manifests to their published JSON schemas (#141,
2ae9b7b)
-
agents: Add OpenCode harness with opt-in [opencode] extra (#115,
5532b87) -
opencode: Add require_token_telemetry, an escape hatch for the zero-token guard (#115,
5532b87)
-
checker-context: Typed model, reject litellm+agent_judge/sim, live test (#137,
727bb7b) -
eval-routing: Restore DEFAULT_JUDGE_MODEL floor + add litellm judge transport (#137,
727bb7b) -
evalboard: Exclude carried-forward passes from the wall-clock aggregates (#125,
dd918e6) -
release: Pin python-semantic-release + GitPython to unbreak version bump (#139,
ef74014) -
routing: Make _resolve_backend_route's match exhaustive (#137,
727bb7b)
-
Retrigger checks (GH Actions appeared stalled repo-wide) (#137,
727bb7b) -
Retrigger checks (previous push did not trigger CI) (#137,
727bb7b) -
fix: Install litellm extra for pyright, address CodeQL findings (#137,
727bb7b)
-
eval-routing: Decouple judge/agent_judge backend+model from the agent's own route (#137,
727bb7b) -
eval-routing: Decouple judge/agent_judge backend+model from the agent's route (#137,
727bb7b) -
evalboard: Chart seconds per passed task on the overview (#125,
dd918e6) -
evalboard: Read the time ratio without hovering (#125,
dd918e6) -
evalboard: Replace the turn-budget signal with time per passed task (#125,
dd918e6) -
evalboard: Run the wall-clock signal beside the turn budget, behind tabs (#125,
dd918e6) -
litellm-judge: Support arbitrary litellm kwargs via params/auth (#137,
727bb7b)
-
docker: Expand ~ and $VAR in extra_mounts destinations (#128,
02dbf69) -
docker: Reject expansions that inject a ':' into a mount path (#128,
02dbf69)
-
ci: Pin pip floor to 26.2 to close PYSEC-2026-3721 (#133,
beecedd) -
codex: Store command output whole in result_summary (+ code-review fixes) (#127,
3d09f0f) -
deps: Address PR #133 review — anthropic 1.0.0 dropped temperature kwarg (#133,
beecedd) -
deps: Bump anthropic to 1.0.0 and migrate Bedrock judge path to httpx2 (#133,
beecedd) -
deps: Bump anthropic to 1.0.0, migrate Bedrock judge path to httpx2 (#133,
beecedd) -
deps: Re-lock pip to 26.2.1 to actually close PYSEC-2026-3721 (#133,
beecedd)
-
docker: Restore container access after the DAC cap drop, shield the task dir (#106,
56bffad) -
reference: Address code review and CodeQL findings (#106,
56bffad) -
reference: Address PR review — scoring correctness, fail-closed anti-cheat (#106,
56bffad)
-
early-stop: CE036 enforces the live_verdict determinism + monotonicity contract (#126,
d854004) -
reference: Skip host-side chmod assertions on Windows (#106,
56bffad)
-
Code review fixes for the Claude Code plugin marketplace (#82,
02e5151) -
Code review fixes for the plugin generic-adopter plan (#82,
02e5151) -
Code review fixes for the plugin-audit P0/P1 plan (#82,
02e5151) -
Derive the poll loop's exit bound from the turn's actual timeout (#111,
d3f1432) -
Make lint-tasks' read-only rule outlive the frontmatter deny (#82,
02e5151) -
Reconcile the plugin branch with main after rebase (#82,
02e5151) -
Remove redundant asyncio re-import flagged by CodeQL (#111,
d3f1432) -
agent: Address PR review — simulator replace mode, preset-aware reports, replace validator (#92,
fddd3c5) -
agent: Address PR review — unconditional preset, judge replace seam (#92,
fddd3c5) -
agent: Align Codex system_prompt with the append-only contract (#92,
fddd3c5) -
agent: Append system_prompt to the Claude Code preset instead of replacing (#92,
fddd3c5) -
agent: Reject a blank system_prompt_file under replace mode (#92,
fddd3c5) -
agent: Resolve system_prompt_file atomically and reject blank prompts (#92,
fddd3c5) -
antigravity: Poll for backgrounded work instead of grading it incomplete (#111,
d3f1432) -
ci: Address PR #81 review — undefined step output, dead job gate, CE035 (#81,
9cc45da) -
ci: Close the five gate-correctness findings from the code review (#81,
9cc45da) -
ci: Code review fixes for the published-action verification (#81,
9cc45da) -
ci: Fail the published-action gate on a run-limit breach (#81,
9cc45da) -
ci: Harden promote ordering and stop preflight misdiagnosing healthy lag (#81,
9cc45da) -
claude: Make Claude speak English, not Claudish (#118,
5006c91) -
cli-called: Move alternation to verb_any_of, close review findings (#103,
f7e9fda) -
codex: Fold sub-agent tokens on a turn-cap stop (#110,
12c5031) -
codex: Report per-turn tokens instead of the thread-cumulative total (#113,
80f3523) -
deps: Bump sqlparse 0.5.5 -> 0.6.0 to clear the pip-audit gate (#120,
ea5a3fc) -
deps: Bump sqlparse to 0.6.0 and close pre/post-run subprocess transports (#120,
ea5a3fc) -
evalboard: Address review findings on the Scribe source layer (#116,
d6f8d7b) -
orchestrator: Close pre/post-run subprocess transports so Windows CI stops leaking (#120,
ea5a3fc) -
plugin: Address PR #82 review — reachable activation suites, least-privilege CI (#82,
02e5151) -
plugin: Address the three PR #82 findings left open (#82,
02e5151) -
plugin: Make analyze compute its numbers, and weight smoke criteria honestly (#82,
02e5151) -
plugin: Name the tautological-criterion trap in init and task (#82,
02e5151) -
plugin: Wire the skill source into the ci skill's scheduled drift run (#82,
02e5151) -
tasks: Armed positives must require success, guarded by CE034 (#82,
02e5151) -
test: Resolve bash by absolute path so Windows CI stops hitting WSL (#81,
9cc45da)
-
Reconcile the published-action verification with main (#81,
9cc45da) -
Renumber a rebase-collided lint-rule candidate (#81,
9cc45da) -
agents: Drop config_support and the Antigravity tool mapping (#110,
12c5031) -
agents: Drop the system_prompt changes from this PR (#110,
12c5031) -
plugin: Pin plugin.json to 0.9.6 after rebasing onto the release (#82,
02e5151) -
plugin: Regenerate the bundled criteria reference (#103,
f7e9fda)
-
Add Tutorial 07 for the plugin, and fix two gaps in PLUGIN.md (#82,
02e5151) -
Defer one harness candidate from the plugin-audit run (#82,
02e5151) -
Rework Tutorial 07 after review — accuracy fixes and far less narration (#82,
02e5151) -
Stop teaching the recursive task glob the ci skill forbids (#82,
02e5151) -
claude: Drop the config_support contract from the repo guide (#110,
12c5031) -
cli-called: Drop the hardcoded subcommand count (#103,
f7e9fda) -
cli-called: Fix the guide's orphaned exact_positional prerequisites (#103,
f7e9fda) -
cli-called: Fix two unclear field descriptions (#103,
f7e9fda) -
cli-called: Name the bug the detail renderer's source avoids (#103,
f7e9fda) -
cli-called: Put the offset comment's two reasons on their own lines (#103,
f7e9fda) -
parity: State the real final_status of a capped run (#110,
12c5031) -
run-limits: Keep the contract on the page, the measurements in the PR (#110,
12c5031) -
run-limits: Re-measure the antigravity timeout case after the poll loop (#110,
12c5031) -
run-limits: Record the measured cross-harness parity results (#110,
12c5031)
-
agent: Emit system_prompt_semantics from the Agent base (#92,
fddd3c5) -
agent: Record system_prompt_semantics marker in environment_info (#92,
fddd3c5) -
agents: Honor run_limits.max_turns on codex and antigravity (#110,
12c5031) -
agents: Make a base-config field mean the same thing on every harness (#110,
12c5031) -
cli-called: Accept alternative verbs via verb_any_of (#103,
f7e9fda) -
cli-called: Add exact_positional to pin the argument tail (#103,
f7e9fda) -
evalboard: Add a Scribe tab, reading the Autopilot suite's own blob container (#116,
d6f8d7b) -
evaluation: Preserve criteria after agent failures (#119,
636e87d) -
lint: CE026 clause 4 — snippet
with:keys must be real action inputs (#82,02e5151) -
plugin: 1/5 — the bundled criteria reference explains optional fields (#82,
02e5151) -
plugin: 1/6 — discover the eval tree instead of assuming tasks/ and runs/latest (#82,
02e5151) -
plugin: 1/6 — marketplace + coder-eval plugin skeleton (#82,
02e5151) -
plugin: 2/5 — shared adversarial task rubric, applied by
task(#82,02e5151) -
plugin: 2/6 — analyze reads the run's actual schema, not one generation's (#82,
02e5151) -
plugin: 2/6 — generate the bundled criteria reference, guard it with CE032 (#82,
02e5151) -
plugin: 3/5 — a real run becomes part of done in
task(#82,02e5151) -
plugin: 3/6 — resolve the version a project pins before validating anything (#82,
02e5151) -
plugin: 3/6 — skill-check skill and the canonical activation suite (#82,
02e5151) -
plugin: 4/5 —
/coder-eval:lint-tasks, a read-only reviewer of existing tasks (#82,02e5151) -
plugin: 4/6 — look before you write: skill-check, init, lint-tasks (#82,
02e5151) -
plugin: 5/5 — activation budgets in
skill-check, layer routing inanalyze(#82,02e5151) -
plugin: 5/6 — repo convention wins, and the CI gate stops measuring the wrong set (#82,
02e5151) -
plugin: 6/6 — document the criterion aliases the loader accepts, from the models (#82,
02e5151) -
plugin: 6/6 — validate the plugin in CI, extend CE026, document it (#82,
02e5151) -
plugin: CLI-driving skills offer to install coder-eval, asking first (#82,
02e5151) -
plugin: Ship coder_eval as a Claude Code plugin + marketplace (#82,
02e5151)
-
Retire the repo-local twins of the plugin's authoring skills (#82,
02e5151) -
agent: Resolve the system prompt to a value, not a mode string (#92,
fddd3c5) -
cli-called: Collapse the pairwise verb check to itertools.combinations (#103,
f7e9fda) -
plugin: Rename skill-check to check-skill, and write the naming rule down (#82,
02e5151) -
reports: Read the recorded prompt regime instead of sniffing its shape (#92,
fddd3c5)
-
Defer the plugin-skill repo-file containment guard (#82,
02e5151) -
antigravity: Clear the CodeQL findings on the fake SDK helper (#110,
12c5031) -
antigravity: Pin the SDK half of the env seam contract (#110,
12c5031) -
plugin: Make a new skill declare whether it needs eval-root discovery (#82,
02e5151) -
run-limits: Add cross-harness max_turns / turn_timeout fixtures (#110,
12c5031) -
run-limits: Make the max_turns fixture assert the cap bound (#110,
12c5031)
-
Code review fixes for evalboard path-to-ga stale tags and mature passes (#94,
ce006c1) -
Review round 2 — gate evalboard in CI, correct GPT-5.6 rates (#94,
ce006c1) -
ci: Correct defects in the uipath runner migration (#86,
57556af) -
criteria: Make glob path resolution literal-first and ignore-filtered (#65,
b3bba2b) -
criteria: Split clustered short flags and keep negative numbers positional (#73,
a7ec3ea) -
evalboard: 1/3 — drop de-tagged tasks and score only executed runs (#94,
ce006c1) -
evalboard: Path-to-GA shows only still-tagged tasks, scored on runs that executed (#94,
ce006c1) -
evalboard: Resync lib/pricing.ts with the authoritative pricing.py table (#94,
ce006c1) -
sandbox: Close the remaining record_cli findings from #73 (#73,
a7ec3ea) -
sandbox: Repair record_cli defects found reviewing #73 (#73,
a7ec3ea) -
sandbox: Stop the recorder dir defeating the PLUGIN_TOOLS_DIR pin (#73,
a7ec3ea)
-
criteria: Accept glob patterns in criterion path fields (#65,
b3bba2b) -
evalboard: 2/3 — surface last-seen and maturity on the Path-to-GA table (#94,
ce006c1) -
sandbox: Generate CLI recording shims via record_cli (#73,
a7ec3ea)
-
command-executed: Keep whole argv-joined payload in shell unwrap (#77,
7abd080) -
command-executed: Match patterns against shell-normalized commands (#77,
7abd080) -
command-executed: Narrow command param to str before shell-normalizing (#77,
7abd080) -
command-executed: Recognize shell wrappers by predicate, not allowlist (#77,
7abd080) -
criteria: Add present predicate so asserting a switch cannot weaken a guard (#72,
8574ded) -
criteria: Make cli_called guards fail loud instead of vacuously passing (#72,
8574ded) -
criteria: Stop ignore_flags re-opening the guard false-PASS (#72,
8574ded) -
early-stop: Address PR review — trajectory parity, reason determinism, doc restore (#78,
4cf8092) -
lint: Derive CE030 criteria from the source union literal, not runtime (#77,
7abd080) -
lint: Enumerate in-tree criteria by module attribute, not the union (#77,
7abd080) -
lint: Scope CE030 criterion parity to in-tree criteria only (#77,
7abd080) -
reports: Explicit return on every early_stop_gate_note path (CodeQL py/mixed-returns) (#78,
4cf8092) -
reports: Pre-initialize the gate note so CodeQL sees it bound on every path (#78,
4cf8092)
-
Surface the Marketplace listing and make the Action quickstarts self-sufficient (#80,
401245a) -
command-executed: Document shell-normalization contract + gate it (CE030) (#77,
7abd080)
-
criteria: Add cli_called for structured invocation matching (#72,
8574ded) -
criteria: Match a flag across spellings with FlagMatch.aliases (#72,
8574ded) -
early-stop: Per-criterion arming via stop_early blocks on live criteria (#78,
4cf8092)
-
litellm: Pin litellm[proxy]==1.95.0 + fastapi==0.140.0 for proxy startup (#76,
f2f8580) -
litellm: Pin proxy deps (litellm 1.95.0 + fastapi 0.140.0) to fix startup crash (#76,
f2f8580)
- action: Rename Marketplace listing to coder_eval, add author
(
a9c274d)
- early-stop: Address PR review — polarity-blind budget, pass_threshold displacement,
gate-semantic split (#74,
800ac77)
-
cost: A task timeout with no preserved turn is unrecorded spend, not free (#63,
93c7fc0) -
cost: Book spend on the error and timeout paths, flag what is unpriced (#63,
93c7fc0) -
cost: Flag every hard-killed task as a cost floor, not just the empty ones (#63,
93c7fc0) -
evalboard: Honest scoped counts, and one definition of a run's scope (#69,
0bdac0b) -
litellm: Gate cost_log_tags on agent capability, not route (fixes non-Claude crash) (#66,
4131a2a) -
litellm: Make the orphaned-spend warning actually fire (#66,
4131a2a) -
litellm: Per-attempt cost-log scoping + single run-id accessor + no-match warning (#66,
4131a2a) -
litellm: Pin each open-weight model to a vetted provider set (no silent fallback) (#66,
4131a2a) -
litellm: Proxy-authoritative token buckets + all-priced gate + transactional join (#66,
4131a2a) -
litellm: Sanitize cost headers, reject non-finite cost, drop debug scaffolding (#66,
4131a2a) -
orchestrator: Recover the in-flight turn's spend on a hard kill (#63,
93c7fc0) -
pricing: Add the claude-opus-5 rate so killed turns stop booking zero (#63,
93c7fc0) -
pricing: Add the five unpriced codex tiers still on OpenAI's rate card (#63,
93c7fc0) -
pricing: Correct every wrong rate-card entry and close the alias gaps (#63,
93c7fc0) -
pricing: Refresh the rate card and correct gemini-3-flash-preview (#63,
93c7fc0) -
reports: Count errors as misses and stop losing cost on error paths (#63,
93c7fc0) -
reports: Count errors as misses in one canonical pass rate (#63,
93c7fc0)
-
cost: Describe the per-turn backfill as the net it is (#63,
93c7fc0) -
cost: Describe the unpriced-crash mechanism accurately and keep comments framework-general (#63,
93c7fc0) -
litellm: Correct the cost contract after cutting per-message distribution (#66,
4131a2a) -
litellm: Document LITELLM_COST_LOG wiring + correct the reconciliation-cost contract (#66,
4131a2a)
-
cost: Publish one accurate total on every reporting surface (#63,
93c7fc0) -
docker: Bind-mount the LiteLLM cost log so --driver docker joins actual cost (#66,
4131a2a) -
evalboard: Compare every harness on the overview, and scope the whole page to one (#69,
0bdac0b) -
evalboard: Compare harnesses on the overview, and identify each run (#69,
0bdac0b) -
evalboard: Lift the harness scope to the page header, in vendor colors (#69,
0bdac0b) -
evalboard: Make each turn's provider-call table a collapsed dropdown (#66,
4131a2a) -
evalboard: Mark a partly-priced run total as a floor, not the bill (#63,
93c7fc0) -
evalboard: One set of pass-rate cutoffs, and a run table that pages through all history (#69,
0bdac0b) -
evalboard: Per-call cost/cache table from provider_call_costs (replaces inline) (#66,
4131a2a) -
evalboard: Read the canonical pass rate and surface incomplete cost (#63,
93c7fc0) -
evalboard: Say which harness, model, and framework version a run used (#69,
0bdac0b) -
litellm: Actual per-call cost + cache accounting for the open-weight backend (#66,
4131a2a)
-
cost: Correct the simulator-cost bound and drop the unread variant error share (#63,
93c7fc0) -
cost: Cut the commentary and drop unreachable rate-card keys (#63,
93c7fc0) -
cost: Define the unpriced-row test once, and only for new runs (#63,
93c7fc0) -
cost: Total_cost_usd means the whole bill everywhere (#63,
93c7fc0) -
litellm: Cut per-message distribution; turn-level join + per-call audit record (#66,
4131a2a) -
litellm: Drop the provider field/column — unavailable on the streaming path (#66,
4131a2a) -
litellm: Stream the cost log + de-duplicate the OpenRouter config comment (#66,
4131a2a)
- agents: Extend cooperative early stop to codex and antigravity
(
b849421)
- criteria: Async-primary BaseCriterion contract with CheckerMisuseError escalation (#60)
-
Remove dead SimulationConfig.parallel_trials; add CE031 to guard the class (
947acd3) -
early-stop: Add stop_when 'auto' — per-instance arming + pass-armed-subset stop rule (#51,
08e21e6) -
early-stop: Defer fail-stop while a pass-armed criterion is undecided (#51,
08e21e6) -
early-stop: Return assert_never explicitly to satisfy CodeQL (#51,
08e21e6) -
reports: Code review fixes for the JUnit CI gate (#37,
74db6fa) -
reports: JUnit CI-gate review fixes + CE027 env-var lint (#37,
74db6fa) -
reports: Make skipped-task JUnit names platform-independent (#37,
74db6fa)
-
Re-trigger CI (GitHub dropped the force-push event) (#37,
74db6fa) -
deps: Lock defusedxml (dev-only, test-side XML parsing) (#37,
74db6fa)
- Disable Docs gh-pages auto-publish on push (Pages not enabled yet)
(
c289d46)
-
1/8 — add DATASETS.md and a task-schema dataset: section (
e3d37ac) -
2/8 — retire BYOD.md into DOCKER_ISOLATION.md (
1a4a2a5) -
3/8 — one complete run_limits reference; document skip (
bcb7e24) -
4/8 — add DIALOG_MODE.md and correct four stale simulation claims (
18dcbe9) -
5/8 — fix prompt_mutations example; add CE029 (
b524009) -
7/8 — generate flat indexes from the mkdocs nav; add CE028 (
1514bcb) -
Add CI Gate reference (GitHub Action + JUnit) and wire into indexes (
b8c6301) -
Agent guides, extending & report-schema references, and fixes (
ce74824) -
Fold nav long tail into one Advanced group; align index ordering (
e2ff053) -
Point docs links to coder-eval.com/docs; drop Ruff badge (
60430e5) -
Point pyproject Documentation URL to coder-eval.com/docs (
be39df2) -
Reword CODER_EVAL_RAW_SDK_LOG to prose form (satisfy CE027) (
d7d5b59) -
Use the brand name "Coder Eval" in prose and titles (
821f11b)
-
Packaged CI gate — JUnit XML output + composite GitHub Action (#37,
74db6fa) -
action: Generic env passthrough + minimum-task-score gate (#37,
74db6fa) -
ci: 3/3 — publish composite action, release automation, PR dogfood (#37,
74db6fa) -
cli: 2/3 — wire run --junit-xml and report -f junit (#37,
74db6fa) -
reports: 1/3 — add reports_junit.py disk-driven JUnit XML writer (#37,
74db6fa)
- 6/8 — CE030 documents-or-exempts model fields
(
c9a3b16)
-
Render weight:0 criteria as informational on every display surface (#34,
9a34e90) -
Weight:0 un-gates criteria (informational criteria) (#34,
9a34e90) -
Weight:0 un-gates criteria and renders as informational (#34,
9a34e90) -
early-stop: Decide skill activation on the tool call, not its result (#43,
d34aa97) -
early-stop: Latch skill activation on any engagement, not first (#43,
d34aa97) -
evalboard: Match watchlist skeleton header to avoid layout shift (#45,
acc1c86) -
reports: 1/2 — exact Student-t p-values in welch_t_test (#38,
6df6e9b) -
reports: Exact Student-t p-values and a paired comparison section (#38,
6df6e9b) -
reports: Fail loud on t* overflow; surface excluded paired tasks (#38,
6df6e9b) -
reports: One source of truth for variant series and paired stats (#38,
6df6e9b) -
reports: Validate confidence and n_resamples in bootstrap_mean_ci (#38,
6df6e9b)
-
evalboard: Make all pages harness-aware and stream tables (#45,
acc1c86) -
evalboard: Make analytics surfaces harness-aware and stream tables (#45,
acc1c86) -
reports: 2/2 — add a Paired Comparison section to experiment reports (#38,
6df6e9b)
-
agents: Run codex + antigravity harnesses on the tempdir/host path (#33,
04498a9) -
antigravity: Pin google-antigravity 0.1.7 to load on glibc 2.35 (#33,
04498a9) -
codex: Always run full-access; drop the in-process OS sandbox (#33,
04498a9) -
codex: Cover zsh login shells (macOS default) in the mock-PATH home shim (#26,
73f0db2) -
codex: Keep mock CLIs shadowed in bash login shells (#26,
73f0db2) -
codex: Make login-shell temp-home lifecycle exception-safe incl. kill_sync (#26,
73f0db2) -
codex: Restore mock-CLI PATH prepend in login shells via per-task HOME profile (#26,
73f0db2) -
codex: Restore original HOME inside generated login-shell profile (#26,
73f0db2) -
codex: Use full-access sandbox on coder_eval-managed tempdir path (#33,
04498a9) -
deps: Bump pyasn1 to 0.6.4 for GHSA-8ppf-4f7h-5ppj / GHSA-hm4w-wwcw-mr6r (#33,
04498a9)
-
deps: Bump pyasn1 0.6.3 -> 0.6.4 (GHSA-8ppf-4f7h-5ppj, GHSA-hm4w-wwcw-mr6r) (#32,
f1f4f9f) -
deps: Upgrade agent SDKs to latest (claude 0.2.124, codex 0.144.4) (#33,
04498a9)
-
Add MkDocs docs site, comparison page, and SEO metadata (#32,
f1f4f9f) -
Address PR review — Coder Eval naming, cleaner table, gh-deploy workflow (#32,
f1f4f9f) -
Address review — git identity for gh-pages, single strict deploy, table glyphs, uv tool install in Tutorial 01 (#32,
f1f4f9f)
-
Accept sandbox_managed kwarg in agent test doubles (#33,
04498a9) -
codex: Pin start-to-CodexConfig env composition and tidy test imports (#26,
73f0db2)
-
deps: Bump mcp to >=1.28.1 for CVE-2026-52869/52870/59950 (#25,
fb1ad4c) -
orchestration: Isolate per-task config-resolution failures from the suite (#25,
fb1ad4c) -
orchestration: Normalize the all-fail config-resolution abort to ValueError (#25,
fb1ad4c) -
orchestration: Re-raise ValueError verbatim in the all-fail abort (#25,
fb1ad4c) -
orchestrator: Interrupt-proof teardown so a timeout can't drop task.json (#29,
89ec0d0) -
sandbox: Prune capture-ignored entries on every preservation path; interrupt-proof teardown (#29,
89ec0d0)
-
deps: Upgrade click to 8.4.2 to resolve PYSEC-2026-2132 (#19,
a240d4e) -
deps: Upgrade click to 8.4.2 to resolve PYSEC-2026-2132 (#20,
bba8645) -
evalboard: Address search-box code review comments (#13,
382d80d) -
evalboard: Prevent search bar from resetting mid-type (#13,
382d80d) -
sandbox: Exclude home-dir dotfiles from capture_to artifacts (#19,
a240d4e) -
sandbox: Extend capture_to denylist with credential stores (#19,
a240d4e)
-
evalboard: Add a Harness (RunConfig) column to the runs tables (#10,
59a240c) -
evalboard: Add Gemini rates to the frontend pricing table (#10,
59a240c) -
evalboard: Show harness as a vendor logo (internal, main table only) (#10,
59a240c) -
evalboard: Show the harness (RunConfig) column on the runs tables (#10,
59a240c)
-
antigravity: Apply env_path_prepend so mock CLIs shadow real ones (#487,
412d28f) -
antigravity: Take harness-spawn lock unconditionally so no-prepend spawns wait out mutated-PATH windows (#487,
412d28f) -
ci: Restore claude-pr-review git-fetch auth broken by persist-credentials (#485,
155ecb2)
-
Add Docker isolation tutorial (04) + venv activation note (#482,
47b3e0b) -
Optimize README + pyproject for discoverability (#485,
155ecb2) -
OSS discoverability + packaging metadata, and a claude-pr-review CI fix (#485,
155ecb2) -
OSS-readiness follow-ups — version drift, positioning, PyPI install (#485,
155ecb2) -
antigravity: Describe the PATH-mutation window as the full harness context-entry, not just the Popen (#487,
412d28f) -
tutorials: Add a section about docker instalation (#482,
47b3e0b) -
tutorials: Changed some texts to make more clear (#482,
47b3e0b) -
tutorials: Fix stale uipath-credentials claim + review nits (#482,
47b3e0b)
- antigravity: Type the spawn-lock loop global as AbstractEventLoop | None instead of Any
(#487,
412d28f)
-
Address PR #468 review — OSS-prep docs/CI/packaging papercuts (#468,
7e75cd8) -
Code review fixes for open-source-docs-cleanup (#481,
1810741) -
Results of Fable code review: discriminated unions, judge retry/ERROR escalation, DEFAULT_* removal, resilience fixes (#483,
6d2564f) -
agents: Address Antigravity review — loud turn-status conversion + test/lint hardening (#461,
74e4774) -
agents: Let Antigravity read skill files inside workspace_only sandbox (#461,
74e4774) -
agents: Normalize Antigravity tool-call params to canonical keys (#461,
74e4774) -
agents: Silence pyright on optional google-antigravity import (#461,
74e4774) -
ci: Close token-exfil + comment-injection gaps in claude-pr-review (#472,
9e374de) -
ci: Harden claude-pr-review with tool + comment allowlists (#472,
9e374de) -
codex: Apply env_path_prepend to app-server PATH so mock CLIs shadow real ones (#480,
8a53bbe) -
docker: Real multi-stage runtime kit + drop unused label (review) (#466,
15712a6) -
errors: 6/6 — categorization fall-through + sandbox-cleanup guard (#483,
6d2564f) -
evalboard: Collapsed replicate row shows Passed if any replicate passed (#474,
caa9cee) -
evalboard: Colour all failure statuses red, not grey (#474,
caa9cee) -
evalboard: Reconcile window-summary Runs denominator and share passClass (#477,
936731b)
-
2/3 — purge all dangling references to the deleted docs/features tree (#481,
1810741) -
Add Contributor License Agreement file and reference it in CONTRIBUTING (#468,
7e75cd8) -
Address PR #468 review round 2 — remove workflow residue from public tree (#468,
7e75cd8) -
Address PR #481 review — scrub residues, purge dangling refs (#481,
1810741) -
Defer type-Literal-default lint candidate from top5-review-fixes run (#483,
6d2564f) -
Open-source prep — community-health files, docs, and review fixes (#468,
7e75cd8) -
Open-source prep — delete internal doc trees, purge references, add tutorials 04/05 (#481,
1810741) -
Sweep stale .env DEFAULT_* references after layer-5 removal (#483,
6d2564f) -
samples: Add step-by-step run guide to SkillsBench README (#473,
6bc706f) -
tasks: Lead the tasks README with samples, sentinels below (#473,
6bc706f) -
tutorials: 3/3 — add 04 Writing a task and 05 Comparing two models (#481,
1810741)
-
agents: Add Antigravity (Gemini) backend via google-antigravity SDK (#461,
74e4774) -
agents: Wire Antigravity skill discovery via native skills_paths (#461,
74e4774) -
config: 5/6 — remove the .env DEFAULT_* layer-5 knobs (#483,
6d2564f) -
docker: Make docker-images — build agent + runtime kit in one command (#466,
15712a6) -
docker: Relocatable runtime kit for inject-mode tasks (#466,
15712a6) -
evalboard: K/N ✓ pass-count badge on replicated task rows (#474,
caa9cee) -
evalboard: Per-task pass rate + k/N ✓ badge for replicated runs (#474,
caa9cee) -
evalboard: Per-task pass rate across replicates (#474,
caa9cee) -
evalboard: Window cost + run summary on the front page (#477,
936731b) -
evaluation: 4/6 — Bedrock judge retry + JudgeInfrastructureError escalation (#483,
6d2564f) -
lint: 2/6 — CE024 discriminated-union rule + TemplateSource wrap (#483,
6d2564f) -
models: 1/6 — discriminated SuccessCriterion union + fail-loud validate_registry (#483,
6d2564f) -
models: 3/6 — extra=forbid on mutation models + CE009 scope extension (#483,
6d2564f) -
samples: Add vendored SkillsBench sample tasks (#473,
6bc706f) -
samples: Swap pddl-tpp-planning for court-form-filling (#473,
6bc706f)
-
evalboard: Address PR review — consistent replicate rollup (#474,
caa9cee) -
tasks: Group agent feature-tests under tasks/agents/ + add index (#473,
6bc706f)
-
agents: Drop redundant module import in Antigravity timeout test (#461,
74e4774) -
codex: Cover env_path_prepend PATH shadowing in _build_codex_env and start() (#480,
8a53bbe) -
config: Isolate defaults test from ambient TELEMETRY_ENABLED (#483,
6d2564f) -
docker: Drift guards + harden runtime kit (review feedback) (#466,
15712a6)
-
ci: Disable telemetry workflow-wide in pr-checks (baked-in default would emit) (#456,
fe27c5f) -
codex: Fall back to full-access sandbox under the docker driver (#459,
371ab28) -
codex: Run end-to-end under the docker driver (#459,
371ab28) -
deps: Bump python-socketio/engineio to clear pip-audit CVEs (#456,
fe27c5f) -
docker: Copy ~/.claude symlinks verbatim so a plugin-cache loop can't abort setup (#460,
b192ca2) -
enums: Make FinalStatus.category exhaustive (no silent "failed" fall-through) (#456,
fe27c5f) -
evalboard: Align run-view count + trends with replicate collapse (#458,
a9f93e5) -
pricing: Inline proxy-shim deprecation message to satisfy pyright (#463,
75a07f9) -
pricing: Keep coder_eval.proxy.pricing as a deprecated alias (#463,
75a07f9) -
sandbox: Root Windows tempdir sandboxes off the user temp tree (#465,
f857c64) -
telemetry: First-run disclosure notice, README, caller-settable Source, SchemaVersion (#456,
fe27c5f) -
telemetry: Keep docker container silent to avoid double-counted Task.End (#456,
fe27c5f) -
telemetry: Never emit real telemetry from the test suite (#456,
fe27c5f) -
telemetry: One CoderEval.Task.End event per task + Category dim (drop divergent buckets) (#456,
fe27c5f) -
telemetry: Single Task.End event + Category dim, test isolation, exhaustiveness, baked-in default connection string (#456,
fe27c5f)
-
Bump socketio/engineio (CVE fixes) + address review feedback (#462,
80991c8) -
Delete coder_eval.proxy.pricing shim + finish gateway-identifier cleanup (#467,
32e4e63) -
Finish LLMGW residual cleanup left by the proxy removal (#467,
32e4e63) -
Finish LLMGW residual cleanup left by the proxy removal (#463) (#467,
32e4e63)
-
Keep proxy design docs as historical records (#463,
75a07f9) -
telemetry: Drop remaining stale Task.End/.Failed references (#456,
fe27c5f)
-
BUILD_FAILED status + capture build log for failed image builds (#462,
80991c8) -
evalboard: Per-replicate task detail with a sticky run selector (#458,
a9f93e5) -
telemetry: Bake in a default App Insights connection string (env overrides) (#456,
fe27c5f)
-
Address PR #450 review — CE020→CE021 rename, else-clause guard, round-trip + ctor-leak tests (#450,
faa91bd) -
Address PR #451 review findings (CodeQL + type/diagnostic nits) (#451,
5c285d6) -
Address PR #451 round-2 review findings (CodeQL + Docker test gap) (#451,
5c285d6) -
Code review fixes for decouple-base-agent-config-sdk-types (#452,
6283346) -
Code review fixes for malformed-task-json/sim-leak plan (#450,
faa91bd) -
Degrade-not-crash on malformed task.json + simulator scratch-dir leak (#450,
faa91bd) -
Full code review fixes for harness-lint-improvements (#440,
6dd7fae) -
High-priority code-review fixes (review 260622-1009) (#439,
3e2d668) -
PR-gate + review fixes (merge main, finalize hardening) (#440,
6dd7fae) -
cli: Drop codex+backend rejection — --backend routes the judges too (#440,
6dd7fae) -
docker: 1/3 — degrade-not-crash on malformed task.json (#450,
faa91bd) -
evalboard: Don't let mature-source scan crash the run page (#455,
b5d838d) -
lint: Full review fix — CE018 also flags list/set membership denylists (#440,
6dd7fae) -
models: Correct misleading type:none validation suggestion (#440,
6dd7fae) -
reports: Avoid implicit string concat in denominator render line (#453,
c2412a7) -
reports: Re-evaluate thresholds on missing-aggregator path so completion_rate is consistent (#453,
c2412a7) -
reports-html: Phase 3 — _status_badge dispatches on FinalStatus.category (#439,
3e2d668) -
results: Phase 1 — calculate_weighted_score fails loud on length mismatch (#439,
3e2d668) -
sampling: Keep --sample-per-stratum nondeterministic by default (#439,
3e2d668) -
scoring: Phase 7 — fail-loud weighted score + single-sourced all_criteria_passed gate (#440,
6dd7fae) -
simulation: 2/3 — self-cleaning UserSimulator.start() (no scratch leak) (#450,
faa91bd) -
task-loader: Phase 2 — CLI --sample-per-stratum reproducible-by-default (#439,
3e2d668) -
telemetry: Full code review fixes — docker per-task events (#441,
2f84789) -
telemetry: Review fixes + reshape to generic product telemetry (#441,
2f84789) -
tests: Reset deadline-break scenario state between replays (#451,
5c285d6) -
types: Phase 1 — promote pyright string-concat + missing-type-arg to error (#440,
6dd7fae)
-
Harness & lint improvements (code-review 2026-06-22) (#440,
6dd7fae) -
Remove suite-aggregate-error-row-denominator scratch docs from repo root (#453,
c2412a7) -
claude: Shared axis catalog, cr-workflow scoring fix + 3-way change class (#445,
21b7386) -
evalboard: Gitignore the coverage/ output dir (#439,
3e2d668) -
evalboard: Remove deploy plumbing (moving to coder_eval_uipath) (#447,
50f570a) -
lint: Phase 5 — enable PLR0915/PLR0912 ceiling with tracked debt markers (#440,
6dd7fae) -
lint: Renumber dialog-loop statement-cap rule CE019 -> CE020 (#451,
5c285d6)
-
claude-review: Drop dead cross-repo skills-YAML check from prompt (#444,
6c4bc35) -
claude-review: Harden PR-review workflow for public open-sourcing (#444,
6c4bc35) -
claude-review: Harden PR-review workflow for public repo (#444,
6c4bc35)
-
Mark suite-aggregate-error-row-denominator plan complete (#453,
c2412a7) -
evalboard: Make public README local-only, drop deploy/Azure details (#447,
50f570a) -
evalboard: Repoint README deploy section after plumbing move (#447,
50f570a) -
orchestrator: Full code review fix — clarify finalize score-wrap asymmetry (#439,
3e2d668)
-
cli: Phase 12 — reject codex+bedrock/proxy, document sampling reproducibility (#440,
6dd7fae) -
cli: Phase 6 — constrain proxy --vendor with click.Choice (#440,
6dd7fae) -
docker: Mount a lean RW copy of ~/.claude instead of host dir read-only (#446,
dfcbf5b) -
evalboard: Add EVALBOARD_EDITION edition flag (#447,
50f570a) -
evalboard: Add OSS edition and move deploy plumbing to coder_eval_uipath (#447,
50f570a) -
evalboard: Clickable mature test links + trends maturity (#455,
b5d838d) -
evalboard: Gate more internal-only surfaces in OSS edition (#447,
50f570a) -
evalboard: Hide internal-only nav links in OSS edition (#447,
50f570a) -
evalboard: Mature-link popover + simpler trends maturity (#455,
b5d838d) -
lint: Phase 3 — CE018 no-final-status-name-denylist + dispatch _status_badge on category (#440,
6dd7fae) -
reports: 1/2 — record ERROR-row exclusion + gateable completion_rate in suite aggregates (#453,
c2412a7) -
reports: Record ERROR-row exclusion & expose gateable completion_rate in suite aggregates (#453,
c2412a7) -
telemetry: Opt-out usage telemetry via OpenTelemetry → Azure App Insights customEvents (#441,
2f84789) -
telemetry: Phase 1 — telemetry module + config + core deps (#441,
2f84789) -
telemetry: Phase 2 — emit run-start + task-end/failed events (#441,
2f84789) -
telemetry: Phase 3 — per-command CoderEval.Cli. events (#441,
2f84789) -
telemetry: Phase 4 — CE018 lint rule + feature doc (#441,
2f84789)
-
Decompose agent turn-loops + the five remaining god-functions (#451,
5c285d6) -
agents: Phase 4 — shared finalize/raise/partial-record kernels on Agent (#451,
5c285d6) -
claude-agent: Phase 2 — extract _ClaudeTurnState + _build_claude_query (#451,
5c285d6) -
cli: Phase 9 — extract pure heartbeat_is_fresh watchdog predicate (#440,
6dd7fae) -
codex-agent: Drop dead prev_prompt_tokens + fix post-watchdog timeout state (#451,
5c285d6) -
codex-agent: Phase 3 — extract _CodexTurnState (#451,
5c285d6) -
docker: 4/5 — decompose DockerRunner.run into staged helpers (#451,
5c285d6) -
errors: 1/5 — decompose categorize_error into group classifiers (#451,
5c285d6) -
errors: Phase 8 — delete dead ErrorCategory members + add liveness contract test (#440,
6dd7fae) -
models: 1/2 — decouple BaseAgentConfig from claude-code-sdk types (#452,
6283346) -
models: Decouple BaseAgentConfig from claude-code-sdk types (+ CE020 guard) (#452,
6283346) -
models: Phase 4 — declare type on BaseSuccessCriterion, drop getattr type-holes (#440,
6dd7fae) -
orchestrator: 5/5 — decompose _simulation_dialog_loop + add CE019 ratchet (#451,
5c285d6) -
orchestrator: Phase 2 — drop run_batch wrapper, promote reportImportCycles to error (#440,
6dd7fae) -
orchestrator: Phase 6 — DRY simulation telemetry via one builder (#439,
3e2d668) -
reports: 2/5 — decompose generate_markdown into section helpers (#451,
5c285d6) -
reports: 3/5 — decompose generate_experiment_report into section helpers (#451,
5c285d6)
-
Address review — guard SettingSource mirror + freeze codex task.json shape (#452,
6283346) -
Move setting_sources merge-strategy assertion to Claude subclass (#452,
6283346) -
Resolve CodeQL mixed-import alert on agent_config (#452,
6283346) -
Stop codex CLI test leaking settings.api_backend into later tests (#440,
6dd7fae) -
agents: Phase 1 review fix — keep Claude goldens collectable without codex (#451,
5c285d6) -
agents: Phase 1 — golden-master characterization harness (#451,
5c285d6) -
agents: Phase 4 review fix — lock _state-before-finalize ordering (#451,
5c285d6) -
cli: Phase 11 — CliRunner coverage for the report command (#440,
6dd7fae) -
cli: Phase 4 — report CLI tests + extract heartbeat_is_alive helper (#439,
3e2d668) -
codex: Split SDK-independent unit tests into an ungated file (#451,
5c285d6) -
lint: 2/2 — add CE020 banning SDK-typed BaseAgentConfig fields (#452,
6283346) -
lint: 3/3 — add CE020 guarding EvaluationResult.model_validate_json (#450,
faa91bd) -
lint: Code review fix for phase 3 — pin CE018 denylist to FinalStatus (#440,
6dd7fae) -
models: Code review fix for phase 4 — accurate union test name + bogus-tag assertion (#440,
6dd7fae) -
proxy: Phase 10 — TokenManager._acquire_token HTTP unit test via httpx.MockTransport (#440,
6dd7fae) -
reports: 2/2 — ERROR-row denominator + completion_rate gate tests (#453,
c2412a7) -
telemetry: Strip ANSI before asserting on --help output (#441,
2f84789)
-
cli: Harden aggregate recovery path and disclose rebuild scope (#438,
5f32a13) -
deps: Bump pydantic-settings and msgpack to patch CVEs (#435,
24a0e06) -
evalboard: Date and sort ad-hoc runs by run start_time (#432,
f058ea8) -
evalboard: Exclude mature-skipped tasks from run-page cost/duration metrics (#427,
18f303c)
-
Restore HTML report generation and the report command (#430,
1abf449) -
cli: Add aggregate to rebuild run.json/run.md from task.json (#438,
5f32a13) -
cli: Add summarize to rebuild run.json/run.md from task.json (#438,
5f32a13) -
evalboard: Cap ad-hoc runs with a show-all toggle (#432,
f058ea8) -
evalboard: Date, sort, and paginate the ad-hoc runs section (#432,
f058ea8) -
evalboard: Mark skipped-mature tasks with a non-clickable badge (#427,
18f303c) -
evalboard: Show mature tasks as green passes, keep trend averages honest (#427,
18f303c)
-
Remove dormant features ahead of open-sourcing (#430,
1abf449) -
Remove dormant features ahead of open-sourcing (#422) (#430,
1abf449)
-
Address pr:424 review nits (pricing seam + error-category contract) (#424,
b24010e) -
Code review findings for pricing seam + error category (#424,
b24010e) -
Full code review fixes for delegate-sdk base primitives (phases 1-3) (#424,
b24010e)
-
Correct the is-not-None lookup rationale (review follow-up) (#424,
b24010e) -
Surface register_pricing extension point + BYOA worked example (#424,
b24010e) -
pricing: Add feature spec for the pricing registration seam (#424,
b24010e)
-
Delegate-sdk base primitives — AgentConfigError + pricing registration seam (#424,
b24010e) -
errors: Add generic AgentConfigError + AGENT_CONFIG_ERROR category (#424,
b24010e) -
pricing: Add register_pricing seam for plugin-contributed model rates (#424,
b24010e)
-
release: Drop dry-run input; dispatch always cuts a release (#426,
633628c) -
release: Make releases manual via dispatch with dry-run preview (#426,
633628c) -
release: Manual dispatch-triggered release + versioned agent image (#426,
633628c) -
release: Pick bump level at dispatch, not from commit messages (#426,
633628c) -
release: Publish versioned agent image from the release job (#426,
633628c)
- review: Add workflow-orchestrated code-review skill + harden the full review command
(#423,
ac7b2e0)
-
Move dashboard CI to coder_eval_uipath (as eval-runner) (#419,
b862794) -
Remove dashboard CI (moved to coder_eval_uipath as eval-runner) (#419,
b862794)
-
agents: Break plugins<->registry import cycle (CodeQL) (#416,
df9ac29) -
agents: Code review fixes for phase 2 — uniform ResolvedAgentConfig (#416,
df9ac29) -
agents: Guard duplicate plugin-kind registration + review polish (#416,
df9ac29) -
ci: Byoa fixture plugin uses hatchling only-include, not setuptools py-modules (#416,
df9ac29)
-
agents: Bring-your-own-agent (BYOA) plugin SPI (#416,
df9ac29) -
agents: Phase 1 — entry-point plugin discovery + string-keyed registry (#416,
df9ac29) -
agents: Phase 2 — registry-driven agent config dispatch (#416,
df9ac29) -
agents: Phase 3 — worked BYOA plugin, live test + CI, docs (#416,
df9ac29)
-
docker: Pin framework entrypoint via docker run --entrypoint (#418,
a86ad2b) -
docker: Restore actionable runtime-image guard; reconcile docs (#418,
a86ad2b)
-
Add --strict so no-op release exits non-zero (#417,
4f6876b) -
Add conventional commits checker for PR titles and commits (#406,
ff73e9a) -
Configure git identity before amending release commit (#420,
37923e8) -
Fix false-positive release on non-releasable commits (#417,
4f6876b)
-
Skill-activation nightly — stratified sampling, per-skill recall, evalboard view (#391,
ecc0403) -
activation: Compute activation score + add Slack line (#391,
ecc0403) -
activation: Enrich case rows + rework the evalboard activation views (#391,
ecc0403) -
activation: Keep activation rows out of run-level metrics (#391,
ecc0403) -
activation: Merge skills + activation into one nightly run (#391,
ecc0403) -
activation: Per-skill recall aggregation + evalboard view (#391,
ecc0403) -
activation: Run 20/skill nightly via --sample-per-stratum (#391,
ecc0403) -
activation: Run as a nested sub-run instead of merging into run.json (#391,
ecc0403) -
activation: Surface activation in the evalboard (front page, run card, dedicated page) (#391,
ecc0403) -
dataset: Stratified random sampling; make --sample random (#391,
ecc0403) -
evalboard: Hide activation rows from run view by default (#391,
ecc0403)
-
Remove UiPath eval content moved to coder-eval-uipath (#401,
6405391) -
Sweep remaining references to moved eval content (#401,
6405391)
-
sandbox: Address PR review — restore -p alias, close test gaps, doc deltas (#398,
cbcfb85) -
sandbox: Clear stale artifacts on resume re-run; drop CLI aliases; review polish (#398,
cbcfb85)
- Initial Release