Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
26 commits
Select commit Hold shift + click to select a range
938a791
benchmakrs added
rishabhraj36 Aug 10, 2026
b7e4d5e
fix: removed missing refs from readme
rishabhraj36 Aug 10, 2026
bfa4d4b
feat(benchmarks): support OpenAI as judge provider
rishabhraj36 Aug 10, 2026
2c5c778
fix: screenshot upload
rishabhraj36 Aug 10, 2026
c861319
Merge branch 'main' into feat/evals
rishabhraj36 Aug 13, 2026
134c8a0
feat: added cost support
rishabhraj36 Aug 13, 2026
712bc99
docs: design benchmark-specific judge prompts
beubax Aug 13, 2026
332db0f
feat(benchmarks): split general and stealth judge prompts
beubax Aug 13, 2026
57fb122
Delete docs/superpowers/specs directory
beubax Aug 13, 2026
c69011f
chore: bump version metadata for 0.7.0 release
beubax Aug 13, 2026
ac8d5fa
Revert "chore: bump version metadata for 0.7.0 release"
beubax Aug 13, 2026
dce482c
feat(benchmarks): expose full Webcmd skill pack to Codex evals
rishabhraj36 Aug 14, 2026
4234b94
Merge remote-tracking branch 'origin/main' into feat/evals
beubax Aug 14, 2026
d7549e7
feat(benchmarks): clarify stealth captcha handling
beubax Aug 14, 2026
0cb090c
Merge remote-tracking branch 'origin/feat/evals' into feat/evals
beubax Aug 14, 2026
317e66c
fix: added support for webcmd browser skills only
rishabhraj36 Aug 14, 2026
268d8b8
feat(benchmarks): run eval attempts in parallel
beubax Aug 14, 2026
9af8437
fix(benchmarks): isolate parallel webcmd sessions
beubax Aug 14, 2026
251f4ef
test(benchmarks): cover parallel webcmd sessions
beubax Aug 14, 2026
3f50d78
pi support (#313)
rishabhraj36 Aug 14, 2026
6cdf4f2
chore(benchmarks): pin webcmd 0.7.1
beubax Aug 14, 2026
59581e0
refactor: disable concurrency in benchmark evaluation and remove inte…
beubax Aug 14, 2026
8d99e67
feat(benchmarks): support Codex subscription evals
beubax Aug 14, 2026
637e4ba
docs(skills): improve browser diff and form guidance
rishabhraj36 Aug 18, 2026
50ad60e
Merge branch 'main' into feat/evals
rishabhraj36 Aug 18, 2026
d63dac2
docs(skills): sync browser skill sources
rishabhraj36 Aug 18, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 10 additions & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -36,3 +36,13 @@ sitemaps/
# Database files
*.db
autoresearch-results.tsv

# webcmd benchmarks (dataset comparison harness)
benchmarks/results/
benchmarks/.venv/
benchmarks/node_modules/
benchmarks/**/__pycache__/
benchmarks/.pytest_cache/
benchmarks/pytest-cache-files-*
benchmarks/datasets/*.json
!benchmarks/datasets/Stealth_Webcmd.json
275 changes: 275 additions & 0 deletions benchmarks/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,275 @@
# Browser Agent Benchmark Comparison

5 agents Β· 99 tasks each Β· 5 categories Β· accuracy, speed & token efficiency

## Summary table

| Agent | Accuracy | Avg time/task | Total time | Total tokens | Tokens/correct | Steps | Tool calls |
|---|---:|---:|---:|---:|---:|---:|---:|
| **dev-browser** | **75.8%** | 162.1s | 267 min | **4.49M** | **59.8K** | 1,710 | 1,504 |
| **libretto** | 73.7% | 141.9s | 234 min | 6.10M | 83.6K | **1,452** | **1,249** |
| **webcmd** | 70.7% | 188.8s | 312 min | 7.11M | 101.6K | 2,936 | 2,709 |
| **agent-browser** | 63.6% | 151.9s | 251 min | 6.39M | 101.4K | 2,815 | 2,614 |
| **chrome-devtools-axi** | 62.6% | **136.6s** | **225 min** | 6.10M | 95.3K | 2,317 | 2,116 |

**Best in class:** accuracy β†’ dev-browser Β· speed β†’ chrome-devtools-axi Β· tokens β†’ dev-browser Β· efficiency β†’ dev-browser Β· fewest steps β†’ libretto

## Overall accuracy (% of 99 passed)

```
dev-browser β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 75.8%
libretto β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 73.7%
webcmd β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 70.7%
agent-browser β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 63.6%
chrome-devtools-axi β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 62.6%
```

## Avg time per task (seconds β€” lower is better)

```
chrome-devtools-axi β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 136.6s ← fastest
libretto β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 141.9s
agent-browser β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘ 151.9s
dev-browser β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘ 162.1s
webcmd β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ 188.8s
```

## Total tokens (millions β€” lower is better)

```
dev-browser β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 4.49M ← leanest
chrome-devtools-axi β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 6.10M
libretto β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 6.10M
agent-browser β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘ 6.39M
webcmd β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ 7.11M
```

## Tokens per correct answer (lower = more efficient)

```
dev-browser β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 59.8K ← best
libretto β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘β–‘ 83.6K
chrome-devtools-axi β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘β–‘β–‘β–‘ 95.3K
agent-browser β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–‘ 101.4K
webcmd β–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆβ–ˆ 101.6K
```

## Accuracy by category (% pass rate)

| Category | webcmd | chrome-devtools-axi | agent-browser | dev-browser | libretto |
|---|---:|---:|---:|---:|---:|
| BrowseComp | 75% | 80% | 75% | **90%** | 80% |
| GAIA | 55% | 35% | **65%** | 60% | 40% |
| InteractionTests | **100%** | 84.2% | 84.2% | 89.5% | **100%** |
| OM2W2 | 65% | 60% | 45% | **80%** | 70% |
| WebBenchREAD | 60% | 55% | 50% | 60% | **80%** |

## Steps & tool calls (total actions)

| Agent | Steps | Tool calls |
|---|---:|---:|
| libretto | 1,452 | 1,249 |
| dev-browser | 1,710 | 1,504 |
| chrome-devtools-axi | 2,317 | 2,116 |
| agent-browser | 2,815 | 2,614 |
| webcmd | 2,936 | 2,709 |

---

*All five agents completed 99 tasks. Time and token totals are cumulative across the run. Avg time = total time Γ· 99. Tokens/correct = total tokens Γ· tasks passed.*

---


# Browser Bench Eval

Run a controlled, sequential browser-tool benchmark. Keep task data and local evidence private.

## Workflow

1. Confirm `uv`, the selected controller CLI, selected browser tool, and judge authentication are available (`GOOGLE_API_KEY` for `google`, `OPENAI_API_KEY` for `openai`, or a ChatGPT-authenticated `codex login` for `codex`). AXI, agent-browser, and dev-browser runs also require CloakBrowser.
2. Ask the user to choose a controller, model, benchmark, and task selection if any is missing.
3. Start with one task unless the user explicitly requests a larger or full run.
4. Run `scripts/run_eval.py` with the explicit choices.
5. Report selected task count, overall accuracy, category accuracy, terminal statuses, controller time, steps, tool calls, token usage, and the ignored local result path.
6. Compare runs only when manifest metadata (benchmark, dataset hash, controller, model, and tools) match.

The legacy `tokens` metric remains non-cached controller input plus controller
output. Task results also record the controller's ordinary input, cached reads,
cache writes, output, and reasoning-output detail. For `gpt-5.6-sol`,
`estimated_api_cost_usd` applies the documented API rates to each completed
controller turn before summing them, including the long-context multiplier.
This is an API-equivalent estimate; ChatGPT-authenticated Codex runs are not
necessarily billed through the API. Judge usage is excluded.

## Commands

```bash
uv run python benchmarks/scripts/run_eval.py \
--controller codex \
--model gpt-5 \
--benchmark BU_Bench_V1 \
--tasks 1 \
--tools webcmd
```

Use OpenAI instead of Gemini for judging:

```bash
uv run python benchmarks/scripts/run_eval.py \
--controller codex \
--model gpt-5 \
--benchmark BU_Bench_V1 \
--tasks 1 \
--tools webcmd \
--judge-provider openai \
--judge-model gpt-4o-mini
```

Use the signed-in Codex subscription for judging:

```bash
uv run python benchmarks/scripts/run_eval.py \
--controller pi \
--model openai-codex/gpt-5.6-sol \
--reasoning-effort low \
--benchmark BU_Bench_V1 \
--tasks 1 \
--tools webcmd \
--judge-provider codex \
--judge-model gpt-5.4
```

Codex judging runs ephemerally in an isolated, read-only temporary directory.
Its token usage and cost are excluded, like the other judge providers.

Pi uses the pinned SDK sidecar from the repository's Node lockfile:

```bash
npm --prefix benchmarks ci --ignore-scripts
webcmd skills add --provider codex --scope user
./benchmarks/node_modules/.bin/pi
# In Pi, run /login and select OpenAI Codex, then exit.

uv run python benchmarks/scripts/run_eval.py \
--controller pi \
--model openai-codex/gpt-5.6-sol \
--reasoning-effort low \
--benchmark BU_Bench_V1 \
--tasks 1 \
--tools webcmd
```

The Pi sidecar reads the resulting OAuth credentials from
`~/.pi/agent/auth.json`, and preflight verifies them before creating a task.
Its `estimated_api_cost_usd` is an API-equivalent comparison metric, not a
subscription charge; the manifest records `billing_mode` as
`chatgpt_subscription`.

For AXI with one dedicated CloakBrowser process and profile per task:

```bash
uv run python benchmarks/scripts/run_eval.py \
--controller codex \
--model gpt-5.6-sol \
--reasoning-effort high \
--benchmark BU_Bench_V1 \
--tasks 1 \
--tools chrome-devtools-axi
```

For agent-browser with one dedicated CloakBrowser process and profile per task:

```bash
uv run python benchmarks/scripts/run_eval.py \
--controller codex \
--model gpt-5.6-sol \
--reasoning-effort high \
--benchmark BU_Bench_V1 \
--tasks 1 \
--tools agent-browser
```

For dev-browser with one dedicated CloakBrowser process and profile per task:

```bash
npm install -g dev-browser
dev-browser install
dev-browser install-skill --codex

uv run python benchmarks/scripts/run_eval.py \
--controller codex \
--model gpt-5.6-sol \
--benchmark BU_Bench_V1 \
--tasks all \
--tools dev-browser
```

Use `--controller pi --model openai-codex/gpt-5.6-sol` to run the same lane through
Pi. The Pi sidecar mounts the installed `dev-browser` skill and enables only
its `bash` and `read` tools.

`dev-browser install` is required because it installs the daemon's Playwright
and QuickJS dependencies. It may also download dev-browser's Chromium, but the
benchmark does not use that browser: the task-private shim always connects
dev-browser to the task's dedicated CloakBrowser CDP endpoint.

For Libretto Browser Tools with Codex or Pi and one dedicated CloakBrowser per
task:

```bash
npm ci

uv run python benchmarks/scripts/run_eval.py \
--controller codex \
--model gpt-5.6-sol \
--reasoning-effort low \
--benchmark BU_Bench_V1 \
--task-indices 0,1 \
--tools libretto
```

With Codex, the harness injects its pinned stdio MCP server per attempt. With
Pi, it registers the equivalent native Pi custom tools directly; use
`--controller pi --model openai-codex/gpt-5.6-sol`. Both expose only `browser_open`,
`browser_exec`, `browser_snapshot`, `browser_status`, and `browser_close`.
`browser_connect` is disabled so the agent cannot leave the task's dedicated
CloakBrowser.

Use `--stealth-view official` only with `Stealth_Bench_V1`. Never add a parallel flag or publish `results/`.

Webcmd attempts use the dedicated `benchmark` Profile. The harness creates one
opaque Session per task before starting the controller, passes that Session to
the controller, and closes it only after the controller exits. The agent does
not create or close Sessions. Its browser surface is only
`webcmd --profile benchmark --session <session-id> browser tabs` β†’ optional
`bind --page PAGE` β†’ optional `snapshot` β†’ one or more
`run --stdin <<'JS' ... JS` calls. Profile-level cookies, cache, and storage are
shared across the Webcmd run; tabs and browser workspace are task-specific.
Removed browser primitives (`open`, `state`,
`click`, `type`, `screenshot`, `wait`, `eval`, `observe`, and `tab`) are not
allowed. Do not run `webcmd browser --help`; the allowed surface is complete.
Run programs must use one quoted heredoc; never invoke `run --stdin` with an
empty stdin body. Avoid shell-fragile JavaScript such as
`$`, template literals, regex end anchors, and mixed-quote one-liners. The global `page` is already
available; do not call `browser.currentPage()` or `page.snapshotForAI()`. `run`
returns a snapshot diff by default. Only positive `--timeout` and `--max-output`
values, `--snapshot-mode act|tree`, and the boolean `--no-snapshot-diff` flag are
accepted:

```bash
webcmd --profile benchmark \
--session session_7d8f2c10-4a11-4f3e-9c22-1b6de0a91f45 \
browser run --stdin --snapshot-mode act <<'JS'
return await page.title()
JS
```

The separate `snapshot` command accepts `--snapshot-mode act|tree|read`.

BU Bench loads `datasets/BU_Bench_V1.json`. To skip a task without shifting its raw index, set its optional field to `"enabled": false`; leave all 100 entries in their original order.

## References

- Read `references/judge-contract.md` when auditing judge decisions.
- Read `references/dataset-provenance.md` before copying, updating, or publishing datasets.
1 change: 1 addition & 0 deletions benchmarks/datasets/BU_Bench_V1.enc

Large diffs are not rendered by default.

1 change: 1 addition & 0 deletions benchmarks/datasets/Stealth_Bench_V1.enc

Large diffs are not rendered by default.

Loading
Loading