English | 简体中文
A usability lab for tools built for coding agents.
One public pilot found coding agents chose a richer semantic tool over grep only 0–6% of the time.
Until now you had no way to measure that on your tool.
One command, no permanent install. Node.js 22.6+ required (the bundled demo server uses built-in TypeScript stripping).
npx ergolab run demo-mcpErgoLab spins a fresh sandbox per agent-task pair and drives a panel — every agent CLI installed on your machine (Claude Code, Codex, Gemini CLI; each resolved by probing PATH first, then known install locations such as the ChatGPT.app-bundled Codex), plus a deterministic keyless mock agent as the baseline. An agent that cannot run keeps its row as skipped (reason). Then it prints the panel and writes the artifacts:
┌──────────┬─────────────────────┬──────────────┬───────────────────┬──────────────────┬─────────────────┬────────────────┐
│ agent │ find-latest-version │ find-license │ list-dependencies │ count-by-keyword │ find-deprecated │ find-publisher │
├──────────┼─────────────────────┼──────────────┼───────────────────┼──────────────────┼─────────────────┼────────────────┤
│ mock │ fail 8ms │ fail 6ms │ fail 7ms │ fail 6ms │ fail 7ms │ fail 5ms │
└──────────┴─────────────────────┴──────────────┴───────────────────┴──────────────────┴─────────────────┴────────────────┘
Scorecard: mock 0/6 — panel 0/6 (0%)
With no agent CLI installed you still get a full report (the mock always runs). To try it on your tool:
cd your-repo
npx ergolab init # scaffolds ergolab.yaml — three runnable starter tasks
# edit the three prompts to name your tool's operations (the one step over 60 seconds)
npx ergolab run . # a panel report on your own toolCommit the report, or host badge.json in your repo for a README badge — reports stay in your repo, there is no registry.
Recorded from node dist/cli.js run suites/showcase-grep-vs-tool --agents mock,gemini-cli (the grep arm of the bundled A/B pair; gemini-cli is not installed on the recording machine, so its row is kept as skipped):
The bundled showcase is an A/B pair over the same six tasks and the same dataset:
| suite | where the answers live | mock agent (grep-only) |
|---|---|---|
showcase-grep-vs-tool |
a plain file in the sandbox | 2/6 — the two substring lookups pass, exact extraction and aggregation fail |
demo-mcp |
behind the bundled stdio MCP registry server | 0/6 — nothing greppable, no tool access |
That gap is the product in miniature: a tool an agent cannot use reads as a wall of failures, and you see exactly which cell broke and at which stage. Real agent CLIs join the panel automatically when installed — on the recording machine discovery found claude at ~/.local/bin/claude and codex at /Applications/ChatGPT.app/Contents/Resources/codex (an install that never touches PATH). Full recorded numbers, commands, and limitations: docs/demo-results.json. The gif is re-renderable with vhs docs/demo.tape.
You shipped an MCP server, an agent-facing CLI, or docs written for agents. The only instrument you had was trying it in your own Claude Code session — one agent, one day, no baseline, no cost column, no way to compare Claude Code against Codex against Gemini CLI on the same tasks.
This is not a hypothetical problem. In August 2026 a tool-platform author hand-rolled exactly this measurement and found agents picked the richer semantic tool over grep only 0–6% of the time on location tasks — a finding that sparked a long Hacker News discussion (97 points, 69 comments). The reaction was "why do agents grep"; the actionable question for tool authors is "can agents use my surface, and where exactly do they fail?"
ErgoLab transplants the classic usability-lab protocol: recruit a panel (here: coding agents), give standard tasks, report where they fail, how long they take, and what they cost. HCI solved this for humans decades ago; agents get the same treatment:
- Reproducible — the suite is a file (
ergolab.yaml), verification is exit-code deterministic (no LLM judge, no oracle drift), every run gets fresh sandboxes. - Cross-agent — the same tasks run against Claude Code, Codex, Gemini CLI, and a keyless mock baseline, sequentially, one sample per pair.
- Actionable — the failure table names the agent, the task, and the stage (setup / agent / verify); the report is a matrix you can diff across releases.
One process, one ergolab binary; agent CLIs are child processes. No services, no database, no sandboxing VMs — agents run in fresh temp workdirs with your normal CLI auth, the environment you ship for.
ergolab.yaml ──► SuiteLoader (yaml + zod)
│
▼
Runner ──► sandbox/<agent>/<task>/ (fresh temp dir each)
1. setup script
2. agent subprocess (driver-built invocation)
3. verify script (exit code decides pass)
│
▼
Drivers: claude-code · codex · gemini-cli · mock
(resolve CLI: PATH first, then fallbacks · register mcp_servers per driver · parse token usage)
│
▼
Reporter ──► terminal table · report.md · ergolab-report.json · badge.json
The driver interface is the only extension point: detect(), buildInvocation(task, sandbox), parseUsage(rawOutput). detect() never assumes PATH alone — on macOS, a Codex that ships only inside ChatGPT.app still runs, and a missing binary is reported as skipped (reason: not-found) or skipped (reason: off-path-at <path>), never silently dropped.
# no install
npx ergolab run demo-mcp
# or globally
npm install -g ergolabRequirements: Node.js 22.6+ (LTS), macOS or Linux. To run real agents you also need the corresponding CLI installed and authenticated — claude (Claude Code), codex, or gemini (Gemini CLI). None are required: the mock agent always produces a report.
From source:
git clone https://github.com/SuperMarioYL/ergolab.git
cd ergolab
npm install
npm run build
./dist/cli.js run demo-mcp --agents mockergolab init scaffold ergolab.yaml (three runnable starter tasks) in the cwd
ergolab run <suite> run a suite against a panel of agents
--agents <a,b> panel selection; default: mock,claude-code,codex,gemini-cli
--out <dir> artifact directory (default: cwd)
ergolab list-suites bundled suites plus any in the cwd
ergolab showcase run the bundled grep-vs-tool showcase suite
ergolab report [file] re-render a saved ergolab-report.json
--out <dir> where to re-write report/report.md and badge.json
<suite> is a directory containing ergolab.yaml, a path to the file itself, or a bundled suite name (demo-mcp, showcase-grep-vs-tool).
tool: leftpad # the tool under test — your library, CLI, or docs
mcp_servers: # optional: stdio servers registered for every agent
registry: # paths resolve against this suite directory
command: node
args: ["--experimental-strip-types", "server/index.ts"]
tasks:
- name: find-latest-version # unique; results are keyed by it
prompt: >- # the ask, handed to the agent verbatim
Find the latest stable version of `leftpad` with the registry MCP
server; write it to answer.txt.
setup: "echo seed-data > data.txt" # optional shell, run before the agent
verify: "grep -q '1.4.2' answer.txt" # exit 0 = pass — the only oracle
timeout_s: 120 # optional, default 120| driver | what runs | token capture | MCP registration |
|---|---|---|---|
claude-code |
claude -p <prompt> --output-format json in the sandbox, permissions bypassed for the non-interactive run |
input, output, and cache fields from the result's usage |
--mcp-config file, the ~/.claude.json shape |
codex |
codex exec --json with workspace-write, working root pinned to the sandbox |
usage events from the JSONL stream | -c mcp_servers.<name>=… config overrides |
gemini-cli |
gemini -y --output-format json -p <prompt> |
best effort (the CLI's output shape is not yet verified against a live binary) | --mcp-config file, the ~/.gemini/settings.json shape |
mock |
no model, no key, no network — a grep-only stand-in agent | deterministic prompt-length estimate | none, by design (it models the agent without the affordance) |
The mock driver models the grep-first agent: it finds the output file the prompt names, greps the sandbox for the entity the prompt mentions, and copies the first matching line — writing nothing when grep finds nothing. Tasks whose answers are greppable pass; tasks whose answers live behind a tool fail. That determinism is the baseline every real agent is read against.
-
The panel report — one row per agent, one cell per task:
pass/fail/timeout/error, the stage the run ended in (setup/agent/verify), wall time, and token cost where the CLI reports it. -
Skipped, never dropped — an agent whose binary cannot run keeps its row with the reason, so panel coverage stays visible.
-
Exit-code verification — deterministic and free; no LLM judge, no fuzzy matching, no oracle drift between runs.
-
The A/B two-suite pattern — run the same tasks with and without your MCP server registered (with and without a plain-file mirror of the data) to measure what tool access changes. See the bundled
demo-mcp/showcase-grep-vs-toolpair. -
Self-hosted badge —
badge.jsonis a shields.io endpoint document; host it anywhere raw (e.g. a gist or your repo) and point a badge at it:https://img.shields.io/endpoint?url=https://raw.githubusercontent.com/<you>/<repo>/main/badge.jsonOn the recorded showcase run it reads
agent-usability 33/100(orange).
Honest limits of one run: sequential, one sample per agent-task pair — the report surfaces cliffs (2 of 6, 0 of 6), not fine differences. Repeats and variance reporting are roadmap items, not shipped behavior.
| field | meaning | default |
|---|---|---|
tool |
the tool under test; names the report context | required |
mcp_servers.<name>.command |
stdio command for an MCP server; suite-relative paths are absolutized per agent | optional |
mcp_servers.<name>.args |
argument list (same path resolution) | [] |
tasks[].name |
unique task id, keyed in results | required |
tasks[].prompt |
the ask, verbatim to the agent | required |
tasks[].setup |
shell run in the sandbox before the agent | optional |
tasks[].verify |
shell whose exit code decides pass/fail | required |
tasks[].timeout_s |
per-task timeout, applied to each stage | 120 |
Artifacts land in --out (default: the cwd): ergolab-report.json, report/report.md, badge.json, and sandbox/<agent>/<task>-*/ workdirs (kept for debugging; add sandbox/ to your .gitignore). Agents inherit your normal CLI auth and environment — v0.1 deliberately runs in the environment you ship for, not a container.
v0.1 is deliberately narrow. Not in it, in rough order of likely arrival:
- Statistical repeats, variance reporting, and parallel panel runs
- Docker/VM sandboxing for hermetic agent execution
- Windows support (macOS and Linux first)
- Transcript-level tool-choice instrumentation (today: the A/B two-suite pattern)
- Prebuilt GitHub Action packaging (today: call the CLI from your own CI)
MIT — see LICENSE.
MIT © 2026 SuperMarioYL
