Skip to content

docs(readme): measured comparison of the judging backends - #4

Merged
ChenYCL merged 1 commit into
mainfrom
docs/readme-measurements
Sep 26, 2026
Merged

ChenYCL merged 1 commit into
mainfrom
docs/readme-measurements

Conversation

@ChenYCL

@ChenYCL ChenYCL commented Sep 26, 2026

Copy link
Copy Markdown
Owner

What

One new README section — ## Judging backends: measured comparison in README.md, mirrored as
## 判定后端实测对比 in README.zh-CN.md — carrying the measured comparison of the judging
backends that until now lived only in experiments/** and docs/**.

It contains: the 20-item accuracy table (hosted Jev, Kev 4B at both row limits, the GGUF 4B
local-readout default, the two other 4B candidates, Kev 0.8B, the GGUF 0.8B and the trivial
positional baseline); what each costs in per-item latency, disk, memory and money; a short "how this
was measured"; the 15 real browser runs; and the three goal_done bars with their measured basis.

No code, tests or config changed. The existing "hosted Jev is the default" statement is untouched —
the new section restates it. The only edit outside the new section is lives in SKILL.md →
also lives in SKILL.md (and 见 → 同时见), so the pointer above the new tables no longer reads as
if the numbers were only in SKILL.md.

Sources for every figure

figure source
hosted Jev 19/20 = 0.95 · browser 14/15 · noul 5/5 · action+click 5/5 experiments/gguf-provider/results/ceiling-jev-20items.json (overall, by_kind); restated in experiments/gguf-provider/RESULTS.md and results/local-models-4b.md §0
hosted Jev 621 / 1,270 ms per item, $0.0035 for the 20 items results/ceiling-jev-20items.json (latency_ms, cost_usd_total)
Kev 4B @16384: 19/20 · 14/15 · 5/5 · 5/5 · 2,223 / 12,183 ms · 36 GB footprint · 23,769 ms packed step experiments/kev-4b/README.md (Measured: Kev 4B / Memory); cross-checked in results/local-models-4b.md §0
Kev 4B @8192: 18/20 · 13/15 · 5/5 · 4/5 · 1,965 / 8,995 ms · 12.2 GiB RSS peak same two files
Kev 4B disk 9.34 GB base + 152 MiB checkpoint; Kev 0.8B 1.72 GiB download docs/local-kev-bringup.md §0, experiments/kev-4b/README.md (verification log)
Kev 0.8B 14/20 (14 of 19 answerable) · 9/15 · 5/5 · 3/5 · 590 / 6,353 ms · 3.6 GiB peak experiments/kev-4b/README.md (Measured: Kev 0.8B, Row limit §3)
GGUF 4B 16/20 · 12/15 · 4/5 · 2/5 · 3,072 / 13,895 ms · 3,362 MiB RSS · 2,740,937,888 B experiments/gguf-provider/RESULTS.md, results/local-models-4b.md §0/§2
Qwen3-4B-Instruct-2507 15/20 · 10/15 · 5/5 · 1/5 · 2,514 / 15,386 ms · 2,497,281,120 B results/local-models-4b.md §0/§2
gemma-3-4b-it 10/20 · 7/15 · 3/5 · 1/5 · 2,189 / 9,725 ms · 2,489,894,016 B results/local-models-4b.md §0/§2
GGUF 0.8B 10/20 · 7/15 · 3/5 · 1/5 · 777 / 4,178 ms · 811,843,840 B experiments/gguf-provider/RESULTS.md
baseline 11/20 · 9/15 · 2/5 · 2/5 results/baseline-first-option.json + RESULTS.md
clean GGUF 4B step 18.3 s cold / 9.0 s warm, 6.1 s of it click_target results/local-models-4b.md §0.2, RESULTS.md
rotation k=0/3/7 → e8/e5/e1, P≈0.99 experiments/gguf-provider/RESULTS.md
15 runs: 4 success / 8 stuck / 2 needs_user / 1 max_steps; 39 step requests; 0 timeouts; 0 LOW_LABEL_MASS; min label mass 0.870 docs/local-backend-run-smoke.md §0/§2/§3
click_target 0.71–0.78 on a 65-element page; type last of four (0.105 vs 0.430); 3 premature stops; 8 of 9 non-success runs stop on action docs/local-backend-run-smoke.md §0/§5
bars 0.174 / 0.482; bands 0.12–0.28 and (0.341, 0.683]; 7/0/0 vs 4/0/3 docs/local-backend-run-smoke.md §4/§9, experiments/kev-4b/README.md (Threshold replay)

Nothing is estimated. Cells the records do not measure say not measured; cells that do not exist
for a backend say —.

Verification

  • Both new tables parse as consistent GFM grids (equal cell counts per row), code fences are
    balanced, and every relative link in the new section resolves to a file that exists.
  • CI: npm test (mock mode, Ubuntu + headless Chrome) — docs-only change.

Add a "Judging backends: measured comparison" section to README.md and its
Chinese mirror, carrying the numbers that used to live only in the experiment
records: the 20-item accuracy table (hosted Jev, both Kev 4B row limits, the
GGUF 4B default, the two 4B candidates, Kev 0.8B, the GGUF 0.8B and the trivial
positional baseline), what each costs in latency, disk, memory and money, how
the set was measured, the 15 real browser runs, and the three goal_done bars
with their measured basis. Every figure is read from the repo's own records;
nothing is estimated and unmeasured cells say so.

The hosted-Jev-default statement is unchanged.
@ChenYCL
ChenYCL merged commit 3432ead into main Sep 26, 2026
1 check passed
@ChenYCL
ChenYCL deleted the docs/readme-measurements branch September 26, 2026 09:56
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant