Skip to content

docs(exp-18): judge-capability rerun on v0.1.0, plus speed and cost - #38

Merged
ghchinoy merged 2 commits into
mainfrom
docs/exp-18-v010
Sep 29, 2026
Merged

ghchinoy merged 2 commits into
mainfrom
docs/exp-18-v010

Conversation

@ghchinoy

Copy link
Copy Markdown
Owner

This updates EXP-18, the mizan Experiment 07 judge study, with a full same-session rerun on serving image v0.1.0. Vertex G4 reported version v0.1.0, revision f241b77; Cloud Run served the same digest. The 2026-09-26 run on the older images is kept as a replication.

  • No regression. Paired on identical items, dgem on G4 went from 67.0% to 67.9% (1,720 verdicts, p = 0.23) and Cloud Run from 68.6% to 69.8%. No suite changed significantly.
  • Verdicts mostly unchanged. dgem is at parity with gemini-3.8-flash on safety and faithfulness, and on RewardBench with the mirror read. LLMBar instruction following still needs an autoregressive judge (79 vs 94, p = 0.0015).
  • MT-Bench and toxicity gaps are no longer significant. They narrowed on v0.1.0 but point the same way on both dates, so the MT-Bench verdict is now 'prefer an autoregressive judge'.
  • Latency. p50 is 105 ms on G4 direct, 169 ms through the gateway (vertex_first), 182 ms on Cloud Run, and 3,125 ms for gemini-3.8-flash.

Checks: docs sync (43/43 pages), link check (0 broken) and make check-public all pass. The full report is mizan docs/experiments/07-judge-capability-rerun.md, landing in ghchinoy/mizan#110.

@ghchinoy ghchinoy changed the title docs(exp-18): judge-capability rerun on v0.1.0 docs(exp-18): judge-capability rerun on v0.1.0, plus speed and cost Sep 29, 2026
@ghchinoy

Copy link
Copy Markdown
Owner Author

Added the speed and cost results (mizan Experiment 07b, ghchinoy/mizan#111), measured on the production G4 endpoint as-is:

  • Throughput: about 83 items/s per G4 replica on short inputs and 68 on long ones. A 3.5-minute sustained run at 32 concurrent held 85 items/s with 0 errors and did not autoscale. Cloud Run reached 78 / 39 items/s.
  • Live cascade: matches the offline estimate within 1–2 items. p50 is 94–133 ms.
  • Thinking budget 0: saves 0.4–1.4 s at p50 on gemini-3.8-flash with no significant accuracy change.
  • Cost per 1,000 judgements: about $0.02 for dgem fully utilized, $1.06–2.25 for 3.8-flash, and ≥ $10 for Vertex predefined metrics. A G4 replica has a $140/day idle floor.

Docs sync, link check and check-public all pass.

@ghchinoy
ghchinoy merged commit e4219cd into main Sep 29, 2026
3 checks passed
@ghchinoy
ghchinoy deleted the docs/exp-18-v010 branch September 29, 2026 17:09
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant