Skip to content

Correct inflated benchmark numbers in docs to match results JSON - #16

Merged
Mark007-R merged 1 commit into
mainfrom
fix/honest-benchmark-numbers
Jun 8, 2026
Merged

Mark007-R merged 1 commit into
mainfrom
fix/honest-benchmark-numbers

Conversation

@Mark007-R

Copy link
Copy Markdown
Owner

Why

A post-project audit found that several headline numbers in the public docs did not match the actual benchmark outputs in results/*.json. The benchmark machinery (token counts, latency, correctness, cost-arithmetic) is real and computed on the real 200-pair / 200-tuple datasets — but the hand-written prose in README and SYSTEM_WRITEUP had drifted to inflated values. This PR reconciles every doc to the data.

Notably, the project's own designated source-of-truth (ui/demo_scenario.py PHASE_FINDINGS dict) and docs/POLICIES.md were already accurate — only the narrative prose had drifted.

Corrections (claimed → actual, from results JSON)

Doc claim Actual value (results file)
naive 12,800 / hybrid 2,480 tokens 1,449.6 / 838.4 (phase3_day18_consolidated)
hybrid "~6× cheaper" 24% cheaper ($0.585 vs $0.769; phase5_*)
hybrid "3× faster" slightly slower (0.48 ms vs 0.24 ms p50)
quality 4.1 / 4.2, "goes up" tied 3.9/5 fact-subset; Phase-3 mock-proxy 3.04 vs hybrid 2.91
naive-LLM "~11 ms" latency 1.8 µs (mock-mode local compute)
policy "~17,000× cheaper" $0 vs ~$0.57/100 projected (undefined ratio — removed)

What's now stated honestly

  • A Measurement honesty note in README + SYSTEM_WRITEUP: the whole project ran in MOCK_LLM mode; quality is a mock-proxy token-overlap heuristic, not a real LLM judge; policy-engine latencies are local compute (no network).
  • The real, defensible findings kept: hybrid uses 42% fewer tokens at parity recall (183/200), 24% cheaper at tied quality; declarative 100% vs naive-LLM 54% correctness at zero marginal cost, naive-LLM misses all 5 reject scenarios.

Files

  • README.md, docs/SYSTEM_WRITEUP.md, docs/DEMO_VIDEO_SCRIPT.md, context_engine/README.md, orchestrator/README.md, results/metrics.json
  • (docs/POLICIES.md, ui/, SYSTEM_DESIGN.md, ARCHITECTURE.md were already accurate — unchanged.)

Tests unaffected (docs-only): suite stays 871 passed / 10 skipped / 90% coverage.

…ts JSON — reconcile README + SYSTEM_WRITEUP + per-component READMEs + DEMO_VIDEO_SCRIPT + metrics summary to the actual computed values (hybrid 838 vs 1450 tokens not 2480 vs 12800; 24% cheaper not ~6x; declarative $0 vs ~$0.57/100 projected not ~17000x; policy latency is mock-mode local compute), add measurement-honesty note flagging mock-mode + mock-proxy quality (post-project doc correction)
Copilot AI review requested due to automatic review settings June 8, 2026 07:48
@Mark007-R
Mark007-R merged commit f242647 into main Jun 8, 2026
@Mark007-R
Mark007-R deleted the fix/honest-benchmark-numbers branch June 8, 2026 07:48

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR corrects previously inflated benchmark headline numbers in the repository’s public documentation so they match the actual values in results/*.json, and adds explicit caveats about MOCK_LLM / mock-proxy quality scoring to prevent misinterpretation.

Changes:

  • Replaced outdated token/cost/latency/quality claims in docs with values from results/phase3_day18_consolidated.json and results/phase5_*.json.
  • Added “measurement honesty” caveats (MOCK_LLM mode, proxy quality judge, local-only latency) to README/writeup and summarized them in results/metrics.json.
  • Updated per-component READMEs / demo script narrative to remove incorrect multipliers and reflect the corrected benchmarks.

Reviewed changes

Copilot reviewed 6 out of 6 changed files in this pull request and generated 3 comments.

Show a summary per file
File Description
results/metrics.json Updates champion strategy summaries and adds a measurement caveat capturing MOCK_LLM/proxy constraints.
README.md Rewrites the phase-champion tables and narrative to match results JSON and adds a measurement honesty note.
orchestrator/README.md Aligns policy-engine headline claims with corrected cost framing and clarifies local/mock latency.
docs/SYSTEM_WRITEUP.md Reconciles the main narrative “headlines” with results JSON and adds an up-front measurement honesty section.
docs/DEMO_VIDEO_SCRIPT.md Updates narrated benchmark claims to match corrected costs/constraints (mock-mode, local latency).
context_engine/README.md Updates the champion retrieval description to match corrected token/cost/quality framing.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread README.md
Comment on lines +75 to +76
**Hybrid uses 42% fewer brief tokens than naive and 41% fewer than
recency, at parity fact recall on 183 of 200 pairs.** Under the
Comment thread docs/SYSTEM_WRITEUP.md
Comment on lines +111 to +113
and a summarised digest for the cold tail. **It uses 42% fewer brief
tokens than naive and 41% fewer than recency, at parity fact recall on
183 of 200 pairs.** Note the honest wrinkle: under the mock-proxy
Comment thread context_engine/README.md
Comment on lines +53 to +55
+ summary (cold tail). Uses 42% fewer brief tokens than naive at
parity fact recall (183/200 pairs); ~24% cheaper per 100q at tied
mock-proxy quality. Champion on the cost / quality-per-token
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants