Correct inflated benchmark numbers in docs to match results JSON - #16
Merged
Merged
Conversation
…ts JSON — reconcile README + SYSTEM_WRITEUP + per-component READMEs + DEMO_VIDEO_SCRIPT + metrics summary to the actual computed values (hybrid 838 vs 1450 tokens not 2480 vs 12800; 24% cheaper not ~6x; declarative $0 vs ~$0.57/100 projected not ~17000x; policy latency is mock-mode local compute), add measurement-honesty note flagging mock-mode + mock-proxy quality (post-project doc correction)
There was a problem hiding this comment.
Pull request overview
This PR corrects previously inflated benchmark headline numbers in the repository’s public documentation so they match the actual values in results/*.json, and adds explicit caveats about MOCK_LLM / mock-proxy quality scoring to prevent misinterpretation.
Changes:
- Replaced outdated token/cost/latency/quality claims in docs with values from
results/phase3_day18_consolidated.jsonandresults/phase5_*.json. - Added “measurement honesty” caveats (MOCK_LLM mode, proxy quality judge, local-only latency) to README/writeup and summarized them in
results/metrics.json. - Updated per-component READMEs / demo script narrative to remove incorrect multipliers and reflect the corrected benchmarks.
Reviewed changes
Copilot reviewed 6 out of 6 changed files in this pull request and generated 3 comments.
Show a summary per file
| File | Description |
|---|---|
results/metrics.json |
Updates champion strategy summaries and adds a measurement caveat capturing MOCK_LLM/proxy constraints. |
README.md |
Rewrites the phase-champion tables and narrative to match results JSON and adds a measurement honesty note. |
orchestrator/README.md |
Aligns policy-engine headline claims with corrected cost framing and clarifies local/mock latency. |
docs/SYSTEM_WRITEUP.md |
Reconciles the main narrative “headlines” with results JSON and adds an up-front measurement honesty section. |
docs/DEMO_VIDEO_SCRIPT.md |
Updates narrated benchmark claims to match corrected costs/constraints (mock-mode, local latency). |
context_engine/README.md |
Updates the champion retrieval description to match corrected token/cost/quality framing. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Comment on lines
+75
to
+76
| **Hybrid uses 42% fewer brief tokens than naive and 41% fewer than | ||
| recency, at parity fact recall on 183 of 200 pairs.** Under the |
Comment on lines
+111
to
+113
| and a summarised digest for the cold tail. **It uses 42% fewer brief | ||
| tokens than naive and 41% fewer than recency, at parity fact recall on | ||
| 183 of 200 pairs.** Note the honest wrinkle: under the mock-proxy |
Comment on lines
+53
to
+55
| + summary (cold tail). Uses 42% fewer brief tokens than naive at | ||
| parity fact recall (183/200 pairs); ~24% cheaper per 100q at tied | ||
| mock-proxy quality. Champion on the cost / quality-per-token |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Why
A post-project audit found that several headline numbers in the public docs did not match the actual benchmark outputs in
results/*.json. The benchmark machinery (token counts, latency, correctness, cost-arithmetic) is real and computed on the real 200-pair / 200-tuple datasets — but the hand-written prose in README and SYSTEM_WRITEUP had drifted to inflated values. This PR reconciles every doc to the data.Notably, the project's own designated source-of-truth (
ui/demo_scenario.pyPHASE_FINDINGSdict) anddocs/POLICIES.mdwere already accurate — only the narrative prose had drifted.Corrections (claimed → actual, from results JSON)
phase3_day18_consolidated)phase5_*)What's now stated honestly
MOCK_LLMmode; quality is a mock-proxy token-overlap heuristic, not a real LLM judge; policy-engine latencies are local compute (no network).Files
README.md,docs/SYSTEM_WRITEUP.md,docs/DEMO_VIDEO_SCRIPT.md,context_engine/README.md,orchestrator/README.md,results/metrics.jsondocs/POLICIES.md,ui/,SYSTEM_DESIGN.md,ARCHITECTURE.mdwere already accurate — unchanged.)Tests unaffected (docs-only): suite stays 871 passed / 10 skipped / 90% coverage.