experiment: backfill Phase 6 deliverables (metrics journal + scorecard re-validation) - #14
Merged
Merged
Conversation
…ay 29-32 entries to results/metrics.json (append-only journal per SKILL Step 4) + add Phase 6 weekly takehome scorecard re-validation row (orch 6/6 + ce 5/5 held across all four additive days) (Phase 6 backfill)
There was a problem hiding this comment.
Pull request overview
Backfills Phase 6 “late” SKILL deliverables by extending the experiment artifacts: the append-only daily metrics journal and the weekly takehome scorecard re-validation, keeping Phase 6’s “one phase = one squash” invariant intact while still recording the missing end-of-phase evidence.
Changes:
- Appends Phase 6 Day 29–32 snapshots to
results/metrics.json, including a Day 32 wrap summary block. - Adds a Phase 6 weekly takehome scorecard re-validation section to
results/takehome_scorecard.mdwith the documented eval command and results.
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated 6 comments.
| File | Description |
|---|---|
| results/metrics.json | Adds Day 29–32 journal entries plus a Phase 6 wrap summary. |
| results/takehome_scorecard.md | Adds a Phase 6 weekly re-validation block with evaluator outcomes and run command. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| @@ -1,5 +1,5 @@ | |||
| { | |||
| "_about": "PennyCore \u00e2\u20ac\u201d append-only metrics journal. Each entry is a per-day snapshot; never overwrite past entries. Defined in SKILL Step 4 + the PRIMARY METRICS section. Initialized on Day 7 as a backfill (the journal was missed on Days 5 and 6 when the first measurable code shipped \u00e2\u20ac\u201d those days' numbers are NOT retroactively reconstructed; the day_05 and day_06 entries record only the binary 'feature shipped' fact so future readers know the journal started here). Comparison metrics (quality, cost, quality/$ ratios) start populating Day 12 when Phase 3 benchmarks land.", | |||
| "_about": "PennyCore — append-only metrics journal. Each entry is a per-day snapshot; never overwrite past entries. Defined in SKILL Step 4 + the PRIMARY METRICS section. Initialized on Day 7 as a backfill (the journal was missed on Days 5 and 6 when the first measurable code shipped — those days' numbers are NOT retroactively reconstructed; the day_05 and day_06 entries record only the binary 'feature shipped' fact so future readers know the journal started here). Comparison metrics (quality, cost, quality/$ ratios) start populating Day 12 when Phase 3 benchmarks land.", | |||
| "naive_tokens": 1059, | ||
| "ratio": 1.483, | ||
| "interpretation": "at budget=2000 with 50-message history, the recency brief is LARGER than naive concat because it adds channel markers, timestamps, and action segments that naive concat omits. The ratio<1 win the Phase 3 hybrid is supposed to deliver only kicks in when the borrower history exceeds the budget \u00e2\u20ac\u201d the 200-history-pair benchmark on Day 12 quantifies that." | ||
| "interpretation": "at budget=2000 with 50-message history, the recency brief is LARGER than naive concat because it adds channel markers, timestamps, and action segments that naive concat omits. The ratio<1 win the Phase 3 hybrid is supposed to deliver only kicks in when the borrower history exceeds the budget — the 200-history-pair benchmark on Day 12 quantifies that." |
| "event_id_received": "evt_963ff24cf2914a90b1318adf", | ||
| "round_trip_latency_observed_sec": "<1 (sleep 1 between publish and recent-read; not microbenchmarked today)", | ||
| "tenant_seeded": "INSERT INTO tenants (id, name, slug) VALUES ('acme', 'Acme Bank', 'acme-bank') \u00e2\u20ac\u201d required by Day-3 schema FK" | ||
| "tenant_seeded": "INSERT INTO tenants (id, name, slug) VALUES ('acme', 'Acme Bank', 'acme-bank') — required by Day-3 schema FK" |
| "anomaly_detected": "notify_loan_officer", | ||
| "system_event": "no_op", | ||
| "unknown": "notify_loan_officer (conservative default \u00e2\u20ac\u201d escalate to human)" | ||
| "unknown": "notify_loan_officer (conservative default — escalate to human)" |
Comment on lines
+182
to
+183
| "make_buffered_handler(_recent_events) — Day 8", | ||
| "make_planning_handler(_recent_proposals) — Day 9" |
| } | ||
| }, | ||
| "notes": "Day 9 planner is mock-mode-only in the live system (placeholder ANTHROPIC_API_KEY); LLM path covered by fake-client tests. Real-LLM exercise deferred to user-driven key setup or Phase 3 benchmark days. docker-compose up verification re-run after initial commit to close SKILL Step 4 'actually RUN the system' loop \u00e2\u20ac\u201d the bind-mounted orchestrator/ directory picks up planner.py without rebuild." | ||
| "notes": "Day 9 planner is mock-mode-only in the live system (placeholder ANTHROPIC_API_KEY); LLM path covered by fake-client tests. Real-LLM exercise deferred to user-driven key setup or Phase 3 benchmark days. docker-compose up verification re-run after initial commit to close SKILL Step 4 'actually RUN the system' loop — the bind-mounted orchestrator/ directory picks up planner.py without rebuild." |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Follow-up to PR #13 — appends two SKILL deliverables that landed late.
What was missed
results/metrics.jsonappend-only journal — SKILL Step 4 mandates a per-day snapshot inresults/metrics.json(append, don't overwrite). Phase 6 (Days 29-32) shipped without those four entries. Last entry was Day 28 (2026-05-31).results/takehome_scorecard.mdis tracked weekly from Day 11. Last update was Day 26; a Phase 6 re-validation was due at the end of Day 32.What this PR adds
results/metrics.json— Day 29 / 30 / 31 / 32 withcomponent,feature,tests_passing,tests_added_today,takehome_score, and notes. Day 32 carries aphase_wrap: trueflag and aphase_6_summaryblock (4 deliverables, 79 tests, suite 766 → 845, branch + squash commit).results/takehome_scorecard.md— confirms both evaluators held: orchestrator 6/6, context-engine 5/5 non-LLM (scenario 6 LLM-gated, 401 on placeholder key, same as Day 26). Adapter modules (takehome/context-engine/memory_system.py+takehome/orchestrator/orchestrator_impl.py) untouched in Phase 6 —evaluate.pyfiles unmodified per rule 17.Verified by running
bash scripts/run_takehome_evals.sh.Why a backfill PR
Mirrors the established Phase-3 / Phase-4 backfill pattern (commits 4cd9ba2, 82afa29, 75158a6 — all post-merge follow-ups for deliverables missed in their main phase PR). Keeps the "one phase = one squash" invariant intact while letting late-noticed SKILL deliverables land.
Test plan
results/metrics.jsonis valid JSON;len(entries)is now 28bash scripts/run_takehome_evals.shprints orchestrator 6/6 + context-engine 5/6 raw (5/5 non-LLM)