Skip to content

experiment: backfill Phase 6 deliverables (metrics journal + scorecard re-validation) - #14

Merged
Mark007-R merged 1 commit into
mainfrom
backfill/phase6-metrics-and-scorecard
Jun 4, 2026
Merged

Mark007-R merged 1 commit into
mainfrom
backfill/phase6-metrics-and-scorecard

Conversation

@Mark007-R

Copy link
Copy Markdown
Owner

Follow-up to PR #13 — appends two SKILL deliverables that landed late.

What was missed

  1. results/metrics.json append-only journal — SKILL Step 4 mandates a per-day snapshot in results/metrics.json (append, don't overwrite). Phase 6 (Days 29-32) shipped without those four entries. Last entry was Day 28 (2026-05-31).
  2. Weekly takehome scorecard re-validation — results/takehome_scorecard.md is tracked weekly from Day 11. Last update was Day 26; a Phase 6 re-validation was due at the end of Day 32.

What this PR adds

  • 4 new entries in results/metrics.json — Day 29 / 30 / 31 / 32 with component, feature, tests_passing, tests_added_today, takehome_score, and notes. Day 32 carries a phase_wrap: true flag and a phase_6_summary block (4 deliverables, 79 tests, suite 766 → 845, branch + squash commit).
  • Phase 6 weekly re-validation block in results/takehome_scorecard.md — confirms both evaluators held: orchestrator 6/6, context-engine 5/5 non-LLM (scenario 6 LLM-gated, 401 on placeholder key, same as Day 26). Adapter modules (takehome/context-engine/memory_system.py + takehome/orchestrator/orchestrator_impl.py) untouched in Phase 6 — evaluate.py files unmodified per rule 17.

Verified by running bash scripts/run_takehome_evals.sh.

Why a backfill PR

Mirrors the established Phase-3 / Phase-4 backfill pattern (commits 4cd9ba2, 82afa29, 75158a6 — all post-merge follow-ups for deliverables missed in their main phase PR). Keeps the "one phase = one squash" invariant intact while letting late-noticed SKILL deliverables land.

Test plan

  • results/metrics.json is valid JSON; len(entries) is now 28
  • bash scripts/run_takehome_evals.sh prints orchestrator 6/6 + context-engine 5/6 raw (5/5 non-LLM)
  • Pre-commit hook passes (secret scan + takehome checksums + local-only)

…ay 29-32 entries to results/metrics.json (append-only journal per SKILL Step 4) + add Phase 6 weekly takehome scorecard re-validation row (orch 6/6 + ce 5/5 held across all four additive days) (Phase 6 backfill)
Copilot AI review requested due to automatic review settings June 4, 2026 02:52
@Mark007-R
Mark007-R merged commit 0518b2a into main Jun 4, 2026
1 check passed
@Mark007-R
Mark007-R deleted the backfill/phase6-metrics-and-scorecard branch June 4, 2026 02:53

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Backfills Phase 6 “late” SKILL deliverables by extending the experiment artifacts: the append-only daily metrics journal and the weekly takehome scorecard re-validation, keeping Phase 6’s “one phase = one squash” invariant intact while still recording the missing end-of-phase evidence.

Changes:

  • Appends Phase 6 Day 29–32 snapshots to results/metrics.json, including a Day 32 wrap summary block.
  • Adds a Phase 6 weekly takehome scorecard re-validation section to results/takehome_scorecard.md with the documented eval command and results.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 6 comments.

File Description
results/metrics.json Adds Day 29–32 journal entries plus a Phase 6 wrap summary.
results/takehome_scorecard.md Adds a Phase 6 weekly re-validation block with evaluator outcomes and run command.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread results/metrics.json
@@ -1,5 +1,5 @@
{
"_about": "PennyCore \u00e2\u20ac\u201d append-only metrics journal. Each entry is a per-day snapshot; never overwrite past entries. Defined in SKILL Step 4 + the PRIMARY METRICS section. Initialized on Day 7 as a backfill (the journal was missed on Days 5 and 6 when the first measurable code shipped \u00e2\u20ac\u201d those days' numbers are NOT retroactively reconstructed; the day_05 and day_06 entries record only the binary 'feature shipped' fact so future readers know the journal started here). Comparison metrics (quality, cost, quality/$ ratios) start populating Day 12 when Phase 3 benchmarks land.",
"_about": "PennyCore — append-only metrics journal. Each entry is a per-day snapshot; never overwrite past entries. Defined in SKILL Step 4 + the PRIMARY METRICS section. Initialized on Day 7 as a backfill (the journal was missed on Days 5 and 6 when the first measurable code shipped — those days' numbers are NOT retroactively reconstructed; the day_05 and day_06 entries record only the binary 'feature shipped' fact so future readers know the journal started here). Comparison metrics (quality, cost, quality/$ ratios) start populating Day 12 when Phase 3 benchmarks land.",
Comment thread results/metrics.json
"naive_tokens": 1059,
"ratio": 1.483,
"interpretation": "at budget=2000 with 50-message history, the recency brief is LARGER than naive concat because it adds channel markers, timestamps, and action segments that naive concat omits. The ratio<1 win the Phase 3 hybrid is supposed to deliver only kicks in when the borrower history exceeds the budget \u00e2\u20ac\u201d the 200-history-pair benchmark on Day 12 quantifies that."
"interpretation": "at budget=2000 with 50-message history, the recency brief is LARGER than naive concat because it adds channel markers, timestamps, and action segments that naive concat omits. The ratio<1 win the Phase 3 hybrid is supposed to deliver only kicks in when the borrower history exceeds the budget — the 200-history-pair benchmark on Day 12 quantifies that."
Comment thread results/metrics.json
"event_id_received": "evt_963ff24cf2914a90b1318adf",
"round_trip_latency_observed_sec": "<1 (sleep 1 between publish and recent-read; not microbenchmarked today)",
"tenant_seeded": "INSERT INTO tenants (id, name, slug) VALUES ('acme', 'Acme Bank', 'acme-bank') \u00e2\u20ac\u201d required by Day-3 schema FK"
"tenant_seeded": "INSERT INTO tenants (id, name, slug) VALUES ('acme', 'Acme Bank', 'acme-bank') — required by Day-3 schema FK"
Comment thread results/metrics.json
"anomaly_detected": "notify_loan_officer",
"system_event": "no_op",
"unknown": "notify_loan_officer (conservative default \u00e2\u20ac\u201d escalate to human)"
"unknown": "notify_loan_officer (conservative default — escalate to human)"
Comment thread results/metrics.json
Comment on lines +182 to +183
"make_buffered_handler(_recent_events) — Day 8",
"make_planning_handler(_recent_proposals) — Day 9"
Comment thread results/metrics.json
}
},
"notes": "Day 9 planner is mock-mode-only in the live system (placeholder ANTHROPIC_API_KEY); LLM path covered by fake-client tests. Real-LLM exercise deferred to user-driven key setup or Phase 3 benchmark days. docker-compose up verification re-run after initial commit to close SKILL Step 4 'actually RUN the system' loop \u00e2\u20ac\u201d the bind-mounted orchestrator/ directory picks up planner.py without rebuild."
"notes": "Day 9 planner is mock-mode-only in the live system (placeholder ANTHROPIC_API_KEY); LLM path covered by fake-client tests. Real-LLM exercise deferred to user-driven key setup or Phase 3 benchmark days. docker-compose up verification re-run after initial commit to close SKILL Step 4 'actually RUN the system' loop — the bind-mounted orchestrator/ directory picks up planner.py without rebuild."
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants