From 448b03660ec1c02a3fa4751d9f362ebbddb86225 Mon Sep 17 00:00:00 2001 From: Brad Burch Date: Mon, 3 Aug 2026 10:01:53 -0700 Subject: [PATCH] Refresh docs to match what shipped: tap capture, round types, three LLM clients MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Several CLAUDE.md claims were not merely stale but wrong, which is worse than missing — a future session would have acted on them: - "CoachingLLM has two implementations" (three now, and only the Anthropic one can enforce a JSON schema; the other two share one prose contract) - system audio described as ScreenCaptureKit, which is exactly the thing that recorded an entire Continuity call as silence - deployment floor still 14.0 (14.2 since CoreAudio taps) Also records what actually cost time this session: - the documented AppleScript for opening the main window silently no-ops without a preceding `activate` — it reports success and leaves 0 windows - `entire contents` doesn't just return empty on the Settings pane, it errors; counting the Form's section groups works, and `screencapture` is not a fallback because the shell lacks Screen Recording - comparing mic vs sys RMS across the recordings directory is the fastest capture diagnostic, and `audioop` is gone in Python 3.13 - never run make-app.sh while a recording is live; it overwrites the running executable - a Finder-launched .app has a minimal PATH, so spawning a helper binary needs absolute-path search rather than /usr/bin/env Adds the converse of the existing "don't split what shares a contract" rule: don't narrow something that already serves two callers. That bug shipped twice this session — a shared query filtered for one caller, and a second reader on one file descriptor. README and docs/local-llm.md now describe all three coaching providers, with the subscription path's tradeoffs stated where someone comparing costs will actually see them. Co-Authored-By: Claude Opus 5 --- CLAUDE.md | 44 +++++++++++++++++++++++++++++++++++++------- README.md | 17 ++++++++++++----- docs/local-llm.md | 18 ++++++++++++++++++ 3 files changed, 67 insertions(+), 12 deletions(-) diff --git a/CLAUDE.md b/CLAUDE.md index 41215fb..eaff891 100644 --- a/CLAUDE.md +++ b/CLAUDE.md @@ -25,7 +25,8 @@ Calibration notes specific to this repo: - **Verification is not mechanical.** Judging *what would falsify this* deserves high effort — see "Verify against reality" below. A passing mocked test means very little here. - **Prompt and rubric changes are the highest-judgment work in the codebase** and the least checkable by tests. Never delegate them cheaply; the failure mode is fluent, plausible, and wrong. - **Never let the author verify their own work.** Spawn a separate agent to check a rubric change or a risky diff. Both real bugs this rubric work shipped (an API-rejected schema, a stored all-zero debrief) were invisible to the code that produced them. -- **Don't split what shares a contract.** The scored dimensions live in `DefaultPrompts` + `PromptStore.dimensions(for:)` + both LLM clients + the DB schema. Parallel agents each editing one end produce a diff that compiles and doesn't work — give one agent the whole contract. +- **Don't split what shares a contract.** The scored dimensions live in `DefaultPrompts` + `PromptStore.dimensions(for:)` + all three LLM clients + the DB schema. Parallel agents each editing one end produce a diff that compiles and doesn't work — give one agent the whole contract. +- **And don't overload what already serves two masters.** The converse bug, shipped twice: `sessionsWithTranscript()` feeds both the re-coach sweep and `exportAll`, so narrowing it for one silently broke the other. Before adding a filter to a shared query — or a second reader to a file descriptor — grep its callers and ask whether they want the same thing. ## Toolchain gotcha (read first) @@ -35,6 +36,8 @@ Every `swift` command **must** run under the full Xcode toolchain — the Comman DEVELOPER_DIR=/Applications/Xcode.app/Contents/Developer swift build ``` +The deployment floor is **macOS 14.2** (`Package.swift` and the bundle's `LSMinimumSystemVersion`, which must stay in sync) — CoreAudio process taps need it. SourceKit in the editor may still report "compiling for macOS 14.0" against a stale index; trust `swift build`. + ## Commands ```sh @@ -55,31 +58,53 @@ DEBRIEF_RUN_INTEGRATION=1 … swift test --filter WhisperIntegrationTests # Real Anthropic call — the ONLY check that the built JSON schema is one the API accepts ANTHROPIC_API_KEY=… DEBRIEF_RUN_INTEGRATION=1 … swift test --filter CoachingIntegrationTests + +# Real system-audio capture (needs an output device + audible volume; plays audio via `say`) +DEBRIEF_RUN_INTEGRATION=1 … swift test --filter SystemAudioIntegrationTests + +# Real `claude -p` call — spends Claude subscription quota +DEBRIEF_RUN_INTEGRATION=1 … swift test --filter ClaudeCodeCLIIntegrationTests ``` -**Why `make-app.sh` instead of `swift run`:** macOS attaches Microphone/Screen-Recording TCC prompts to the *bundle*, and `CallAlerts` touches `UNUserNotificationCenter`, which traps in an unbundled binary. Always run the app via the `.app`. The script signs with a stable self-signed identity (`Debrief Local Signing`) so TCC (Microphone/Screen-Recording) grants survive rebuilds — the header comment in the script explains why ad-hoc signing loses grants every build. +**Never run `make-app.sh` while a recording is in progress** — it overwrites the running +executable. Check for a live session first (`manifest.json` with `"finalized":false`, or a +`sys-*.wav` written seconds ago) and stop it before rebuilding. + +**Why `make-app.sh` instead of `swift run`:** macOS attaches Microphone/system-audio TCC prompts to the *bundle*, and `CallAlerts` touches `UNUserNotificationCenter`, which traps in an unbundled binary. Always run the app via the `.app`. The script signs with a stable self-signed identity (`Debrief Local Signing`) so TCC grants survive rebuilds — the header comment in the script explains why ad-hoc signing loses grants every build. + +Hardware capture paths can't be fully unit-tested; verify them against `docs/manual-test-checklist.md`. -Hardware capture paths (mic/screen) can't be unit-tested; verify them against `docs/manual-test-checklist.md`. +**A Finder-launched `.app` inherits a minimal PATH** (`/usr/bin:/bin:/usr/sbin:/sbin`). Anything that spawns a helper binary must search absolute paths — `/usr/bin/env ` resolves in your shell and fails in the bundle. See `ClaudeCodeCLIClient.defaultSearchPaths()`. ## Verify against reality — green tests prove less here than usual -Two whole classes of bug in this app are invisible to `swift test`, and both have shipped: +Three whole classes of bug in this app are invisible to `swift test`, and all three have shipped: 1. **The API rejects a schema the mocks accept.** `outputSchema(dimensions:)` is hand-built JSON; every unit test mocks `URLSession`, so a schema the Messages API 400s on passes the suite cleanly and fails every real debrief. Run `CoachingIntegrationTests` after touching the schema, the clients, or the scored dimensions. 2. **A prompt change that "reads better" and scores worse.** Nothing in the suite can tell you a rubric got *dumber*. Run a real transcript through it and compare against the old prompts before believing a rubric improvement. +3. **Capture that runs, writes correctly-sized files, and records silence.** Every mocked recorder test passed while ScreenCaptureKit wrote full-rate zeros for an entire call. Nothing short of measuring real audio catches it. + +**The fastest capture diagnostic is the chunks already on disk.** Compare mic vs sys RMS across `~/Library/Application Support/Debrief/recordings/*` — the pattern instantly separates "capture is broken" (every session silent) from "this call type is excluded" (one session silent, twenty fine), and a full-length chunk of pure zeros is a different bug from a missing or short chunk. Item 16 of the checklist has a ready-to-paste script. Note `audioop` was **removed in Python 3.13**, so compute RMS with `array` + `math`, not the obvious one-liner. The UI is drivable — don't stop at "it builds." `Debrief` is a `MenuBarExtra` plus `Window(id: "main")`, so it launches with **zero windows**; open the main window with: ```applescript +tell application "Debrief" to activate +delay 1 tell application "System Events" to tell process "Debrief" to click menu item "Debrief" of menu 1 of menu bar item "Window" of menu bar 1 ``` +**The `activate` is load-bearing.** Without it the click reports success and the window +count stays 0 — a silent no-op, not an error. + Then drive it via System Events accessibility. Useful paths (macOS 15, verified): - sidebar tabs: `outline 1 of scroll area 1 of group 1 of splitter group 1 of group 1 of window 1` → `select row N` (1=Sessions, 2=Pipeline, 3=Trends, 4=Settings) - detail pane: `group 2 of splitter group 1 of group 1 of window 1` AppleScript gotchas that cost real time: `right`, `st`, and other reserved words silently break scripts with confusing syntax errors; `entire contents` is flaky on large trees and returns empty rather than erroring — prefer explicit paths and `count of (groups of x)` style probes; compare a known-good state against the state under test rather than trusting one absolute reading. +On the Settings pane specifically, `entire contents` **errors outright** (`-1700`, "can't make into type specifier"). What works is counting the Form's sections — `count of UI elements of scroll area 1 of group 2 of splitter group 1 of group 1 of window 1` returns one `group` per `Section`, so an added section is visible as the count changing. `screencapture` is not a fallback: the shell lacks Screen Recording, so it fails with "could not create image from display". + ## Architecture Swift Package, no `.xcodeproj`. Four library targets + one executable; the executable is the only place they're wired together: @@ -98,6 +123,10 @@ Library targets depend only on `Store` (or nothing) and take injected protocols Mic stream = **you**, system-audio stream = **them**. Two separate WAV streams give perfect two-party attribution with no ML diarization. `TranscriptMerger.merge(you:them:)` interleaves them by timestamp. Everything downstream assumes this mapping. +System audio comes from a **CoreAudio process tap**, not ScreenCaptureKit. SCK scopes capture to audio from *windows on a display*, so a Continuity/cellular call — voiced by a windowless daemon — delivered full-rate digital silence for the whole call. A tap scopes by what reaches the output device instead. Two invariants follow, both easy to undo by accident: +- **The aggregate device has no output sub-device, only the tap.** Binding it to a device UID would go stale when the output switches mid-call. +- **A tap delivers *nothing* while the output is idle** — zero callbacks until audio plays. `SystemAudioRecorder.padSilenceToNow` reconstructs the missing wall-clock time as silence. Remove it and the mic track (which always streams) keeps real time while the system track compresses, so every "them" segment after a gap lands early in the merge. Adding an output sub-device does **not** fix the idling; it was measured. + ### Crash-safety: chunks on disk are the source of truth Audio is flushed to disk in ~30s WAV chunks *during* capture (`WavChunkWriter`). The transcript and debrief are always re-derivable from those chunks, so a mid-interview crash loses nothing and Debrief offers recovery on next launch (`RecordingStore.unfinalizedSessions()`). @@ -111,7 +140,7 @@ Phase state machine: `.idle → .recording → .finalizing(status:) → .idle` ( ### LLM abstraction -`CoachingLLM` protocol has two implementations: `AnthropicClient` (Claude API) and `OpenAICompatibleClient` (Ollama / LM Studio / any OpenAI-compatible server — see `docs/local-llm.md`). `AppEnvironment.resolveLLM()` picks based on `UserDefaults`. The API key lives in a 0600 `secrets.json` under Application Support (`SecretStore`), falling back to `ANTHROPIC_API_KEY`. It is deliberately *not* the Keychain: the self-signed, no-Team-ID signing pins each keychain item to the app's cdhash, which changes every rebuild and re-prompts for the keychain password on every launch/save — the header comment in `SecretStore` records the full dead-end (data-protection keychain and null-ACL both ruled out). Sign with a Team ID cert if you want keychain-grade at-rest encryption back. `coordinator.coaching` is reassignable so a key/model changed in Settings applies to the next debrief without relaunch (`rebuildCoaching()`). +`CoachingLLM` protocol has three implementations: `AnthropicClient` (Claude API), `OpenAICompatibleClient` (Ollama / LM Studio / any OpenAI-compatible server — see `docs/local-llm.md`), and `ClaudeCodeCLIClient` (shells out to `claude -p`, billing a Claude subscription instead of an API key). `AppEnvironment.resolveLLM()` picks based on `UserDefaults`, falling back to `AnthropicClient` when the CLI isn't found. Only the Anthropic path can enforce a JSON schema; the other two share one prose contract — `OpenAICompatibleClient.formatAppendix` states the rules and `decodeCoaching(from:dimensions:)` enforces them, so a change to either half needs the other. The API key lives in a 0600 `secrets.json` under Application Support (`SecretStore`), falling back to `ANTHROPIC_API_KEY`. It is deliberately *not* the Keychain: the self-signed, no-Team-ID signing pins each keychain item to the app's cdhash, which changes every rebuild and re-prompts for the keychain password on every launch/save — the header comment in `SecretStore` records the full dead-end (data-protection keychain and null-ACL both ruled out). Sign with a Team ID cert if you want keychain-grade at-rest encryption back. `coordinator.coaching` is reassignable so a key/model changed in Settings applies to the next debrief without relaunch (`rebuildCoaching()`). Coaching runs with `try?` inside `runFinalize` — a failed debrief leaves the session **retryable** (status `pending`/`failed`), never blocks finalization. Settings → "Retry pending debriefs" calls `retryAllPending()`; Settings → "Re-run debriefs on current rubric" calls `recoachAll()`, which re-coaches **already-complete** sessions too (the only way a prompt change reaches existing debriefs). @@ -124,13 +153,14 @@ Schema gotcha: the Messages API **rejects `minimum`/`maximum` on integer types** ### Prompts: the rubric is data, not code - **Global** prompts are plain markdown in `~/Library/Application Support/Debrief/prompts/` (`PromptStore`, seeded from `DefaultPrompts`): `base.md` + per-round-type overlays. Editing them retunes every debrief without rebuilding. **`ensureDefaults()` only writes a file that doesn't exist** — editing `DefaultPrompts` alone will NOT update an existing install. -- **Scored dimensions are parsed out of the markdown.** Each file's `## Scored dimensions` section (`- key: description`) becomes the JSON-schema keys via `PromptStore.dimensions(for:)` = base's shared delivery dimensions + the overlay's round-specific ones. A new round type is a new `.md` file — no Swift change. +- **Scored dimensions are parsed out of the markdown.** Each file's `## Scored dimensions` section (`- key: description`) becomes the JSON-schema keys via `PromptStore.dimensions(for:)` = base's shared delivery dimensions + the overlay's round-specific ones. A new round type is a new `.md` file — no Swift change. Settings → **Interview types** is a convenience over that folder (create/duplicate/edit/delete), never a second source of truth; editing the files by hand still works. +- **A round type can opt out of scoring** with `transcript-only: true` in its overlay's leading metadata block (only that block is scanned, so rubric prose discussing it can't switch scoring off). `mock_interview` ships this way. The guard lives in `CoachingService.coach()` — all three paths (finalize, `retryAllPending`, `recoachAll`) funnel through there, so one check keeps a practice round from ever billing a call. - **The verdict is the headline, and it is not an average.** `advancement` (`Advancement`, a 4-point forced choice) is elicited from the model *directly* and must never be derived from `scores` — real scorecards co-record the verdict and the ratings. `overallScore` is a secondary trend line only; it is comparable *within* a round type, not across (dimension sets differ), and LLM judges compress toward the top of a 1–5 scale, so a flat mean discriminates poorly. See the provenance comment atop `DefaultPrompts` for what the rubric design is and isn't evidence-backed by. - **Per-interview** grading criteria (`session.customInstructions`, added in migration v2) override the global prompt for one session only. ### Store -GRDB with a `DatabaseMigrator` (`AppDatabase.migrator`) — **schema changes go in a new `registerMigration` block, never by editing an existing one.** In-memory DB for tests (`AppDatabase.inMemory()`), on-disk for the app. LLM feedback (scores, highlights, action items, process notes) is stored as JSON strings in columns; weakness tags are a separate indexed table for trend queries. +GRDB with a `DatabaseMigrator` (`AppDatabase.migrator`) — **schema changes go in a new `registerMigration` block, never by editing an existing one.** (`coachingStatus` is a plain TEXT column with no CHECK constraint, so adding a `CoachingStatus` case needs no migration — `skipped` was added that way. It is terminal, unlike `failed`: both coaching sweeps exclude it, or a transcript-only session would be offered for retry forever.) In-memory DB for tests (`AppDatabase.inMemory()`), on-disk for the app. LLM feedback (scores, highlights, action items, process notes) is stored as JSON strings in columns; weakness tags are a separate indexed table for trend queries. `insertSegments` runs `TranscriptArtifacts.clean` on every write, so the transcript table holds speech and nothing else. WhisperKit narrates non-speech as `[BLANK_AUDIO]`, `[ Silence ]`, diff --git a/README.md b/README.md index c92abc0..d223796 100644 --- a/README.md +++ b/README.md @@ -47,9 +47,9 @@ launch. ## Requirements -- macOS 14 or later (Apple Silicon recommended — transcription uses CoreML/Metal) +- macOS 14.2 or later (Apple Silicon recommended — transcription uses CoreML/Metal) - Full Xcode installed (the build needs XCTest, which the Command Line Tools alone don't ship) -- A Claude API key for the coaching step (transcription is free and local) — or run coaching fully offline against a local model; see [docs/local-llm.md](docs/local-llm.md) +- Something to generate the coaching debrief — transcription is always free and local. Pick one: a **Claude API key** (recommended), a **Claude subscription** via the Claude Code CLI, or a **local model** for fully offline coaching; see [docs/local-llm.md](docs/local-llm.md) ## Build & run @@ -149,8 +149,15 @@ to *Debrief* instead of your terminal. not across (different round types score different dimensions). - **Choose your coaching model.** Settings → Coaching model lets you pick Claude Opus (best quality, default), Sonnet (balanced), or Haiku (fastest, - cheapest) — or switch the provider to a local/OpenAI-compatible server; see - [docs/local-llm.md](docs/local-llm.md). + cheapest) — or switch provider entirely: your **Claude subscription** via the + Claude Code CLI (no API key; costs more tokens per debrief and can hit + subscription rate limits), or a **local/OpenAI-compatible server** for fully + offline coaching; see [docs/local-llm.md](docs/local-llm.md). +- **Practice rounds don't get scored.** Tag a session **Mock Interview** and + Debrief records and transcribes it but never sends it to an LLM — a friend + improvising isn't a calibrated interviewer, and scoring it would put noise in + the same trend lines your real rounds feed. Settings → Interview types lets you + add your own round types and mark any of them transcript-only. - **Export for Claude Cowork.** Settings → Cowork export writes one Markdown file per session to a folder you choose, and keeps writing one on every new debrief from then on. @@ -196,7 +203,7 @@ Swift Package, no `.xcodeproj`. Five targets: | --------------- | ----------------------------------------------------- | | `CaptureKit` | Call detection, mic + system-audio recorders, WAV chunking | | `Transcriber` | WhisperKit wrapper and two-stream transcript merge | -| `CoachingEngine`| Prompt assembly, LLM clients (Claude API and local/OpenAI-compatible), coaching service | +| `CoachingEngine`| Prompt assembly, LLM clients (Claude API, Claude Code CLI, local/OpenAI-compatible), coaching service | | `Store` | GRDB/SQLite schema, records, and trend/pipeline queries | | `DebriefApp` | SwiftUI menu-bar app wiring it all together | diff --git a/docs/local-llm.md b/docs/local-llm.md index 5894b01..9d96a02 100644 --- a/docs/local-llm.md +++ b/docs/local-llm.md @@ -5,6 +5,24 @@ everything on your machine (or use another provider), point Debrief at any OpenAI-compatible server. This guide uses Ollama; LM Studio and remote providers are covered at the end. +## If you're here to avoid API cost, read this first + +Two options avoid a metered API key, and they trade off differently: + +- **A local model** (this guide) — genuinely free, fully offline, no rate + limits, and the only option that keeps interview transcripts entirely on your + machine. Costs you coaching quality; see the expectations below. +- **Your Claude subscription** (Settings → Coaching model → "Claude + subscription") — runs debriefs through the Claude Code CLI, so you get real + Claude quality with no API key. It carries roughly 15–20k extra tokens of the + CLI's own prompt per debrief, subscription rate limits can throttle "Re-run + debriefs on current rubric", and because Claude Code is a coding tool this + isn't a supported integration — a CLI update can break it. Requires the CLI + installed and signed in. + +Both use the same prompt-based response contract described below, since neither +can enforce a JSON schema the way the Claude API path does. + ## Honest expectations Local models produce noticeably weaker coaching than Claude: shallower