From 535facf9d75df7375746401222244b03434f9a1a Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Sat, 19 Sep 2026 01:19:11 -0400 Subject: [PATCH 1/8] chore(harness): drive the e2e routes over the control socket The snapshot route and the wizard-ci MCP server spawn the real binary with `--ci --control-socket` in a PTY and drive it through the control API: the run is released with POST /run, state is long polled, and every commit goes through POST /actions. The in-process host, its store driver, and the action registry are gone; the harness keeps the launcher, the detection picks, the profiles, and the result payload. Harness tests run against the store's ControlDriver and per-flow actions with unchanged goldens, and two process specs exercise each surface end to end. Docs describe the API and the launcher; the workbench environment contract is unchanged. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- .claude/skills/exploring-the-wizard/SKILL.md | 61 +- docs/benchmarking.md | 167 ++-- docs/local-dev.md | 115 +-- e2e-harness/ARCHITECTURE.md | 286 ++++--- .../__tests__/control-actions-parity.test.ts | 54 -- ...-driver.test.ts => control-driver.test.ts} | 59 +- .../__tests__/control-socket-headless.test.ts | 77 ++ .../__tests__/control-socket-tui.test.ts | 82 ++ .../__tests__/e2e-flow-snapshot.test.ts | 7 +- e2e-harness/__tests__/e2e-profile-ask.test.ts | 10 +- e2e-harness/__tests__/e2e-result.test.ts | 4 +- .../__tests__/keyboard-equivalence.test.tsx | 4 +- e2e-harness/__tests__/launch.test.ts | 99 +++ e2e-harness/action-registry.ts | 402 ---------- e2e-harness/e2e-profile.ts | 10 +- e2e-harness/e2e-result.ts | 33 +- e2e-harness/launch.ts | 109 +++ e2e-harness/picks.ts | 101 +++ e2e-harness/wizard-ci-driver.ts | 220 ----- scripts/README.md | 75 +- scripts/tui-host.no-jest.ts | 754 ------------------ scripts/tui-snapshots.no-jest.ts | 401 +++++++++- scripts/wizard-ci-mcp.no-jest.ts | 144 ++-- src/store/control/__tests__/state.test.ts | 10 +- src/store/control/state.ts | 1 + src/store/control/types.ts | 1 + src/store/programs/index.ts | 1 + src/store/state/store-api.ts | 2 - 28 files changed, 1343 insertions(+), 1946 deletions(-) delete mode 100644 e2e-harness/__tests__/control-actions-parity.test.ts rename e2e-harness/__tests__/{wizard-ci-driver.test.ts => control-driver.test.ts} (89%) create mode 100644 e2e-harness/__tests__/control-socket-headless.test.ts create mode 100644 e2e-harness/__tests__/control-socket-tui.test.ts create mode 100644 e2e-harness/__tests__/launch.test.ts delete mode 100644 e2e-harness/action-registry.ts create mode 100644 e2e-harness/launch.ts create mode 100644 e2e-harness/picks.ts delete mode 100644 e2e-harness/wizard-ci-driver.ts delete mode 100644 scripts/tui-host.no-jest.ts diff --git a/.claude/skills/exploring-the-wizard/SKILL.md b/.claude/skills/exploring-the-wizard/SKILL.md index 4310f08d1..23d64733f 100644 --- a/.claude/skills/exploring-the-wizard/SKILL.md +++ b/.claude/skills/exploring-the-wizard/SKILL.md @@ -9,7 +9,7 @@ compatibility: wizard-ci MCP server. metadata: author: posthog - version: '5.0' + version: '6.0' --- # Exploring the wizard as an agent @@ -23,8 +23,8 @@ sequence, and gateway policy. For new exploration, launch the server with `SNAP_HARNESS=pi`; prefer `SNAP_SEQUENCE=orchestrator` for the integration flow. These are server environment variables, not MCP arguments. Restart an existing server to change its environment. See the -[host architecture](../../../e2e-harness/ARCHITECTURE.md) for other programs, -overrides, and current limitations. +[harness architecture](../../../e2e-harness/ARCHITECTURE.md) for the control +API, other programs, and overrides. ## Prepare the run @@ -35,16 +35,14 @@ wizard, so finish recording one app before opening another. - **Detection only:** pass `appDir` and `projectId` (both required strings), with no key. Stop at `auth` without calling `run_agent`. - **Full integration:** reuse the authorized phx key file path, separate gateway - token file path, and project id; - ask only for missing inputs. Prefer `keyFile` so the key stays out of tool - arguments. Set `WIZARD_CI_GATEWAY_TOKEN_FILE` in the MCP server environment - before launch (restart an existing server); it is not an `open_app` argument. - The file must contain an already-issued gateway bearer, not the phx key. - CI does not mint or refresh it. Never print or commit either secret. See - [local credential setup](../../../docs/local-dev.md#credentials-for-local-ci-and-headless-runs). Read the - [credential and region limitations](../../../e2e-harness/ARCHITECTURE.md#current-host-limitations) - before starting: an inherited key can shadow `keyFile`, and the host currently - hardcodes the US region. + token file path, and project id; ask only for missing inputs. Prefer `keyFile` + so the key stays out of tool arguments. Set `WIZARD_CI_GATEWAY_TOKEN_FILE` in + the MCP server environment before launch (restart an existing server); it is + not an `open_app` argument. The file must contain an already-issued gateway + bearer, not the phx key. CI does not mint or refresh it. Never print or commit + either secret. See + [local credential setup](../../../docs/local-dev.md#credentials-for-local-ci-and-headless-runs). + `region` selects the PostHog region for auth and the gateway. - **Questions during the run:** launch the server with `E2E_ASK=true` to keep `wizard_ask` available in this CI session. Handle questions yourself through the actions below; fixed-route answer profiles do not drive the MCP route. @@ -79,31 +77,33 @@ own decisions through the same state and action contract. judging detection. Inspect `session.integration` and `setupQuestions`. 2. Capture `render_screen` before each decision and during task or phase changes. Save numbered frames such as `/tmp/wz-explore-snaps/01-intro.txt`. -3. Commit only actions currently offered. Common choices are below; the - [action registry](../../../e2e-harness/action-registry.ts) defines the full - set. +3. Commit only actions currently offered. Common choices are below; the generic + set lives in + [`src/store/control/actions.ts`](../../../src/store/control/actions.ts) and a + program adds its own through `controlActions` on its steps. 4. For a full run, confirm setup and call `run_agent` at `auth`. Continue reading state and handling overlays while it runs; polling alone cannot answer them. 5. Check `runPhase` (`idle`, `running`, `completed`, `error`), background status, and the rendered outro. On error, capture the frame and reason before dismissing it. An error outro can wait for dismissal while `integration` - still says `running`; a host exit can instead surface as a socket error. + still says `running`; a wizard exit can instead surface as a socket error. 6. After successful agent completion, finish the offered outro and follow-up actions. For the integration flow, `session.skillsComplete` marks the tail's completion. Other programs can have a terminal outro or exit screen. -| Decision | Action and `params` | -| ---------------------------------------- | -------------------------------------------------------------------------------------------- | -| Confirm intro / dismiss blocking outage | `confirm_setup` / `dismiss_outage` | -| Answer setup question | `choose`, `{ key, value }` from `setupQuestions` | -| Answer every question in a pending batch | `answer_question`, `{ answers: { questionId: value } }`; values are strings or string arrays | -| Cancel a question batch | `cancel_question` | -| Accept or decline an optional task | `resolve_notice`, `{ keep: true }` or `{ keep: false }` | -| Finish outro | `dismiss_outro` | -| Record MCP outcome | `set_mcp_outcome`, `{ outcome: "skipped" }` or `{ outcome: "installed", clients: [...] }` | -| Dismiss suggested prompts / Slack step | `dismiss` / `dismiss_slack` | -| Record keep-skills choice | `keep_skills`, `{ kept: true }` or `{ kept: false }` | +| Decision | Action and `params` | +| ---------------------------------------- | ---------------------------------------------------------------------------------------------------------------- | +| Confirm intro / dismiss blocking outage | `confirm_setup` / `dismiss_outage` | +| Answer setup question | `choose`, `{ key, value }` from `setupQuestions` | +| Answer every question in a pending batch | `answer_question`, `{ answers: { questionId: value } }`; values are strings or string arrays | +| Cancel a question batch | `cancel_question` | +| Accept or decline an optional task | `resolve_notice`, `{ keep: true }` or `{ keep: false }` | +| Finish outro | `dismiss_outro` | +| Record MCP outcome | `set_mcp_outcome`, `{ outcome: "skipped" }` or `{ outcome: "installed", clients: [...] }` | +| Dismiss suggested prompts / Slack step | `dismiss` / `dismiss_slack` | +| Record keep-skills choice | `keep_skills`, `{ kept: true }` or `{ kept: false }` | +| Pick the project on a detect screen | `pick_integration_target`, `{ path, integration }`; source maps: `pick_source_maps_project`, `{ variant, path }` | MCP and keep-skills actions commit store state; recording an outcome does not perform the corresponding installation or cleanup. Report which outcomes were @@ -129,5 +129,6 @@ The shared log is `/tmp/posthog-wizard.log`. Record its byte count before a run and read from that count plus one afterward. Run sweeps serially so their logs remain attributable. `read_state` omits `frameworkContext`; an empty `setupQuestions` list alone does not prove a router mode. When necessary, -inspect the detector under [`src/store/frameworks/`](../../../src/store/frameworks/) against -the same fixture. +inspect the detector under +[`src/store/frameworks/`](../../../src/store/frameworks/) against the same +fixture. diff --git a/docs/benchmarking.md b/docs/benchmarking.md index e5fadf4f7..17a4bc6e8 100644 --- a/docs/benchmarking.md +++ b/docs/benchmarking.md @@ -29,29 +29,29 @@ Selection criteria, checked in this order: 1. **Real product, in production** — an open-source app people actually run (stars are a proxy; a hosted instance is better evidence). 2. **Single-app repo** — reject monorepos: fetch the repo's top-level listing - and reject on `pnpm-workspace.yaml`, `turbo.json`, `lerna.json`, or - top-level `apps/`/`packages/` directories. + and reject on `pnpm-workspace.yaml`, `turbo.json`, `lerna.json`, or top-level + `apps/`/`packages/` directories. 3. **Greenfield** — grep the repo for `posthog` (manifest and source). An app that already integrates PostHog measures augmentation discipline, not integration quality; keep at most one such app and exclude it from quality scoring. 4. **Framework coverage** — spread picks across the frameworks the wizard supports; results do not transfer between them. -5. **Locally installable** — its toolchain (node/python/php/ruby/gradle) - exists on the bench machine, or its runs will fail for reasons that are - yours, not the model's. +5. **Locally installable** — its toolchain (node/python/php/ruby/gradle) exists + on the bench machine, or its runs will fail for reasons that are yours, not + the model's. Apps used in the 2026-07 benchmark, as worked examples of the spread: -| app | upstream | stack | -|---|---|---| -| Maybe | `maybe-finance/maybe` | Rails | -| Outline | `outline/outline` | React + Koa / TS | -| WordPress-Android | `wordpress-mobile/WordPress-Android` | native Kotlin | -| healthchecks | `healthchecks/healthchecks` | Django | -| Firefly III | `firefly-iii/firefly-iii` | Laravel, server-rendered | -| Monica | `monicahq/monica` | Laravel + Inertia/Vue | -| Papermark | `mfts/papermark` | Next.js — already shipped posthog-js; kept as the augment-existing case, excluded from quality scoring | +| app | upstream | stack | +| ----------------- | ------------------------------------ | ------------------------------------------------------------------------------------------------------ | +| Maybe | `maybe-finance/maybe` | Rails | +| Outline | `outline/outline` | React + Koa / TS | +| WordPress-Android | `wordpress-mobile/WordPress-Android` | native Kotlin | +| healthchecks | `healthchecks/healthchecks` | Django | +| Firefly III | `firefly-iii/firefly-iii` | Laravel, server-rendered | +| Monica | `monicahq/monica` | Laravel + Inertia/Vue | +| Papermark | `mfts/papermark` | Next.js — already shipped posthog-js; kept as the augment-existing case, excluded from quality scoring | ## Setup (from a bare machine) @@ -90,21 +90,21 @@ Define each config as a `WIZARD_CI_FLAG_OVERRIDES` JSON, plus the baseline as Everything below ships in this repo (`wizard/`) and its workbench (`wizard-workbench/`); paths are from each repo's root. -- **Headless run (snapshotting CI harness):** `wizard/scripts/tui-snapshots.no-jest.ts` - spawns the real TUI (`wizard/scripts/tui-host.no-jest.ts`) in a PTY via +- **Headless run (snapshotting CI harness):** + `wizard/scripts/tui-snapshots.no-jest.ts` spawns the real wizard + (`bin.ts --ci --control-socket`) in a PTY via `wizard/e2e-harness/tui-capture.ts`, self-drives the fixed e2e profile - (`wizard/e2e-harness/wizard-ci-driver.ts`, `wizard/e2e-harness/profiles.ts`) - through auth, the agent run, and the outro, and writes each screen as an - `NN-.ans` frame. An `NN-outro.ans` frame is the flow-completion - signal. `tsx` runs source — no build step. Invocation: see the run-cell - recipe below. + (`wizard/e2e-harness/profiles.ts`) over the control socket through auth, the + agent run, and the outro, and writes each screen as an `NN-.ans` + frame. An `NN-outro.ans` frame is the flow-completion signal. `tsx` runs + source — no build step. Invocation: see the run-cell recipe below. - **Config selection:** the flag axis is `wizard-orchestrator` (on → the orchestrator on pi, per-task models from context-mill frontmatter; off → the linear anthropic default). Per-stage variations ride - `wizard-orchestrator-override` payloads (`{stage: {model?, effort?}}`, - variant keys in `wizard/src/agent/runner/switchboard/flags/schemes.ts`). - The baseline is `{"wizard-orchestrator":"false"}` — never an empty override, - or live remote flags leak into the baseline. + `wizard-orchestrator-override` payloads (`{stage: {model?, effort?}}`, variant + keys in `wizard/src/agent/runner/switchboard/flags/schemes.ts`). The baseline + is `{"wizard-orchestrator":"false"}` — never an empty override, or live remote + flags leak into the baseline. ## Running one cell @@ -147,19 +147,19 @@ files=$(git -C "$WORK" diff --name-only main integ | wc -l | tr -d ' ')" \ | tee "$OUT/result.txt" ``` -Both commits are `--no-verify` (see Traps). The diff, frames, stdout, and -result line are the cell's complete artifact set — everything else (the shared -debug log) is unreliable under parallelism. +Both commits are `--no-verify` (see Traps). The diff, frames, stdout, and result +line are the cell's complete artifact set — everything else (the shared debug +log) is unreliable under parallelism. ## Running the matrix -- One app at a time; per app, launch its configs in parallel (≤4 on one - machine) and `wait`. Contention inflates absolute times roughly uniformly. +- One app at a time; per app, launch its configs in parallel (≤4 on one machine) + and `wait`. Contention inflates absolute times roughly uniformly. - Cost: anthropic-harness cells report `modelUsage.costUSD` in - `/tmp/posthog-wizard.log` — zero the log before each app's wave and slice - the block per baseline run. pi-harness cells do not persist token totals; - add a temporary hook in the pi harness success path that writes the session - token stats to a per-run file, and price them at list rates. + `/tmp/posthog-wizard.log` — zero the log before each app's wave and slice the + block per baseline run. pi-harness cells do not persist token totals; add a + temporary hook in the pi harness success path that writes the session token + stats to a per-run file, and price them at list rates. - Rerun any anomalous cell solo (zeroed log, no parallelism) before drawing a conclusion from it. @@ -167,61 +167,61 @@ debug log) is unreliable under parallelism. - **Target-app git hooks.** Your `git commit` runs the app's husky/lint-staged hooks if a prior install activated them; a failing hook silently rolls the - tree back and the run measures as zero-diff. Always commit `--no-verify`. - On any zero-diff run, check `git stash list` before believing it. -- **Zero-diff has many causes.** Distinguish: the agent honestly declined - (read its setup report), the agent's tool calls failed, your harness ate the - work, or `.gitignore` hid it (env files never show in diffs). Attribute - before you blame the model. + tree back and the run measures as zero-diff. Always commit `--no-verify`. On + any zero-diff run, check `git stash list` before believing it. +- **Zero-diff has many causes.** Distinguish: the agent honestly declined (read + its setup report), the agent's tool calls failed, your harness ate the work, + or `.gitignore` hid it (env files never show in diffs). Attribute before you + blame the model. - **"Reached the outro" is not success.** The flow completes even when nothing was integrated. Treat completion as outro + a non-trivial diff. - **Parallel runs interleave shared state.** The shared debug log cannot be attributed per-run; capture everything per-run or run solo when attribution matters. - **Sandbox/allowlist gaps look like model failures.** If a config produces - empty or thin work, check whether a blocked command (package-manager - install, formatter) caused it, and whether other models worked around the - same block. File the gap; exclude the affected cells. -- **Repo-wide format scripts.** An agent running the app's `format`/`lint - --fix` buries its real diff under hundreds of churn files. Count "real - files" excluding scaffolding, lockfiles, env files — and read a sample of - the churn before scoring. -- **A stale credential fails silently mid-batch.** Read the key per run, not - per session. + empty or thin work, check whether a blocked command (package-manager install, + formatter) caused it, and whether other models worked around the same block. + File the gap; exclude the affected cells. +- **Repo-wide format scripts.** An agent running the app's `format`/`lint --fix` + buries its real diff under hundreds of churn files. Count "real files" + excluding scaffolding, lockfiles, env files — and read a sample of the churn + before scoring. +- **A stale credential fails silently mid-batch.** Read the key per run, not per + session. ## Judging Use the wizard-workbench PR evaluator's rubric — do not invent your own. It lives at `wizard-workbench/services/pr-evaluator/`: -- **Rubric criteria:** `wizard-workbench/services/pr-evaluator/prompts/evaluation.md` - — per-item YES/NO/N-A checks grouped into four dimensions. -- **Scoring math:** `wizard-workbench/services/pr-evaluator/evaluator.ts` — - each dimension scores `max(1, round(pass_rate × 5))` over its applicable - items; confidence = `min(app_sanity, round(mean of the four))`. +- **Rubric criteria:** + `wizard-workbench/services/pr-evaluator/prompts/evaluation.md` — per-item + YES/NO/N-A checks grouped into four dimensions. +- **Scoring math:** `wizard-workbench/services/pr-evaluator/evaluator.ts` — each + dimension scores `max(1, round(pass_rate × 5))` over its applicable items; + confidence = `min(app_sanity, round(mean of the four))`. - **Automated run:** from `wizard-workbench/`, `pnpm run evaluate --branch --base --test-run` (needs `POSTHOG_PERSONAL_API_KEY`; judge model via `EVALUATOR_MODEL`). Output lands in `wizard-workbench/test-evaluations//` as `rubric.json` + `scores.json`. -- **Manual run:** an agent applies the same rubric directly to each cell's - diff — faster for many cells, and what the 2026-07 benchmark did. Either - way, report the four dimensions under their full names, 1–5 each - (5 production-ready, 3 works with real issues, 1 broken or empty): +- **Manual run:** an agent applies the same rubric directly to each cell's diff + — faster for many cells, and what the 2026-07 benchmark did. Either way, + report the four dimensions under their full names, 1–5 each (5 + production-ready, 3 works with real issues, 1 broken or empty): -- **Files** (`file_analysis`) — right files touched, nothing unrelated, - imports valid +- **Files** (`file_analysis`) — right files touched, nothing unrelated, imports + valid - **App** (`app_sanity`) — nothing broken: builds, existing code and configs preserved, changes minimal - **PostHog** (`posthog_implementation`) — SDK installed, initialized at the - right entry points, env-based keys, real distinct id, identify, error - tracking + right entry points, env-based keys, real distinct id, identify, error tracking - **Events** (`event_quality`) — real user actions, useful properties, no PII, consistent names -Verify claims against the diff (grep for `capture`/`identify` call sites, -check the init file, check the manifest), and build or typecheck where cheap. -Judge the same subset of apps for every config you compare. +Verify claims against the diff (grep for `capture`/`identify` call sites, check +the init file, check the manifest), and build or typecheck where cheap. Judge +the same subset of apps for every config you compare. ## Publishing evidence @@ -230,48 +230,55 @@ inspectable: 1. Fork each app to the operator's account. 2. Pin a `bench-base` branch at the exact commit the runs used. -3. Per cell: branch from `bench-base`, apply the **sanitized** patch, push, - open a draft PR against `bench-base`. +3. Per cell: branch from `bench-base`, apply the **sanitized** patch, push, open + a draft PR against `bench-base`. 4. Sanitize before anything touches a public fork: drop env files, wizard - scaffolding, and lockfiles from the patch; redact every token literal. - Verify zero secrets in the pushed diff before opening the PR. + scaffolding, and lockfiles from the patch; redact every token literal. Verify + zero secrets in the pushed diff before opening the PR. ## Report template ```markdown # — model benchmark - + ## Summary — configs that completed everywhere -| config | completed | median time | median cost | quality (judged on) | - ## Results - -| config | | | … | -