diff --git a/.claude/skills/exploring-the-wizard/SKILL.md b/.claude/skills/exploring-the-wizard/SKILL.md index 4310f08d1..d7a302517 100644 --- a/.claude/skills/exploring-the-wizard/SKILL.md +++ b/.claude/skills/exploring-the-wizard/SKILL.md @@ -9,7 +9,7 @@ compatibility: wizard-ci MCP server. metadata: author: posthog - version: '5.0' + version: '6.0' --- # Exploring the wizard as an agent @@ -23,8 +23,8 @@ sequence, and gateway policy. For new exploration, launch the server with `SNAP_HARNESS=pi`; prefer `SNAP_SEQUENCE=orchestrator` for the integration flow. These are server environment variables, not MCP arguments. Restart an existing server to change its environment. See the -[host architecture](../../../e2e-harness/ARCHITECTURE.md) for other programs, -overrides, and current limitations. +[harness architecture](../../../e2e-harness/ARCHITECTURE.md) for the control +API, other programs, and overrides. ## Prepare the run @@ -35,16 +35,14 @@ wizard, so finish recording one app before opening another. - **Detection only:** pass `appDir` and `projectId` (both required strings), with no key. Stop at `auth` without calling `run_agent`. - **Full integration:** reuse the authorized phx key file path, separate gateway - token file path, and project id; - ask only for missing inputs. Prefer `keyFile` so the key stays out of tool - arguments. Set `WIZARD_CI_GATEWAY_TOKEN_FILE` in the MCP server environment - before launch (restart an existing server); it is not an `open_app` argument. - The file must contain an already-issued gateway bearer, not the phx key. - CI does not mint or refresh it. Never print or commit either secret. See - [local credential setup](../../../docs/local-dev.md#credentials-for-local-ci-and-headless-runs). Read the - [credential and region limitations](../../../e2e-harness/ARCHITECTURE.md#current-host-limitations) - before starting: an inherited key can shadow `keyFile`, and the host currently - hardcodes the US region. + token file path, and project id; ask only for missing inputs. Prefer `keyFile` + so the key stays out of tool arguments. Set `WIZARD_CI_GATEWAY_TOKEN_FILE` in + the MCP server environment before launch (restart an existing server); it is + not an `open_app` argument. The file must contain an already-issued gateway + bearer, not the phx key. CI does not mint or refresh it. Never print or commit + either secret. See + [local credential setup](../../../docs/local-dev.md#credentials-for-local-ci-and-headless-runs). + `region` selects the PostHog region for auth and the gateway. - **Questions during the run:** launch the server with `E2E_ASK=true` to keep `wizard_ask` available in this CI session. Handle questions yourself through the actions below; fixed-route answer profiles do not drive the MCP route. @@ -64,11 +62,16 @@ It exposes exactly these tools: | `run_agent` | None | Starts the real program in the background and returns immediately | Use `read_state.actions` for action ids and parameters; there is no -`list_actions` MCP tool. Framework identity is `session.integration`, with +`list_actions` MCP tool. The state mirrors the wizard's store: `session` holds +the detection, setup, run, and follow-up fields (`session.runPhase`, +`session.pendingQuestion`, `session.taskNotice`, `session.outroData`, a redacted +`session.frameworkContext`), beside `tasks`, `setupQuestions`, and `actions`. +Framework identity is `session.integration`, with `session.detectedFrameworkLabel` and `session.detectionComplete`. The separate top-level `integration` field is the background status: `idle`, `running`, -`done`, or `failed`; `integrationError` holds a caught failure. Those two fields -are added by `read_state` and are absent from `perform_action` replies. +`done`, or `failed`, derived from `session.runPhase`; `integrationError` holds a +caught failure. Those two fields are added by `read_state` and are absent from +`perform_action` replies. ## Drive and record @@ -79,31 +82,34 @@ own decisions through the same state and action contract. judging detection. Inspect `session.integration` and `setupQuestions`. 2. Capture `render_screen` before each decision and during task or phase changes. Save numbered frames such as `/tmp/wz-explore-snaps/01-intro.txt`. -3. Commit only actions currently offered. Common choices are below; the - [action registry](../../../e2e-harness/action-registry.ts) defines the full - set. +3. Commit only actions currently offered. Common choices are below; the generic + set lives in + [`src/store/control/actions.ts`](../../../src/store/control/actions.ts) and a + program adds its own through `controlActions` on its steps. 4. For a full run, confirm setup and call `run_agent` at `auth`. Continue reading state and handling overlays while it runs; polling alone cannot answer them. -5. Check `runPhase` (`idle`, `running`, `completed`, `error`), background - status, and the rendered outro. On error, capture the frame and reason before - dismissing it. An error outro can wait for dismissal while `integration` - still says `running`; a host exit can instead surface as a socket error. +5. Check `session.runPhase` (`idle`, `running`, `completed`, `error`), + background status, and the rendered outro. On error, capture the frame and + reason before dismissing it. An error outro can wait for dismissal while + `integration` still says `running`; a wizard exit can instead surface as a + socket error. 6. After successful agent completion, finish the offered outro and follow-up actions. For the integration flow, `session.skillsComplete` marks the tail's completion. Other programs can have a terminal outro or exit screen. -| Decision | Action and `params` | -| ---------------------------------------- | -------------------------------------------------------------------------------------------- | -| Confirm intro / dismiss blocking outage | `confirm_setup` / `dismiss_outage` | -| Answer setup question | `choose`, `{ key, value }` from `setupQuestions` | -| Answer every question in a pending batch | `answer_question`, `{ answers: { questionId: value } }`; values are strings or string arrays | -| Cancel a question batch | `cancel_question` | -| Accept or decline an optional task | `resolve_notice`, `{ keep: true }` or `{ keep: false }` | -| Finish outro | `dismiss_outro` | -| Record MCP outcome | `set_mcp_outcome`, `{ outcome: "skipped" }` or `{ outcome: "installed", clients: [...] }` | -| Dismiss suggested prompts / Slack step | `dismiss` / `dismiss_slack` | -| Record keep-skills choice | `keep_skills`, `{ kept: true }` or `{ kept: false }` | +| Decision | Action and `params` | +| ---------------------------------------- | ---------------------------------------------------------------------------------------------------------------- | +| Confirm intro / dismiss blocking outage | `confirm_setup` / `dismiss_outage` | +| Answer setup question | `choose`, `{ key, value }` from `setupQuestions` | +| Answer every question in a pending batch | `answer_question`, `{ answers: { questionId: value } }`; values are strings or string arrays | +| Cancel a question batch | `cancel_question` | +| Accept or decline an optional task | `resolve_notice`, `{ keep: true }` or `{ keep: false }` | +| Finish outro | `dismiss_outro` | +| Record MCP outcome | `set_mcp_outcome`, `{ outcome: "skipped" }` or `{ outcome: "installed", clients: [...] }` | +| Dismiss suggested prompts / Slack step | `dismiss` / `dismiss_slack` | +| Record keep-skills choice | `keep_skills`, `{ kept: true }` or `{ kept: false }` | +| Pick the project on a detect screen | `pick_integration_target`, `{ path, integration }`; source maps: `pick_source_maps_project`, `{ variant, path }` | MCP and keep-skills actions commit store state; recording an outcome does not perform the corresponding installation or cleanup. Report which outcomes were @@ -127,7 +133,8 @@ copies after recording their results. The shared log is `/tmp/posthog-wizard.log`. Record its byte count before a run and read from that count plus one afterward. Run sweeps serially so their logs -remain attributable. `read_state` omits `frameworkContext`; an empty -`setupQuestions` list alone does not prove a router mode. When necessary, -inspect the detector under [`src/store/frameworks/`](../../../src/store/frameworks/) against -the same fixture. +remain attributable. `read_state` shows `session.frameworkContext` after +redaction; an empty `setupQuestions` list alone does not prove a router mode. +When necessary, inspect the detector under +[`src/store/frameworks/`](../../../src/store/frameworks/) against the same +fixture. diff --git a/.claude/skills/wizard-development/SKILL.md b/.claude/skills/wizard-development/SKILL.md index a751e2a78..d7395f719 100644 --- a/.claude/skills/wizard-development/SKILL.md +++ b/.claude/skills/wizard-development/SKILL.md @@ -19,7 +19,7 @@ infrastructure should consume those boundaries. | Concern | Owner | | ----------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| Framework detection, context, env conventions | [FrameworkConfig](../../../src/store/framework-config.ts) and [framework configs](../../../src/frameworks/) | +| Framework detection, context, env conventions | [FrameworkConfig](../../../src/store/framework-config.ts) and [framework configs](../../../src/store/frameworks/) | | Integration instructions and orchestrator flows/tasks | [context-mill](https://github.com/PostHog/context-mill) | | Programs, steps, prerequisites and outcomes | [programs](../../../src/store/programs/) | | Sequence, harness, model and effort selection | [switchboard](../../../src/agent/runner/switchboard/) | diff --git a/.claude/skills/wizard-development/references/ARCHITECTURE.md b/.claude/skills/wizard-development/references/ARCHITECTURE.md index f13384acc..75183874c 100644 --- a/.claude/skills/wizard-development/references/ARCHITECTURE.md +++ b/.claude/skills/wizard-development/references/ARCHITECTURE.md @@ -19,9 +19,7 @@ lifecycle in [runner/index.ts](../../../../src/agent/runner/index.ts) resolves a program's `run` definition, calls shared bootstrap, selects a binding, dispatches the -sequence, and flushes the scanner report on cleanup. The old -[agent-runner.ts](../../../../src/agent/runner/index.ts) is a compatibility -export. +sequence, and flushes the scanner report on cleanup. | Layer | Source and responsibility | | --------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | @@ -57,7 +55,8 @@ Do not migrate a linear program merely by changing its binding if it depends on these hooks. Inspect the orchestrator's flow and completion path instead. [Metrics](../../../../src/store/programs/metrics/) is a current Pi/orchestrator example. Native command modules still need registration in -[bin.ts](../../../../bin.ts); screen sequences derive from the program registry. +[src/cli/main.ts](../../../../src/cli/main.ts); flows derive from the program +registry. ## Switchboard contract @@ -138,7 +137,7 @@ durable credentials. ## UI state and agent output Business logic uses [WizardUI](../../../../src/store/ui/wizard-ui.ts) through -`getUI()`. [InkUI](../../../../src/store/ui/store-ui.ts) updates the TUI store; +`getUI()`. [StoreUI](../../../../src/store/ui/store-ui.ts) updates the store; [LoggingUI](../../../../src/tui/console/logging-ui.ts) is available for noninteractive callers that select it. A missing TTY does not automatically mean an arbitrary caller uses LoggingUI; snapshot CI drives Ink in a PTY. @@ -152,11 +151,11 @@ session event handlers. Orchestrated tasks also have queue and handoff state. Do not assume all harness output passes through `handleSDKMessage`. Session changes go through explicit store setters. They emit updates, -re-evaluate gates, detect transitions, and refresh rendering. The -[router](../../../../src/tui/router.ts) resolves overlays first, then the first -visible incomplete screen from -[screen-sequences.ts](../../../../src/tui/screen-sequences.ts). Those sequences -are projected from registered program steps. Change the state/predicate that +re-evaluate gates, detect transitions, and refresh rendering. The store's +[flow resolution](../../../../src/store/state/flow-resolution.ts) resolves +interrupts first, then the first visible incomplete step of the program's flow +([flowFor](../../../../src/store/programs/flow-for.ts)). Those flows are +projected from registered program steps. Change the state/predicate that represents progress rather than adding imperative navigation. ## MCP and instrumentation @@ -167,9 +166,8 @@ them differently; inspect the selected harness rather than assuming identical tool names or discovery. Context-mill supplies skills and flow/task prompts. [Middleware](../../../../src/agent/middleware/) provides opt-in message/phase -instrumentation. The linear sequence creates the benchmark pipeline; there is no -pipeline construction in the compatibility `agent-runner.ts`. Inspect the actual -consumer before extending instrumentation to another sequence or harness. +instrumentation. The linear sequence creates the benchmark pipeline. Inspect the +actual consumer before extending instrumentation to another sequence or harness. ## Surfaces and the control API @@ -184,9 +182,10 @@ lists what it owns and may import: | `src/cli` | argv, command tree, runners that sequence runs and pass context, `ControlHooks` | every surface, through its public entries only | Cross-surface imports go through `@store`, `@store/types`, `@store/programs`, -`@agent`, `@agent/types`, `@tui`, `@tui/types`, and `@tui/console`. -`src/__tests__/architecture` enforces the matrix and the public-entry rule; -`tsc -b tsconfig.solution.json` mirrors it with project references. +`@agent`, `@agent/types`, `@tui`, `@tui/types`, and `@tui/console`; the two cli +runners load `@store/control` lazily. `src/__tests__/architecture` enforces the +matrix and the public-entry rule; `tsc -b tsconfig.solution.json` mirrors it +with project references. `--control-socket ` serves an HTTP/1.1 API over a unix socket from `src/store/control`: state with long polling, actions that call one store setter @@ -203,13 +202,14 @@ Run it: ```bash # headless, every build; the key travels in the environment +mkdir -p /tmp/w POSTHOG_WIZARD_API_KEY=phx_... WIZARD_CI_GATEWAY_TOKEN_FILE=/path/to/token \ npx tsx bin.ts --headless-DONOTUSE-EXPERIMENTAL --control-socket /tmp/w/w.sock \ --project-id --region us --install-dir /tmp/app curl -s --unix-socket /tmp/w/w.sock -X POST -H 'content-type: application/json' -d '{}' http://localhost/detect curl -s --unix-socket /tmp/w/w.sock -X POST -H 'content-type: application/json' \ -d '{"programId":"posthog-integration"}' http://localhost/runs -curl -s --unix-socket /tmp/w/w.sock 'http://localhost/state?wait=60000&since=0' | jq '.state.run, .state.tasks' +curl -s --unix-socket /tmp/w/w.sock 'http://localhost/state?wait=60000&since=0' | jq '.state.session.runPhase, .state.tasks' curl -s --unix-socket /tmp/w/w.sock http://localhost/runs curl -s --unix-socket /tmp/w/w.sock -X POST http://localhost/shutdown diff --git a/AGENTS.md b/AGENTS.md index 31ceaa57f..98f9bf562 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -34,11 +34,14 @@ Each domain has a dedicated boundary: - **TUI** → screen components and primitives in `src/tui/` The tree is three surfaces plus a composition root: `src/store` (state and -contracts), `src/agent` (one agent run), `src/tui` (rendering), `src/cli` -(argv and wiring). Surfaces import each other only through `@store`, -`@store/types`, `@store/programs`, `@agent`, `@agent/types`, `@tui`, -`@tui/types`, and `@tui/console`; `src/__tests__/architecture` enforces it. -Each surface's `README.md` lists what it owns and may import. +contracts), `src/agent` (one agent run), `src/tui` (rendering), `src/cli` (argv +and wiring). Surfaces import each other only through `@store`, `@store/types`, +`@store/programs`, `@agent`, `@agent/types`, `@tui`, `@tui/types`, and +`@tui/console`, plus `@store/control` loaded lazily by the two cli runners; +`src/__tests__/architecture` enforces it. Each surface's `README.md` lists what +it owns and may import. To drive a run without a keyboard (snapshots, an agent, +CI), use the control socket described in +[`e2e-harness/ARCHITECTURE.md`](e2e-harness/ARCHITECTURE.md). Adding a new concern means finding the narrowest existing surface, not adding logic to the runner. Keep changes local to the boundary that owns them. @@ -104,8 +107,8 @@ aliases. | Subcommand | What it audits | | ----------------------------- | ---------------------------------------------------- | -| `wizard audit events` | event capture quality + cost | -| `wizard audit all` | comprehensive audit across every area (**default**) | +| `wizard audit events` | event capture quality + cost | +| `wizard audit all` | comprehensive audit across every area (**default**) | | `wizard audit autocapture` | autocapture setup + cost | | `wizard audit feature-flags` | feature flag usage + cost | | `wizard audit identify` | `$identify` implementation | @@ -129,16 +132,16 @@ confuse it with the top-level `wizard skill` command. - **Registration:** [`src/cli/main.ts`](src/cli/main.ts) — the `.use()` chain wires each command. [`bin.ts`](bin.ts) runs the Node preflight, then imports it. -- **Command shape:** [`src/cli/commands/command.ts`](src/cli/commands/command.ts) — the - `Command` interface every command implements. +- **Command shape:** + [`src/cli/commands/command.ts`](src/cli/commands/command.ts) — the `Command` + interface every command implements. - **Flat native commands** (e.g. `revenue-analytics`, `upload-source-maps`) are built with `nativeCommandFactory` ([`src/cli/commands/factories/native-command-factory.ts`](src/cli/commands/factories/native-command-factory.ts)). - **Family commands** (e.g. `audit`) resolve subcommands at runtime against the `cliEntries` in `skill-menu.json`. Logic lives in - [`src/cli/dispatch-family.ts`](src/cli/dispatch-family.ts). - Adding a skill-backed subcommand is a **context-mill** release, not a wizard - change. + [`src/cli/dispatch-family.ts`](src/cli/dispatch-family.ts). Adding a + skill-backed subcommand is a **context-mill** release, not a wizard change. ### Commands vs. programs (don't confuse these) @@ -166,6 +169,7 @@ pnpm try --install-dir= # Run the wizard locally against a test proje pnpm build # Compile TypeScript pnpm test # Unit tests (builds first) pnpm test: # One Vitest project: store, agent, tui, cli, harness, arch + # harness spawns the real binary over its control socket; WIZARD_PTY_TESTS=0 skips the PTY spec pnpm typecheck # tsc -b over the surface projects (tsconfig.solution.json) pnpm typecheck: # One surface project pnpm test:watch # Unit tests in watch mode @@ -182,8 +186,8 @@ nonmutating lint checks, and scope formatting fixes to edited files. Do not add tests for prose, compiler-enforced shapes, or duplicated implementation. Keep new code comments to one line; put longer explanations in linked docs. -Local `--ci`, smoke-test, and full headless runs require two separate secrets: -a PostHog personal API key and an already-issued gateway token supplied through +Local `--ci`, smoke-test, and full headless runs require two separate secrets: a +PostHog personal API key and an already-issued gateway token supplied through `WIZARD_CI_GATEWAY_TOKEN_FILE`, plus the target project ID. Follow the [credential setup](docs/local-dev.md#credentials-for-local-ci-and-headless-runs). diff --git a/README.md b/README.md index 481d49d9d..37fe55b7c 100644 --- a/README.md +++ b/README.md @@ -2,8 +2,8 @@ posthoglogo

- -> have any feedback, please drop an email to **[wizard@posthog.com](mailto:wizard@posthog.com)**. +> have any feedback, please drop an email to +> **[wizard@posthog.com](mailto:wizard@posthog.com)**.

PostHog wizard ✨

@@ -19,22 +19,36 @@ To use the wizard, you can run it directly using: npx @posthog/wizard@latest ``` -Currently the wizard can be used for over 16+ frameworks for frontend, backend, and mobile applications. If you have other integrations you would like the wizard to -support, please open a [GitHub issue](https://github.com/posthog/wizard/issues)! +Currently the wizard can be used for over 16+ frameworks for frontend, backend, +and mobile applications. If you have other integrations you would like the +wizard to support, please open a +[GitHub issue](https://github.com/posthog/wizard/issues)! -Visit our [docs](https://posthog.com/docs/ai-engineering/ai-wizard) to learn more. +Visit our [docs](https://posthog.com/docs/ai-engineering/ai-wizard) to learn +more. ## Privacy & data usage -The wizard uses **AI models from Anthropic or OpenAI**, routed through PostHog's AI gateway, to read your project's source files and integrate PostHog. A few things worth knowing up front: - -- **Source files** are sent to the selected model provider as part of the agent's context. -- **`.env*` files and secrets** stay on your machine. The wizard's security scanner blocks anything it identifies as a secret from being read by the agent. -- **Telemetry** (run metadata — phase, task list, planned events) is sent to PostHog by default. Pass `--no-telemetry` (or set `POSTHOG_WIZARD_NO_TELEMETRY=1`) to disable. -- **AI opt-in**: for existing organizations in interactive runs, the wizard checks `is_ai_data_processing_approved` and waits for approval before agent work. CI and signup runs bypass this interactive gate. -- **Prefer your own AI?** The wizard's integration knowledge ships as a context-mill skill you can download and run inside your own agent. - -The wizard's "Privacy & data usage" menu (intro screen) and the `[I]` shortcut on the auth screen surface the same information in-terminal. +The wizard uses **AI models from Anthropic or OpenAI**, routed through PostHog's +AI gateway, to read your project's source files and integrate PostHog. A few +things worth knowing up front: + +- **Source files** are sent to the selected model provider as part of the + agent's context. +- **`.env*` files and secrets** stay on your machine. The wizard's security + scanner blocks anything it identifies as a secret from being read by the + agent. +- **Telemetry** (run metadata — phase, task list, planned events) is sent to + PostHog by default. Pass `--no-telemetry` (or set + `POSTHOG_WIZARD_NO_TELEMETRY=1`) to disable. +- **AI opt-in**: for existing organizations in interactive runs, the wizard + checks `is_ai_data_processing_approved` and waits for approval before agent + work. CI and signup runs bypass this interactive gate. +- **Prefer your own AI?** The wizard's integration knowledge ships as a + context-mill skill you can download and run inside your own agent. + +The wizard's "Privacy & data usage" menu (intro screen) and the `[I]` shortcut +on the auth screen surface the same information in-terminal. ## MCP Commands @@ -51,33 +65,43 @@ npx @posthog/wizard@latest mcp remove ## Wizard programs -The wizard's commands are grouped into **programs** — self-contained agentic jobs that install, audit, or wire up a specific piece of PostHog. They're powered by skills from the [context mill](https://github.com/PostHog/context-mill). +The wizard's commands are grouped into **programs** — self-contained agentic +jobs that install, audit, or wire up a specific piece of PostHog. They're +powered by skills from the +[context mill](https://github.com/PostHog/context-mill). ### PostHog integration (default) -Running the wizard with no arguments installs PostHog into your project. It detects your framework, wires up initialization, instruments a starter set of events, and walks you through a first dashboard: +Running the wizard with no arguments installs PostHog into your project. It +detects your framework, wires up initialization, instruments a starter set of +events, and walks you through a first dashboard: ```bash npx @posthog/wizard@latest ``` -Powered by the `posthog-integration` program. Most other programs below build on it (they declare `requires: ['posthog-integration']`) and will offer to run it first if PostHog isn't already set up. +Powered by the `posthog-integration` program. Most other programs below build on +it (they declare `requires: ['posthog-integration']`) and will offer to run it +first if PostHog isn't already set up. ### Self-driving -Autonomously sets up PostHog self-driving end-to-end. It connects GitHub, enables Session Replay and Error Tracking, wires up signal sources, and configures a Signals scout troop that watches your project for you. +Autonomously sets up PostHog self-driving end-to-end. It connects GitHub, +enables Session Replay and Error Tracking, wires up signal sources, and +configures a Signals scout troop that watches your project for you. ```bash npx @posthog/wizard@latest self-driving ``` -If PostHog isn't already installed, the wizard runs the default integration first (composed run) before starting the self-driving setup. +If PostHog isn't already installed, the wizard runs the default integration +first (composed run) before starting the self-driving setup. ### Audit Audit an existing PostHog integration for correctness and best practices. The -`audit` command is a **family**. With no subcommand it runs the **events** -audit (the default); pass a subcommand to run a specific one: +`audit` command is a **family**. With no subcommand it runs the **events** audit +(the default); pass a subcommand to run a specific one: ```bash # Runs the events audit (the default) — no subcommand needed @@ -96,12 +120,12 @@ npx @posthog/wizard@latest audit web-analytics # web analytics setup Most audit subcommands resolve at runtime from the published skill registry, so new audits appear without a wizard release (`web-analytics` is wizard-native). -> **`audit ` chooses an audit area — it does not take a skill name.** -> The audit subcommands above *are* context-mill skills promoted to commands (via -> a `cli: role: command` block); [`wizard skill `](#run-a-single-skill) -> runs a skill that hasn't been promoted. Same machinery, two surfaces. -> (`wizard audit --help` still labels the positional `[skill]` — read it as "pick -> a subcommand.") +> **`audit ` chooses an audit area — it does not take a skill +> name.** The audit subcommands above _are_ context-mill skills promoted to +> commands (via a `cli: role: command` block); +> [`wizard skill `](#run-a-single-skill) runs a skill that hasn't +> been promoted. Same machinery, two surfaces. (`wizard audit --help` still +> labels the positional `[skill]` — read it as "pick a subcommand.") ### Revenue Analytics @@ -129,7 +153,8 @@ OAuth sources open the PostHog app's new-source flow in your browser. ### Upload source maps -Upload JavaScript source maps to PostHog error tracking so stack traces are symbolicated back to your original code: +Upload JavaScript source maps to PostHog error tracking so stack traces are +symbolicated back to your original code: ```bash npx @posthog/wizard@latest upload-source-maps @@ -149,36 +174,35 @@ npx @posthog/wizard@latest skill # run one by name Reviews are auto-requested via [`.github/CODEOWNERS`](.github/CODEOWNERS) — the file is the source of truth; this table just mirrors it for readability. -`team-wizard-docs` is the default reviewer; the team-owned programs below -route review to their owning team instead. - -| Path | Owning team | -|---|---| -| `*` (everything else, including all other programs) | `@PostHog/team-wizard-docs` | -| `src/agent/` | `@PostHog/team-wizard-docs` | -| `src/store/programs/posthog-integration/` | `@PostHog/team-wizard-docs` | -| `src/store/programs/error-tracking-upload-source-maps/` | `@PostHog/team-error-tracking` | -| `src/store/programs/mcp-analytics/` | `@PostHog/team-mcp-analytics` | -| `src/store/programs/revenue-analytics/` | `@PostHog/team-web-analytics` | -| `src/store/programs/self-driving/` | `@PostHog/team-self-driving` | -| `src/store/programs/warehouse-source/` | `@PostHog/team-warehouse-sources` | -| `src/store/programs/web-analytics-doctor/` | `@PostHog/team-web-analytics` | - -Ownership is by directory. Programs not listed above -(`agent-skill`, `audit`, `events-audit`, `mcp`, `migration`, `posthog-doctor`, -`shared`, `slack`) fall through the default and are owned by -`team-wizard-docs`. Today CODEOWNERS only auto-requests review — approval is -not a merge gate. +`team-wizard-docs` is the default reviewer; the team-owned programs below route +review to their owning team instead. + +| Path | Owning team | +| ------------------------------------------------------- | --------------------------------- | +| `*` (everything else, including all other programs) | `@PostHog/team-wizard-docs` | +| `src/agent/` | `@PostHog/team-wizard-docs` | +| `src/store/programs/posthog-integration/` | `@PostHog/team-wizard-docs` | +| `src/store/programs/error-tracking-upload-source-maps/` | `@PostHog/team-error-tracking` | +| `src/store/programs/mcp-analytics/` | `@PostHog/team-mcp-analytics` | +| `src/store/programs/revenue-analytics/` | `@PostHog/team-web-analytics` | +| `src/store/programs/self-driving/` | `@PostHog/team-self-driving` | +| `src/store/programs/warehouse-source/` | `@PostHog/team-warehouse-sources` | +| `src/store/programs/web-analytics-doctor/` | `@PostHog/team-web-analytics` | + +Ownership is by directory. Programs not listed above (`agent-skill`, `audit`, +`events-audit`, `mcp`, `migration`, `posthog-doctor`, `shared`, `slack`) fall +through the default and are owned by `team-wizard-docs`. Today CODEOWNERS only +auto-requests review — approval is not a merge gate. ## Headless signup + install (agents / CI) -> ⚠️ `--ci` is **not currently supported in published builds** (see [CI Mode](#ci-mode)). -> This flow works in development builds only. +> ⚠️ `--ci` is **not currently supported in published builds** (see +> [CI Mode](#ci-mode)). This flow works in development builds only. -For a fully non-interactive first-run (no existing PostHog account, no TTY, -no browser), combine `--ci --signup --email`. The wizard provisions a new -account, uses the returned personal API key to run the normal CI install, -and wires PostHog into the project at `--install-dir`: +For a fully non-interactive first-run (no existing PostHog account, no TTY, no +browser), combine `--ci --signup --email`. The wizard provisions a new account, +uses the returned personal API key to run the normal CI install, and wires +PostHog into the project at `--install-dir`: ```bash npx @posthog/wizard@latest --ci --signup \ @@ -204,38 +228,37 @@ npx @posthog/wizard@latest provision --email user@example.com --region eu --json ``` Success prints the full `ProvisioningResult` (`projectApiKey`, `host`, -`projectId`, `accountId`, `accessToken`, `refreshToken`, and -`personalApiKey` if present). Failure exits 1; in `--json` mode the error -is emitted to stderr as `{"error":"...","code":"..."}`, with `code` set to -`email_exists` when the address is already registered. +`projectId`, `accountId`, `accessToken`, `refreshToken`, and `personalApiKey` if +present). Failure exits 1; in `--json` mode the error is emitted to stderr as +`{"error":"...","code":"..."}`, with `code` set to `email_exists` when the +address is already registered. -> ⚠️ **Output contains live credentials.** Pipe it into a secrets store — -> do not let it be captured by shared CI logs. Mask the step output or -> redirect stdout to a file your job reads and discards. +> ⚠️ **Output contains live credentials.** Pipe it into a secrets store — do not +> let it be captured by shared CI logs. Mask the step output or redirect stdout +> to a file your job reads and discards. # Options The following CLI arguments are available: -| Option | Description | Type | Default | Choices | Environment Variable | -| ----------------- | ---------------------------------------------------------------- | ------- | ------- | ---------------------------------------------------- | ------------------------------ | -| `--help` | Show help | boolean | | | | -| `--version` | Show version number | boolean | | | | -| `--debug` | Enable verbose logging | boolean | `false` | | `POSTHOG_WIZARD_DEBUG` | -| `--signup` | Create a new PostHog account during setup | boolean | `false` | | `POSTHOG_WIZARD_SIGNUP` | -| `--install-dir` | Directory to install PostHog in | string | | | `POSTHOG_WIZARD_INSTALL_DIR` | -| `--ci` | Enable CI mode for non-interactive execution | boolean | `false` | | `POSTHOG_WIZARD_CI` | -| `--api-key` | PostHog personal API key (phx_xxx) for authentication | string | | | `POSTHOG_WIZARD_API_KEY` | -| `--no-telemetry` | Disable wizard run-state telemetry | boolean | `false` | | `POSTHOG_WIZARD_NO_TELEMETRY` | - +| Option | Description | Type | Default | Choices | Environment Variable | +| ---------------- | ----------------------------------------------------- | ------- | ------- | ------- | ----------------------------- | +| `--help` | Show help | boolean | | | | +| `--version` | Show version number | boolean | | | | +| `--debug` | Enable verbose logging | boolean | `false` | | `POSTHOG_WIZARD_DEBUG` | +| `--signup` | Create a new PostHog account during setup | boolean | `false` | | `POSTHOG_WIZARD_SIGNUP` | +| `--install-dir` | Directory to install PostHog in | string | | | `POSTHOG_WIZARD_INSTALL_DIR` | +| `--ci` | Enable CI mode for non-interactive execution | boolean | `false` | | `POSTHOG_WIZARD_CI` | +| `--api-key` | PostHog personal API key (phx_xxx) for authentication | string | | | `POSTHOG_WIZARD_API_KEY` | +| `--no-telemetry` | Disable wizard run-state telemetry | boolean | `false` | | `POSTHOG_WIZARD_NO_TELEMETRY` | # CI Mode **CI mode is available only in development/test builds.** Published builds reject `--ci`; use an interactive terminal for `npx @posthog/wizard@latest`. -Local CI runs require a PostHog personal API key **and a separate gateway -token file**, plus the target project ID. See +Local CI runs require a PostHog personal API key **and a separate gateway token +file**, plus the target project ID. See [local credentials](docs/local-dev.md#credentials-for-local-ci-and-headless-runs) for setup and the CI secret names. With both secrets configured: @@ -257,8 +280,10 @@ The CLI args override environment variables in CI mode. ### Required Flags for CI Mode -- `--api-key`: Personal API key (`phx_xxx`) from your [PostHog settings](https://app.posthog.com/settings/user-api-keys) -- `--install-dir`: Directory to install PostHog in (e.g., `.` for current directory) +- `--api-key`: Personal API key (`phx_xxx`) from your + [PostHog settings](https://app.posthog.com/settings/user-api-keys) +- `--install-dir`: Directory to install PostHog in (e.g., `.` for current + directory) ### Required API Key Scopes @@ -270,8 +295,8 @@ dashboard:write insight:write notebook:write event_definition:write health_issue:read wizard_session:read wizard_session:write ``` -The source of truth is `WIZARD_OAUTH_SCOPES` in `src/store/shared/constants.ts`, which -documents why each scope is needed — if this block drifts, trust the code. +The source of truth is `WIZARD_OAUTH_SCOPES` in `src/store/shared/constants.ts`, +which documents why each scope is needed — if this block drifts, trust the code. Some programs request more on top (`PROGRAM_SCOPE_ADDITIONS` in `src/store/services/oauth/program-scopes.ts`); the default integration flow adds `integration:read` and `external_data_source:read` / @@ -279,23 +304,23 @@ Some programs request more on top (`PROGRAM_SCOPE_ADDITIONS` in ### OAuth app scope ceiling -The wizard's OAuth app on the PostHog side caps the scopes its tokens may -carry (`OAuthApplication.scopes`). Any scope requested in this repo (see -`src/store/services/oauth/program-scopes.ts`) must be grantable under that ceiling, or -`/authorize` drops it and the call that needs it 403s. +The wizard's OAuth app on the PostHog side caps the scopes its tokens may carry +(`OAuthApplication.scopes`). Any scope requested in this repo (see +`src/store/services/oauth/program-scopes.ts`) must be grantable under that +ceiling, or `/authorize` drops it and the call that needs it 403s. **A granted token can be narrower than the request even with a correct ceiling.** The consent screen lets the user deselect any scope the app doesn't mark required (`OAuthApplication.required_scopes`), and out-of-ceiling scopes are clamped silently (`clamp_scopes_to_ceiling`) — neither path errors; -`/oauth/token` just returns a smaller `scope`. So never assume the token -carries what was requested: the token response's `scope` field is the truth. -The wizard diffs granted vs requested at login (`missingOAuthScopes` in -`src/store/shared/oauth.ts`), warns the user which permissions are missing, and emits -`wizard: oauth grant narrowed` so narrowed runs are countable in analytics. -The diff also rides on the session (`credentials.missingScopes`), so when a -run does fail on a scope-gated step, the error names the missing permission -and the fix instead of the generic report-a-bug line. +`/oauth/token` just returns a smaller `scope`. So never assume the token carries +what was requested: the token response's `scope` field is the truth. The wizard +diffs granted vs requested at login (`missingOAuthScopes` in +`src/store/shared/oauth.ts`), warns the user which permissions are missing, and +emits `wizard: oauth grant narrowed` so narrowed runs are countable in +analytics. The diff also rides on the session (`credentials.missingScopes`), so +when a run does fail on a scope-gated step, the error names the missing +permission and the fix instead of the generic report-a-bug line. **To make scopes impossible to deselect, list them explicitly in the app's `scopes`.** `required_scopes` is not a separate field — it is derived @@ -305,8 +330,8 @@ consent POST 400s with `invalid_scope` if the grant misses one), while scopes covered only by `@default` stay deselectable and `optional_scopes` are declinable extras. That is why `llm_gateway:read` and `wizard_session:*` are already un-deselectable today, and everything else is not. To pin the base set -the wizard cannot run without, seed each region's app with `@default` plus -every scope in `WIZARD_OAUTH_SCOPES`: +the wizard cannot run without, seed each region's app with `@default` plus every +scope in `WIZARD_OAUTH_SCOPES`: ``` python manage.py seed_oauth_app_scopes --client-id --dry-run \ @@ -314,10 +339,10 @@ python manage.py seed_oauth_app_scopes --client-id --dry-run \ ``` then re-run without `--dry-run`. Keep `@default` in the list — dropping it -narrows the ceiling to only the explicit entries and strips the -program-specific additions. The wizard has no client-side lever for any of -this; the login diff and prompt-threaded degrade above handle a narrowed -grant, but only pinning prevents one. +narrows the ceiling to only the explicit entries and strips the program-specific +additions. The wizard has no client-side lever for any of this; the login diff +and prompt-threaded degrade above handle a narrowed grant, but only pinning +prevents one. **The live wizard apps use the `@default` sentinel, so most net-new scopes need no ceiling edit.** The prod US app's `scopes` is: @@ -362,18 +387,18 @@ used to silently add permissions. The CLI was overhauled to consolidate commands into a smaller, extensible surface. If you used an older command, here's where it went: -| Old command | New command | What changed | -|---|---|---| -| `wizard integrate` | `wizard` (default flow) | Command removed; the default flow runs the integration | -| `wizard events-audit` | `wizard audit events` | Now an `audit`-family subcommand | -| `wizard audit` (single audit) | `wizard audit ` | Now a family; see [Audit](#audit) for the subcommands | -| `wizard audit-3000` | *removed* | Retired | -| `wizard revenue` | `wizard revenue-analytics` | Renamed (old `revenue` removed) | -| `wizard upload-sourcemaps` | `wizard upload-source-maps` | Renamed; `upload-sourcemaps` still works as an alias | - -> **Commands vs. programs:** `integrate` was the *command*; the program behind it -> is `posthog-integration`, which still exists and now powers the default flow. -> Other commands depend on it via `requires: ['posthog-integration']`. The +| Old command | New command | What changed | +| ----------------------------- | --------------------------- | ------------------------------------------------------ | +| `wizard integrate` | `wizard` (default flow) | Command removed; the default flow runs the integration | +| `wizard events-audit` | `wizard audit events` | Now an `audit`-family subcommand | +| `wizard audit` (single audit) | `wizard audit ` | Now a family; see [Audit](#audit) for the subcommands | +| `wizard audit-3000` | _removed_ | Retired | +| `wizard revenue` | `wizard revenue-analytics` | Renamed (old `revenue` removed) | +| `wizard upload-sourcemaps` | `wizard upload-source-maps` | Renamed; `upload-sourcemaps` still works as an alias | + +> **Commands vs. programs:** `integrate` was the _command_; the program behind +> it is `posthog-integration`, which still exists and now powers the default +> flow. Other commands depend on it via `requires: ['posthog-integration']`. The > program id is internal — it was never a command you typed. # Steal this code @@ -396,8 +421,8 @@ and set up the general flow of the application. ## Analytics Did you know you can capture PostHog events even for smaller, supporting -products like a command line tool? `src/store/shared/analytics.ts` is a great example -of how to do it. +products like a command line tool? `src/store/shared/analytics.ts` is a great +example of how to do it. This file wraps `posthog-node` with some convenience functions to set up an analytics session and log events. We can see the usage and outcomes of this @@ -409,8 +434,8 @@ When the user authenticates, the wizard also streams live run state — current phase, task list, planned events — to `POST /api/projects/{id}/wizard/sessions/` so the PostHog web app can render real-time progress. Updates are debounced (250ms) with phase changes flushed immediately; failures fall back silently to -the wizard's debug log without disturbing the TUI. Pass `--no-telemetry` (or -set `POSTHOG_WIZARD_NO_TELEMETRY=1`) to disable. +the wizard's debug log without disturbing the TUI. Pass `--no-telemetry` (or set +`POSTHOG_WIZARD_NO_TELEMETRY=1`) to disable. ## Leave rules behind @@ -445,17 +470,16 @@ users of the wizard, no training delays or other ambiguity. ## Keep secrets out of the LLM -The wizard somtimes needs to move a secret. The agent -orchestrates that journey, but the raw value should _never_ enter the LLM -conversation, where it would be sent to the model provider, written to -transcripts, and captured in logs. +The wizard somtimes needs to move a secret. The agent orchestrates that journey, +but the raw value should _never_ enter the LLM conversation, where it would be +sent to the model provider, written to transcripts, and captured in logs. -`src/store/session/secret-vault.ts` is a small, reusable pattern for exactly this. It's a -session-scoped, in-memory vault: a tool that handles a secret calls `put()` to -store the raw value and hands the agent an opaque `secret:` reference -instead. The agent passes that ref between tools as if it were the value; the -host resolves it back to the real secret only at the last moment, inside the -process, when it writes the file. +`src/store/session/secret-vault.ts` is a small, reusable pattern for exactly +this. It's a session-scoped, in-memory vault: a tool that handles a secret calls +`put()` to store the raw value and hands the agent an opaque `secret:` +reference instead. The agent passes that ref between tools as if it were the +value; the host resolves it back to the real secret only at the last moment, +inside the process, when it writes the file. Two tools in `src/store/tools/tools.ts` form the ends of that pipe: @@ -471,48 +495,59 @@ drive the work end to end, but the only thing it ever sees is an opaque handle. ## Build system -Built with [tsdown](https://tsdown.dev/) (Rolldown). `pnpm build` bundles `bin.ts` into ESM chunks in `dist/`, inlining all local source and keeping npm dependencies external. +Built with [tsdown](https://tsdown.dev/) (Rolldown). `pnpm build` bundles +`bin.ts` into ESM chunks in `dist/`, inlining all local source and keeping npm +dependencies external. ### Environment variables -**Build-time (locked).** `NODE_ENV` is replaced with `"production"` at compile time. It cannot be overridden at runtime. All URLs, OAuth client IDs, and dev-mode code paths resolve to their production values unconditionally. +**Build-time (locked).** `NODE_ENV` is replaced with `"production"` at compile +time. It cannot be overridden at runtime. All URLs, OAuth client IDs, and +dev-mode code paths resolve to their production values unconditionally. -To add a new build-time constant, add it to `env` in `tsdown.config.ts` and export it from `src/env.ts`. +To add a new build-time constant, add it to `env` in `tsdown.config.ts` and +export it from `src/env.ts`. -**Runtime (allowlisted).** Runtime env reads go through `runtimeEnv()` in `src/env.ts`, which only accepts keys in the `RuntimeEnvKey` union: +**Runtime (allowlisted).** Runtime env reads go through `runtimeEnv()` in +`src/env.ts`, which only accepts keys in the `RuntimeEnvKey` union: -| Variable | Purpose | -|---|---| -| `POSTHOG_WIZARD_BENCHMARK_CONFIG` | Path to benchmark config file | -| `POSTHOG_WIZARD_BENCHMARK_FILE` | Output path for benchmark results | -| `POSTHOG_WIZARD_LOG_DIR` | Log directory override | -| `POSTHOG_WIZARD_DEBUG` / `DEBUG` | Enable debug output | -| `MCP_URL` | Override MCP server URL | -| `POSTHOG_API_KEY` | API key for MCP subprocess auth | -| `TERM`, `TERM_PROGRAM`, `CI`, etc. | Terminal/platform detection | -| `APPDATA`, `XDG_CONFIG_HOME` | Platform path resolution | +| Variable | Purpose | +| ---------------------------------- | --------------------------------- | +| `POSTHOG_WIZARD_BENCHMARK_CONFIG` | Path to benchmark config file | +| `POSTHOG_WIZARD_BENCHMARK_FILE` | Output path for benchmark results | +| `POSTHOG_WIZARD_LOG_DIR` | Log directory override | +| `POSTHOG_WIZARD_DEBUG` / `DEBUG` | Enable debug output | +| `MCP_URL` | Override MCP server URL | +| `POSTHOG_API_KEY` | API key for MCP subprocess auth | +| `TERM`, `TERM_PROGRAM`, `CI`, etc. | Terminal/platform detection | +| `APPDATA`, `XDG_CONFIG_HOME` | Platform path resolution | To add a new runtime env var, add its key to `RuntimeEnvKey` in `src/env.ts`. -**Direct `process.env` access** is only used for subprocess environment writes (e.g. `agent-interface.ts` setting `ANTHROPIC_BASE_URL`), vendored code, and tests. +**Direct `process.env` access** is only used for subprocess environment writes +(e.g. `agent-interface.ts` setting `ANTHROPIC_BASE_URL`), vendored code, and +tests. ### Import aliases Path aliases defined in `tsconfig.build.json`, resolved by tsdown: -| Alias | Maps to | -|---|---| -| `@env` | `src/env.ts` | -| `@store`, `@store/types`, `@store/programs` | `src/store/{index,types,programs/index}.ts` | -| `@agent`, `@agent/types` | `src/agent/{index,types}.ts` | -| `@tui`, `@tui/types`, `@tui/console` | `src/tui/{index,types,console/index}.ts` | -| `@cli/*` | `src/cli/*` (composition root, internal) | -| `@store/*`, `@agent/*`, `@tui/*` | surface internals; tests only, never across surfaces | +| Alias | Maps to | +| ------------------------------------------- | -------------------------------------------------------------------------- | +| `@env` | `src/env.ts` | +| `@store`, `@store/types`, `@store/programs` | `src/store/{index,types,programs/index}.ts` | +| `@agent`, `@agent/types` | `src/agent/{index,types}.ts` | +| `@tui`, `@tui/types`, `@tui/console` | `src/tui/{index,types,console/index}.ts` | +| `@cli/*` | `src/cli/*` (composition root, internal) | +| `@store/control` | `src/store/control/index.ts`; the cli runners only, by dynamic import | +| `@e2e-harness/*` | `e2e-harness/*` (harness, scripts, tests) | +| `@store/*`, `@agent/*`, `@tui/*` | surface internals; harness, scripts, and tests only, never across surfaces | ## Running locally For `--ci`, smoke tests, and full headless runs, configure both the personal API -key and gateway token file first: [local credentials](docs/local-dev.md#credentials-for-local-ci-and-headless-runs). +key and gateway token file first: +[local credentials](docs/local-dev.md#credentials-for-local-ci-and-headless-runs). Interactive runs mint their gateway token after authentication. ### Quick test without linking @@ -527,7 +562,8 @@ pnpm try --install-dir=[a path] pnpm run dev ``` -This builds, links globally, and watches for changes. Leave it running - any `.ts` file changes will auto-rebuild. Then from any project: +This builds, links globally, and watches for changes. Leave it running - any +`.ts` file changes will auto-rebuild. Then from any project: ```bash wizard --integration=nextjs @@ -538,8 +574,8 @@ wizard --integration=nextjs --local-mcp # MCP from localhost:8787 wizard --integration=nextjs --local-dev # context-mill + MCP + PostHog ``` -See [`docs/local-dev.md`](docs/local-dev.md) for the full catalog. -`--local-mcp` selects the MCP server; `--local-context-mill` selects the skills server. +See [`docs/local-dev.md`](docs/local-dev.md) for the full catalog. `--local-mcp` +selects the MCP server; `--local-context-mill` selects the skills server. ### Testing @@ -564,8 +600,8 @@ You can hand the wizard to an AI agent and have it drive the real flow itself deciding each screen and snapshotting the TUI to see what happened. The agent drives through the `wizard-ci` MCP tools (`open_app` / `read_state` / `perform_action` / `render_screen` / `run_agent`), which are registered in this -repo's `.mcp.json` and bound in every session here — approve `wizard-ci` the first -time you're prompted. The how-to is the `exploring-the-wizard` skill +repo's `.mcp.json` and bound in every session here — approve `wizard-ci` the +first time you're prompted. The how-to is the `exploring-the-wizard` skill (`.claude/skills/exploring-the-wizard/SKILL.md`), which an agent discovers automatically. @@ -573,12 +609,12 @@ Example prompt — explore against [open-saas](https://github.com/wasp-lang/open-saas): > Explore the PostHog wizard against open-saas, following the -> `exploring-the-wizard` skill. Reuse my phx key file path, gateway token file path, and project id, -> asking only for missing inputs. Launch the MCP server with -> `WIZARD_CI_GATEWAY_TOKEN_FILE` set to the gateway token file path; -> then clone `https://github.com/wasp-lang/open-saas` into a throwaway `/tmp` -> copy. Drive the whole flow yourself through the `wizard-ci` MCP tools, deciding -> each screen: +> `exploring-the-wizard` skill. Reuse my phx key file path, gateway token file +> path, and project id, asking only for missing inputs. Launch the MCP server +> with `WIZARD_CI_GATEWAY_TOKEN_FILE` set to the gateway token file path; then +> clone `https://github.com/wasp-lang/open-saas` into a throwaway `/tmp` copy. +> Drive the whole flow yourself through the `wizard-ci` MCP tools, deciding each +> screen: > > 1. `open_app` on the `/tmp` copy, then `read_state` to see the screen and the > actions legal right now. @@ -605,23 +641,24 @@ To make your version of a tool usable with a one-line `npx` command: # Health checks -`src/store/health-checks/` checks skills download origins before the wizard runs. -The entry point is `evaluateWizardReadiness()`, which only blocks on skill downloads: +`src/store/health-checks/` checks skills download origins before the wizard +runs. The entry point is `evaluateWizardReadiness()`, which only blocks on skill +downloads: -| Decision | Meaning | -| ------------------- | --------------------------------------------------------------- | -| `yes` | Skills are reachable — proceed without outage warnings. | -| `no` | Neither skills origin is reachable — do not run. | +| Decision | Meaning | +| -------- | ------------------------------------------------------- | +| `yes` | Skills are reachable — proceed without outage warnings. | +| `no` | Neither skills origin is reachable — do not run. | ### Module layout -| File | Responsibility | -| --- | --- | -| `types.ts` | Enums, interfaces (`ServiceHealthStatus`, `AllServicesHealth`, etc.) | +| File | Responsibility | +| -------------- | ----------------------------------------------------------------------- | +| `types.ts` | Enums, interfaces (`ServiceHealthStatus`, `AllServicesHealth`, etc.) | | `endpoints.ts` | Direct gateway (`/readyz`) and skills origin (`skill-menu.json`) checks | | `readiness.ts` | `checkAllExternalServices`, `evaluateWizardReadiness`, readiness config | -| `index.ts` | Barrel re-export | -| `testme.md` | Test running instructions and endpoint reference | +| `index.ts` | Barrel re-export | +| `testme.md` | Test running instructions and endpoint reference | ## What blocks a run @@ -643,29 +680,35 @@ The same policy applies during signup. Third-party status pages are not queried. After minting a token, `gateway-session.ts` checks `/readyz` on the returned gateway URL and reports an unavailable gateway through the existing error path. -`skillsOrigin` is one entry covering two origins: skills are published to -GitHub Releases and an AWS mirror under the same filenames, and downloads fail -over between them (`src/store/fetch-retry.ts`). Both are probed in parallel, so -the key only reports **Down** when neither origin answers — a GitHub Releases -outage on its own doesn't block a run, including a 403 or 404, which is as -often about the origin (expired asset redirect, blocked region, a publish that -reached one origin and not the other) as about the asset. +`skillsOrigin` is one entry covering two origins: skills are published to GitHub +Releases and an AWS mirror under the same filenames, and downloads fail over +between them (`src/store/fetch-retry.ts`). Both are probed in parallel, so the +key only reports **Down** when neither origin answers — a GitHub Releases outage +on its own doesn't block a run, including a 403 or 404, which is as often about +the origin (expired asset redirect, blocked region, a publish that reached one +origin and not the other) as about the asset. ## Smoke test helper (`scripts/smoke-test-ci.sh`) -This repo includes a helper script to run a full end‑to‑end smoke test of the wizard packaged in a tarball against a real app from [`posthog/wizard-workbench`](https://github.com/PostHog/wizard-workbench). This will catch certain packaging issues that might not be caught by other tests. +This repo includes a helper script to run a full end‑to‑end smoke test of the +wizard packaged in a tarball against a real app from +[`posthog/wizard-workbench`](https://github.com/PostHog/wizard-workbench). This +will catch certain packaging issues that might not be caught by other tests. **Prerequisites** - Point to a `wizard-workbench` checkout either by: - Setting `WIZARD_WORKBENCH_ROOT=/absolute/path/to/wizard-workbench`, or - - Cloning `wizard-workbench` next to this repo (so it lives at `../wizard-workbench`). -- Set `POSTHOG_PERSONAL_API_KEY` either in your shell or in `../wizard-workbench/.env`. + - Cloning `wizard-workbench` next to this repo (so it lives at + `../wizard-workbench`). +- Set `POSTHOG_PERSONAL_API_KEY` either in your shell or in + `../wizard-workbench/.env`. - Set `WIZARD_CI_GATEWAY_TOKEN_FILE` to an absolute path containing the separate - AI gateway token. See [local credentials](docs/local-dev.md#credentials-for-local-ci-and-headless-runs). + AI gateway token. See + [local credentials](docs/local-dev.md#credentials-for-local-ci-and-headless-runs). - Set `POSTHOG_WIZARD_PROJECT_ID` to the intended test project and - `POSTHOG_WIZARD_REGION` to `us` or `eu` (CI uses `us`). The helper also accepts - `POSTHOG_PROJECT_ID` and `POSTHOG_REGION` as fallback names. + `POSTHOG_WIZARD_REGION` to `us` or `eu` (CI uses `us`). The helper also + accepts `POSTHOG_PROJECT_ID` and `POSTHOG_REGION` as fallback names. **Usage** @@ -695,8 +738,10 @@ The script will: - Copy the selected app into a temp directory - Install dependencies for the app - Install the packed wizard tarball into an isolated temp project -- Run `wizard` in `--ci` mode against the copied app and perform basic post‑install checks +- Run `wizard` in `--ci` mode against the copied app and perform basic + post‑install checks ## Contributing -Start with [AGENTS.md](AGENTS.md) for the development skills and execution policy. +Start with [AGENTS.md](AGENTS.md) for the development skills and execution +policy. diff --git a/docs/benchmarking.md b/docs/benchmarking.md index e5fadf4f7..17a4bc6e8 100644 --- a/docs/benchmarking.md +++ b/docs/benchmarking.md @@ -29,29 +29,29 @@ Selection criteria, checked in this order: 1. **Real product, in production** — an open-source app people actually run (stars are a proxy; a hosted instance is better evidence). 2. **Single-app repo** — reject monorepos: fetch the repo's top-level listing - and reject on `pnpm-workspace.yaml`, `turbo.json`, `lerna.json`, or - top-level `apps/`/`packages/` directories. + and reject on `pnpm-workspace.yaml`, `turbo.json`, `lerna.json`, or top-level + `apps/`/`packages/` directories. 3. **Greenfield** — grep the repo for `posthog` (manifest and source). An app that already integrates PostHog measures augmentation discipline, not integration quality; keep at most one such app and exclude it from quality scoring. 4. **Framework coverage** — spread picks across the frameworks the wizard supports; results do not transfer between them. -5. **Locally installable** — its toolchain (node/python/php/ruby/gradle) - exists on the bench machine, or its runs will fail for reasons that are - yours, not the model's. +5. **Locally installable** — its toolchain (node/python/php/ruby/gradle) exists + on the bench machine, or its runs will fail for reasons that are yours, not + the model's. Apps used in the 2026-07 benchmark, as worked examples of the spread: -| app | upstream | stack | -|---|---|---| -| Maybe | `maybe-finance/maybe` | Rails | -| Outline | `outline/outline` | React + Koa / TS | -| WordPress-Android | `wordpress-mobile/WordPress-Android` | native Kotlin | -| healthchecks | `healthchecks/healthchecks` | Django | -| Firefly III | `firefly-iii/firefly-iii` | Laravel, server-rendered | -| Monica | `monicahq/monica` | Laravel + Inertia/Vue | -| Papermark | `mfts/papermark` | Next.js — already shipped posthog-js; kept as the augment-existing case, excluded from quality scoring | +| app | upstream | stack | +| ----------------- | ------------------------------------ | ------------------------------------------------------------------------------------------------------ | +| Maybe | `maybe-finance/maybe` | Rails | +| Outline | `outline/outline` | React + Koa / TS | +| WordPress-Android | `wordpress-mobile/WordPress-Android` | native Kotlin | +| healthchecks | `healthchecks/healthchecks` | Django | +| Firefly III | `firefly-iii/firefly-iii` | Laravel, server-rendered | +| Monica | `monicahq/monica` | Laravel + Inertia/Vue | +| Papermark | `mfts/papermark` | Next.js — already shipped posthog-js; kept as the augment-existing case, excluded from quality scoring | ## Setup (from a bare machine) @@ -90,21 +90,21 @@ Define each config as a `WIZARD_CI_FLAG_OVERRIDES` JSON, plus the baseline as Everything below ships in this repo (`wizard/`) and its workbench (`wizard-workbench/`); paths are from each repo's root. -- **Headless run (snapshotting CI harness):** `wizard/scripts/tui-snapshots.no-jest.ts` - spawns the real TUI (`wizard/scripts/tui-host.no-jest.ts`) in a PTY via +- **Headless run (snapshotting CI harness):** + `wizard/scripts/tui-snapshots.no-jest.ts` spawns the real wizard + (`bin.ts --ci --control-socket`) in a PTY via `wizard/e2e-harness/tui-capture.ts`, self-drives the fixed e2e profile - (`wizard/e2e-harness/wizard-ci-driver.ts`, `wizard/e2e-harness/profiles.ts`) - through auth, the agent run, and the outro, and writes each screen as an - `NN-.ans` frame. An `NN-outro.ans` frame is the flow-completion - signal. `tsx` runs source — no build step. Invocation: see the run-cell - recipe below. + (`wizard/e2e-harness/profiles.ts`) over the control socket through auth, the + agent run, and the outro, and writes each screen as an `NN-.ans` + frame. An `NN-outro.ans` frame is the flow-completion signal. `tsx` runs + source — no build step. Invocation: see the run-cell recipe below. - **Config selection:** the flag axis is `wizard-orchestrator` (on → the orchestrator on pi, per-task models from context-mill frontmatter; off → the linear anthropic default). Per-stage variations ride - `wizard-orchestrator-override` payloads (`{stage: {model?, effort?}}`, - variant keys in `wizard/src/agent/runner/switchboard/flags/schemes.ts`). - The baseline is `{"wizard-orchestrator":"false"}` — never an empty override, - or live remote flags leak into the baseline. + `wizard-orchestrator-override` payloads (`{stage: {model?, effort?}}`, variant + keys in `wizard/src/agent/runner/switchboard/flags/schemes.ts`). The baseline + is `{"wizard-orchestrator":"false"}` — never an empty override, or live remote + flags leak into the baseline. ## Running one cell @@ -147,19 +147,19 @@ files=$(git -C "$WORK" diff --name-only main integ | wc -l | tr -d ' ')" \ | tee "$OUT/result.txt" ``` -Both commits are `--no-verify` (see Traps). The diff, frames, stdout, and -result line are the cell's complete artifact set — everything else (the shared -debug log) is unreliable under parallelism. +Both commits are `--no-verify` (see Traps). The diff, frames, stdout, and result +line are the cell's complete artifact set — everything else (the shared debug +log) is unreliable under parallelism. ## Running the matrix -- One app at a time; per app, launch its configs in parallel (≤4 on one - machine) and `wait`. Contention inflates absolute times roughly uniformly. +- One app at a time; per app, launch its configs in parallel (≤4 on one machine) + and `wait`. Contention inflates absolute times roughly uniformly. - Cost: anthropic-harness cells report `modelUsage.costUSD` in - `/tmp/posthog-wizard.log` — zero the log before each app's wave and slice - the block per baseline run. pi-harness cells do not persist token totals; - add a temporary hook in the pi harness success path that writes the session - token stats to a per-run file, and price them at list rates. + `/tmp/posthog-wizard.log` — zero the log before each app's wave and slice the + block per baseline run. pi-harness cells do not persist token totals; add a + temporary hook in the pi harness success path that writes the session token + stats to a per-run file, and price them at list rates. - Rerun any anomalous cell solo (zeroed log, no parallelism) before drawing a conclusion from it. @@ -167,61 +167,61 @@ debug log) is unreliable under parallelism. - **Target-app git hooks.** Your `git commit` runs the app's husky/lint-staged hooks if a prior install activated them; a failing hook silently rolls the - tree back and the run measures as zero-diff. Always commit `--no-verify`. - On any zero-diff run, check `git stash list` before believing it. -- **Zero-diff has many causes.** Distinguish: the agent honestly declined - (read its setup report), the agent's tool calls failed, your harness ate the - work, or `.gitignore` hid it (env files never show in diffs). Attribute - before you blame the model. + tree back and the run measures as zero-diff. Always commit `--no-verify`. On + any zero-diff run, check `git stash list` before believing it. +- **Zero-diff has many causes.** Distinguish: the agent honestly declined (read + its setup report), the agent's tool calls failed, your harness ate the work, + or `.gitignore` hid it (env files never show in diffs). Attribute before you + blame the model. - **"Reached the outro" is not success.** The flow completes even when nothing was integrated. Treat completion as outro + a non-trivial diff. - **Parallel runs interleave shared state.** The shared debug log cannot be attributed per-run; capture everything per-run or run solo when attribution matters. - **Sandbox/allowlist gaps look like model failures.** If a config produces - empty or thin work, check whether a blocked command (package-manager - install, formatter) caused it, and whether other models worked around the - same block. File the gap; exclude the affected cells. -- **Repo-wide format scripts.** An agent running the app's `format`/`lint - --fix` buries its real diff under hundreds of churn files. Count "real - files" excluding scaffolding, lockfiles, env files — and read a sample of - the churn before scoring. -- **A stale credential fails silently mid-batch.** Read the key per run, not - per session. + empty or thin work, check whether a blocked command (package-manager install, + formatter) caused it, and whether other models worked around the same block. + File the gap; exclude the affected cells. +- **Repo-wide format scripts.** An agent running the app's `format`/`lint --fix` + buries its real diff under hundreds of churn files. Count "real files" + excluding scaffolding, lockfiles, env files — and read a sample of the churn + before scoring. +- **A stale credential fails silently mid-batch.** Read the key per run, not per + session. ## Judging Use the wizard-workbench PR evaluator's rubric — do not invent your own. It lives at `wizard-workbench/services/pr-evaluator/`: -- **Rubric criteria:** `wizard-workbench/services/pr-evaluator/prompts/evaluation.md` - — per-item YES/NO/N-A checks grouped into four dimensions. -- **Scoring math:** `wizard-workbench/services/pr-evaluator/evaluator.ts` — - each dimension scores `max(1, round(pass_rate × 5))` over its applicable - items; confidence = `min(app_sanity, round(mean of the four))`. +- **Rubric criteria:** + `wizard-workbench/services/pr-evaluator/prompts/evaluation.md` — per-item + YES/NO/N-A checks grouped into four dimensions. +- **Scoring math:** `wizard-workbench/services/pr-evaluator/evaluator.ts` — each + dimension scores `max(1, round(pass_rate × 5))` over its applicable items; + confidence = `min(app_sanity, round(mean of the four))`. - **Automated run:** from `wizard-workbench/`, `pnpm run evaluate --branch --base --test-run` (needs `POSTHOG_PERSONAL_API_KEY`; judge model via `EVALUATOR_MODEL`). Output lands in `wizard-workbench/test-evaluations//` as `rubric.json` + `scores.json`. -- **Manual run:** an agent applies the same rubric directly to each cell's - diff — faster for many cells, and what the 2026-07 benchmark did. Either - way, report the four dimensions under their full names, 1–5 each - (5 production-ready, 3 works with real issues, 1 broken or empty): +- **Manual run:** an agent applies the same rubric directly to each cell's diff + — faster for many cells, and what the 2026-07 benchmark did. Either way, + report the four dimensions under their full names, 1–5 each (5 + production-ready, 3 works with real issues, 1 broken or empty): -- **Files** (`file_analysis`) — right files touched, nothing unrelated, - imports valid +- **Files** (`file_analysis`) — right files touched, nothing unrelated, imports + valid - **App** (`app_sanity`) — nothing broken: builds, existing code and configs preserved, changes minimal - **PostHog** (`posthog_implementation`) — SDK installed, initialized at the - right entry points, env-based keys, real distinct id, identify, error - tracking + right entry points, env-based keys, real distinct id, identify, error tracking - **Events** (`event_quality`) — real user actions, useful properties, no PII, consistent names -Verify claims against the diff (grep for `capture`/`identify` call sites, -check the init file, check the manifest), and build or typecheck where cheap. -Judge the same subset of apps for every config you compare. +Verify claims against the diff (grep for `capture`/`identify` call sites, check +the init file, check the manifest), and build or typecheck where cheap. Judge +the same subset of apps for every config you compare. ## Publishing evidence @@ -230,48 +230,55 @@ inspectable: 1. Fork each app to the operator's account. 2. Pin a `bench-base` branch at the exact commit the runs used. -3. Per cell: branch from `bench-base`, apply the **sanitized** patch, push, - open a draft PR against `bench-base`. +3. Per cell: branch from `bench-base`, apply the **sanitized** patch, push, open + a draft PR against `bench-base`. 4. Sanitize before anything touches a public fork: drop env files, wizard - scaffolding, and lockfiles from the patch; redact every token literal. - Verify zero secrets in the pushed diff before opening the PR. + scaffolding, and lockfiles from the patch; redact every token literal. Verify + zero secrets in the pushed diff before opening the PR. ## Report template ```markdown # — model benchmark - + ## Summary — configs that completed everywhere -| config | completed | median time | median cost | quality (judged on) | - ## Results - -| config | | | … | -