From acc5824473c74eded1355431a4935a586cb5e833 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Tue, 22 Sep 2026 17:42:57 -0400 Subject: [PATCH 01/29] docs(programs): add non-TUI developer API reference Document callable agent and program contracts, development CI invocation, outcomes, cancellation, and current headless limits. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- README.md | 5 ++ docs/developer-interfaces.md | 124 +++++++++++++++++++++++++++ docs/local-dev.md | 3 + src/agent/README.md | 87 +++++++++++++++---- src/programs/README.md | 157 +++++++++++++++++++++++++++++++++++ 5 files changed, 359 insertions(+), 17 deletions(-) create mode 100644 docs/developer-interfaces.md create mode 100644 src/programs/README.md diff --git a/README.md b/README.md index be133cf7b..2cb1129b1 100644 --- a/README.md +++ b/README.md @@ -388,6 +388,11 @@ that conventional code implies. If you want to use this code as a starting place for your own project, here's a quick explainer on its structure. +For code that runs without the terminal UI, see the +[non-interactive developer interfaces](docs/developer-interfaces.md), including +the [standalone agent](src/agent/README.md) and +[callable programs](src/programs/README.md). + ## Entrypoint: `run.ts` The entrypoint for this tool is `run.ts`. Use this file to interpret arguments diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md new file mode 100644 index 000000000..0a9e2c924 --- /dev/null +++ b/docs/developer-interfaces.md @@ -0,0 +1,124 @@ +# Non-interactive developer interfaces + +Wizard has two repository-local TypeScript call surfaces and one development CLI +mode for running without a terminal UI. The TypeScript aliases below are +internal to this repository; `@posthog/wizard` currently publishes a CLI, not +these functions as a stable package API. + +| Surface | Use it for | Detailed contract | +| --------------------------------------- | ----------------------------------------------- | ------------------------------------------------------------------------------------------ | +| `runAgent(config, input, options)` | One already-configured AI run | [Agent reference](../src/agent/README.md) | +| `runProgram(programId, input, options)` | A registered program with invocation-owned data | [Programs reference](../src/programs/README.md) | +| Development `--ci` | A process-owned, non-interactive CLI run | [Local CI credentials and recipe](local-dev.md#credentials-for-local-ci-and-headless-runs) | + +## Standalone agent + +`runAgent` takes a resolved `RunConfig` (run definition, binding, tools and +policy), `RunInput` (project, credentials, optional inference-auth provider, +flags and host), and optional `onProgress`, `interaction`, and `signal` options. +Import the function from `@agent` and types from `@agent/types`. It returns a +`RunResult` with a `success`, `aborted`, `failed`, or `crashed` outcome and a +final task/status/usage snapshot. Progress is delivered in emission order; +observer errors do not fail the run. Without `interaction`, questions have no +answer bridge and optional task notices are declined. + +```ts +import { runAgent } from '@agent'; +import type { RunConfig, RunInput } from '@agent/types'; + +export async function runStandalone( + config: RunConfig, + input: RunInput, + signal?: AbortSignal, +) { + return runAgent(config, input, { + signal, + onProgress: (event) => { + if (event.kind === 'status') console.log(event.message); + }, + }); +} +``` + +The caller prepares `config` and `input`; the agent does not authenticate the +PostHog user or detect the project. Supply `input.inferenceAuth` when the host +owns gateway authentication. A legacy fallback remains when it is absent. There +is no session control protocol on this API. An aborted signal returns an +`aborted` result; it does not pause the run. + +## Callable program + +`runProgram` takes a registered ID, `ProgramInput` with at least `installDir`, +and optional `ProgramOptions`. The host supplies resolved credentials or a +credential provider, prepared detection and framework context where needed, and +callbacks for questions, approvals, progress, or program-specific effects. An +optional `signal` cancels an active agent run. It returns a `ProgramRunOutcome`: +outcome and failure, final progress, actual settled agent runs, program-specific +data, artifacts, and invocation data (including a captured event plan). The +latter contains credentials and should not be logged. + +```ts +import { runProgram } from '@programs'; +import type { ProgramOptions } from '@programs/types'; + +export async function runAudit( + installDir: string, + credentials: NonNullable, + awaitAiApproval: NonNullable, + signal?: AbortSignal, +) { + const result = await runProgram( + 'audit', + { installDir }, + { + credentials, + awaitAiApproval, + signal, + onProgress: ({ runId, event }) => { + if (event.kind === 'tasks') console.log(runId, event.tasks); + }, + }, + ); + if (result.outcome !== 'success') { + throw new Error(result.failure?.message ?? `Audit ${result.outcome}`); + } + return result.artifacts.reportFile; +} +``` + +The caller implements the credential and approval callbacks. Some programs +require additional prepared inputs or host effects; the +[program reference](../src/programs/README.md#inputs-and-capabilities) lists +them. Host callbacks such as credential resolution, approval, and MCP work do +not receive the signal. There is no live store or step-control handle. + +## Development CI and experimental headless runner + +Development/test builds accept `--ci`. This is a whole-process CLI path, not an +awaitable function returning `ProgramRunOutcome`. It requires an install +directory, a PostHog personal API key, a project ID, and an already-issued +gateway token in the file named by `WIZARD_CI_GATEWAY_TOKEN_FILE`: + +```bash +WIZARD_CI_GATEWAY_TOKEN_FILE="$HOME/.config/posthog/wizard-gateway-token" \ +pnpm try --ci --api-key "$POSTHOG_PERSONAL_API_KEY" \ + --project-id "$POSTHOG_WIZARD_PROJECT_ID" \ + --region us --install-dir /absolute/path/to/test-app +``` + +The runner logs progress and writes a local task-stream JSONL dump. Callers +observe the process exit and its logs, rather than a returned result. The +gateway token file is read directly for CI; this path does not mint or refresh +that token. Published builds reject `--ci`. The internal +`runWizardCI(config, options): void` entry point still uses the legacy session +adapter, which now calls `runProgram` for agent execution. + +An experimental published-build headless path exists internally as +`runWizardHeadless(config, options): void`. It shares the process-owned runner, +logs progress, and can push task-stream updates to PostHog when telemetry is +enabled. Its selector is deliberately hidden and is not a supported invocation +recipe. Neither internal function returns a structured, awaitable outcome. + +Controlled headless and socket control APIs are not available yet. There is no +supported route, command, or event protocol for pausing a run, supplying an +answer later, or reading its live state from another process. diff --git a/docs/local-dev.md b/docs/local-dev.md index 93f949cdd..5fc65f934 100644 --- a/docs/local-dev.md +++ b/docs/local-dev.md @@ -3,6 +3,9 @@ Running the wizard against local servers. Four things can independently be local, and this doc is the catalog of how to control each. +For the callable agent and program contracts, see the +[non-interactive developer interfaces](developer-interfaces.md). + ## Credentials for local CI and headless runs Local `--ci` runs, smoke tests, and full headless/snapshot agent runs need diff --git a/src/agent/README.md b/src/agent/README.md index 3aff9c1ce..c5fc2757e 100644 --- a/src/agent/README.md +++ b/src/agent/README.md @@ -1,29 +1,68 @@ # Agent -The agent runs one program's AI pipeline against a project directory. It takes resolved data in, reports through progress events, asks through an injected answerer, and returns a result. It never reads a session, a store or a UI. +The agent runs one program's AI pipeline against a project directory. It takes +resolved data in, reports through progress events, asks through an injected +answerer, and returns a result. It never reads a session, a store or a UI. For +the callable program host and development CI runner, see the +[non-interactive developer interfaces](../../docs/developer-interfaces.md). ## Signatures -Import runtime values from `@agent` and types from `@agent/types`. Nothing outside `src/agent` imports deeper; lint and the architecture test reject it. +Import runtime values from `@agent` and types from `@agent/types`. Nothing +outside `src/agent` imports deeper; lint and the architecture test reject it. ```ts import { runAgent, RunOutcome } from '@agent'; -import type { RunConfig, RunInput, RunResult, AgentProgress } from '@agent/types'; +import type { RunConfig, RunInput, RunResult, AgentProgress, AgentInteraction } from '@agent/types'; runAgent(config: RunConfig, input: RunInput, options?: { + signal?: AbortSignal; onProgress?: (event: AgentProgress) => void; interaction?: AgentInteraction; }): Promise ``` -- `RunConfig`: the opaque program id, its `AgentRunDefinition` (prompt, skill, tools, copy), the resolved `binding` (sequence, harness, model and task-role routes), supplied program commandments and stage policy, the skills origin, flag snapshot, trace tags, tool allow and deny lists, seed tasks and bound completion `hooks`. -- `RunInput`: install directory, resolved credentials, project and user payloads, skill id, detected integration, `flags` (`ci`, `signup`, `debug`, `e2eAsk`, `localMcp`, `captureAio`, `benchmark`, `yaraReport`) and the host the CLI was told. -- `RunResult`: `outcome` is `RunOutcome.Success | Aborted | Failed | Crashed`. Success may carry an `outro`; the other three carry a `failure` (`AgentFailure`: message, outro data, error, exit code, error code, detail). Every result carries `skillId` and a `snapshot` of what the run reported: tasks, status lines, stage, token usage totals, final cost, dashboard and notebook URLs, handoff text. -- `AgentProgress`: one event per thing the run reports, in emission order. Kinds: `lifecycle`, `spinner`, `log`, `status`, `tasks`, `stage`, `url`, `usage`, `finalCost`, `authError`, `handoff`, `completion`. Payloads are copies, never live objects. -- `AgentInteraction`: every member optional. `ask(question)` resolves with answers, `cancelAsk()` dismisses the open question, `taskNotice(notice)` resolves with whether to keep an optional task, `cancelTaskNotice()` declines it. -- Errors: the agent does not exit the process and does not throw for a decided failure. An unexpected throw becomes `outcome: Crashed` with the error attached. A gateway 401 emits `authError` and then fails. - -Other runtime exports: `DEFAULT_AGENT_BINDING` for standalone callers, the generic `resolveBinding` and `resolveHarness` helpers, `shouldDisableAsk`, `initializeAgent`, `executeAgent`, `buildRunTags`, `AgentSignals`, `configureGatewayFromCIEnvironment`, `downloadSkill`, `WIZARD_TOOL_NAMES`, `LONGER_ASK_TIMEOUT_MS`, `flushScanReport`, and `runMcpPromptViaSdk`, which loads the streaming module on first call. +- `RunConfig`: the opaque program id, its `AgentRunDefinition` (prompt, skill, + tools, copy), the resolved `binding` (sequence, harness, model and task-role + routes), supplied program commandments and stage policy, the skills origin, + flag snapshot, trace tags, tool allow and deny lists, seed tasks and bound + completion `hooks`. +- `RunInput`: install directory, resolved PostHog credentials, optional + `inferenceAuth`, project and user payloads, skill id, detected integration, + `flags` (`ci`, `signup`, `debug`, `e2eAsk`, `localMcp`, `captureAio`, + `benchmark`, `yaraReport`) and the host the CLI was told. The caller can + supply an `InferenceAuthProvider` whose `resolve()` returns gateway + authentication. The agent resolves it before execution and again when the + harness needs refreshed auth; if absent, the legacy gateway-auth fallback + remains. +- `RunResult`: `outcome` is `RunOutcome.Success | Aborted | Failed | Crashed`. + Success may carry an `outro`; the other three carry a `failure` + (`AgentFailure`: message, outro data, error, exit code, error code, detail). + Every result carries `skillId` and a `snapshot` of what the run reported: + tasks, status lines, stage, token usage totals, final cost, dashboard and + notebook URLs, handoff text. +- `AgentProgress`: one event per thing the run reports, in emission order. + Kinds: `lifecycle`, `spinner`, `log`, `status`, `tasks`, `stage`, `url`, + `usage`, `finalCost`, `authError`, `handoff`, `completion`. Payloads are + copies, never live objects. +- `AgentInteraction`: every member optional. `ask(question)` resolves with + answers, `cancelAsk()` dismisses the open question, `taskNotice(notice)` + resolves with whether to keep an optional task, `cancelTaskNotice()` declines + it. +- `signal`: an optional `AbortSignal` from the host. A pre-aborted signal + returns `Aborted` before execution; aborting during execution is passed to the + active harness and returns `Aborted` with the current snapshot. It does not + pause or resume a run. +- Errors: the agent does not exit the process and does not throw for a decided + failure. An unexpected throw becomes `outcome: Crashed` with the error + attached. A gateway 401 emits `authError` and then fails. + +Other runtime exports: `DEFAULT_AGENT_BINDING` for standalone callers, the +generic `resolveBinding` and `resolveHarness` helpers, `shouldDisableAsk`, +`initializeAgent`, `executeAgent`, `buildRunTags`, `AgentSignals`, +`configureGatewayFromCIEnvironment`, `downloadSkill`, `WIZARD_TOOL_NAMES`, +`LONGER_ASK_TIMEOUT_MS`, `flushScanReport`, and `runMcpPromptViaSdk`, which +loads the streaming module on first call. Minimal invocation: @@ -41,21 +80,32 @@ if (result.outcome !== RunOutcome.Success) { } ``` -`src/agent/__tests__/run-agent-standalone.test.ts` runs this with no UI, no store and no registry. +`src/agent/__tests__/run-agent-standalone.test.ts` runs this with no UI, no +store and no registry. ## Intent -Programs call the agent to do the work a skill describes. The TUI and the headless runner observe the run through `onProgress` and answer it through `interaction`; today `src/programs/run-agent-legacy.ts` does both on top of the session. +Programs call the agent to do the work a skill describes. A standalone host can +observe the run through `onProgress` and answer it through `interaction`; the +legacy TUI and non-interactive runner still use +`src/programs/run-agent-legacy.ts` for session gates and UI translation, then +call the same `runProgram` host. -Without `onProgress` the run completes and its snapshot still comes back in the result. Without `interaction` the agent installs no ask bridge: `wizard_ask` returns its "not available" error and optional task notices are declined, which is what a `--ci` run does. A throwing observer is logged and the run continues. +Without `onProgress` the run completes and its snapshot still comes back in the +result. Without `interaction` the agent installs no ask bridge: `wizard_ask` +returns its "not available" error and optional task notices are declined, which +is what a `--ci` run does. A throwing observer is logged and the run continues. ## Architecture -The agent owns run state for one invocation: the task queue, phase, status, resolved skill, handoff text, usage and the final result. It depends on `src/shared` and on `src/env.ts`, and on program types only until the bindings table moves to programs. +The agent owns run state for one invocation: the task queue, phase, status, +resolved skill, handoff text, usage and the final result. It depends on +`src/shared` and `src/env.ts`; callers supply program routing and policy through +`RunConfig`. ```text caller ── RunConfig + RunInput ──▶ runAgent - │ prepareRun: gateway mint, triage provider + │ prepareRun: resolve inference auth, triage provider ▼ sequence (linear | orchestrator) │ @@ -66,4 +116,7 @@ caller ── RunConfig + RunInput ──▶ runAgent RunResult ``` -`runner/` holds the dispatcher, sequences, harnesses and the switchboard. `tools/` holds the wizard tools shared by both harnesses. `middleware/` holds the benchmark pipeline. `progress.ts` defines the event and interaction contracts; `yara-hooks.ts` scans what the run installs. +`runner/` holds the dispatcher, sequences, harnesses and the switchboard. +`tools/` holds the wizard tools shared by both harnesses. `middleware/` holds +the benchmark pipeline. `progress.ts` defines the event and interaction +contracts; `yara-hooks.ts` scans what the run installs. diff --git a/src/programs/README.md b/src/programs/README.md new file mode 100644 index 000000000..73c73d3d1 --- /dev/null +++ b/src/programs/README.md @@ -0,0 +1,157 @@ +# Programs + +`runProgram` runs a registered Wizard program from caller-supplied data. The +program owns its invocation state, resolves its run definition and binding, and +passes one or more agent runs to [`runAgent`](../agent/README.md). Callers do +not create a `WizardSession` or a UI store. This is a repository-local +TypeScript interface; the npm package does not currently export it as a public +library API. + +## Signature + +Import the runtime function from `@programs` and types from `@programs/types`: + +```ts +import { runProgram } from '@programs'; +import type { + ProgramInput, + ProgramOptions, + ProgramRunOutcome, +} from '@programs/types'; + +runProgram( + programId: string, + input: ProgramInput, + options?: ProgramOptions, +): Promise +``` + +`programId` must resolve through the runtime registry. Unknown IDs return a +failed outcome. The host supplies an absolute or relative `installDir`; report +paths in the result are resolved against it. + +## Intent and ownership + +The host authenticates, detects the project, and chooses any workflow decisions +it needs before calling. It passes those facts as plain data. `runProgram` +resolves the program recipe, applies the program binding and policy, owns the +progress projection, and invokes the agent. An agent run receives resolved +inputs and does not read the host's session. + +```text +host ── program id + input + capabilities ──▶ runProgram + │ │ + │ registry + data-only run resolver + binding + │ │ + ◀── attributed progress ── ProgramStore ◀── runAgent + ◀── outcome + settled runs + final data ───┘ +``` + +## Inputs and capabilities + +| Surface | What the caller provides | +| ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `ProgramInput.installDir` | Project directory for agent work and report paths. | +| `ProgramInput.credentials` or `ProgramOptions.credentials` | A resolved `{ posthog, inferenceAuth, project, apiUser }` object, or a provider with `resolve(programId)` that returns one. When used, the provider is called once for the program ID. Agent programs require credentials; a few non-agent programs do not. The host owns authentication and token storage. | +| `ProgramInput.runId` | Optional stable attribution ID; generated when absent. Composed child runs receive their own IDs. | +| Detection data | `integration`, `typescript`, `frameworkConfig`, `frameworkContext`, `warehouseSources`, `detectedTools`, `sourceMapsSelection`, and related data required by the chosen program. Dynamic recipes use these values instead of a session. | +| Run selection | Optional `run` override, `binding`, `skillId`, `flags`, `wizardFlags`, `wizardFlagPayloads`, `wizardMetadata`, `seedTasks`, `hooks`, `allowedTools`, `disallowedTools`, `agentFlow`, and `host`. These are resolved snapshots for this invocation. | +| Composition | `composition.integration` can supply a prepared child integration for `self-driving`; `compositionWorkflow` can confirm the handoff and GitHub steps. Boolean decisions in `composition` are also accepted. | +| Host capabilities | `interaction` answers agent questions; `onProgress` observes events. `signal` cancels an active agent run. `mcp` and `workflow` serve programs without an agent. `integrationEffects` supplies the integration recipe's host effects. `awaitAiApproval` resolves the AI-processing approval gate when needed. | + +The exact shapes are in [`ProgramInput` and `ProgramOptions`](run-program.ts). +`ProgramOptions['credentials']` and `ProgramInput['credentials']` expose the +provider and resolved-credential types to callers; their underlying named types +live in [`credentials.ts`](credentials.ts). Do not log the outcome's `data`: it +includes PostHog credentials. + +`posthog-integration` needs a prepared `frameworkConfig` and +`integrationEffects` unless the caller supplies an explicit `run` override. +`self-driving` can compose that integration before its own run. It requires a +confirmed GitHub connection, supplied as `composition.githubConnected: true` or +through `compositionWorkflow`; a composed integration also requires a confirmed +handoff. An AI program whose organization lacks AI-processing approval needs +`awaitAiApproval` for a normal invocation; declining aborts the program. + +## Outcomes and progress + +`ProgramRunOutcome.outcome` is `success`, `aborted`, `failed`, or `crashed`. +Decided pre-run failures, such as an unknown program or missing credentials, +return `failed` with `failure.message`. An agent crash appears as `crashed`. +External host capabilities can still reject, so callers should also handle a +rejected promise. + +`options.signal` accepts an `AbortSignal`. A signal aborted before the run +starts returns `aborted` with an agent-abort failure code. During an agent run, +the signal reaches the active harness and the result is `aborted`. It is not a +general cancellation protocol for host-supplied credential, approval, MCP, or +workflow callbacks; those callbacks should manage their own lifetime. + +The result includes: + +- `runResults`: completed agent results in registration order. +- `settledRuns`: actual completed agent invocations in finish order, each with + `runId`, optional `stepId`, and its `RunResult`. This is separate from the + progress projection. +- `progress`: the final per-run task, status, stage, usage, and outcome + projection, plus bounded observer diagnostics. +- `data`: invocation-owned credentials, project and user data, detection + context, captured event plan, and completed composition steps. +- `programData`: program-specific data for flows without an agent, such as + doctor issues or MCP client results. +- `artifacts.reportFile`: the resolved report path when an agent run definition + provides one; it does not assert that the file was written. + +`onProgress` receives `{ runId, stepId?, event }`. The `event` is a copied +[`AgentProgress`](../agent/README.md#signatures) value, including `tasks` and +`status` changes. The store applies the event before calling the observer. +Observers are not awaited; thrown or rejected observers are isolated and +recorded as bounded diagnostics when observed. A late asynchronous rejection may +arrive after the returned progress snapshot. Keep observers short and use +`runId` and `stepId` to attribute composed runs. The final result remains +available when no observer is supplied. + +```ts +import { runProgram } from '@programs'; +import type { ProgramOptions } from '@programs/types'; + +export async function runMetrics( + installDir: string, + credentials: NonNullable, + awaitAiApproval: NonNullable, + signal?: AbortSignal, +) { + const result = await runProgram( + 'metrics', + { installDir }, + { + credentials, + awaitAiApproval, + signal, + onProgress: ({ runId, event }) => { + if (event.kind === 'tasks') console.log(runId, event.tasks); + if (event.kind === 'status') console.log(runId, event.message); + }, + }, + ); + + if (result.outcome !== 'success') { + throw new Error(result.failure?.message ?? `Metrics ${result.outcome}`); + } + return result; +} +``` + +The caller implements `credentials` and `awaitAiApproval` through its own auth +and consent flow. No token is embedded in this example. + +## Current limits + +The callable host does not discover credentials, choose a project, or render +questions itself. The host supplies those data and capabilities. Some legacy +recipes still need an explicit data-only run definition; unsupported +combinations return a failed outcome. Existing terminal and CI callers route +their agent work through this host via `run-agent-legacy.ts`, which still owns +session gates and UI translation. There is no socket controller or step-by-step +control API. For the existing process-owned CI runner, see the +[non-interactive developer interfaces](../../docs/developer-interfaces.md). From 2902f31225690ac3070c2aa0b52e6388c600b90b Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Tue, 22 Sep 2026 18:12:52 -0400 Subject: [PATCH 02/29] docs(programs): require caller-owned inference authentication Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- docs/developer-interfaces.md | 15 ++++++++------- src/agent/README.md | 10 +++++----- 2 files changed, 13 insertions(+), 12 deletions(-) diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index 0a9e2c924..0a398338a 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -14,7 +14,7 @@ these functions as a stable package API. ## Standalone agent `runAgent` takes a resolved `RunConfig` (run definition, binding, tools and -policy), `RunInput` (project, credentials, optional inference-auth provider, +policy), `RunInput` (project, credentials, required inference-auth provider, flags and host), and optional `onProgress`, `interaction`, and `signal` options. Import the function from `@agent` and types from `@agent/types`. It returns a `RunResult` with a `success`, `aborted`, `failed`, or `crashed` outcome and a @@ -41,10 +41,10 @@ export async function runStandalone( ``` The caller prepares `config` and `input`; the agent does not authenticate the -PostHog user or detect the project. Supply `input.inferenceAuth` when the host -owns gateway authentication. A legacy fallback remains when it is absent. There -is no session control protocol on this API. An aborted signal returns an -`aborted` result; it does not pause the run. +PostHog user or detect the project. The caller must supply +`input.inferenceAuth`, whose `resolve()` returns gateway authentication and can +refresh it during a long run. There is no session control protocol on this API. +An aborted signal returns an `aborted` result; it does not pause the run. ## Callable program @@ -108,8 +108,9 @@ pnpm try --ci --api-key "$POSTHOG_PERSONAL_API_KEY" \ The runner logs progress and writes a local task-stream JSONL dump. Callers observe the process exit and its logs, rather than a returned result. The -gateway token file is read directly for CI; this path does not mint or refresh -that token. Published builds reject `--ci`. The internal +gateway token file is read into a fixed provider for CI. Pre-run detection and +composed child runs use that same provider; this path does not mint or refresh +the token. Published builds reject `--ci`. The internal `runWizardCI(config, options): void` entry point still uses the legacy session adapter, which now calls `runProgram` for agent execution. diff --git a/src/agent/README.md b/src/agent/README.md index 6a3dbd532..3e9130f40 100644 --- a/src/agent/README.md +++ b/src/agent/README.md @@ -31,9 +31,9 @@ runAgent(config: RunConfig, input: RunInput, options?: { `inferenceAuth`, project and user payloads, skill id, detected integration, `flags` (`ci`, `signup`, `debug`, `e2eAsk`, `localMcp`, `captureAio`, `benchmark`, `yaraReport`) and the host the CLI was told. The caller supplies - an `InferenceAuthProvider` whose `resolve()` returns gateway - authentication. The agent resolves it before execution and again when the - harness needs refreshed auth. + an `InferenceAuthProvider` whose `resolve()` returns gateway authentication. + The agent resolves it before execution and again when the harness needs + refreshed auth. - `RunResult`: `outcome` is `RunOutcome.Success | Aborted | Failed | Crashed`. Success may carry an `outro`; the other three carry a `failure` (`AgentFailure`: message, outro data, error, exit code, error code, detail). @@ -60,8 +60,8 @@ Other runtime exports: `DEFAULT_AGENT_BINDING` for standalone callers, the generic `resolveBinding` and `resolveHarness` helpers, `shouldDisableAsk`, `initializeAgent`, `executeAgent`, `buildRunTags`, `AgentSignals`, `downloadSkill`, `WIZARD_TOOL_NAMES`, `LONGER_ASK_TIMEOUT_MS`, -`flushScanReport`, and `runMcpPromptViaSdk`, which -loads the streaming module on first call. +`flushScanReport`, and `runMcpPromptViaSdk`, which loads the streaming module on +first call. Minimal invocation: From 855bcfdc1daba98c91bdf735524be7f12cffd713 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Tue, 22 Sep 2026 21:20:48 -0400 Subject: [PATCH 03/29] Clarify callable Wizard API contracts --- docs/developer-interfaces.md | 30 ++++++++++---- src/agent/README.md | 10 +++-- src/programs/README.md | 77 ++++++++++++++++++++++++------------ 3 files changed, 81 insertions(+), 36 deletions(-) diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index 0a398338a..79f852dac 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -18,9 +18,10 @@ policy), `RunInput` (project, credentials, required inference-auth provider, flags and host), and optional `onProgress`, `interaction`, and `signal` options. Import the function from `@agent` and types from `@agent/types`. It returns a `RunResult` with a `success`, `aborted`, `failed`, or `crashed` outcome and a -final task/status/usage snapshot. Progress is delivered in emission order; -observer errors do not fail the run. Without `interaction`, questions have no -answer bridge and optional task notices are declined. +final task/status/usage snapshot. Progress is delivered in emission order. +Synchronous observer throws are logged without failing the run; asynchronous +observers must handle their own rejected promises. Without `interaction`, +questions have no answer bridge and optional task notices are declined. ```ts import { runAgent } from '@agent'; @@ -46,13 +47,23 @@ PostHog user or detect the project. The caller must supply refresh it during a long run. There is no session control protocol on this API. An aborted signal returns an `aborted` result; it does not pause the run. +### Inference authentication + +For first-party inference authentication, import +`createPosthogInferenceAuthProvider` from `@programs` and pass authenticated +PostHog credentials and the run's program ID (for example, `config.programId`, +`'audit'`, or `'metrics'`). Its provider mints a gateway token and refreshes it +near expiry; the host still handles user login and project selection. +Development [CI](#development-ci-and-experimental-headless-runner) instead uses +an already-issued fixed token. + ## Callable program `runProgram` takes a registered ID, `ProgramInput` with at least `installDir`, and optional `ProgramOptions`. The host supplies resolved credentials or a credential provider, prepared detection and framework context where needed, and callbacks for questions, approvals, progress, or program-specific effects. An -optional `signal` cancels an active agent run. It returns a `ProgramRunOutcome`: +optional `signal` requests cancellation. It returns a `ProgramRunOutcome`: outcome and failure, final progress, actual settled agent runs, program-specific data, artifacts, and invocation data (including a captured event plan). The latter contains credentials and should not be logged. @@ -88,9 +99,10 @@ export async function runAudit( The caller implements the credential and approval callbacks. Some programs require additional prepared inputs or host effects; the -[program reference](../src/programs/README.md#inputs-and-capabilities) lists -them. Host callbacks such as credential resolution, approval, and MCP work do -not receive the signal. There is no live store or step-control handle. +[program reference](../src/programs/README.md#inputs-and-capabilities) describes +the available fields and capabilities. Host callbacks such as credential +resolution, approval, and MCP work do not receive the signal. There is no live +store or step-control handle. ## Development CI and experimental headless runner @@ -112,7 +124,9 @@ gateway token file is read into a fixed provider for CI. Pre-run detection and composed child runs use that same provider; this path does not mint or refresh the token. Published builds reject `--ci`. The internal `runWizardCI(config, options): void` entry point still uses the legacy session -adapter, which now calls `runProgram` for agent execution. +adapter, which calls `runProgram` for each program's main agent run. Agentic +detection can call the agent separately before that run; MCP suggested prompts +also use a separate agent path with their own progress and cancellation. An experimental published-build headless path exists internally as `runWizardHeadless(config, options): void`. It shares the process-owned runner, diff --git a/src/agent/README.md b/src/agent/README.md index c3fa28f94..3a0b78532 100644 --- a/src/agent/README.md +++ b/src/agent/README.md @@ -33,7 +33,9 @@ runAgent(config: RunConfig, input: RunInput, options?: { `benchmark`, `yaraReport`) and the host the CLI was told. The caller supplies an `InferenceAuthProvider` whose `resolve()` returns gateway authentication. The agent resolves it before execution and again when the harness needs - refreshed auth. + refreshed auth. See the + [first-party provider](../../docs/developer-interfaces.md#inference-authentication) + for gateway token minting and refresh. - `RunResult`: `outcome` is `RunOutcome.Success | Aborted | Failed | Crashed`. Success may carry an `outro`; the other three carry a `failure` (`AgentFailure`: message, outro data, error, exit code, error code, detail). @@ -87,13 +89,15 @@ store and no registry. Programs call the agent to do the work a skill describes. A standalone host can observe the run through `onProgress` and answer it through `interaction`; the legacy TUI and non-interactive runner still use -`src/lib/runners/run-program-agent.ts` for session gates and UI translation, then -call the same `runProgram` host. +`src/lib/runners/run-program-agent.ts` for session gates and UI translation, +then call the same `runProgram` host. Without `onProgress` the run completes and its snapshot still comes back in the result. Without `interaction` the agent installs no ask bridge: `wizard_ask` returns its "not available" error and optional task notices are declined, which is what a `--ci` run does. A throwing observer is logged and the run continues. +Progress callbacks are not awaited, so an asynchronous observer must handle its +own rejected promises. ## Architecture diff --git a/src/programs/README.md b/src/programs/README.md index 7d6046e21..704bb570b 100644 --- a/src/programs/README.md +++ b/src/programs/README.md @@ -2,10 +2,10 @@ `runProgram` runs a registered Wizard program from caller-supplied data. The program owns its invocation state, resolves its run definition and binding, and -passes one or more agent runs to [`runAgent`](../agent/README.md). Callers do -not create a `WizardSession` or a UI store. This is a repository-local -TypeScript interface; the npm package does not currently export it as a public -library API. +passes any main agent run and composed child runs to +[`runAgent`](../agent/README.md). Callers do not create a `WizardSession` or a +UI store. This is a repository-local TypeScript interface; the npm package does +not currently export it as a public library API. ## Signature @@ -49,15 +49,15 @@ host ── program id + input + capabilities ──▶ runProgram ## Inputs and capabilities -| Surface | What the caller provides | -| ---------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `ProgramInput.installDir` | Project directory for agent work and report paths. | -| `ProgramInput.credentials` or `ProgramOptions.credentials` | A resolved `{ posthog, inferenceAuth, project, apiUser }` object, or a provider with `resolve(programId)` that returns one. When used, the provider is called once for the program ID. Agent programs require credentials; a few non-agent programs do not. The host owns authentication and token storage. | -| `ProgramInput.runId` | Optional stable attribution ID; generated when absent. Composed child runs receive their own IDs. | -| Detection data | `integration`, `typescript`, `frameworkConfig`, `frameworkContext`, `warehouseSources`, `detectedTools`, `sourceMapsSelection`, and related data required by the chosen program. Dynamic recipes use these values instead of a session. | -| Run selection | Optional `run` override, `binding`, `skillId`, `flags`, `wizardFlags`, `wizardFlagPayloads`, `wizardMetadata`, `seedTasks`, `hooks`, `allowedTools`, `disallowedTools`, `agentFlow`, and `host`. These are resolved snapshots for this invocation. | -| Composition | `composition.integration` can supply a prepared child integration for `self-driving`; `compositionWorkflow` can confirm the handoff and GitHub steps. Boolean decisions in `composition` are also accepted. | -| Host capabilities | `interaction` answers agent questions; `onProgress` observes events. `signal` cancels an active agent run. `mcp` and `workflow` serve programs without an agent. `integrationEffects` supplies the integration recipe's host effects. `awaitAiApproval` resolves the AI-processing approval gate when needed. | +| Surface | What the caller provides | +| ---------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `ProgramInput.installDir` | Project directory for agent work and report paths. | +| `ProgramInput.credentials` or `ProgramOptions.credentials` | A resolved `{ posthog, inferenceAuth, project, apiUser }` object, or a provider with `resolve(programId)` that returns one. When used, the provider is called once for the program ID. Agent programs require credentials; a few non-agent programs do not. The host owns authentication and token storage. | +| `ProgramInput.runId` | Optional stable attribution ID; generated when absent. Composed child runs receive their own IDs. | +| Detection data | `integration`, `typescript`, `frameworkConfig`, `frameworkContext`, `warehouseSources`, `detectedTools`, `sourceMapsSelection`, and related data required by the chosen program. Dynamic recipes use these values instead of a session. | +| Run selection | Optional `run` override, `binding`, `skillId`, `flags`, `wizardFlags`, `wizardFlagPayloads`, `wizardMetadata`, `seedTasks`, `hooks`, `allowedTools`, `disallowedTools`, `agentFlow`, and `host`. These are resolved snapshots for this invocation. | +| Composition | `composition.integration` can supply a prepared child integration for `self-driving`; `compositionWorkflow` can confirm the handoff and GitHub steps. Boolean decisions in `composition` are also accepted. | +| Host capabilities | `interaction` answers agent questions; `onProgress` observes events. `signal` requests cancellation. `mcp` and `workflow` serve programs without an agent. `integrationEffects` supplies the integration recipe's host effects. `awaitAiApproval` resolves the AI-processing approval gate when needed. | The exact shapes are in [`ProgramInput` and `ProgramOptions`](run-program.ts). `ProgramOptions['credentials']` and `ProgramInput['credentials']` expose the @@ -65,27 +65,50 @@ provider and resolved-credential types to callers; their underlying named types live in [`credentials.ts`](credentials.ts). Do not log the outcome's `data`: it includes PostHog credentials. +For first-party inference auth, use +[`createPosthogInferenceAuthProvider`](../../docs/developer-interfaces.md#inference-authentication) +with authenticated PostHog credentials and the program ID. The development +`--ci` runner instead uses an already-issued fixed gateway token. + `posthog-integration` needs a prepared `frameworkConfig` and -`integrationEffects` unless the caller supplies an explicit `run` override. +`integrationEffects` unless the caller supplies an explicit `run` override. For +a normal source-map upload, pass a detected `sourceMapsSelection.variant`; +without it, the fallback prompt still starts an agent before asking it to abort. `self-driving` can compose that integration before its own run. It requires a confirmed GitHub connection, supplied as `composition.githubConnected: true` or through `compositionWorkflow`; a composed integration also requires a confirmed handoff. An AI program whose organization lacks AI-processing approval needs `awaitAiApproval` for a normal invocation; declining aborts the program. +`flags.ci` and `flags.signup` bypass that approval check and do not call +`awaitAiApproval`. Hosts should use those flags only when consent has already +been handled by their CI authorization or signup flow; the flags do not prove +consent. + +When `composition.integration` is supplied, its agent runs before the handoff +and GitHub confirmations, then the self-driving agent runs. A declined later +confirmation returns `aborted` without rolling back integration edits. Hosts +that need both decisions before any project write must establish them before +calling `runProgram`; `compositionWorkflow` takes precedence over boolean +decisions and is queried at those later checkpoints. ## Outcomes and progress `ProgramRunOutcome.outcome` is `success`, `aborted`, `failed`, or `crashed`. Decided pre-run failures, such as an unknown program or missing credentials, return `failed` with `failure.message`. An agent crash appears as `crashed`. -External host capabilities can still reject, so callers should also handle a -rejected promise. +Handled credential-resolution, approval, and composition callback rejections +also resolve as `failed` with a message, not the callback's original error +class. An unexpected invocation throw, such as a duplicate composed `runId`, +rejects the promise. Hosts should inspect the outcome and separately catch +rejected promises. `options.signal` accepts an `AbortSignal`. A signal aborted before the run -starts returns `aborted` with an agent-abort failure code. During an agent run, -the signal reaches the active harness and the result is `aborted`. It is not a -general cancellation protocol for host-supplied credential, approval, MCP, or -workflow callbacks; those callbacks should manage their own lifetime. +starts returns `aborted` with an agent-abort failure code. Agent programs check +again before startup and pass the signal to the active harness. Programs without +an agent recheck after credential resolution and host work; an abort returns +`aborted` even if a callback has completed. Workflow requests receive the +signal, but in-flight host effects must cooperate with cancellation and +completed external effects are not rolled back. The result includes: @@ -99,8 +122,10 @@ The result includes: context, captured event plan, and completed composition steps. - `programData`: program-specific data for flows without an agent, such as doctor issues or MCP client results. -- `artifacts.reportFile`: the resolved report path when an agent run definition - provides one; it does not assert that the file was written. +- `artifacts.reportFile`: the intended absolute report path once the run + definition resolves and the agent is about to run. Pre-run failure or abort + leaves it absent, and the path does not prove that a file was written. A + failed composed child can return its own report path. `onProgress` receives `{ runId, stepId?, event }`. The `event` is a copied [`AgentProgress`](../agent/README.md#signatures) value, including `tasks` and @@ -151,8 +176,10 @@ The callable host does not discover credentials, choose a project, or render questions itself. The host supplies those data and capabilities. Some legacy recipes still need an explicit data-only run definition; unsupported combinations return a failed outcome. Existing terminal and CI callers route -their agent work through this host via +their main program runs through this host via `src/lib/runners/run-program-agent.ts`, which still owns session gates and UI -translation. There is no socket controller or step-by-step -control API. For the existing process-owned CI runner, see the +translation. Agentic detection and MCP suggested-prompt streaming call the agent +separately, with their own progress and cancellation contracts. There is no +socket controller or step-by-step control API. For the existing process-owned CI +runner, see the [non-interactive developer interfaces](../../docs/developer-interfaces.md). From aba2628fe3c6f74643eaad32fe89e326699095f3 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Tue, 22 Sep 2026 21:46:51 -0400 Subject: [PATCH 04/29] docs(agent): clarify B4 result and error ownership Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- docs/developer-interfaces.md | 15 ++++- src/agent/README.md | 34 ++++++++--- src/agent/runner/README.md | 109 ++++++++++++++++++++++++++--------- src/programs/README.md | 16 +++-- 4 files changed, 130 insertions(+), 44 deletions(-) diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index 0a398338a..43ca1a0b1 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -20,7 +20,10 @@ Import the function from `@agent` and types from `@agent/types`. It returns a `RunResult` with a `success`, `aborted`, `failed`, or `crashed` outcome and a final task/status/usage snapshot. Progress is delivered in emission order; observer errors do not fail the run. Without `interaction`, questions have no -answer bridge and optional task notices are declined. +answer bridge and optional task notices are declined. Non-success results carry +a `failure`; its `error` may be available when an `Error` was caught, while +`code` and `message` are optional. The caller chooses how to log or present a +failure. ```ts import { runAgent } from '@agent'; @@ -44,7 +47,10 @@ The caller prepares `config` and `input`; the agent does not authenticate the PostHog user or detect the project. The caller must supply `input.inferenceAuth`, whose `resolve()` returns gateway authentication and can refresh it during a long run. There is no session control protocol on this API. -An aborted signal returns an `aborted` result; it does not pause the run. +An aborted signal returns an `aborted` result; it does not pause the run. The +agent catches unexpected errors in its run body and reports `crashed` with the +caught `Error` (or an `Error` wrapper for a non-`Error` throw); the host can +rethrow that object when it needs exception semantics. ## Callable program @@ -55,7 +61,9 @@ callbacks for questions, approvals, progress, or program-specific effects. An optional `signal` cancels an active agent run. It returns a `ProgramRunOutcome`: outcome and failure, final progress, actual settled agent runs, program-specific data, artifacts, and invocation data (including a captured event plan). The -latter contains credentials and should not be logged. +latter contains credentials and should not be logged. Agent failures retain an +attached `Error` when one exists. External host callbacks may still reject the +promise, so callers handle those exceptions as well as returned outcomes. ```ts import { runProgram } from '@programs'; @@ -80,6 +88,7 @@ export async function runAudit( }, ); if (result.outcome !== 'success') { + if (result.failure?.error) throw result.failure.error; throw new Error(result.failure?.message ?? `Audit ${result.outcome}`); } return result.artifacts.reportFile; diff --git a/src/agent/README.md b/src/agent/README.md index c3fa28f94..06af2fe55 100644 --- a/src/agent/README.md +++ b/src/agent/README.md @@ -36,10 +36,13 @@ runAgent(config: RunConfig, input: RunInput, options?: { refreshed auth. - `RunResult`: `outcome` is `RunOutcome.Success | Aborted | Failed | Crashed`. Success may carry an `outro`; the other three carry a `failure` - (`AgentFailure`: message, outro data, error, exit code, error code, detail). - Every result carries `skillId` and a `snapshot` of what the run reported: - tasks, status lines, stage, token usage totals, final cost, dashboard and - notebook URLs, handoff text. + (`AgentFailure`: optional message, outro data, `Error`, exit code, error code, + and detail). `failure.error` may be present when the agent caught an `Error`; + `Crashed` requires one. A missing `Error` object does not mean the outcome + succeeded, and not every failed result has a code or message. Every result + carries `skillId` and a `snapshot` of what the run reported: tasks, status + lines, stage, token usage totals, final cost, dashboard and notebook URLs, + handoff text. - `AgentProgress`: one event per thing the run reports, in emission order. Kinds: `lifecycle`, `spinner`, `log`, `status`, `tasks`, `stage`, `url`, `usage`, `finalCost`, `authError`, `handoff`, `completion`. Payloads are @@ -52,9 +55,13 @@ runAgent(config: RunConfig, input: RunInput, options?: { returns `Aborted` before execution; aborting during execution is passed to the active harness and returns `Aborted` with the current snapshot. It does not pause or resume a run. -- Errors: the agent does not exit the process and does not throw for a decided - failure. An unexpected throw becomes `outcome: Crashed` with the error - attached. A gateway 401 emits `authError` and then fails. +- Errors: the agent does not exit the process or throw for a decided failure. It + catches unexpected errors in its run body, logs them, and returns + `outcome: Crashed` with the caught `Error` attached (or an `Error` wrapper for + a non-`Error` throw). A gateway 401 emits `authError` and then fails. The host + decides how to present a returned failure, set an exit code, or rethrow an + attached error. Final scan-report flushing runs outside that catch and can + still reject the promise; hosts requiring a hard promise boundary catch it. Other runtime exports: `DEFAULT_AGENT_BINDING` for standalone callers, the generic `resolveBinding` and `resolveHarness` helpers, `shouldDisableAsk`, @@ -75,10 +82,19 @@ const result = await runAgent(config, input, { }, }); if (result.outcome !== RunOutcome.Success) { + console.error( + result.failure.error ?? result.failure.message ?? `Agent ${result.outcome}`, + ); process.exitCode = result.failure.exitCode ?? 1; } ``` +The result is the agent's termination report. Check `outcome` first and then +read `failure`; a failed result can have no attached `Error`. If a higher layer +uses exceptions, it can rethrow `failure.error` when present and construct an +error from `failure.message` otherwise. Preserve the original `Error` object +when rethrowing so its stack and cause remain available. + `src/agent/__tests__/run-agent-standalone.test.ts` runs this with no UI, no store and no registry. @@ -87,8 +103,8 @@ store and no registry. Programs call the agent to do the work a skill describes. A standalone host can observe the run through `onProgress` and answer it through `interaction`; the legacy TUI and non-interactive runner still use -`src/lib/runners/run-program-agent.ts` for session gates and UI translation, then -call the same `runProgram` host. +`src/lib/runners/run-program-agent.ts` for session gates and UI translation, +then call the same `runProgram` host. Without `onProgress` the run completes and its snapshot still comes back in the result. Without `interaction` the agent installs no ask bridge: `wizard_ask` diff --git a/src/agent/runner/README.md b/src/agent/runner/README.md index d98975871..b95f87405 100644 --- a/src/agent/runner/README.md +++ b/src/agent/runner/README.md @@ -7,12 +7,12 @@ which model) and the pieces that then actually run it. ``` ┌──────────────┐ ┌─────────────┐ ┌────────────────────────────┐ │ │ │ │────▶│ sequence (query shape) │ - │ programs │────▶│ switchboard │ │ linear | orchestrator │ + │ programs │────▶│ binding │ │ linear | orchestrator │ │ │ │ │ └────────────────────────────┘ - │ integration │ │ binds each │ - │ audit │ │ program to │ ┌────────────────────────────┐ - │ migration │ │ a pair │────▶│ harness (SDK adapter) │ - │ ... │ │ │ │ anthropic | pi | ... │ + │ integration │ │ selects the │ + │ audit │ │ sequence, │ ┌────────────────────────────┐ + │ migration │ │ harness and │────▶│ harness (SDK adapter) │ + │ ... │ │ model │ │ anthropic | pi | ... │ └──────────────┘ └─────────────┘ └────────────────────────────┘ ``` @@ -40,23 +40,24 @@ for the coordinated change checklist. Five layers, each with its own job. Nothing crosses layers unless it has to. **The entry point** (`index.ts`) is the front door: -`runAgent(config, input, {onProgress?, interaction?}) → RunResult`. It takes -resolved execution data and an invocation snapshot (`shared/types.ts`), reports -through `onProgress` and asks through `interaction` (`../progress.ts`), and -returns every ending as a result. It never renders, reads a session or exits. -The gates, OAuth, flags and binding lookup that used to run here live in -`src/lib/runners/run-program-agent.ts`, which also maps progress back onto -`getUI()` for today's runners. +`runAgent(config, input, {onProgress?, interaction?, signal?}) → RunResult`. It +takes resolved execution data and an invocation snapshot (`shared/types.ts`), +reports through `onProgress` and asks through `interaction` (`../progress.ts`), +and returns decided outcomes and caught run-body crashes as results. It never +renders, reads a session or exits. `src/programs/run-program.ts` resolves the +binding from caller data. The legacy `src/lib/runners/run-program-agent.ts` owns +session gates and maps progress back onto `getUI()`. **Prepare** (`shared/bootstrap.ts`) is the on-ramp inside the agent: logging -targets, the gateway mint and the scan-triage classifier. Whether the run turns -out to be linear or orchestrator, anthropic or pi, the setup is the same. +targets, caller-supplied inference auth and the scan-triage classifier. Whether +the run turns out to be linear or orchestrator, anthropic or pi, the setup is +the same. -**The switchboard** (`switchboard/`) is the router. Given a program id + the -fetched flags + any CLI overrides, it returns a `ProgramBinding` — which query -shape (sequence), which agent SDK (harness), which model. Two independent -middleware chains, one per axis, apply precedence rules (CLI > flag > program -config > default). This is the only layer that makes routing decisions. +**The switchboard** (`switchboard/`) contains the sequence, harness and model +resolution helpers. The program layer turns its program ID, validated flag route +and CLI overrides into a resolved binding before calling `runAgent`. Agent code +uses that binding to select a sequence and harness; it does not read the program +registry or parse feature flags. **Sequences** (`sequence/`) are LLM query shapes. Once the switchboard has picked one, that sequence takes over the run and owns _how the LLM's work is @@ -77,24 +78,80 @@ gateway. ## How they connect -- Programs supply inference auth; prepare resolves it and builds triage for the resolved harness. -- The switchboard knows which sequences and harnesses exist (via its two - registries), but not what they do. +- Programs supply inference auth; prepare resolves it and builds triage for the + resolved harness. +- The program layer resolves the binding with the switchboard helpers. Agent + code dispatches the selected sequence and harness. - A sequence knows how to shape a conversation, but delegates the actual model call to a harness. - A harness adapts its SDK, gateway transport, security hooks, and tool surface. Each layer is replaceable. +## Ownership map + +```mermaid +%%{init: {"block": {"padding": 20}}}%% +block-beta + columns 11 + hostBand["Host: program or caller"]:11 + runProgram["runProgram"]:3 space:1 programOutcome["ProgramRunOutcome"]:3 space:1 hostSignal["Host AbortSignal"]:3 + space:11 + runnerBand["Agent runner"]:11 + runAgent["runAgent"]:3 space:1 runResult["RunResult"]:3 space:1 runnerSignal["RunAgentOptions.signal"]:3 + space:11 + sequenceBand["Selected sequence"]:11 + sequence["linear | orchestrator"]:3 space:1 sequenceResult["SequenceResult"]:3 space:1 sequenceSignal["signal"]:3 + space:11 + harnessBand["Selected harness"]:11 + agentHarness["AgentHarness"]:3 space:1 agentResult["AgentResult"]:3 space:1 harnessSignal["harness input signal"]:3 + space:11 + sdkBand["External model SDK"]:11 + sdk["Selected SDK"]:3 space:8 + + runProgram --> runAgent + runProgram --> programOutcome + runAgent --> sequence + runAgent --> runResult + sequence --> agentHarness + agentHarness --> sdk + agentHarness --> agentResult + agentResult --> sequenceResult + sequenceResult --> runResult + runResult --> programOutcome + hostSignal --> runnerSignal + runnerSignal --> sequenceSignal + sequenceSignal --> harnessSignal + harnessSignal --> agentHarness + + classDef owner fill:#9ca3af1f,stroke:#9ca3af,stroke-width:1.5px + classDef contract fill:#3b82f626,stroke:#3b82f6,stroke-width:2px + class hostBand,runnerBand,sequenceBand,harnessBand,sdkBand owner + class programOutcome,runResult,sequenceResult,agentResult contract +``` + +Calls descend on the left, results return through the middle, and a host-owned +abort signal descends on the right. A standalone caller invokes `runAgent` +without `runProgram`. A program may return a pre-run failure without starting +the agent, and a caught preparation error produces `RunResult` without a +`SequenceResult`. The orchestrator stops starting new tasks after a fatal run +error and waits for active siblings; that error does not itself abort those +siblings. A host signal can cancel active harness work. + ## Flow 1. The caller runs its gates, authenticates, fetches PostHog flags and resolves a `ProgramBinding { sequence, harness, model }`; analytics tags the run. -2. `runAgent(config, input, options)` resolves the supplied inference auth and prepares triage. +2. `runAgent(config, input, options)` resolves the supplied inference auth and + prepares triage. 3. Sequence takes over — shapes the LLM's work into one conversation (linear) or many (orchestrator), reporting through `onProgress`. 4. Harness drives each conversation through its SDK, using the bound model, on the PostHog LLM gateway. -5. The scan report flushes; `runAgent` returns a `RunResult`. -6. The caller applies it: a decided failure goes to `wizardAbort`, a crash is - rethrown for the runner's own handling. +5. The scan report flushes as the run ends. `runAgent` normally resolves a + `RunResult` with an outcome and progress snapshot. A non-success result + carries a failure whose `error` may be available when the agent caught an + `Error`; an error in final report flushing can still reject. +6. The caller applies it. The legacy runner sends a decided failure to + `wizardAbort`; for a crash it rethrows the attached `Error` when present. + Other hosts can log, present, or rethrow the failure as they need. diff --git a/src/programs/README.md b/src/programs/README.md index 7d6046e21..dbd17b037 100644 --- a/src/programs/README.md +++ b/src/programs/README.md @@ -78,8 +78,12 @@ handoff. An AI program whose organization lacks AI-processing approval needs `ProgramRunOutcome.outcome` is `success`, `aborted`, `failed`, or `crashed`. Decided pre-run failures, such as an unknown program or missing credentials, return `failed` with `failure.message`. An agent crash appears as `crashed`. -External host capabilities can still reject, so callers should also handle a -rejected promise. +Non-success agent outcomes carry the agent's `failure`, including any attached +`Error`. `failure.error` is optional, as are its code and message. Read the +outcome to decide how the run ended, and use the attached error for diagnostics +or an upstream rethrow. The host owns logging and user-facing error messages. +External host capabilities and failures outside the agent's run-body catch can +still reject, so callers should also handle a rejected promise. `options.signal` accepts an `AbortSignal`. A signal aborted before the run starts returns `aborted` with an agent-abort failure code. During an agent run, @@ -136,6 +140,7 @@ export async function runMetrics( ); if (result.outcome !== 'success') { + if (result.failure?.error) throw result.failure.error; throw new Error(result.failure?.message ?? `Metrics ${result.outcome}`); } return result; @@ -151,8 +156,7 @@ The callable host does not discover credentials, choose a project, or render questions itself. The host supplies those data and capabilities. Some legacy recipes still need an explicit data-only run definition; unsupported combinations return a failed outcome. Existing terminal and CI callers route -their agent work through this host via -`src/lib/runners/run-program-agent.ts`, which still owns session gates and UI -translation. There is no socket controller or step-by-step -control API. For the existing process-owned CI runner, see the +their agent work through this host via `src/lib/runners/run-program-agent.ts`, +which still owns session gates and UI translation. There is no socket controller +or step-by-step control API. For the existing process-owned CI runner, see the [non-interactive developer interfaces](../../docs/developer-interfaces.md). From 0cc620dfb9f5e5fcae8ad25925cbb14774e08680 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Tue, 22 Sep 2026 21:53:09 -0400 Subject: [PATCH 05/29] docs: reconcile agent signal contract after stack merge --- src/agent/README.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/src/agent/README.md b/src/agent/README.md index 87bb85130..694d1b2ec 100644 --- a/src/agent/README.md +++ b/src/agent/README.md @@ -16,7 +16,6 @@ import { runAgent, RunOutcome } from '@agent'; import type { RunConfig, RunInput, RunResult, AgentProgress, AgentInteraction } from '@agent/types'; runAgent(config: RunConfig, input: RunInput, options?: { - signal?: AbortSignal; onProgress?: (event: AgentProgress) => void; interaction?: AgentInteraction; signal?: AbortSignal; @@ -85,7 +84,8 @@ if (result.outcome !== RunOutcome.Success) { `src/agent/__tests__/run-agent-standalone.test.ts` runs this with no UI, no store and no registry. -Pass an `AbortController` signal in the options and call `controller.abort()` to cancel an active run. The result then has `RunOutcome.Aborted`. +Pass an `AbortController` signal in the options and call `controller.abort()` to +cancel an active run. The result then has `RunOutcome.Aborted`. ## Intent From ade12f5365289b4243757a0fb2ccb184cf63dba2 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Tue, 22 Sep 2026 22:17:47 -0400 Subject: [PATCH 06/29] Align runner comment with scan flush contract --- src/agent/runner/index.ts | 7 ++++--- 1 file changed, 4 insertions(+), 3 deletions(-) diff --git a/src/agent/runner/index.ts b/src/agent/runner/index.ts index 744a563b2..cba412d7a 100644 --- a/src/agent/runner/index.ts +++ b/src/agent/runner/index.ts @@ -13,13 +13,14 @@ * [skill install] → agent init → prompt → run → errors → [postRun] → outro * * The agent reports and asks, it never renders, never reads a session, never - * exits the process and never rejects. A decided failure comes back in + * exits the process. A decided failure comes back in * `RunResult.failure` with the same fields `wizardAbort` takes; an error the * agent did not decide (a refused mint, an SDK crash) comes back as * `outcome: RunOutcome.Crashed` with the original error attached, so a caller can keep * handling it the way it always did. The legacy adapter in * `src/lib/runners/run-program-agent.ts` rebuilds today's session-driven - * behavior on top of this call for every existing caller. + * behavior on top of this call for existing program callers. Final scan-report + * flushing can still reject after the run body has settled. */ import { Sequence } from '@shared/constants'; @@ -149,7 +150,7 @@ export async function runAgent( }; } // Not a decision the agent made. Hand it back whole rather than throw, so - // every ending of a run is a result the caller reads the same way. + // run-body endings are results the caller reads the same way. const failure = classifyRunFailure(error); logToFile('[agent-runner] run crashed:', error); cleanFailedRun(); From 249a2a37fa6b66939c7a5e5cf84f20790db33a1e Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Tue, 22 Sep 2026 22:28:17 -0400 Subject: [PATCH 07/29] docs: align agent observer contract with A3 --- docs/developer-interfaces.md | 12 ++++++------ 1 file changed, 6 insertions(+), 6 deletions(-) diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index 687c44cb3..6ccbe6a3b 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -19,12 +19,12 @@ flags and host), and optional `onProgress`, `interaction`, and `signal` options. Import the function from `@agent` and types from `@agent/types`. It returns a `RunResult` with a `success`, `aborted`, `failed`, or `crashed` outcome and a final task/status/usage snapshot. Progress is delivered in emission order. -Synchronous observer throws are logged without failing the run; asynchronous -observers must handle their own rejected promises. Without `interaction`, -questions have no answer bridge and optional task notices are declined. -Non-success results carry a `failure`; its `error` may be available when an -`Error` was caught, while `code` and `message` are optional. The caller chooses -how to log or present a failure. +Observer throws and rejections from thenables returned by `onProgress` are +logged without failing the run. Callbacks must handle errors from detached +asynchronous work they start. Without `interaction`, questions have no answer +bridge and optional task notices are declined. Non-success results carry a +`failure`; its `error` may be available when an `Error` was caught, while `code` +and `message` are optional. The caller chooses how to log or present a failure. ```ts import { runAgent } from '@agent'; From a845d1fbe7c6ef62582d583201404d69231cccb3 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Tue, 22 Sep 2026 22:44:19 -0400 Subject: [PATCH 08/29] docs: align B4 contracts with A3 terminal outcomes --- docs/developer-interfaces.md | 11 ++++++----- src/agent/README.md | 34 +++++++++++++++++----------------- src/agent/runner/README.md | 14 +++++++------- src/programs/README.md | 11 +++++------ 4 files changed, 35 insertions(+), 35 deletions(-) diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index 6ccbe6a3b..eb03e361c 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -23,8 +23,8 @@ Observer throws and rejections from thenables returned by `onProgress` are logged without failing the run. Callbacks must handle errors from detached asynchronous work they start. Without `interaction`, questions have no answer bridge and optional task notices are declined. Non-success results carry a -`failure`; its `error` may be available when an `Error` was caught, while `code` -and `message` are optional. The caller chooses how to log or present a failure. +`failure` with a code and message, and may have an attached `Error`. The caller +chooses how to log or present a failure. ```ts import { runAgent } from '@agent'; @@ -49,9 +49,10 @@ PostHog user or detect the project. The caller must supply `input.inferenceAuth`, whose `resolve()` returns gateway authentication and can refresh it during a long run. There is no session control protocol on this API. An aborted signal returns an `aborted` result; it does not pause the run. The -agent catches unexpected errors in its run body and reports `crashed` with the -caught `Error` (or an `Error` wrapper for a non-`Error` throw); the host can -rethrow that object when it needs exception semantics. +agent returns a caught coded error as `failed` and an uncoded throw as +`crashed`. Both retain the caught `Error` (or an `Error` wrapper for a +non-`Error` throw), which the host can rethrow when it needs exception +semantics. ### Inference authentication diff --git a/src/agent/README.md b/src/agent/README.md index 7eec0dbca..c0fecd73f 100644 --- a/src/agent/README.md +++ b/src/agent/README.md @@ -16,7 +16,7 @@ import { runAgent, RunOutcome } from '@agent'; import type { RunConfig, RunInput, RunResult, AgentProgress, AgentInteraction } from '@agent/types'; runAgent(config: RunConfig, input: RunInput, options?: { - onProgress?: (event: AgentProgress) => void; + onProgress?: (event: AgentProgress) => unknown; interaction?: AgentInteraction; signal?: AbortSignal; }): Promise @@ -38,13 +38,12 @@ runAgent(config: RunConfig, input: RunInput, options?: { for gateway token minting and refresh. - `RunResult`: `outcome` is `RunOutcome.Success | Aborted | Failed | Crashed`. Success may carry an `outro`; the other three carry a `failure` - (`AgentFailure`: optional message, outro data, `Error`, exit code, error code, - and detail). `failure.error` may be present when the agent caught an `Error`; - `Crashed` requires one. A missing `Error` object does not mean the outcome - succeeded, and not every failed result has a code or message. Every result - carries `skillId` and a `snapshot` of what the run reported: tasks, status - lines, stage, token usage totals, final cost, dashboard and notebook URLs, - handoff text. + (`AgentFailure`: required code and message, optional outro data, `Error`, exit + code, detail, and authentication detail). `failure.error` may be attached; + `Crashed` requires one. A failed result need not have an attached `Error`. + Every result carries `skillId` and a `snapshot` of what the run reported: + tasks, status lines, stage, token usage totals, final cost, dashboard and + notebook URLs, handoff text. - `AgentProgress`: one event per thing the run reports, in emission order. Kinds: `lifecycle`, `spinner`, `log`, `status`, `tasks`, `stage`, `url`, `usage`, `finalCost`, `authError`, `handoff`, `completion`. Payloads are @@ -57,13 +56,13 @@ runAgent(config: RunConfig, input: RunInput, options?: { returns `Aborted` before execution; aborting during execution is passed to the active harness and returns `Aborted` with the current snapshot. It does not pause or resume a run. -- Errors: the agent does not exit the process or throw for a decided failure. It - catches unexpected errors in its run body, logs them, and returns - `outcome: Crashed` with the caught `Error` attached (or an `Error` wrapper for - a non-`Error` throw). A gateway 401 emits `authError` and then fails. The host - decides how to present a returned failure, set an exit code, or rethrow an - attached error. Final scan-report flushing runs outside that catch and can - still reject the promise; hosts requiring a hard promise boundary catch it. +- Errors: the agent does not exit the process or throw for a decided failure. A + caught coded error returns `Failed`; an uncoded throw returns `Crashed`. Both + retain the caught `Error` (or an `Error` wrapper for a non-`Error` throw). A + gateway 401 returns an authentication failure with detail for the host to + present. The host decides how to present a returned failure, set an exit code, + or rethrow an attached error. Final scan-report flushing is best effort and + does not replace the run result. Other runtime exports: `DEFAULT_AGENT_BINDING` for standalone callers, the generic `resolveBinding` and `resolveHarness` helpers, `shouldDisableAsk`, @@ -115,8 +114,9 @@ Without `onProgress` the run completes and its snapshot still comes back in the result. Without `interaction` the agent installs no ask bridge: `wizard_ask` returns its "not available" error and optional task notices are declined, which is what a `--ci` run does. A throwing observer is logged and the run continues. -Progress callbacks are not awaited, so an asynchronous observer must handle its -own rejected promises. +Progress callbacks are not awaited. Throws and rejections from returned +thenables are logged; observers must handle errors from detached work they +start. ## Architecture diff --git a/src/agent/runner/README.md b/src/agent/runner/README.md index b95f87405..9ec78687c 100644 --- a/src/agent/runner/README.md +++ b/src/agent/runner/README.md @@ -134,9 +134,10 @@ Calls descend on the left, results return through the middle, and a host-owned abort signal descends on the right. A standalone caller invokes `runAgent` without `runProgram`. A program may return a pre-run failure without starting the agent, and a caught preparation error produces `RunResult` without a -`SequenceResult`. The orchestrator stops starting new tasks after a fatal run -error and waits for active siblings; that error does not itself abort those -siblings. A host signal can cancel active harness work. +`SequenceResult`. The orchestrator stops scheduling on the first fatal task +result, cancels active siblings and pending asks, and waits for them to settle +before returning that failure. A host signal can also cancel active harness +work. ## Flow @@ -148,10 +149,9 @@ siblings. A host signal can cancel active harness work. many (orchestrator), reporting through `onProgress`. 4. Harness drives each conversation through its SDK, using the bound model, on the PostHog LLM gateway. -5. The scan report flushes as the run ends. `runAgent` normally resolves a - `RunResult` with an outcome and progress snapshot. A non-success result - carries a failure whose `error` may be available when the agent caught an - `Error`; an error in final report flushing can still reject. +5. The scan report flushes on a best-effort basis as the run ends. `runAgent` + resolves a `RunResult` with an outcome and progress snapshot. A non-success + result carries a code and message; a caught error remains attached. 6. The caller applies it. The legacy runner sends a decided failure to `wizardAbort`; for a crash it rethrows the attached `Error` when present. Other hosts can log, present, or rethrow the failure as they need. diff --git a/src/programs/README.md b/src/programs/README.md index f54c37eff..d0fc85073 100644 --- a/src/programs/README.md +++ b/src/programs/README.md @@ -99,12 +99,11 @@ return `failed` with `failure.message`. An agent crash appears as `crashed`. Handled credential-resolution, approval, and composition callback rejections also resolve as `failed` with a message, not the callback's original error class. An unexpected invocation throw, such as a duplicate composed `runId`, -rejects the promise. Non-success agent outcomes carry the agent's `failure`, -including any attached `Error`. `failure.error` is optional, as are its code and -message. Read the outcome to decide how the run ended, and use the attached -error for diagnostics or an upstream rethrow. The host owns logging and -user-facing error messages. Hosts should inspect the outcome and separately -catch rejected promises. +rejects the promise. Non-success agent outcomes carry a failure code and message +and may have an attached `Error`. Read the outcome to decide how the run ended, +and use the attached error for diagnostics or an upstream rethrow. The host owns +logging and user-facing error messages. Hosts should inspect the outcome and +separately catch rejected promises. `options.signal` accepts an `AbortSignal`. A signal aborted before the run starts returns `aborted` with an agent-abort failure code. Agent programs check From 98c3e78ec7701d3363afe47815a25598ec45d6c5 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Wed, 23 Sep 2026 12:19:13 -0400 Subject: [PATCH 09/29] docs: correct the callable program's rejection and signal notes, and point to the reference hosts - The developer interfaces said external host callbacks may reject the runProgram promise. Credential, approval and composition callback rejections resolve as `failed`. Only an unexpected invocation error rejects. - The programs reference said workflow requests receive the signal. Only the no-agent `workflow` does. `compositionWorkflow.confirmStep` does not, and its rejection during an abort returns `failed`. - The agent and program sections now name their runnable reference hosts, `pnpm test:e2e:agent` and `pnpm test:e2e:programs`, and the environment they read. - The agent reference lists `OutroKind` among the runtime exports. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- docs/developer-interfaces.md | 18 ++++++++++++++++-- src/agent/README.md | 2 +- src/programs/README.md | 7 ++++--- 3 files changed, 21 insertions(+), 6 deletions(-) diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index eb03e361c..fe54c8677 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -54,6 +54,13 @@ agent returns a caught coded error as `failed` and an uncoded throw as non-`Error` throw), which the host can rethrow when it needs exception semantics. +A runnable reference host is `scripts/e2e-agent.no-jest.ts`, run by +`pnpm test:e2e:agent`. It runs a `quack` skill from a loopback skills server in +an empty directory. Its environment is described in +`e2e-harness/surface-e2e.ts`: `PROJECT_ID`, a PostHog key from +`POSTHOG_PERSONAL_API_KEY` or `POSTHOG_KEY_FILE`, and a gateway token from +`WIZARD_CI_GATEWAY_TOKEN_FILE`. + ### Inference authentication For first-party inference authentication, import @@ -74,8 +81,10 @@ optional `signal` requests cancellation. It returns a `ProgramRunOutcome`: outcome and failure, final progress, actual settled agent runs, program-specific data, artifacts, and invocation data (including a captured event plan). The latter contains credentials and should not be logged. Agent failures retain an -attached `Error` when one exists. External host callbacks may still reject the -promise, so callers handle those exceptions as well as returned outcomes. +attached `Error` when one exists. Rejections from the credential, approval and +composition callbacks resolve as `failed`. The promise rejects only on an +unexpected invocation error, such as a duplicate composed `runId` or a run +definition that throws, so callers read the outcome and still catch a rejection. ```ts import { runProgram } from '@programs'; @@ -114,6 +123,11 @@ the available fields and capabilities. Host callbacks such as credential resolution, approval, and MCP work do not receive the signal. There is no live store or step-control handle. +A runnable reference host is `scripts/e2e-programs.no-jest.ts`, run by +`pnpm test:e2e:programs`. It runs posthog-integration against the app in +`APP_DIR`, with the same environment as the agent route +(`e2e-harness/surface-e2e.ts`). + ## Development CI and experimental headless runner Development/test builds accept `--ci`. This is a whole-process CLI path, not an diff --git a/src/agent/README.md b/src/agent/README.md index 8f543be56..49315fa39 100644 --- a/src/agent/README.md +++ b/src/agent/README.md @@ -67,7 +67,7 @@ runAgent(config: RunConfig, input: RunInput, options?: { Other runtime exports: `DEFAULT_AGENT_BINDING` for standalone callers, the generic `resolveBinding` and `resolveHarness` helpers, `shouldDisableAsk`, `initializeAgent`, `executeAgent`, `buildRunTags`, `AgentSignals`, -`downloadSkill`, `WIZARD_TOOL_NAMES`, `LONGER_ASK_TIMEOUT_MS`, +`downloadSkill`, `WIZARD_TOOL_NAMES`, `LONGER_ASK_TIMEOUT_MS`, `OutroKind`, `flushScanReport`, and `runMcpPromptViaSdk`, which loads the streaming module on first call. diff --git a/src/programs/README.md b/src/programs/README.md index 10d740bc2..43978a31b 100644 --- a/src/programs/README.md +++ b/src/programs/README.md @@ -110,9 +110,10 @@ separately catch rejected promises. starts returns `aborted` with an agent-abort failure code. Agent programs check again before startup and pass the signal to the active harness. Programs without an agent recheck after credential resolution and host work; an abort returns -`aborted` even if a callback has completed. Workflow requests receive the -signal, but in-flight host effects must cooperate with cancellation and -completed external effects are not rolled back. +`aborted` even if a callback has completed. The no-agent `workflow` receives the +signal. `compositionWorkflow.confirmStep` does not, and a rejection from it +during an abort returns `failed`. In-flight host effects must cooperate with +cancellation, and completed external effects are not rolled back. The result includes: From ba93d813b824286d875ca435abb90dd03ac52310 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Wed, 23 Sep 2026 18:17:36 -0400 Subject: [PATCH 10/29] docs: describe the post-WP3 agent and program contracts Bring the developer interfaces guide and the agent and runner READMEs in line with the code. The callable program section now covers the two ways to supply credentials, with the provider's resolve(programId, { signal }) called once per invocation. It says runProgram identifies the user, stamps the AI SDK evidence and refreshes a token near expiry, resolves the binding from input.overrides, and copies its input. It lists which capabilities receive the signal, including the workflow connector's step, and the example narrows on the ProgramProgress kind. A new Preflight section describes the shared readiness and settings checks. Both surfaces now say they send no terminal analytics, and the CI section says detection runs through its own runAgent call. The agent README adds the collectTranscript tail and requestRemark options, spells out scanReport: 'defer', lists the ten @agent runtime names, passes the per-request signal in the ask example, and names agentic detection as a direct runAgent caller. The runner README moves credential resolution, stamping, refresh, overrides and the switchboard decision into runProgram and starts the flow with preflight. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- docs/developer-interfaces.md | 109 +++++++++++++++++++++++++---------- src/agent/README.md | 85 +++++++++++++++++---------- src/agent/runner/README.md | 55 ++++++++++-------- 3 files changed, 161 insertions(+), 88 deletions(-) diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index b57c48dc9..48cf6ad4f 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -2,8 +2,8 @@ Wizard has two repository-local TypeScript call surfaces and one development CLI mode for running without a terminal UI. The TypeScript aliases below are -internal to this repository; `@posthog/wizard` currently publishes a CLI, not -these functions as a stable package API. +internal to this repository. `@posthog/wizard` publishes a CLI, not these +functions as a stable package API. | Surface | Use it for | Detailed contract | | --------------------------------------- | ----------------------------------------------- | ------------------------------------------------------------------------------------------ | @@ -44,15 +44,16 @@ export async function runStandalone( } ``` -The caller prepares `config` and `input`; the agent does not authenticate the +The caller prepares `config` and `input`. The agent doesn't authenticate the PostHog user or detect the project. The caller must supply `input.inferenceAuth`, whose `resolve()` returns gateway authentication and can refresh it during a long run. There is no session control protocol on this API. -An aborted signal returns an `aborted` result; it does not pause the run. The +An aborted signal returns an `aborted` result. It doesn't pause the run. The agent returns a caught coded error as `failed` and an uncoded throw as `crashed`. Both retain the caught `Error` (or an `Error` wrapper for a non-`Error` throw), which the host can rethrow when it needs exception -semantics. +semantics. `runAgent` never sends the terminal `setup wizard finished` event. +The host sends it when its process is done. A runnable reference host is `scripts/e2e-agent.no-jest.ts`, run by `pnpm test:e2e:agent`. It runs a `quack` skill from a loopback skills server in @@ -67,24 +68,52 @@ For first-party inference authentication, import `createPosthogInferenceAuthProvider` from `@programs` and pass authenticated PostHog credentials and the run's program ID (for example, `config.programId`, `'audit'`, or `'metrics'`). Its provider mints a gateway token and refreshes it -near expiry; the host still handles user login and project selection. +near expiry. The host still handles user login and project selection. +`runProgram` builds this provider itself when resolved credentials leave +`inferenceAuth` out, so you need it only when you call `runAgent` directly. Development [CI](#development-ci-and-experimental-headless-runner) instead uses an already-issued fixed token. ## Callable program -`runProgram` takes a registered ID, `ProgramInput` with at least `installDir`, -and optional `ProgramOptions`. The host supplies resolved credentials or a -credential provider, prepared detection and framework context where needed, and -callbacks for questions, approvals, progress, or program-specific effects. An -optional `signal` requests cancellation. It returns a `ProgramRunOutcome`: -outcome and failure, final progress, actual settled agent runs, program-specific -data, artifacts, and invocation data (including a captured event plan). The -latter contains credentials and should not be logged. Agent failures retain an -attached `Error` when one exists. Rejections from the credential, approval and -composition callbacks resolve as `failed`. The promise rejects only on an -unexpected invocation error, such as a duplicate composed `runId` or a run -definition that throws, so callers read the outcome and still catch a rejection. +`runProgram` takes a registered ID, a `ProgramInput` with at least `installDir`, +and optional `ProgramOptions`. Import it from `@programs` and types from +`@programs/types`. It returns a `ProgramRunOutcome`: outcome and failure, final +progress, actual settled agent runs, program-specific data, artifacts, and +invocation data (including a captured event plan). The invocation data contains +credentials, so don't log it. Agent failures retain an attached `Error` when one +exists. + +The host supplies credentials in one of two ways: + +- **Resolved.** `input.credentials` carries + `{ posthog, inferenceAuth?, project, apiUser }`. +- **A provider.** `options.credentials.resolve(programId, { signal })`. + `runProgram` calls it once per invocation, and composed child runs reuse the + result. + +Either way, `runProgram` identifies the user for analytics, stamps the +organization's AI SDK evidence, and refreshes an OAuth token that is close to +expiry before each agent run. Launch choices go in `input.overrides` as +`{ harness?, sequence?, model? }`. `runProgram` resolves the binding from them +and the flag snapshot, and captures the switchboard decision once for each agent +run. It copies the input when it receives it, so a later host write can't reach +the run. + +Awaited host capabilities receive the invocation's signal: +`credentials.resolve`, `awaitAiApproval({ programId, signal })`, and +`workflow.step(request, { signal })`. The workflow connector answers the +post-auth, child-run and confirm requests that gated and composed programs make. +A rejection from any of them resolves as `failed`, or as `aborted` once the +signal has aborted. `featureFlags`, the MCP port and the integration effects +don't receive the signal. The promise rejects only on an invocation error, such +as a duplicate composed `runId`, a run definition that throws, or input that +can't be copied. Read the outcome, and still catch a rejection. + +`onProgress` receives two kinds of `ProgramProgress`. A run event is +`{ kind: 'run', runId, stepId?, event }`, where `event` is the agent's progress. +A program-data event is `{ kind: 'program', data }`, a copy of the invocation +data after each write. Narrow on `kind` first: ```ts import { runProgram } from '@programs'; @@ -103,8 +132,11 @@ export async function runAudit( credentials, awaitAiApproval, signal, - onProgress: ({ runId, event }) => { - if (event.kind === 'tasks') console.log(runId, event.tasks); + onProgress: (progress) => { + if (progress.kind !== 'run') return; + if (progress.event.kind === 'tasks') { + console.log(progress.runId, progress.event.tasks); + } }, }, ); @@ -117,17 +149,29 @@ export async function runAudit( ``` The caller implements the credential and approval callbacks. Some programs -require additional prepared inputs or host effects; the +require additional prepared inputs or host effects. The [program reference](../src/programs/README.md#inputs) describes the available -fields and capabilities. Host callbacks such as credential resolution, approval, -and MCP work do not receive the signal. There is no live store or step-control -handle. +fields and capabilities. There is no live store or step-control handle. +`runProgram` never sends the terminal `setup wizard finished` event. A +long-lived host decides when to send it, from the outcome. A runnable reference host is `scripts/e2e-programs.no-jest.ts`, run by `pnpm test:e2e:programs`. It runs posthog-integration against the app in `APP_DIR`, with the same environment as the agent route (`e2e-harness/surface-e2e.ts`). +### Preflight + +Hosts call `preflight(programId, host)` from `@programs` before `runProgram`. +`runProgram` doesn't call it. It runs the readiness check, then the Claude +settings check, and returns `{ kind: 'proceed', restoreSettings }` or +`{ kind: 'abort', failure }`. The host supplies the presentation (`showOutage`, +`setReadinessWarnings` and `showSettingsOverride`) and its policy +(`interactive`, `signup`, and any readiness it already computed). An outage +aborts only an interactive host. An unfixable settings conflict aborts only a +non-interactive host. Call `restoreSettings()` when the run ends. See the +[program reference](../src/programs/README.md#preflight) for the details. + ## Development CI and experimental headless runner Development/test builds accept `--ci`. This is a whole-process CLI path, not an @@ -145,12 +189,13 @@ pnpm try --ci --api-key "$POSTHOG_PERSONAL_API_KEY" \ The runner logs progress and writes a local task-stream JSONL dump. Callers observe the process exit and its logs, rather than a returned result. The gateway token file is read into a fixed provider for CI. Pre-run detection and -composed child runs use that same provider; this path does not mint or refresh +composed child runs use that same provider. This path doesn't mint or refresh the token. Published builds reject `--ci`. The internal -`runWizardCI(config, options): void` entry point still uses the legacy session -adapter, which calls `runProgram` for each program's main agent run. Agentic -detection can call the agent separately before that run; MCP suggested prompts -also use a separate agent path with their own progress and cancellation. +`runWizardCI(config, options): void` entry point uses the session adapter, +`src/lib/runners/run-program-agent.ts`. The adapter runs `preflight`, then calls +`runProgram` for each program's main agent run. Agentic detection runs before +that call, through its own `runAgent` call. MCP suggested prompts use a separate +SDK path with their own progress and cancellation. An experimental published-build headless path exists internally as `runWizardHeadless(config, options): void`. It shares the process-owned runner, @@ -158,6 +203,6 @@ logs progress, and can push task-stream updates to PostHog when telemetry is enabled. Its selector is deliberately hidden and is not a supported invocation recipe. Neither internal function returns a structured, awaitable outcome. -Controlled headless and socket control APIs are not available yet. There is no -supported route, command, or event protocol for pausing a run, supplying an -answer later, or reading its live state from another process. +There is no controlled headless mode and no socket control API. No supported +route, command, or event protocol pauses a run, supplies an answer later, or +reads its live state from another process. diff --git a/src/agent/README.md b/src/agent/README.md index c112b36ac..cb70cdd01 100644 --- a/src/agent/README.md +++ b/src/agent/README.md @@ -9,7 +9,7 @@ the callable program host and development CI runner, see the ## Signatures Import runtime values from `@agent` and types from `@agent/types`. Nothing -outside `src/agent` imports deeper; lint and the architecture test reject it. +outside `src/agent` imports deeper. Lint and the architecture test reject it. ```ts import { runAgent, RunOutcome } from '@agent'; @@ -26,8 +26,15 @@ runAgent(config: RunConfig, input: RunInput, options?: { tools, copy), the resolved `binding` (sequence, harness, model and task-role routes), supplied program commandments and stage policy, the skills origin, flag snapshot, trace tags, tool allow and deny lists, seed tasks, bound - completion `hooks` and `scanReport` (`defer` leaves the scan report to the - host run). + completion `hooks` and `scanReport`. `scanReport: 'defer'` leaves this run's + scans to the host run's report. The default, `'flush'`, writes the report when + this run ends. +- Two `AgentRunDefinition` options shape the run's output. + `collectTranscript: true` keeps the last 256 KiB of assistant text as + `snapshot.transcriptTail` and reports each step as `activity` progress. It + works on the linear sequence with the Anthropic harness. + `requestRemark: false` skips the end-of-run reflection remark, which is on by + default. - `RunInput`: install directory, resolved PostHog credentials, required `inferenceAuth`, project and user payloads, skill id, detected integration, `flags` (`ci`, `signup`, `debug`, `e2eAsk`, `localMcp`, `captureAio`, @@ -38,9 +45,9 @@ runAgent(config: RunConfig, input: RunInput, options?: { [first-party provider](../../docs/developer-interfaces.md#inference-authentication) for gateway token minting and refresh. - `RunResult`: `outcome` is `RunOutcome.Success | Aborted | Failed | Crashed`. - Success may carry an `outro`; the other three carry a `failure` + Success may carry an `outro`. The other three carry a `failure` (`AgentFailure`: required code and message, optional outro data, `Error`, exit - code, detail, and authentication detail). `failure.error` may be attached; + code, detail, and authentication detail). `failure.error` may be attached, and `Crashed` requires one. A failed result need not have an attached `Error`. Every result carries a `snapshot` of what the run reported: tasks, status lines, stage, token usage totals, final cost, dashboard and notebook URLs, @@ -55,29 +62,39 @@ runAgent(config: RunConfig, input: RunInput, options?: { resolves with answers, and `taskNotice(notice, { signal })` resolves with whether to keep an optional task. Each request has its own signal, which aborts when that request times out, the host aborts the run, or another task - fails the run; on abort the host dismisses that request alone, without + fails the run. On abort the host dismisses that request alone, without throwing. - `signal`: an optional `AbortSignal` from the host. A pre-aborted signal - returns `Aborted` before execution; aborting during execution is passed to the + returns `Aborted` before execution. An abort during execution reaches the active harness and returns `Aborted` with the current snapshot. It does not pause or resume a run. `Aborted` means only that the host's signal cancelled the run. - Errors: the agent does not exit the process or throw for a decided failure. A - caught coded error returns `Failed`; an uncoded throw returns `Crashed`. Both - retain the caught `Error` (or an `Error` wrapper for a non-`Error` throw). An - agent that stops itself with `[ABORT]` returns `Failed` with its abort code. A - gateway 401 returns an authentication failure with detail for the host to - present. The host decides how to present a returned failure, set an exit code, - or rethrow an attached error. Final scan-report flushing is best effort and - does not replace the run result. + caught coded error returns `Failed`, and an uncoded throw returns `Crashed`. + Both retain the caught `Error` (or an `Error` wrapper for a non-`Error` + throw). An agent that stops itself with `[ABORT]` returns `Failed` with its + abort code. A gateway 401 returns an authentication failure with detail for + the host to present. The host decides how to present a returned failure, set + an exit code, or rethrow an attached error. Final scan-report flushing is best + effort and does not replace the run result. - Analytics shutdown is host-owned: the agent never sends the terminal - `setup wizard finished` event. The host sends it from the outcome: `Success` - is `success`, `Aborted` is `cancelled`, `Failed` and `Crashed` are `error`. - -Other runtime exports: `DEFAULT_AGENT_BINDING` for standalone callers, -`resolveHarness` and `harnessRunsTasks`, which programs resolve a binding with, -`AgentSignals`, `WIZARD_TOOL_NAMES`, `OutroKind`, `downloadSkill`, and -`runMcpPromptViaSdk`, which loads the streaming module on first call. + `setup wizard finished` event, and `runProgram` doesn't either. The host sends + it from the outcome: `Success` is `success`, `Aborted` is `cancelled`, + `Failed` and `Crashed` are `error`. + +`@agent` exports ten runtime names, and +`src/agent/__tests__/public-entry.test.ts` holds that list: + +- **`runAgent` and `RunOutcome`.** The run and its outcome enum. +- **`OutroKind`.** The kind of an outro in `completion` progress and in failure + outro data. +- **`DEFAULT_AGENT_BINDING`.** The Pi and linear binding for standalone callers. +- **`resolveHarness` and `harnessRunsTasks`.** What programs resolve a binding + with. `harnessRunsTasks` says which harnesses the orchestrator can drive. +- **`AgentSignals` and `WIZARD_TOOL_NAMES`.** The marker strings that program + prompts embed, and the tool ids that go in tool allow and deny lists. +- **`downloadSkill` and `runMcpPromptViaSdk`.** The skill installer and the + suggested-prompts stream. Each loads its module on first call. Minimal invocation: @@ -87,7 +104,7 @@ const result = await runAgent(config, input, { if (event.kind === 'log') console.log(event.message); }, interaction: { - ask: async (question) => answersFor(question), + ask: async (question, { signal }) => answersFor(question, signal), }, }); if (result.outcome !== RunOutcome.Success) { @@ -99,7 +116,7 @@ if (result.outcome !== RunOutcome.Success) { ``` The result is the agent's termination report. Check `outcome` first and then -read `failure`; a failed result can have no attached `Error`. If a higher layer +read `failure`. A failed result can have no attached `Error`. If a higher layer uses exceptions, it can rethrow `failure.error` when present and construct an error from `failure.message` otherwise. Preserve the original `Error` object when rethrowing so its stack and cause remain available. @@ -112,25 +129,31 @@ cancel an active run. The result then has `RunOutcome.Aborted`. ## Intent -Programs call the agent to do the work a skill describes. A standalone host can -observe the run through `onProgress` and answer it through `interaction`; the -legacy TUI and non-interactive runner still use -`src/lib/runners/run-program-agent.ts` for session gates and UI translation, -then call the same `runProgram` host. +Programs call the agent to do the work a skill describes. `runProgram` builds +the `RunConfig` and `RunInput` for every program run. The TUI and the `--ci` +runner reach it through `src/lib/runners/run-program-agent.ts`, which supplies +session capabilities and maps progress back onto `getUI()`. + +Agentic detection calls `runAgent` itself, before the program runs. It uses a +linear Haiku run on the Anthropic harness, with `collectTranscript`, +`requestRemark: false` and `scanReport: 'defer'`. It reads its report from the +transcript tail and makes up to two attempts, with deadlines of 60 and 90 +seconds. A standalone host builds the config and input itself, as +`scripts/e2e-agent.no-jest.ts` does. Without `onProgress` the run completes and its snapshot still comes back in the result. Without `interaction` the agent installs no ask bridge: `wizard_ask` returns its "not available" error and optional task notices are declined. Plain `--ci` runs also disable the ask bridge and decline notices. A throwing observer is logged and the run continues. Progress callbacks are not awaited. Throws and -rejections from returned thenables are logged; observers must handle errors from +rejections from returned thenables are logged. Observers must handle errors from detached work they start. ## Architecture The agent owns run state for one invocation: the task queue, phase, status, resolved skill, handoff text, usage and the final result. It depends on -`src/shared` and `src/env.ts`; callers supply program routing and policy through +`src/shared` and `src/env.ts`. Callers supply program routing and policy through `RunConfig`. ```text @@ -149,4 +172,4 @@ caller ── RunConfig + RunInput ──▶ runAgent `runner/` holds the dispatcher, sequences, harnesses and the switchboard. `tools/` holds the wizard tools shared by both harnesses. `middleware/` holds the benchmark pipeline. `progress.ts` defines the event and interaction -contracts; `yara-hooks.ts` scans what the run installs. +contracts. `yara-hooks.ts` scans what the run installs. diff --git a/src/agent/runner/README.md b/src/agent/runner/README.md index 26dc49418..f170c058e 100644 --- a/src/agent/runner/README.md +++ b/src/agent/runner/README.md @@ -23,7 +23,7 @@ retained for very simple tasks and legacy support. The Anthropic Agent SDK is a supported legacy fallback, deprecated as the default, retained for major Pi vulnerabilities or gaps in support for new Anthropic models. -`DEFAULT_AGENT_BINDING`, the standalone default, is Pi + linear; explicit +`DEFAULT_AGENT_BINDING`, the standalone default, is Pi + linear. Explicit program bindings and flags determine actual behavior. Both harnesses implement `run` and `runTask`. Composed sub-runs are clamped to linear, and linear-only post-run/outro hooks do not automatically transfer to an orchestrated flow. @@ -44,9 +44,11 @@ Five layers, each with its own job. Nothing crosses layers unless it has to. takes resolved execution data and an invocation snapshot (`shared/types.ts`), reports through `onProgress` and asks through `interaction` (`../progress.ts`), and returns decided outcomes and caught run-body crashes as results. It never -renders, reads a session or exits. `src/programs/run-program.ts` resolves -credentials through a host provider, awaits the host's gates, loads flags and -resolves the binding from caller data. The legacy +renders, reads a session or exits. Its caller does the host work. For a program +run, `runProgram` in `src/programs/run-program.ts` calls the host's credentials +provider once, identifies the user, stamps the AI SDK evidence, awaits the +host's approval and workflow connector, loads flags, refreshes an OAuth token +near expiry, and resolves the binding. The legacy `src/lib/runners/run-program-agent.ts` supplies those capabilities from the session and maps progress back onto `getUI()`. @@ -59,10 +61,10 @@ the same. and the harness and model resolution helper, `resolveHarness` (CLI > flag > program config > default). The program layer owns the sequence precedence (`resolveProgramBinding`) and turns its program ID, validated flag route and CLI -overrides into a resolved binding before calling `runAgent`; `harnessRunsTasks` -tells it which harnesses the orchestrator can drive. Agent code uses that -binding to select a sequence and harness; it does not read the program registry -or parse feature flags. +overrides into a resolved binding before it calls `runAgent`. The CLI overrides +arrive as `ProgramInput.overrides`. `harnessRunsTasks` tells it which harnesses +the orchestrator can drive. Agent code uses that binding to select a sequence +and harness. It doesn't read the program registry or parse feature flags. **Sequences** (`sequence/`) are LLM query shapes. Once the binding has picked one, that sequence takes over the run and owns _how the LLM's work is shaped_. @@ -83,7 +85,7 @@ gateway. ## How they connect -- Programs supply inference auth; prepare resolves it and builds triage for the +- Programs supply inference auth. Prepare resolves it and builds triage for the resolved harness. - The program layer resolves the binding with the switchboard helpers. Agent code dispatches the selected sequence and harness. @@ -137,30 +139,33 @@ block-beta Calls descend on the left, results return through the middle, and a host-owned abort signal descends on the right. A standalone caller invokes `runAgent` -without `runProgram`. A program may return a pre-run failure without starting -the agent, and a caught preparation error produces `RunResult` without a -`SequenceResult`. The orchestrator stops scheduling on the first fatal task -result, cancels active siblings and pending asks, and waits for them to settle -before returning that failure. A host signal can also cancel active harness -work. +without `runProgram`, and so does agentic detection. A program may return a +pre-run failure without starting the agent, and a caught preparation error +produces `RunResult` without a `SequenceResult`. The orchestrator stops +scheduling on the first fatal task result, cancels active siblings and pending +asks, and waits for them to settle before returning that failure. A host signal +can also cancel active harness work. ## Flow -1. The caller runs its gates, authenticates, fetches PostHog flags and resolves - a `ProgramBinding { sequence, harness, model }`; analytics tags the run. +1. The host runs `preflight`. `runProgram` then resolves credentials, awaits the + host's gates, loads PostHog flags, refreshes a token near expiry and resolves + a `ProgramBinding { sequence, harness, model }` from `input.overrides` and + the flags. It tags the run and captures the switchboard decision. 2. `runAgent(config, input, options)` resolves the supplied inference auth and prepares triage. -3. Sequence takes over — shapes the LLM's work into one conversation (linear) or - many (orchestrator), reporting through `onProgress`. +3. The sequence takes over. It shapes the LLM's work into one conversation + (linear) or many (orchestrator), and reports through `onProgress`. 4. Harness drives each conversation through its SDK, using the bound model, on the PostHog LLM gateway. 5. The scan report flushes once, on a best-effort basis: as the run ends, or earlier when a process drain runs the cleanups, unless `RunConfig.scanReport` defers it to the host run. Its line arrives as `log` progress. `runAgent` resolves a `RunResult` with an outcome and progress snapshot. A non-success - result carries a code and message; a caught error remains attached. -6. The caller applies it. The legacy runner sends a decided failure to - `wizardAbort` with the terminal status its outcome names; for a crash it - rethrows the attached `Error` when present; after a non-composed success it - sends the terminal success analytics. Other hosts can log, present, or - rethrow the failure as they need. The agent sends no terminal analytics. + result carries a code and message, and a caught error stays attached. +6. `runProgram` records the result in its outcome, and the host applies it. The + legacy runner sends a decided failure to `wizardAbort` with the terminal + status its outcome names. For a crash, it rethrows the attached `Error` when + present. After a non-composed success, it sends the terminal success + analytics. Other hosts can log, present, or rethrow the failure as they need. + Neither `runAgent` nor `runProgram` sends terminal analytics. From 49d53ea5cafe5938ad623ad9f6a2380b210c4f32 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Thu, 24 Sep 2026 16:24:58 -0400 Subject: [PATCH 11/29] docs: describe the caller-built program run and the host-owned gates Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- docs/developer-interfaces.md | 145 ++++++++++++++++++++--------------- src/agent/README.md | 58 +++++++++----- src/agent/runner/README.md | 20 +++-- src/agent/runner/index.ts | 5 +- src/programs/README.md | 6 +- 5 files changed, 138 insertions(+), 96 deletions(-) diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index b4c51d9c3..105ce3490 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -5,11 +5,11 @@ mode for running without a terminal UI. The TypeScript aliases below are internal to this repository. `@posthog/wizard` publishes a CLI, not these functions as a stable package API. -| Surface | Use it for | Detailed contract | -| --------------------------------------- | ----------------------------------------------- | ------------------------------------------------------------------------------------------ | -| `runAgent(config, input, options)` | One already-configured AI run | [Agent reference](../src/agent/README.md) | -| `runProgram(programId, input, options)` | A registered program with invocation-owned data | [Programs reference](../src/programs/README.md) | -| Development `--ci` | A process-owned, non-interactive CLI run | [Local CI credentials and recipe](local-dev.md#credentials-for-local-ci-and-headless-runs) | +| Surface | Use it for | Detailed contract | +| --------------------------------------- | -------------------------------------------------------- | ------------------------------------------------------------------------------------------ | +| `runAgent(config, input, options)` | One already-configured AI run | [Agent reference](../src/agent/README.md) | +| `runProgram(programId, input, options)` | One program run the caller builds from a `ProgramConfig` | [Programs reference](../src/programs/README.md) | +| Development `--ci` | A process-owned, non-interactive CLI run | [Local CI credentials and recipe](local-dev.md#credentials-for-local-ci-and-headless-runs) | ## Standalone agent @@ -74,60 +74,83 @@ an already-issued fixed token. ## Callable program -`runProgram` takes a registered ID, a `ProgramInput` with at least `installDir`, -and optional `ProgramOptions`. Import it from `@programs` and types from -`@programs/types`. It returns a `ProgramRunOutcome`: outcome and failure, the -settled agent runs, observer diagnostics, the data a program with no agent -returned, artifacts, and invocation data (including a captured event plan). The -invocation data contains credentials, so don't log it. Agent failures retain an -attached `Error` when one exists. +`runProgram` takes a program ID, a `ProgramInput` and optional `ProgramOptions`. +Import it from `@programs` and types from `@programs/types`. It doesn't look the +ID up in a registry. The ID picks the binding policy, the commandments, the +stage overrides and the analytics attribution. The caller builds the rest from +the program's `ProgramConfig`: + +- **`input.run`.** The required `AgentRunDefinition`. A static + `ProgramConfig.run` passes as is. A dynamic one reads a session and a + `ProgramRunHost`, so the caller resolves it first. +- **`input.program`.** The `ProgramSettings`: `requiresAi`, `agentFlow`, the + tool allow and deny lists, `excludedTaskTypes`, the audit ledger and its seed + checks, the event plan file and the post-auth gates. The host supplies credentials in one of two ways: - **Resolved.** `input.credentials` carries `{ posthog, inferenceAuth?, project, apiUser }`. - **A provider.** `options.credentials.resolve(programId, { signal })`. - `runProgram` calls it once per invocation, and composed child runs reuse the - result. + `runProgram` calls it once, only when `input.credentials` is absent. Either way, `runProgram` identifies the user for analytics, stamps the organization's AI SDK evidence, and refreshes an OAuth token that is close to -expiry before each agent run. Launch choices go in `input.overrides` as +expiry before the agent starts. Launch choices go in `input.overrides` as `{ harness?, sequence?, model? }`. `runProgram` resolves the binding from them -and the flag snapshot, and captures the switchboard decision once for each agent -run. It copies the input when it receives it, so a later host write can't reach -the run. - -Awaited host capabilities receive the invocation's signal: -`credentials.resolve`, `awaitAiApproval({ programId, signal })`, -`workflow.step(request, { signal })` and `noAgentWorkflow(request)`. The -workflow connector answers the post-auth, child-run and confirm requests that -gated and composed programs make. `noAgentWorkflow` runs the programs with no -agent, such as `posthog-doctor`, `mcp-add` and `slack`. A rejection from any of -them, or from `featureFlags` or a run definition that throws, resolves as -`failed`, or as `aborted` once the signal has aborted. `featureFlags` and the -integration effects don't receive the signal. The promise rejects only on an -invocation error, such as input that can't be copied. Read the outcome, and -still catch a rejection. +and the flag snapshot, and captures the switchboard decision. It copies the +input when it receives it, so a later host write can't reach the run. + +The other options are host capabilities: + +- **`awaitAiApproval({ programId, signal })`.** Answers the AI-processing + approval. Without it, a run that needs approval fails. +- **`awaitPostAuthGates({ programId, gates, signal })`.** Settles the post-auth + gates, such as a project picker. +- **`featureFlags()`.** Evaluates flags when the input has none. It gets no + signal. +- **`interaction`, `onProgress`, `deferSkillCommit` and `signal`.** Answer the + agent's questions, observe the run, leave new skills for the host to commit, + and cancel the invocation. + +A rejection from `credentials`, `awaitAiApproval`, `awaitPostAuthGates`, +`featureFlags` or the token refresh resolves as `failed`, or as `aborted` once +the signal has aborted. The promise rejects only when the invocation itself +breaks, such as input that can't be copied. Read the outcome, and still catch a +rejection. + +`runProgram` returns a `ProgramRunOutcome`: the outcome and failure, +`settledRuns` with the agent run's result, `diagnostics` for observer failures +and late events, `artifacts.reportFile`, and the invocation `data`, including a +captured event plan. The data contains credentials, so don't log it. Agent +failures retain an attached `Error` when one exists. `onProgress` receives two kinds of `ProgramProgress`. A run event is -`{ kind: 'run', runId, stepId?, event }`, where `event` is the agent's progress. -A program-data event is `{ kind: 'program', data }`, a copy of the invocation -data after each write. Narrow on `kind` first: +`{ kind: 'run', runId, event }`, where `event` is the agent's progress. A +program-data event is `{ kind: 'program', data }`, a copy of the invocation data +after each write. Narrow on `kind` first: ```ts -import { runProgram } from '@programs'; +import { getProgramConfig, runProgram } from '@programs'; import type { ProgramOptions } from '@programs/types'; -export async function runAudit( +export async function runMetrics( installDir: string, credentials: NonNullable, awaitAiApproval: NonNullable, signal?: AbortSignal, ) { + const config = getProgramConfig('metrics'); + if (!config.run || typeof config.run === 'function') { + throw new Error('metrics has a static run definition'); + } const result = await runProgram( - 'audit', - { installDir }, + config.id, + { + installDir, + run: config.run, + program: { agentFlow: config.agentFlow, requiresAi: config.requiresAi }, + }, { credentials, awaitAiApproval, @@ -142,34 +165,31 @@ export async function runAudit( ); if (result.outcome !== 'success') { if (result.failure?.error) throw result.failure.error; - throw new Error(result.failure?.message ?? `Audit ${result.outcome}`); + throw new Error(result.failure?.message ?? `Metrics ${result.outcome}`); } return result.artifacts.reportFile; } ``` -The caller implements the credential and approval callbacks. Some programs -require additional prepared inputs or host effects. The -[program reference](../src/programs/README.md#inputs) describes the available -fields and capabilities. There is no live store or step-control handle. -`runProgram` never sends the terminal `setup wizard finished` event. A -long-lived host decides when to send it, from the outcome. +The [program reference](../src/programs/README.md#inputs) lists every input, +setting and option. There is no live store or step-control handle. `runProgram` +never sends the terminal `setup wizard finished` event. A long-lived host +decides when to send it, from the outcome. A runnable reference host is the workbench harness's `pnpm wizard-program`. It -runs one program against the app in `APP_DIR`, with resolved credentials and no -TUI. +builds the run and settings from the `ProgramConfig`, the way the session +adapter does, and runs one program against the app in `APP_DIR`, with resolved +credentials and no TUI. -### Preflight +### What stays with the host -Hosts call `preflight(programId, host)` from `@programs` before `runProgram`. -`runProgram` doesn't call it. It runs the readiness check, then the Claude -settings check, and returns `{ kind: 'proceed', restoreSettings }` or -`{ kind: 'abort', failure }`. The host supplies the presentation (`showOutage`, -`setReadinessWarnings` and `showSettingsOverride`) and its policy -(`interactive`, `signup`, and any readiness it already computed). An outage -aborts only an interactive host. An unfixable settings conflict aborts only a -non-interactive host. Call `restoreSettings()` when the run ends. See the -[program reference](../src/programs/README.md#preflight) for the details. +`runProgram` runs one agent. It doesn't check service readiness or Claude +settings, walk composed steps, or run a program with no agent. The session +adapter, `src/lib/runners/run-program-agent.ts`, runs the readiness and settings +gates before it calls `runProgram`. The TUI walks each composed step as its own +call with `composed: true`, and runs the steps of programs with no agent, such +as `posthog-doctor`, `mcp-add` and `slack`. See the +[program reference](../src/programs/README.md#current-limits). ## Development CI and experimental headless runner @@ -188,13 +208,14 @@ pnpm try --ci --api-key "$POSTHOG_PERSONAL_API_KEY" \ The runner logs progress and writes a local task-stream JSONL dump. Callers observe the process exit and its logs, rather than a returned result. The gateway token file is read into a fixed provider for CI. Pre-run detection and -composed child runs use that same provider. This path doesn't mint or refresh -the token. Published builds reject `--ci`. The internal +the program's agent run use that same provider. This path doesn't mint or +refresh the token. Published builds reject `--ci`. The internal `runWizardCI(config, options): void` entry point uses the session adapter, -`src/lib/runners/run-program-agent.ts`. The adapter runs `preflight`, then calls -`runProgram` for each program's main agent run. Agentic detection runs before -that call, through its own `runAgent` call. MCP suggested prompts use a separate -SDK path with their own progress and cancellation. +`src/lib/runners/run-program-agent.ts`. The adapter builds the run from the +`ProgramConfig`, runs the readiness and settings gates, then calls `runProgram` +once. Agentic detection runs before that call, through its own `runAgent` call. +MCP suggested prompts use a separate SDK path with their own progress and +cancellation. An experimental published-build headless path exists internally as `runWizardHeadless(config, options): void`. It shares the process-owned runner, diff --git a/src/agent/README.md b/src/agent/README.md index 2f30f472f..a4f387d90 100644 --- a/src/agent/README.md +++ b/src/agent/README.md @@ -12,14 +12,24 @@ Import runtime values from `@agent` and types from `@agent/types`. Nothing outside `src/agent` imports deeper. Lint and the architecture test reject it. ```ts -import { runAgent, RunOutcome } from '@agent'; -import type { RunConfig, RunInput, RunResult, AgentProgress, AgentInteraction } from '@agent/types'; - -runAgent(config: RunConfig, input: RunInput, options?: { - onProgress?: (event: AgentProgress) => unknown; - interaction?: AgentInteraction; - signal?: AbortSignal; -}): Promise +import type { + AgentInteraction, + AgentProgress, + RunConfig, + RunInput, + RunResult, +} from '@agent/types'; + +// The shape of `runAgent`, exported from `@agent`. +declare function runAgent( + config: RunConfig, + input: RunInput, + options?: { + onProgress?: (event: AgentProgress) => unknown; + interaction?: AgentInteraction; + signal?: AbortSignal; + }, +): Promise; ``` - `RunConfig`: the opaque program id, its `AgentRunDefinition` (prompt, skill, @@ -104,19 +114,25 @@ runAgent(config: RunConfig, input: RunInput, options?: { Minimal invocation: ```ts -const result = await runAgent(config, input, { - onProgress: (event) => { - if (event.kind === 'log') console.log(event.message); - }, - interaction: { - ask: async (question, { signal }) => answersFor(question, signal), - }, -}); -if (result.outcome !== RunOutcome.Success) { - console.error( - result.failure.error ?? result.failure.message ?? `Agent ${result.outcome}`, - ); - process.exitCode = result.failure.exitCode ?? 1; +import { runAgent, RunOutcome } from '@agent'; +import type { AgentInteraction, RunConfig, RunInput } from '@agent/types'; + +export async function runOnce( + config: RunConfig, + input: RunInput, + ask: NonNullable, +) { + const result = await runAgent(config, input, { + onProgress: (event) => { + if (event.kind === 'log') console.log(event.message); + }, + interaction: { ask }, + }); + if (result.outcome !== RunOutcome.Success) { + console.error(result.failure.error ?? result.failure.message); + process.exitCode = result.failure.exitCode ?? 1; + } + return result; } ``` diff --git a/src/agent/runner/README.md b/src/agent/runner/README.md index f170c058e..6dba64608 100644 --- a/src/agent/runner/README.md +++ b/src/agent/runner/README.md @@ -47,10 +47,12 @@ and returns decided outcomes and caught run-body crashes as results. It never renders, reads a session or exits. Its caller does the host work. For a program run, `runProgram` in `src/programs/run-program.ts` calls the host's credentials provider once, identifies the user, stamps the AI SDK evidence, awaits the -host's approval and workflow connector, loads flags, refreshes an OAuth token -near expiry, and resolves the binding. The legacy -`src/lib/runners/run-program-agent.ts` supplies those capabilities from the -session and maps progress back onto `getUI()`. +host's approval and post-auth gates, loads flags, refreshes an OAuth token near +expiry, and resolves the binding. Its caller builds the run definition and the +program settings from the `ProgramConfig`. The legacy +`src/lib/runners/run-program-agent.ts` does that from the session, supplies the +capabilities, and maps progress back onto `getUI()`. Agentic detection builds +its own config and input, and calls `runAgent` directly. **Prepare** (`shared/bootstrap.ts`) is the on-ramp inside the agent: logging targets, caller-supplied inference auth and the scan-triage classifier. Whether @@ -148,10 +150,12 @@ can also cancel active harness work. ## Flow -1. The host runs `preflight`. `runProgram` then resolves credentials, awaits the - host's gates, loads PostHog flags, refreshes a token near expiry and resolves - a `ProgramBinding { sequence, harness, model }` from `input.overrides` and - the flags. It tags the run and captures the switchboard decision. +1. The host builds the run and settings from the `ProgramConfig`, and runs any + readiness or settings gates it needs. `runProgram` then resolves credentials, + awaits the host's gates, loads PostHog flags, refreshes a token near expiry, + starts the file watchers and resolves a `ResolvedBinding` (sequence, harness + and model) from `input.overrides` and the flags. It tags the run and captures + the switchboard decision. 2. `runAgent(config, input, options)` resolves the supplied inference auth and prepares triage. 3. The sequence takes over. It shapes the LLM's work into one conversation diff --git a/src/agent/runner/index.ts b/src/agent/runner/index.ts index d1310c853..df43d548c 100644 --- a/src/agent/runner/index.ts +++ b/src/agent/runner/index.ts @@ -16,8 +16,9 @@ * sends the process's terminal analytics. * Coded errors return Failed, uncoded throws return Crashed, and both retain * the caught error. Scan-report flushing is best effort after the result is - * decided. Every host reaches this call through programs' `runProgram`, which - * builds the config and input; a standalone caller builds them itself. + * decided. A program run reaches this call through programs' `runProgram`, + * which builds the config and input. Agentic detection and a standalone caller + * build them and call it directly. */ import { Sequence } from '@shared/constants'; diff --git a/src/programs/README.md b/src/programs/README.md index 8ef7f8d94..2b1db5fb1 100644 --- a/src/programs/README.md +++ b/src/programs/README.md @@ -12,18 +12,18 @@ The npm package doesn't export it as a public library API. Import runtime values from `@programs` and types from `@programs/types`: ```ts -import { runProgram } from '@programs'; import type { ProgramInput, ProgramOptions, ProgramRunOutcome, } from '@programs/types'; -runProgram( +// The shape of `runProgram`, exported from `@programs`. +declare function runProgram( programId: string, input: ProgramInput, options?: ProgramOptions, -): Promise +): Promise; ``` `programId` picks the binding policy, the commandments, the stage overrides and From 0d2c2c91a54b0a8f7b41634cef42ad17b466df57 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Thu, 24 Sep 2026 17:56:43 -0400 Subject: [PATCH 12/29] docs: list the restored jest e2e suite next to the live TUI run Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- AGENTS.md | 1 + README.md | 10 +++++++++- 2 files changed, 10 insertions(+), 1 deletion(-) diff --git a/AGENTS.md b/AGENTS.md index ca391ebec..bff00c6db 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -163,6 +163,7 @@ pnpm try --install-dir= # Run the wizard locally against a test proje pnpm build # Compile TypeScript pnpm test # Unit tests (builds first) pnpm test:watch # Unit tests in watch mode +pnpm test:e2e # Jest E2E suite on recorded fixtures (builds first) pnpm test:e2e:tui # Live, credentialed: full TUI on a workbench app copy pnpm lint # Prettier + ESLint checks pnpm fix # Auto-fix lint issues diff --git a/README.md b/README.md index af49e3c90..75239f93c 100644 --- a/README.md +++ b/README.md @@ -565,7 +565,15 @@ To run unit tests, run: bin/test ``` -End-to-end runs are live and credentialed. Point `APP_DIR` at an app copy from +To run the jest E2E suite, which replays recorded LLM calls, run: + +```bash +bin/test-e2e +``` + +See [`e2e-tests/README.md`](e2e-tests/README.md) to add or re-record tests. + +Live end-to-end runs are credentialed. Point `APP_DIR` at an app copy from [wizard-workbench](https://github.com/PostHog/wizard-workbench), which owns the fixture apps and the assertions: From bf66ab9d7de6b9a0718d49e8d8a5abb0da74ca08 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Thu, 24 Sep 2026 19:07:08 -0400 Subject: [PATCH 13/29] chore(programs): take the host-capabilities tree for every non-doc file Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- .github/CODEOWNERS | 2 + .../__tests__/e2e-flow-snapshot.test.ts | 12 +- e2e-harness/__tests__/e2e-result.test.ts | 29 - .../__tests__/tui-snapshot-signature.test.ts | 43 -- e2e-harness/e2e-result.ts | 30 +- e2e-harness/tui-snapshot-signature.ts | 30 - package.json | 1 - scripts/a3-fault-probe.no-jest.ts | 15 +- scripts/check-screens.tsx | 2 +- scripts/tui-host.no-jest.ts | 240 +++---- .../architecture/import-boundaries.test.ts | 95 +-- .../architecture/known-violations.json | 90 +++ src/__tests__/cli.test.ts | 2 +- src/__tests__/programs-cli.test.ts | 2 +- src/__tests__/provision-cli.test.ts | 2 +- src/agent/__tests__/agent-interface.test.ts | 2 - .../__tests__/agent-runner-ask.test.ts} | 16 +- src/agent/__tests__/commandments.test.ts | 10 +- src/agent/__tests__/entry-closure.test.ts | 19 - src/agent/__tests__/entry-streaming.test.ts | 28 +- .../__tests__/gateway-session.test.ts | 62 +- .../__tests__/harness-capabilities.test.ts | 16 - src/agent/__tests__/public-entry.test.ts | 18 - .../__tests__/run-agent-standalone.test.ts | 217 ++----- .../__tests__/run-tags.test.ts | 2 +- .../__tests__/variant-gating.test.ts | 2 +- src/agent/agent-interface.ts | 52 +- src/agent/agent-runner.ts | 3 + src/agent/default-binding.ts | 10 - src/{programs => agent}/gateway-session.ts | 152 ++++- src/agent/index.ts | 60 +- src/agent/mcp-prompt-streaming.ts | 12 +- src/agent/progress.ts | 2 +- .../runner}/__tests__/switchboard.test.ts | 12 +- .../__tests__/pending-question.test.ts | 18 +- src/agent/runner/harness/anthropic/index.ts | 4 - .../harness/pi/__tests__/gateway.test.ts | 2 +- src/agent/runner/harness/pi/gateway.ts | 2 +- src/agent/runner/harness/pi/index.ts | 15 +- src/agent/runner/harness/pi/task.ts | 11 +- src/agent/runner/index.ts | 125 ++-- src/agent/runner/sequence/linear.ts | 21 +- .../orchestrator/orchestrator-runner.ts | 58 +- .../runner/sequence/orchestrator/queue.ts | 4 +- src/agent/runner/shared/bootstrap.ts | 47 +- src/agent/runner/shared/errors.ts | 6 - src/agent/runner/shared/types.ts | 34 +- src/agent/runner/switchboard/commandments.ts | 29 +- .../flags}/__tests__/binding-cases.ts | 9 +- .../flags}/__tests__/flags.test.ts | 34 +- .../runner/switchboard/flags}/index.ts | 0 .../runner/switchboard/flags}/orchestrator.ts | 0 .../runner/switchboard/flags}/schemes.ts | 2 +- .../runner/switchboard/flags}/self-driving.ts | 0 src/agent/runner/switchboard/harness.ts | 87 ++- src/agent/runner/switchboard/index.ts | 154 ++++- .../runner/switchboard/resolve-harness.ts | 90 --- src/agent/runner/switchboard/sequence.ts | 97 ++- src/agent/tools/tool-names.ts | 24 - src/agent/tools/tools.ts | 23 +- src/agent/types.ts | 9 +- src/agent/wizard-ask-bridge.ts | 8 + src/commands/ai-observability.ts | 2 - src/commands/audit.ts | 2 - .../factories/family-command-factory.ts | 9 +- .../factories/native-command-factory.ts | 5 +- src/commands/factories/shared.ts | 8 +- src/{shared => lib}/file-watcher.ts | 0 .../__tests__/ci-inference-auth.test.ts | 43 -- .../__tests__/composed-program-step.test.ts | 40 -- .../runners/__tests__/mint-recovery.test.ts | 68 +- src/lib/runners/ci-inference-auth.ts | 25 - src/lib/runners/run-non-interactive.ts | 76 +-- src/lib/runners/run-wizard.ts | 68 +- src/lib/wizard-session.ts | 67 +- src/programs/__tests__/authenticate.test.ts | 56 -- src/programs/__tests__/binding-owner.test.ts | 88 --- .../__tests__/binding-telemetry.test.ts | 36 -- src/programs/__tests__/credentials.test.ts | 33 - src/programs/__tests__/error-tracking.test.ts | 47 +- src/programs/__tests__/flow-traces.test.ts | 2 +- .../__tests__/posthog-cli-preinstall.test.ts | 20 +- .../__tests__/program-file-watchers.test.ts | 143 ----- src/programs/__tests__/program-store.test.ts | 16 +- .../refresh-access-token-if-needed.test.ts | 134 ++++ .../__tests__/run-agent-legacy.test.ts | 590 +++--------------- src/programs/__tests__/run-program.test.ts | 395 +++--------- .../__tests__/self-driving-deck.test.ts | 13 - .../__tests__/self-driving-detect.test.ts | 11 +- src/programs/__tests__/token-refresh.test.ts | 128 ---- .../__tests__/warehouse-ask-timeout.test.ts | 6 +- .../__tests__/warehouse-suggestion.test.ts | 4 +- src/programs/agent-skill/index.ts | 47 +- src/programs/agent-skill/run-definition.ts | 43 -- src/programs/agent-skill/steps.ts | 2 +- src/programs/ai-observability/index.ts | 42 +- src/programs/ai-observability/run.ts | 39 -- src/programs/ai-opt-in-gate.ts | 15 +- src/programs/audit/index.ts | 52 +- .../{watch-ledger.ts => ledger-watcher.ts} | 20 +- src/programs/audit/types.ts | 5 +- src/programs/authenticate.ts | 93 ++- src/programs/binding-telemetry.ts | 50 -- src/programs/binding.ts | 160 ----- src/programs/commandments.ts | 18 - src/programs/credentials.ts | 16 - .../__tests__/agentic-progress.test.ts | 156 +---- .../detection/__tests__/agentic-retry.test.ts | 40 +- .../__tests__/framework-labels.test.ts | 168 ----- .../detection/__tests__/project-scope.test.ts | 25 +- src/programs/detection/agentic.ts | 46 +- src/programs/detection/context.ts | 10 - src/programs/detection/features.ts | 2 +- src/programs/detection/project-scope.ts | 15 +- src/{commands => programs}/dispatch-family.ts | 17 +- .../detect-agentic.ts | 7 +- .../detect.ts | 3 +- .../index.ts | 7 +- .../steps.ts | 7 +- src/programs/error-tracking/detect-agentic.ts | 18 +- src/programs/error-tracking/index.ts | 28 +- src/programs/events-audit/index.ts | 24 +- src/programs/events-audit/steps.ts | 9 +- src/programs/framework-config.ts | 3 - .../frameworks/astro/astro-wizard-agent.ts | 8 +- .../frameworks/django/django-wizard-agent.ts | 12 - src/programs/frameworks/django/utils.ts | 5 + .../fastapi/fastapi-wizard-agent.ts | 17 +- src/programs/frameworks/fastapi/utils.ts | 4 + .../frameworks/flask/flask-wizard-agent.ts | 6 - src/programs/frameworks/flask/utils.ts | 6 + .../laravel/laravel-wizard-agent.ts | 6 - src/programs/frameworks/laravel/utils.ts | 4 + .../frameworks/nextjs/nextjs-wizard-agent.ts | 12 +- .../frameworks/rails/rails-wizard-agent.ts | 6 - src/programs/frameworks/rails/utils.ts | 3 + .../react-native/react-native-wizard-agent.ts | 4 - src/programs/frameworks/react-native/utils.ts | 7 + .../react-router/react-router-wizard-agent.ts | 8 +- .../tanstack-router-wizard-agent.ts | 8 +- src/programs/host-capabilities.ts | 13 +- src/programs/index.ts | 33 +- src/programs/mcp-analytics/index.ts | 62 +- src/programs/mcp-analytics/run.ts | 62 -- src/programs/mcp/index.ts | 2 +- src/programs/metrics/index.ts | 47 +- src/programs/metrics/run.ts | 44 -- src/programs/migration/index.ts | 46 +- src/programs/migration/run.ts | 40 -- src/programs/migration/steps.ts | 2 +- .../__tests__/detect.test.ts | 25 +- .../__tests__/index.test.ts | 72 +-- .../posthog-integration/ai-sdk-stamp.ts | 37 -- src/programs/posthog-integration/detect.ts | 62 +- src/programs/posthog-integration/index.ts | 87 ++- src/programs/posthog-integration/steps.ts | 9 +- src/programs/program-file-watchers.ts | 62 -- src/programs/program-registry.ts | 2 + src/programs/program-run.ts | 25 +- src/programs/program-step.ts | 31 +- src/programs/program-store.ts | 8 - src/programs/replay-vision/index.ts | 65 +- src/programs/replay-vision/run.ts | 47 -- src/programs/revenue-analytics/abort-cases.ts | 25 - src/programs/revenue-analytics/detect.ts | 28 +- src/programs/revenue-analytics/index.ts | 16 +- src/programs/revenue-analytics/run.ts | 14 - src/programs/revenue-analytics/steps.ts | 2 +- .../run-agent-legacy.ts} | 204 +++--- src/programs/run-program.ts | 285 ++++----- src/programs/self-driving/detect-agentic.ts | 13 +- src/programs/self-driving/detect.ts | 9 +- src/programs/self-driving/index.ts | 15 +- src/programs/self-driving/steps.ts | 11 +- src/programs/shared/health-check-step.ts | 10 +- src/programs/shared/posthog-cli-preinstall.ts | 2 +- src/programs/snapshot-program-input.ts | 21 - .../__tests__/event-plan-watcher.test.ts | 163 +++++ .../__tests__/task-stream-push.test.ts | 88 ++- .../task-stream/destinations/posthog.ts | 2 +- .../event-plan-watcher.ts} | 17 +- src/programs/task-stream/task-stream-push.ts | 51 +- src/programs/task-stream/types.ts | 2 +- src/programs/token-refresh.ts | 63 -- src/programs/types.ts | 2 - src/programs/warehouse-source/detect.ts | 9 +- src/programs/warehouse-source/index.ts | 11 +- src/programs/warehouse-source/steps.ts | 2 +- .../web-analytics-doctor/abort-cases.ts | 34 - src/programs/web-analytics-doctor/detect.ts | 37 +- src/programs/web-analytics-doctor/index.ts | 27 +- src/programs/web-analytics-doctor/run.ts | 27 - .../__tests__/claude-settings-backup.test.ts | 13 - .../__tests__/skill-run-cleanup.test.ts | 38 -- src/shared/ask-policy.ts | 31 - src/shared/ci-gateway-auth.ts | 27 - src/shared/claude-settings.ts | 2 +- .../errors/__tests__/run-failure.test.ts | 7 +- src/shared/gateway-auth.ts | 99 --- src/shared/posthog-cli-install.ts | 46 -- src/shared/run-state.ts | 27 - src/shared/run-tags.ts | 26 - src/shared/scan-consent.ts | 33 - src/shared/skill-run-cleanup.ts | 62 -- .../utils/__tests__/environment.test.ts | 4 +- .../utils/__tests__/oauth-refresh.test.ts | 2 +- src/shared/utils/__tests__/oauth.test.ts | 2 +- src/shared/utils/analytics.ts | 4 +- src/shared/utils/cleanup-registry.ts | 22 - src/shared/utils/environment.ts | 2 +- src/shared/utils/oauth-token.ts | 84 --- src/shared/utils/oauth.ts | 81 ++- src/shared/utils/package-manager.ts | 51 ++ src/shared/utils/provisioning.ts | 2 +- src/shared/utils/setup-utils.ts | 52 ++ src/shared/utils/wizard-abort.ts | 25 +- src/steps/install-cli-steering/index.ts | 60 +- .../upload-environment-variables/index.ts | 3 +- src/ui/__tests__/headless-ui.test.ts | 10 - src/ui/headless-ui.ts | 17 +- src/ui/index.ts | 1 - .../__tests__/keyboard-equivalence.test.tsx | 4 - src/ui/tui/__tests__/store-invariants.test.ts | 5 - src/ui/tui/decks/__tests__/registry.test.ts | 55 -- src/ui/tui/decks/registry.ts | 48 -- src/ui/tui/decks/revenue-analytics/index.tsx | 7 + src/ui/tui/decks/self-driving/index.tsx | 2 +- .../tui/decks/self-driving/pricing.ts} | 0 src/ui/tui/decks/self-driving/tips.ts | 2 +- src/ui/tui/decks/warehouse-source/index.tsx | 7 + src/ui/tui/hooks/file-watcher.ts | 12 +- src/ui/tui/playground/demos/LearnDeckDemo.tsx | 20 +- src/ui/tui/playground/demos/RunScreenDemo.tsx | 8 +- .../tui/screens/ErrorTrackingDetectScreen.tsx | 2 - src/ui/tui/screens/RunScreen.tsx | 21 +- .../SelfDrivingIntegrationDetectScreen.tsx | 2 - src/ui/tui/screens/SelfDrivingIntroScreen.tsx | 2 +- src/ui/tui/screens/SourceMapsDetectScreen.tsx | 15 +- src/ui/tui/screens/audit/AuditRunScreen.tsx | 3 +- .../mcp-suggested-prompts-services.test.ts | 70 --- .../mcp-suggested-prompts-services.ts | 9 +- src/ui/tui/store.ts | 5 - src/ui/wizard-ui.ts | 12 +- test/module-graph.ts | 123 ---- test/program-host.ts | 7 - tsconfig.json | 2 - 246 files changed, 3385 insertions(+), 5605 deletions(-) delete mode 100644 e2e-harness/__tests__/tui-snapshot-signature.test.ts delete mode 100644 e2e-harness/tui-snapshot-signature.ts rename src/{shared/__tests__/ask-policy.test.ts => agent/__tests__/agent-runner-ask.test.ts} (76%) delete mode 100644 src/agent/__tests__/entry-closure.test.ts rename src/{programs => agent}/__tests__/gateway-session.test.ts (94%) delete mode 100644 src/agent/__tests__/harness-capabilities.test.ts delete mode 100644 src/agent/__tests__/public-entry.test.ts rename src/{shared => agent}/__tests__/run-tags.test.ts (97%) rename src/{programs => agent}/__tests__/variant-gating.test.ts (89%) delete mode 100644 src/agent/default-binding.ts rename src/{programs => agent}/gateway-session.ts (69%) rename src/{programs => agent/runner}/__tests__/switchboard.test.ts (97%) rename src/{programs/experiments => agent/runner/switchboard/flags}/__tests__/binding-cases.ts (91%) rename src/{programs/experiments => agent/runner/switchboard/flags}/__tests__/flags.test.ts (93%) rename src/{programs/experiments => agent/runner/switchboard/flags}/index.ts (100%) rename src/{programs/experiments => agent/runner/switchboard/flags}/orchestrator.ts (100%) rename src/{programs/experiments => agent/runner/switchboard/flags}/schemes.ts (99%) rename src/{programs/experiments => agent/runner/switchboard/flags}/self-driving.ts (100%) delete mode 100644 src/agent/runner/switchboard/resolve-harness.ts delete mode 100644 src/agent/tools/tool-names.ts rename src/{shared => lib}/file-watcher.ts (100%) delete mode 100644 src/lib/runners/__tests__/ci-inference-auth.test.ts delete mode 100644 src/lib/runners/__tests__/composed-program-step.test.ts delete mode 100644 src/lib/runners/ci-inference-auth.ts delete mode 100644 src/programs/__tests__/authenticate.test.ts delete mode 100644 src/programs/__tests__/binding-owner.test.ts delete mode 100644 src/programs/__tests__/binding-telemetry.test.ts delete mode 100644 src/programs/__tests__/credentials.test.ts delete mode 100644 src/programs/__tests__/program-file-watchers.test.ts create mode 100644 src/programs/__tests__/refresh-access-token-if-needed.test.ts delete mode 100644 src/programs/__tests__/token-refresh.test.ts delete mode 100644 src/programs/agent-skill/run-definition.ts delete mode 100644 src/programs/ai-observability/run.ts rename src/programs/audit/{watch-ledger.ts => ledger-watcher.ts} (52%) delete mode 100644 src/programs/binding-telemetry.ts delete mode 100644 src/programs/binding.ts delete mode 100644 src/programs/commandments.ts delete mode 100644 src/programs/detection/__tests__/framework-labels.test.ts rename src/{commands => programs}/dispatch-family.ts (92%) delete mode 100644 src/programs/mcp-analytics/run.ts delete mode 100644 src/programs/metrics/run.ts delete mode 100644 src/programs/migration/run.ts delete mode 100644 src/programs/posthog-integration/ai-sdk-stamp.ts delete mode 100644 src/programs/program-file-watchers.ts delete mode 100644 src/programs/replay-vision/run.ts delete mode 100644 src/programs/revenue-analytics/abort-cases.ts delete mode 100644 src/programs/revenue-analytics/run.ts rename src/{lib/runners/run-program-agent.ts => programs/run-agent-legacy.ts} (72%) delete mode 100644 src/programs/snapshot-program-input.ts create mode 100644 src/programs/task-stream/__tests__/event-plan-watcher.test.ts rename src/programs/{posthog-integration/watch-event-plan.ts => task-stream/event-plan-watcher.ts} (83%) delete mode 100644 src/programs/token-refresh.ts delete mode 100644 src/programs/web-analytics-doctor/abort-cases.ts delete mode 100644 src/programs/web-analytics-doctor/run.ts delete mode 100644 src/shared/__tests__/skill-run-cleanup.test.ts delete mode 100644 src/shared/ask-policy.ts delete mode 100644 src/shared/ci-gateway-auth.ts delete mode 100644 src/shared/gateway-auth.ts delete mode 100644 src/shared/posthog-cli-install.ts delete mode 100644 src/shared/run-state.ts delete mode 100644 src/shared/run-tags.ts delete mode 100644 src/shared/scan-consent.ts delete mode 100644 src/shared/skill-run-cleanup.ts delete mode 100644 src/shared/utils/cleanup-registry.ts delete mode 100644 src/shared/utils/oauth-token.ts delete mode 100644 src/ui/tui/decks/__tests__/registry.test.ts delete mode 100644 src/ui/tui/decks/registry.ts create mode 100644 src/ui/tui/decks/revenue-analytics/index.tsx rename src/{shared/self-driving-pricing.ts => ui/tui/decks/self-driving/pricing.ts} (100%) create mode 100644 src/ui/tui/decks/warehouse-source/index.tsx delete mode 100644 src/ui/tui/services/__tests__/mcp-suggested-prompts-services.test.ts delete mode 100644 test/module-graph.ts diff --git a/.github/CODEOWNERS b/.github/CODEOWNERS index 9ccfa69f8..7bceb017f 100644 --- a/.github/CODEOWNERS +++ b/.github/CODEOWNERS @@ -21,7 +21,9 @@ # Team-owned program decks /src/ui/tui/decks/error-tracking-upload-source-maps/ @PostHog/team-error-tracking +/src/ui/tui/decks/revenue-analytics/ @PostHog/team-web-analytics /src/ui/tui/decks/self-driving/ @PostHog/team-self-driving +/src/ui/tui/decks/warehouse-source/ @PostHog/team-warehouse-sources # Agent runner (harness, model, and tools) /src/agent/ @PostHog/team-wizard-docs diff --git a/e2e-harness/__tests__/e2e-flow-snapshot.test.ts b/e2e-harness/__tests__/e2e-flow-snapshot.test.ts index 3341cf8f3..ac9d7b85c 100644 --- a/e2e-harness/__tests__/e2e-flow-snapshot.test.ts +++ b/e2e-harness/__tests__/e2e-flow-snapshot.test.ts @@ -21,7 +21,11 @@ import { Integration } from '@shared/constants'; import { HostResolution } from '@shared/host-resolution'; import { FRAMEWORK_REGISTRY } from '@programs/registry'; import { WizardReadiness } from '@shared/health-checks/readiness'; -import { Program, getProgramConfig, type ProgramId } from '@programs'; +import { + Program, + getProgramConfig, + type ProgramId, +} from '@programs'; import { ScreenId } from '@ui/tui/router'; import { SELF_DRIVING_INTEGRATE_PATH_KEY } from '@programs/self-driving/detect'; import { WizardCiDriver } from '../wizard-ci-driver'; @@ -109,8 +113,8 @@ function traceFlow( path: '.', }); } else if (screen === ScreenId.Run) { - // The run screen is shared by composed run steps (a step declaring a - // child program, e.g. self-driving's integrate-run) and the program's own + // The run screen is shared by composed run steps (a step carrying its own + // `run` thunk, e.g. self-driving's integrate-run) and the program's own // run. Complete the active run step the way the runner would: a composed // step via completeRunStep, the main run via runPhase. const steps = getProgramConfig(store.router.activeProgram).steps; @@ -120,7 +124,7 @@ function traceFlow( (!s.show || s.show(store.session)) && (!s.isComplete || !s.isComplete(store.session)), ); - if (runStep?.runProgramId) { + if (runStep?.run) { store.completeRunStep(runStep.id); } else { store.setRunPhase(RunPhase.Completed); diff --git a/e2e-harness/__tests__/e2e-result.test.ts b/e2e-harness/__tests__/e2e-result.test.ts index 96a9eff1b..2757c7530 100644 --- a/e2e-harness/__tests__/e2e-result.test.ts +++ b/e2e-harness/__tests__/e2e-result.test.ts @@ -20,7 +20,6 @@ import { E2eRunRecorder, abortReasonFrom, buildE2eResult, - createE2eResultWriter, detectedSourcesFrom, readReportFile, taskOutcomesFrom, @@ -464,34 +463,6 @@ describe('buildE2eResult', () => { expect(build()).toMatchObject(base); }); - it('replaces an outro result only when the final skills decision is written', () => { - const directory = fs.mkdtempSync(path.join(os.tmpdir(), 'wizard-result-')); - const file = path.join(directory, 'result.json'); - let skillsComplete = false; - const write = createE2eResultWriter(file, () => ({ - ...build(), - skillsComplete, - })); - - try { - write(); - expect(JSON.parse(fs.readFileSync(file, 'utf8'))).toMatchObject({ - skillsComplete: false, - }); - skillsComplete = true; - write(); - expect(JSON.parse(fs.readFileSync(file, 'utf8'))).toMatchObject({ - skillsComplete: false, - }); - write(true); - expect(JSON.parse(fs.readFileSync(file, 'utf8'))).toMatchObject({ - skillsComplete: true, - }); - } finally { - fs.rmSync(directory, { recursive: true, force: true }); - } - }); - it('projects tasks down to label and status', () => { expect(build().tasks).toEqual([ { label: 'Connect your data sources', status: 'completed' }, diff --git a/e2e-harness/__tests__/tui-snapshot-signature.test.ts b/e2e-harness/__tests__/tui-snapshot-signature.test.ts deleted file mode 100644 index 631b4771b..000000000 --- a/e2e-harness/__tests__/tui-snapshot-signature.test.ts +++ /dev/null @@ -1,43 +0,0 @@ -/** - * The snapshot signature changes for the state the real TUI shows: a - * status-only progress event, and an audit ledger update that keeps the same - * context key. - */ -import { createUiReducer } from '@ui/agent-progress'; -import { InkUI } from '@ui/tui/ink-ui'; -import { WizardStore } from '@ui/tui/store'; -import { AUDIT_CHECKS_KEY } from '@programs/audit/types'; -import { tuiSnapshotSignature } from '../tui-snapshot-signature'; - -it('captures status-only host progress through the real TUI store', () => { - const store = new WizardStore(); - const onChange = vi.fn(); - store.subscribe(onChange); - const reduce = createUiReducer(new InkUI(store)); - const before = tuiSnapshotSignature(store); - - reduce({ kind: 'status', message: 'Inspecting the project' }); - const first = tuiSnapshotSignature(store); - reduce({ kind: 'status', message: 'Installing the SDK' }); - const second = tuiSnapshotSignature(store); - - expect(store.statusMessages).toEqual([ - 'Inspecting the project', - 'Installing the SDK', - ]); - expect(onChange).toHaveBeenCalledTimes(2); - expect(first).not.toBe(before); - expect(second).not.toBe(first); -}); - -it('captures changed audit checks under the same context key', () => { - const store = new WizardStore(); - store.setFrameworkContext(AUDIT_CHECKS_KEY, [ - { id: 'first', area: 'Events', label: 'First', status: 'pending' }, - ]); - const first = tuiSnapshotSignature(store); - store.setFrameworkContext(AUDIT_CHECKS_KEY, [ - { id: 'first', area: 'Events', label: 'First', status: 'pass' }, - ]); - expect(tuiSnapshotSignature(store)).not.toBe(first); -}); diff --git a/e2e-harness/e2e-result.ts b/e2e-harness/e2e-result.ts index b4046ebe1..8e6b27171 100644 --- a/e2e-harness/e2e-result.ts +++ b/e2e-harness/e2e-result.ts @@ -22,7 +22,6 @@ import fs from 'fs'; import path from 'path'; import { OutroKind, type WizardSession } from '@lib/wizard-session'; -import { RunPhase } from '@shared/run-state'; import { TASK_OUTCOMES_KEY } from '@agent'; import type { TaskOutcome } from '@agent/types'; import { DETECTED_WAREHOUSE_SOURCES_KEY } from '@programs/warehouse-source/detect'; @@ -240,7 +239,7 @@ export function taskOutcomesFrom( /** The keys the result payload carried before the warehouse work. */ export interface E2eResultBase { - runPhase: RunPhase; + runPhase: string; hasPosthogDep: boolean; newDeps: string[]; envFile: string | null; @@ -248,31 +247,6 @@ export interface E2eResultBase { skillsComplete: boolean; } -export type E2eResultPayload = E2eResultBase & { - asks: E2eAskRecord[]; - unansweredAsks: number; - refusedAsks: number; - notices: E2eNoticeRecord[]; - tasks: Array<{ label: string; status: string }>; - taskOutcomes?: Array>; - detectedSources: DetectedSource[]; - reportFile: E2eReportFile | null; - abort: string | null; -}; - -/** Write the first outcome once, then allow the final skills decision to replace it. */ -export function createE2eResultWriter( - file: string | undefined, - getResult: () => E2eResultPayload, -): (final?: boolean) => void { - let written = false; - return (final = false) => { - if (!file || (written && !final)) return; - fs.writeFileSync(file, JSON.stringify(getResult(), null, 2)); - written = true; - }; -} - /** * Build the `E2E_RESULT_JSON` payload. Additive over {@link E2eResultBase} — * the pre-existing keys are copied through byte-identical. @@ -283,7 +257,7 @@ export function buildE2eResult(args: { session: Pick; tasks: Array<{ label: string; status: string }>; reportFile: E2eReportFile | null; -}): E2eResultPayload { +}): Record { const { base, recorder, session, tasks, reportFile } = args; return { ...base, diff --git a/e2e-harness/tui-snapshot-signature.ts b/e2e-harness/tui-snapshot-signature.ts deleted file mode 100644 index 9c679c78c..000000000 --- a/e2e-harness/tui-snapshot-signature.ts +++ /dev/null @@ -1,30 +0,0 @@ -/** - * The fixed-route snapshot signature. `scripts/tui-host.no-jest.ts` takes a - * frame whenever this string changes: the screen, the overlay, the task list, - * the status lines, the run phase or the framework context. `ctx` hashes the - * context values, not only its keys, because audit ledger updates keep the - * same key and still need a frame. - */ -import type { WizardStore } from '@ui/tui/store'; - -function digest(s: string): string { - let h = 0x811c9dc5; - for (let i = 0; i < s.length; i++) { - h ^= s.charCodeAt(i); - h = Math.imul(h, 0x01000193); - } - return (h >>> 0).toString(36); -} - -/** State changes that should produce a new fixed-route TUI frame. */ -export function tuiSnapshotSignature(store: WizardStore): string { - return JSON.stringify({ - screen: store.currentScreen, - overlay: store.router.hasOverlay, - tasks: store.tasks.map((t) => [t.label, t.status, t.done]), - status: store.statusMessages, - phase: store.session.runPhase, - // Include values: audit ledger updates keep the same context key. - ctx: digest(JSON.stringify(store.session.frameworkContext)), - }); -} diff --git a/package.json b/package.json index 407569a16..4fe8bf62a 100644 --- a/package.json +++ b/package.json @@ -143,7 +143,6 @@ "test:coverage": "pnpm build && vitest run --coverage", "test:e2e": "pnpm build && ./e2e-tests/run.sh", "test:e2e-record": "export RECORD_FIXTURES=true && pnpm build && ./e2e-tests/run.sh", - "test:e2e:tui": "tsx scripts/tui-snapshots.no-jest.ts", "try": "tsx bin.ts", "dev": "pnpm build && pnpm link --global && pnpm build:watch", "test:watch": "vitest", diff --git a/scripts/a3-fault-probe.no-jest.ts b/scripts/a3-fault-probe.no-jest.ts index ebc39c450..91b83e5ac 100644 --- a/scripts/a3-fault-probe.no-jest.ts +++ b/scripts/a3-fault-probe.no-jest.ts @@ -28,7 +28,9 @@ globalThis.fetch = (input, init) => { }; const { runAgent } = await import('@agent/runner'); -const { createCiGatewayAuth } = await import('@shared/ci-gateway-auth'); +const { configureGatewayCredentialsForCI } = await import( + '@agent/gateway-session' +); const { DEFAULT_AGENT_MODEL, Harness, Sequence } = await import( '@shared/constants' ); @@ -41,8 +43,7 @@ analytics.captureException = () => {}; analytics.wizardCapture = () => {}; analytics.shutdown = async () => {}; -const chosenHarness = harness === 'pi' ? Harness.pi : Harness.anthropic; -const ciAuth = createCiGatewayAuth( +configureGatewayCredentialsForCI( 'phe_synthetic_fault_probe', 228144, gatewayUrl, @@ -62,9 +63,14 @@ const config: RunConfig = { composed: false, binding: { sequence: Sequence.linear, - harness: chosenHarness, + harness: harness as Harness, model: DEFAULT_AGENT_MODEL, }, + switchboard: { + program: 'fault-probe', + flags: {}, + cliHarness: harness as Harness, + }, skillsBaseUrl: 'http://127.0.0.1:1', wizardFlags: {}, wizardFlagPayloads: {}, @@ -79,7 +85,6 @@ const input: RunInput = { host: HostResolution.fromApiHost('http://127.0.0.1:1', { localMcp: true }), projectId: 228144, }, - inferenceAuth: { resolve: () => Promise.resolve(ciAuth) }, project: null, apiUser: null, flags: { diff --git a/scripts/check-screens.tsx b/scripts/check-screens.tsx index e82d5fe0f..034106c70 100644 --- a/scripts/check-screens.tsx +++ b/scripts/check-screens.tsx @@ -15,7 +15,7 @@ import { ProgressList } from '@ui/tui/primitives/ProgressList'; import { ManagedSettingsScreen } from '@ui/tui/screens/ManagedSettingsScreen'; import { SettingsOverrideScreen } from '@ui/tui/screens/SettingsOverrideScreen'; import { WizardAskScreen } from '@ui/tui/screens/WizardAskScreen'; -import type { SettingsConflict } from '@shared/claude-settings'; +import type { SettingsConflict } from '@agent/agent-interface'; function fakeStore(session: Record): any { return { diff --git a/scripts/tui-host.no-jest.ts b/scripts/tui-host.no-jest.ts index 3325f2278..c34f718eb 100644 --- a/scripts/tui-host.no-jest.ts +++ b/scripts/tui-host.no-jest.ts @@ -17,20 +17,18 @@ import fs from 'fs'; import net from 'net'; import { spawnSync } from 'child_process'; import { startTUI } from '@ui/tui/start-tui'; -import { getUI } from '@ui'; import { VERSION } from '@shared/version'; -import { Program, getProgramConfig, type ProgramId } from '@programs'; +import { + Program, + getProgramConfig, + type ProgramId, +} from '@programs'; import type { Harness, Sequence } from '@shared/constants'; import { buildSession } from '@lib/wizard-session'; import { initLocalDev } from '@shared/local-dev'; -import { loadCiInferenceAuthProvider } from '@lib/runners/ci-inference-auth'; -import type { InferenceAuthProvider } from '@agent/types'; -import { runProgramAgent } from '@lib/runners/run-program-agent'; -import { commitRegisteredRunSkillCleanups } from '@shared/skill-run-cleanup'; -import { - TaskStreamPush, - createFileDestination, -} from '@programs/task-stream/index'; +import { configureGatewayFromCIEnvironment } from '@agent/gateway-session'; +import { runProgramAgent } from '@programs/run-agent-legacy'; +import { TaskStreamPush, createFileDestination } from '@programs/task-stream/index'; import { getAuditChecks } from '@programs/audit/types'; import { authenticate } from '@programs/authenticate'; import { getOrAskForProjectData } from '@utils/setup-utils'; @@ -56,22 +54,22 @@ import { profileFor, resolveE2eProfile } from '@e2e-harness/profiles'; import { E2eRunRecorder, buildE2eResult, - createE2eResultWriter, readReportFile, } from '@e2e-harness/e2e-result'; -import { tuiSnapshotSignature } from '@e2e-harness/tui-snapshot-signature'; + +/** Cheap 32-bit FNV-1a, to fold framework-context values into a signature. */ +function digest(s: string): string { + let h = 0x811c9dc5; + for (let i = 0; i < s.length; i++) { + h ^= s.charCodeAt(i); + h = Math.imul(h, 0x01000193); + } + return (h >>> 0).toString(36); +} const sleep = (ms: number) => new Promise((r) => setTimeout(r, ms)); const mark = (m: string) => logToFile(`[tui-host] ${m}`); -/** A blank variable counts as unset, so the key file is the fallback. */ -function readPersonalApiKey(env: NodeJS.ProcessEnv): string { - const inline = env.POSTHOG_PERSONAL_API_KEY?.trim(); - if (inline) return inline; - const file = env.POSTHOG_KEY_FILE?.trim(); - return file ? fs.readFileSync(file, 'utf8').trim() : ''; -} - /** Tri-state: absent ⇒ `undefined`, so `resolveLocalDev` can apply the umbrella. */ function envFlag(name: string): boolean | undefined { const raw = process.env[name]; @@ -193,7 +191,12 @@ function runAppBuild(root: string): boolean { } async function main() { - const apiKey = readPersonalApiKey(process.env); + const apiKey = ( + process.env.POSTHOG_PERSONAL_API_KEY ?? + (process.env.POSTHOG_KEY_FILE + ? fs.readFileSync(process.env.POSTHOG_KEY_FILE, 'utf8') + : '') + ).trim(); const projectId = process.env.PROJECT_ID!; // Which program to drive — defaults to the integration flow. Set PROGRAM to // an id (e.g. `self-driving`) to host a different one. @@ -225,7 +228,7 @@ async function main() { // Keep the `wizard_ask` bridge wired despite `ci: true`. The driver loop // below is the answerer — without this the agent-in-the-loop layer of a // flow (credential questions, the orchestrator's seeded warehouse task) is - // never exercised. Only this host sets it; see `isAskDisabled`. + // never exercised. Only this host sets it; see `shouldDisableAsk`. e2eAsk: process.env.E2E_ASK === 'true', apiKey, projectId, @@ -234,6 +237,7 @@ async function main() { // skills (:8765) against the production MCP. localDev: process.env.POSTHOG_WIZARD_LOCAL_DEV === 'true', localMcp: envFlag('POSTHOG_WIZARD_LOCAL_MCP'), + localContextMill: envFlag('POSTHOG_WIZARD_LOCAL_CONTEXT_MILL'), localPosthog: envFlag('POSTHOG_WIZARD_LOCAL_POSTHOG'), // Switchboard variation overrides (see e2e.json `variations`), threaded by // the snapshot driver as one run per variation. Empty ⇒ resolved default. @@ -241,17 +245,6 @@ async function main() { sequence: (process.env.SNAP_SEQUENCE || undefined) as Sequence | undefined, model: process.env.SNAP_MODEL || undefined, }); - // Read the one-use token file only when a route first requests inference. - let ciAuth: InferenceAuthProvider | undefined; - store.setInferenceAuth({ - resolve: async () => { - ciAuth ??= loadCiInferenceAuthProvider( - Number(projectId), - store.session.region ?? 'us', - ); - return ciAuth.resolve(); - }, - }); // Dumped, never pushed: an e2e run is synthetic, like `--ci`. const streamLog = createFileDestination(process.env.TASK_STREAM_LOG ?? ''); if (streamLog) { @@ -259,6 +252,9 @@ async function main() { store, programId, destinations: [streamLog], + eventPlanPath: programConfig.eventPlanFile + ? join(store.session.installDir, programConfig.eventPlanFile) + : undefined, auditChecks: programConfig.auditLedgerFile ? () => getAuditChecks(store.session) : undefined, @@ -296,7 +292,15 @@ async function main() { // Pass the pre-run gates and run the program's real agent. The auth and run // screens never advance on their own; this is what moves them. Mirrors // run-wizard's flow, including in-program run phases. + let gatewayConfigured = false; const runProgram = async () => { + if (!gatewayConfigured) { + configureGatewayFromCIEnvironment( + Number(projectId), + store.session.region ?? 'us', + ); + gatewayConfigured = true; + } await store.getGate('intro'); await store.getGate('integration-check'); await store.getGate('health-check'); @@ -306,12 +310,11 @@ async function main() { // or scope their own run to a picked project (error-tracking). // `authenticate` here resolves the phx key, not OAuth, since the session is // built with ci + apiKey. - if (programConfig.steps.some((s) => s.runProgramId || s.targetDir)) { + if (programConfig.steps.some((s) => s.run || s.targetDir)) { const runSessionFor = async ( step: (typeof programConfig.steps)[number], ) => { const live = store.session; - const previousLabel = live.detectedFrameworkLabel; const runSession = step.targetDir ? { ...live, @@ -320,25 +323,15 @@ async function main() { } : live; if (step.onRunPrep) await step.onRunPrep(runSession); - if ( - runSession.detectedFrameworkLabel && - runSession.detectedFrameworkLabel !== previousLabel - ) { - store.setDetectedFramework(runSession.detectedFrameworkLabel); - } return runSession; }; for (const step of programConfig.steps) { if (step.screenId === 'outro') break; if (step.show && !step.show(store.session)) continue; if (step.screenId === 'auth') { - await authenticate(store.session, programConfig.id, getUI()); - } else if (step.runProgramId) { - await runProgramAgent( - getProgramConfig(step.runProgramId), - await runSessionFor(step), - { composed: true }, - ); + await authenticate(store.session, programConfig.id); + } else if (step.run) { + await step.run(await runSessionFor(step)); store.completeRunStep(step.id); } else if (step.screenId === 'run') { await runProgramAgent(programConfig, await runSessionFor(step)); @@ -349,8 +342,6 @@ async function main() { } else { await runProgramAgent(programConfig, store.session); } - // runProgramAgent leaves new skills armed; a finished run keeps them. - commitRegisteredRunSkillCleanups(); }; if (process.env.MODE === 'serve') return serve(); @@ -452,6 +443,7 @@ async function main() { }; const recorder = new E2eRunRecorder(); const screenPath: string[] = []; + let resultWritten = false; // An abort exits from inside the runner, so hook `exit` too — see writeResult. process.on('exit', () => writeResult()); // Snapshot on key moments — a screen change, a task-list update, or a @@ -462,8 +454,18 @@ async function main() { // signature and serialized. let lastSig = ''; let chain: Promise = Promise.resolve(); + const signature = () => + JSON.stringify({ + screen: store.currentScreen, + overlay: store.router.hasOverlay, + tasks: store.tasks.map((t) => [t.label, t.status, t.done]), + phase: store.session.runPhase, + // Values, not just keys: a screen rerendering from an artifact updated + // in place (the audit ledger) keeps its key and would snap once, empty. + ctx: digest(JSON.stringify(store.session.frameworkContext)), + }); const snap = (): Promise => { - const sig = tuiSnapshotSignature(store); + const sig = signature(); if (sig === lastSig) return chain; lastSig = sig; const screen = store.currentScreen; @@ -625,73 +627,79 @@ async function main() { // integration re-writes it after keep-skills (skillsComplete). Registered // on `exit` too: `wizardAbort` renders the error outro and exits, and an // aborted run would otherwise write nothing at all. - const writeResult = createE2eResultWriter( - process.env.E2E_RESULT_JSON, - () => { - const appDir = process.env.APP_DIR!; - // One dependency-name pattern per ecosystem manifest. A run only needs - // the names, so a line-level scan beats per-format parsers. - const MANIFESTS: Array<[string, RegExp]> = [ - ['pubspec.yaml', /^ {2}([A-Za-z_][A-Za-z0-9_]*)\s*:/gm], - ['go.mod', /^\s*([\w.\/-]+)\s+v[\w.-]+/gm], - ['Cargo.toml', /^([A-Za-z0-9_-]+)\s*=/gm], - ['pom.xml', /([^<]+)<\/artifactId>/g], - ['build.gradle', /['"]([\w.-]+:[\w.-]+)[:'"]/g], - ['mix.exs', /\{:([a-z_]+)\s*,/g], - ]; - const deps: string[] = []; - try { - // package.json needs a real parse: a line scan would also match script - // names, and only the dependency blocks carry dependencies. - const pkg = JSON.parse( - fs.readFileSync(`${appDir}/package.json`, 'utf8'), - ); - deps.push( - ...Object.keys({ ...pkg.dependencies, ...pkg.devDependencies }), - ); - } catch { - /* not a JS project */ - } - for (const [file, pattern] of MANIFESTS) { - try { - const text = fs.readFileSync(`${appDir}/${file}`, 'utf8'); - for (const match of text.matchAll(pattern)) deps.push(match[1]); - } catch { - /* app doesn't use this ecosystem */ - } - } - const posthogDeps = [ - ...new Set(deps.filter((d) => d.toLowerCase().includes('posthog'))), - ]; - let envFile: string | null = null; + const writeResult = (): void => { + if (!process.env.E2E_RESULT_JSON || resultWritten) return; + resultWritten = true; + const appDir = process.env.APP_DIR!; + // One dependency-name pattern per ecosystem manifest. A run only needs + // the names, so a line-level scan beats per-format parsers. + const MANIFESTS: Array<[string, RegExp]> = [ + ['pubspec.yaml', /^ {2}([A-Za-z_][A-Za-z0-9_]*)\s*:/gm], + ['go.mod', /^\s*([\w.\/-]+)\s+v[\w.-]+/gm], + ['Cargo.toml', /^([A-Za-z0-9_-]+)\s*=/gm], + ['pom.xml', /([^<]+)<\/artifactId>/g], + ['build.gradle', /['"]([\w.-]+:[\w.-]+)[:'"]/g], + ['mix.exs', /\{:([a-z_]+)\s*,/g], + ]; + const deps: string[] = []; + try { + // package.json needs a real parse: a line scan would also match script + // names, and only the dependency blocks carry dependencies. + const pkg = JSON.parse( + fs.readFileSync(`${appDir}/package.json`, 'utf8'), + ); + deps.push( + ...Object.keys({ ...pkg.dependencies, ...pkg.devDependencies }), + ); + } catch { + /* not a JS project */ + } + for (const [file, pattern] of MANIFESTS) { try { - const hit = fs - .readdirSync(appDir) - .find( - (f) => - (f.startsWith('.env') || f.endsWith('.env')) && - /posthog/i.test(fs.readFileSync(`${appDir}/${f}`, 'utf8')), - ); - envFile = hit ? `${appDir}/${hit}` : null; + const text = fs.readFileSync(`${appDir}/${file}`, 'utf8'); + for (const match of text.matchAll(pattern)) deps.push(match[1]); } catch { - /* none */ + /* app doesn't use this ecosystem */ } - return buildE2eResult({ - base: { - runPhase: store.session.runPhase, - hasPosthogDep: posthogDeps.length > 0, - newDeps: posthogDeps, - envFile, - screenPath, - skillsComplete: store.session.skillsComplete, - }, - recorder, - session: store.session, - tasks: store.tasks, - reportFile: readReportFile(appDir, programConfig.reportFile), - }); - }, - ); + } + const posthogDeps = [ + ...new Set(deps.filter((d) => d.toLowerCase().includes('posthog'))), + ]; + let envFile: string | null = null; + try { + const hit = fs + .readdirSync(appDir) + .find( + (f) => + (f.startsWith('.env') || f.endsWith('.env')) && + /posthog/i.test(fs.readFileSync(`${appDir}/${f}`, 'utf8')), + ); + envFile = hit ? `${appDir}/${hit}` : null; + } catch { + /* none */ + } + fs.writeFileSync( + process.env.E2E_RESULT_JSON, + JSON.stringify( + buildE2eResult({ + base: { + runPhase: store.session.runPhase, + hasPosthogDep: posthogDeps.length > 0, + newDeps: posthogDeps, + envFile, + screenPath, + skillsComplete: store.session.skillsComplete, + }, + recorder, + session: store.session, + tasks: store.tasks, + reportFile: readReportFile(appDir, programConfig.reportFile), + }), + null, + 2, + ), + ); + }; const unsubResult = store.subscribe(() => { if (store.currentScreen === 'outro') writeResult(); }); @@ -709,7 +717,7 @@ async function main() { unsubResult(); await snap(); // the final screen await chain; // flush any pending snapshots - writeResult(true); // integration: replace the early outro result after keep-skills + writeResult(); // final write (integration: after keep-skills) process.exit(0); } } diff --git a/src/__tests__/architecture/import-boundaries.test.ts b/src/__tests__/architecture/import-boundaries.test.ts index 86f64500f..e4b566abd 100644 --- a/src/__tests__/architecture/import-boundaries.test.ts +++ b/src/__tests__/architecture/import-boundaries.test.ts @@ -1,14 +1,6 @@ import * as fs from 'fs'; import * as path from 'path'; import { fileURLToPath } from 'url'; -import { - aliasTarget, - loadAliases, - probe, - REPO_ROOT, - staticImportClosure, - toRepoRelative, -} from '../../../test/module-graph'; export type Surface = | 'env' @@ -20,6 +12,7 @@ export type Surface = | 'cli'; const HERE = path.dirname(fileURLToPath(import.meta.url)); +const REPO_ROOT = path.resolve(HERE, '../../..'); const SURFACE_RULES: ReadonlyArray boolean]> = [ @@ -208,6 +201,14 @@ function stripComments(source: string): string { return out; } +function toRepoRelative(abs: string): string { + return path.relative(REPO_ROOT, abs).split(path.sep).join('/'); +} + +function isFile(abs: string): boolean { + return fs.statSync(abs, { throwIfNoEntry: false })?.isFile() ?? false; +} + function collectFiles(absDir: string, into: string[]): void { for (const entry of fs.readdirSync(absDir, { withFileTypes: true })) { const abs = path.join(absDir, entry.name); @@ -223,6 +224,51 @@ function collectFiles(absDir: string, into: string[]): void { } } +function loadAliases(): ReadonlyArray { + const tsconfig = JSON.parse( + fs.readFileSync(path.join(REPO_ROOT, 'tsconfig.build.json'), 'utf8'), + ) as { compilerOptions?: { paths?: Record } }; + return Object.entries(tsconfig.compilerOptions?.paths ?? {}).map( + ([pattern, targets]) => [pattern, targets[0]] as const, + ); +} + +function aliasTarget( + spec: string, + aliases: ReadonlyArray, +): string | null { + for (const [pattern, target] of aliases) { + if (pattern.endsWith('*')) { + const prefix = pattern.slice(0, -1); + if (spec.startsWith(prefix)) { + return path.resolve( + REPO_ROOT, + target.slice(0, -1) + spec.slice(prefix.length), + ); + } + } else if (spec === pattern) { + return path.resolve(REPO_ROOT, target); + } + } + return null; +} + +function probe(base: string): string | null { + const candidates: string[] = []; + if (base.endsWith('.js')) { + const stem = base.slice(0, -3); + candidates.push(`${stem}.ts`, `${stem}.tsx`); + } + candidates.push( + `${base}.ts`, + `${base}.tsx`, + path.join(base, 'index.ts'), + path.join(base, 'index.tsx'), + base, + ); + return candidates.find(isFile) ?? null; +} + function specifiersIn(text: string): string[] { const found = new Set(); for (const pattern of SPECIFIER_PATTERNS) { @@ -362,35 +408,6 @@ describe('import boundaries', () => { }); }); -// The program watchers load inside this closure. -it('keeps the callable runProgram closure free of UI, session and legacy imports', () => { - const forbidden = staticImportClosure( - 'src/programs/run-program.ts', - true, - ).filter( - (file) => - file === 'src/programs/program-registry.ts' || - file.startsWith('src/ui/') || - file.startsWith('src/steps/') || - file.startsWith('src/lib/wizard-session') || - file.startsWith('src/lib/runners/') || - file.startsWith('src/commands/') || - file.startsWith('src/programs/task-stream/'), - ); - expect(forbidden).toEqual([]); -}); - -it('keeps program decks and task-stream state behind the TUI boundary', () => { - const forbidden = analysis.edges.filter( - (edge) => - (edge.startsWith('src/programs/') && - edge.includes(' -> src/ui/tui/decks/')) || - (edge.startsWith('src/programs/task-stream/') && - edge.includes(' -> src/ui/')), - ); - expect(forbidden).toEqual([]); -}); - describe('surface classification', () => { it('maps representative paths to their surface', () => { expect(classifySurface('src/env.ts')).toBe('env'); @@ -479,8 +496,8 @@ describe('programs entry modules', () => { 'matrix:programs->tui', ); expect( - rule('src/commands/dispatch-family.ts', 'src/commands/command.ts'), - ).toBe(null); + rule('src/programs/dispatch-family.ts', 'src/commands/command.ts'), + ).toBe('matrix:programs->cli'); }); }); diff --git a/src/__tests__/architecture/known-violations.json b/src/__tests__/architecture/known-violations.json index 07d3ab535..035428ddd 100644 --- a/src/__tests__/architecture/known-violations.json +++ b/src/__tests__/architecture/known-violations.json @@ -1,5 +1,8 @@ { "violations": [ + "src/agent/runner/switchboard/flags/index.ts -> src/programs/types.ts", + "src/agent/runner/switchboard/flags/schemes.ts -> src/programs/types.ts", + "src/agent/runner/switchboard/index.ts -> src/programs/types.ts", "src/commands/ai-observability.ts -> src/programs/ai-observability/index.ts", "src/commands/audit.ts -> src/programs/audit/index.ts", "src/commands/basic-integration/ci-install.ts -> src/programs/posthog-integration/index.ts", @@ -7,6 +10,7 @@ "src/commands/basic-integration/skill.ts -> src/programs/agent-skill/index.ts", "src/commands/doctor.ts -> src/programs/posthog-doctor/index.ts", "src/commands/error-tracking.ts -> src/programs/error-tracking/index.ts", + "src/commands/factories/family-command-factory.ts -> src/programs/dispatch-family.ts", "src/commands/factories/family-picker.tsx -> src/commands/command.ts", "src/commands/mcp-analytics.ts -> src/programs/mcp-analytics/index.ts", "src/commands/metrics.ts -> src/programs/metrics/index.ts", @@ -19,17 +23,102 @@ "src/env.ts -> src/lib/headless-mode.ts", "src/lib/runners/run-non-interactive.ts -> src/programs/audit/types.ts", "src/lib/runners/run-non-interactive.ts -> src/programs/detect-map.ts", + "src/lib/runners/run-non-interactive.ts -> src/programs/run-agent-legacy.ts", "src/lib/runners/run-non-interactive.ts -> src/programs/task-stream/index.ts", "src/lib/runners/run-non-interactive.ts -> src/programs/task-stream/task-stream-push.ts", "src/lib/runners/run-wizard.ts -> src/programs/audit/types.ts", "src/lib/runners/run-wizard.ts -> src/programs/authenticate.ts", "src/lib/runners/run-wizard.ts -> src/programs/posthog-integration/detect.ts", + "src/lib/runners/run-wizard.ts -> src/programs/run-agent-legacy.ts", "src/lib/runners/run-wizard.ts -> src/programs/task-stream/destinations/file.ts", "src/lib/runners/run-wizard.ts -> src/programs/task-stream/destinations/posthog.ts", "src/lib/runners/run-wizard.ts -> src/programs/task-stream/index.ts", "src/lib/runners/run-wizard.ts -> src/programs/task-stream/task-stream-push.ts", "src/lib/wizard-session.ts -> src/agent/progress.ts", + "src/programs/agent-skill/index.ts -> src/ui/tui/decks/agent-skill/index.tsx", + "src/programs/agent-skill/steps.ts -> src/lib/wizard-session.ts", + "src/programs/ai-observability/index.ts -> src/lib/headless-mode.ts", + "src/programs/ai-observability/index.ts -> src/ui/tui/decks/agent-skill/index.tsx", + "src/programs/ai-opt-in-gate.ts -> src/lib/wizard-session.ts", + "src/programs/audit/index.ts -> src/lib/headless-mode.ts", + "src/programs/audit/index.ts -> src/lib/wizard-session.ts", + "src/programs/audit/ledger-watcher.ts -> src/lib/file-watcher.ts", + "src/programs/audit/ledger-watcher.ts -> src/ui/index.ts", + "src/programs/audit/types.ts -> src/lib/wizard-session.ts", + "src/programs/authenticate.ts -> src/lib/wizard-session.ts", + "src/programs/authenticate.ts -> src/ui/index.ts", + "src/programs/detection/agentic.ts -> src/lib/wizard-session.ts", + "src/programs/detection/agentic.ts -> src/ui/agent-progress.ts", + "src/programs/detection/agentic.ts -> src/ui/index.ts", + "src/programs/detection/features.ts -> src/lib/wizard-session.ts", + "src/programs/detection/project-scope.ts -> src/lib/wizard-session.ts", + "src/programs/dispatch-family.ts -> src/commands/command.ts", + "src/programs/dispatch-family.ts -> src/commands/factories/shared.ts", + "src/programs/error-tracking-upload-source-maps/detect-agentic.ts -> src/lib/wizard-session.ts", + "src/programs/error-tracking-upload-source-maps/detect.ts -> src/lib/wizard-session.ts", + "src/programs/error-tracking-upload-source-maps/index.ts -> src/lib/wizard-session.ts", + "src/programs/error-tracking-upload-source-maps/index.ts -> src/ui/tui/decks/error-tracking-upload-source-maps/index.tsx", + "src/programs/error-tracking-upload-source-maps/steps.ts -> src/lib/wizard-session.ts", + "src/programs/error-tracking/detect-agentic.ts -> src/lib/wizard-session.ts", + "src/programs/error-tracking/index.ts -> src/lib/wizard-session.ts", + "src/programs/error-tracking/index.ts -> src/ui/tui/decks/error-tracking/index.tsx", + "src/programs/error-tracking/index.ts -> src/ui/tui/decks/error-tracking/tips.ts", + "src/programs/events-audit/index.ts -> src/lib/wizard-session.ts", + "src/programs/events-audit/steps.ts -> src/lib/wizard-session.ts", + "src/programs/frameworks/astro/astro-wizard-agent.ts -> src/ui/index.ts", + "src/programs/frameworks/django/utils.ts -> src/ui/index.ts", + "src/programs/frameworks/fastapi/utils.ts -> src/ui/index.ts", + "src/programs/frameworks/flask/utils.ts -> src/ui/index.ts", + "src/programs/frameworks/laravel/utils.ts -> src/ui/index.ts", + "src/programs/frameworks/nextjs/nextjs-wizard-agent.ts -> src/ui/index.ts", + "src/programs/frameworks/rails/utils.ts -> src/ui/index.ts", + "src/programs/frameworks/react-native/utils.ts -> src/ui/index.ts", + "src/programs/frameworks/react-router/react-router-wizard-agent.ts -> src/ui/index.ts", + "src/programs/frameworks/tanstack-router/tanstack-router-wizard-agent.ts -> src/ui/index.ts", + "src/programs/mcp/index.ts -> src/lib/wizard-session.ts", + "src/programs/metrics/index.ts -> src/ui/tui/decks/agent-skill/index.tsx", + "src/programs/migration/index.ts -> src/ui/tui/decks/migration/index.tsx", + "src/programs/migration/steps.ts -> src/lib/wizard-session.ts", + "src/programs/posthog-integration/detect.ts -> src/lib/wizard-session.ts", + "src/programs/posthog-integration/index.ts -> src/lib/wizard-session.ts", + "src/programs/posthog-integration/index.ts -> src/steps/index.ts", + "src/programs/posthog-integration/index.ts -> src/ui/tui/decks/posthog-integration/index.tsx", + "src/programs/posthog-integration/steps.ts -> src/lib/wizard-session.ts", + "src/programs/program-registry.ts -> src/ui/tui/decks/agent-skill/index.tsx", + "src/programs/program-run.ts -> src/lib/wizard-session.ts", "src/programs/program-step.ts -> src/lib/wizard-session.ts", + "src/programs/program-step.ts -> src/ui/tui/components/TipsCard.tsx", + "src/programs/program-step.ts -> src/ui/tui/primitives/index.ts", + "src/programs/program-step.ts -> src/ui/tui/store.ts", + "src/programs/replay-vision/index.ts -> src/lib/wizard-session.ts", + "src/programs/revenue-analytics/detect.ts -> src/lib/wizard-session.ts", + "src/programs/revenue-analytics/index.ts -> src/ui/tui/decks/revenue-analytics/index.tsx", + "src/programs/revenue-analytics/steps.ts -> src/lib/wizard-session.ts", + "src/programs/run-agent-legacy.ts -> src/lib/wizard-session.ts", + "src/programs/run-agent-legacy.ts -> src/ui/agent-progress.ts", + "src/programs/run-agent-legacy.ts -> src/ui/index.ts", + "src/programs/run-program.ts -> src/lib/wizard-session.ts", + "src/programs/self-driving/detect-agentic.ts -> src/lib/wizard-session.ts", + "src/programs/self-driving/detect.ts -> src/lib/wizard-session.ts", + "src/programs/self-driving/index.ts -> src/lib/wizard-session.ts", + "src/programs/self-driving/index.ts -> src/ui/tui/decks/self-driving/index.tsx", + "src/programs/self-driving/index.ts -> src/ui/tui/decks/self-driving/pricing.ts", + "src/programs/self-driving/index.ts -> src/ui/tui/decks/self-driving/tips.ts", + "src/programs/self-driving/steps.ts -> src/lib/wizard-session.ts", + "src/programs/shared/health-check-step.ts -> src/lib/wizard-session.ts", + "src/programs/shared/posthog-cli-preinstall.ts -> src/steps/install-cli-steering/index.ts", + "src/programs/task-stream/destinations/posthog.ts -> src/lib/wizard-session.ts", + "src/programs/task-stream/event-plan-watcher.ts -> src/lib/file-watcher.ts", + "src/programs/task-stream/event-plan-watcher.ts -> src/ui/tui/store.ts", + "src/programs/task-stream/task-stream-push.ts -> src/lib/wizard-session.ts", + "src/programs/task-stream/task-stream-push.ts -> src/ui/tui/store.ts", + "src/programs/task-stream/task-stream-push.ts -> src/ui/wizard-ui.ts", + "src/programs/task-stream/types.ts -> src/lib/wizard-session.ts", + "src/programs/warehouse-source/detect.ts -> src/lib/wizard-session.ts", + "src/programs/warehouse-source/index.ts -> src/lib/wizard-session.ts", + "src/programs/warehouse-source/index.ts -> src/ui/tui/decks/warehouse-source/index.tsx", + "src/programs/warehouse-source/steps.ts -> src/lib/wizard-session.ts", + "src/programs/web-analytics-doctor/detect.ts -> src/lib/wizard-session.ts", "src/shared/utils/analytics.ts -> src/lib/wizard-session.ts", "src/shared/utils/oauth.ts -> src/ui/index.ts", "src/shared/utils/setup-utils.ts -> src/lib/wizard-session.ts", @@ -39,6 +128,7 @@ "src/shared/utils/wizard-abort.ts -> src/lib/wizard-session.ts", "src/shared/utils/wizard-abort.ts -> src/ui/index.ts", "src/shared/utils/wizard-abort.ts -> src/ui/logging-ui.ts", + "src/ui/headless-ui.ts -> src/ui/tui/store.ts", "src/ui/tui/decks/error-tracking/tips.ts -> src/programs/replay-vision/index.ts", "src/ui/tui/playground/demos/AuditChecksDemo.tsx -> src/programs/audit/types.ts", "src/ui/tui/playground/demos/DoctorReportDemo.tsx -> src/programs/posthog-doctor/index.ts", diff --git a/src/__tests__/cli.test.ts b/src/__tests__/cli.test.ts index 953e7cdea..c740d3fc8 100644 --- a/src/__tests__/cli.test.ts +++ b/src/__tests__/cli.test.ts @@ -113,7 +113,7 @@ vi.mock('@utils/wizard-abort', async (importOriginal) => ({ ...(await importOriginal()), wizardAbort: vi.fn(), })); -vi.mock('../lib/runners/run-program-agent', () => ({ +vi.mock('../programs/run-agent-legacy', () => ({ runProgramAgent: vi.fn().mockResolvedValue(undefined), })); diff --git a/src/__tests__/programs-cli.test.ts b/src/__tests__/programs-cli.test.ts index cfb523c73..998808992 100644 --- a/src/__tests__/programs-cli.test.ts +++ b/src/__tests__/programs-cli.test.ts @@ -29,7 +29,7 @@ import { selfDrivingCommand } from '../commands/self-driving'; import { dispatchFamily, pickerChildrenToShow, -} from '../commands/dispatch-family'; +} from '@programs/dispatch-family'; import type { Command } from '../commands/command'; import { fetchSkillMenu, type CliEntry } from '@shared/skill-menu'; import { auditConfig } from '@programs/audit/index'; diff --git a/src/__tests__/provision-cli.test.ts b/src/__tests__/provision-cli.test.ts index b901fb59f..0224d25fa 100644 --- a/src/__tests__/provision-cli.test.ts +++ b/src/__tests__/provision-cli.test.ts @@ -64,7 +64,7 @@ vi.mock('@utils/wizard-abort', async (importOriginal) => ({ ...(await importOriginal()), wizardAbort: vi.fn(), })); -vi.mock('../lib/runners/run-program-agent', () => ({ +vi.mock('../programs/run-agent-legacy', () => ({ runProgramAgent: vi.fn().mockResolvedValue(undefined), })); diff --git a/src/agent/__tests__/agent-interface.test.ts b/src/agent/__tests__/agent-interface.test.ts index 6e4ab44e7..c096578e6 100644 --- a/src/agent/__tests__/agent-interface.test.ts +++ b/src/agent/__tests__/agent-interface.test.ts @@ -237,8 +237,6 @@ describe('runAgent', () => { kind: 'abort', classification: 'WIZARD_ABORT', }); - const [{ options }] = mockQuery.mock.calls[0]; - expect(options.abortController.signal.aborted).toBe(true); }); it('returns a failure when the stream ends without a terminal result', async () => { diff --git a/src/shared/__tests__/ask-policy.test.ts b/src/agent/__tests__/agent-runner-ask.test.ts similarity index 76% rename from src/shared/__tests__/ask-policy.test.ts rename to src/agent/__tests__/agent-runner-ask.test.ts index 0e9f7aaac..dfee8067f 100644 --- a/src/shared/__tests__/ask-policy.test.ts +++ b/src/agent/__tests__/agent-runner-ask.test.ts @@ -1,21 +1,21 @@ -import { isAskDisabled } from '@shared/ask-policy'; +import { shouldDisableAsk } from '@agent/agent-runner'; import { buildSession } from '@lib/wizard-session'; -describe('isAskDisabled', () => { +describe('shouldDisableAsk', () => { it('enables wizard_ask in interactive runs by default', () => { - expect(isAskDisabled({ ci: false, signup: false, e2eAsk: false })).toBe( + expect(shouldDisableAsk({ ci: false, signup: false, e2eAsk: false })).toBe( false, ); }); it('auto-disables when running in CI mode', () => { - expect(isAskDisabled({ ci: true, signup: false, e2eAsk: false })).toBe( + expect(shouldDisableAsk({ ci: true, signup: false, e2eAsk: false })).toBe( true, ); }); it('auto-disables during the signup flow (which is non-interactive at the prompt layer)', () => { - expect(isAskDisabled({ ci: false, signup: true, e2eAsk: false })).toBe( + expect(shouldDisableAsk({ ci: false, signup: true, e2eAsk: false })).toBe( true, ); }); @@ -35,14 +35,14 @@ describe('isAskDisabled', () => { ])( 'ci=$ci signup=$signup e2eAsk=$e2eAsk → disabled=$disabled', ({ ci, signup, e2eAsk, disabled }) => { - expect(isAskDisabled({ ci, signup, e2eAsk })).toBe(disabled); + expect(shouldDisableAsk({ ci, signup, e2eAsk })).toBe(disabled); }, ); it('leaves a plain --ci session disabled — buildSession defaults e2eAsk to false', () => { const session = buildSession({ installDir: '/tmp/ask-policy', ci: true }); expect(session.e2eAsk).toBe(false); - expect(isAskDisabled(session)).toBe(true); + expect(shouldDisableAsk(session)).toBe(true); }); it('re-enables the bridge when the harness asks for it', () => { @@ -51,6 +51,6 @@ describe('isAskDisabled', () => { ci: true, e2eAsk: true, }); - expect(isAskDisabled(session)).toBe(false); + expect(shouldDisableAsk(session)).toBe(false); }); }); diff --git a/src/agent/__tests__/commandments.test.ts b/src/agent/__tests__/commandments.test.ts index d4a9722d5..2a694c0f2 100644 --- a/src/agent/__tests__/commandments.test.ts +++ b/src/agent/__tests__/commandments.test.ts @@ -1,6 +1,5 @@ import { WIZARD_COMMANDMENTS } from '@agent/commandments'; import { assembleCommandments } from '@agent/runner/switchboard/commandments'; -import { getProgramCommandments } from '@programs/commandments'; import { Harness, Sequence } from '@shared/constants'; const global = WIZARD_COMMANDMENTS.join('\n'); @@ -11,13 +10,7 @@ const prompt = ( harness: Harness, sequence: Sequence, program = 'posthog-integration', -) => - assembleCommandments({ - programCommandments: getProgramCommandments(program), - sequence, - harness, - caps: CAPS, - }); +) => assembleCommandments({ program, sequence, harness, caps: CAPS }); const COMBOS = [ ['anthropic', Harness.anthropic, Sequence.linear], @@ -155,6 +148,7 @@ describe('commandments by axis', () => { describe('runtime caps gate the pi runtime notes', () => { const withCaps = (caps: { bash: boolean; posthogMcp: boolean }) => assembleCommandments({ + program: 'warehouse-source', sequence: Sequence.linear, harness: Harness.pi, caps, diff --git a/src/agent/__tests__/entry-closure.test.ts b/src/agent/__tests__/entry-closure.test.ts deleted file mode 100644 index 0db2317d4..000000000 --- a/src/agent/__tests__/entry-closure.test.ts +++ /dev/null @@ -1,19 +0,0 @@ -import { staticImportClosure } from '../../../test/module-graph'; - -describe('@agent entry closure', () => { - // The CLI's startup chunk imports @agent, so its static imports load before any work. - it('loads only leaf data at startup; the runner, agent interface and tools load on first call', () => { - const agentFiles = staticImportClosure('src/agent/index.ts').filter( - (file) => file.startsWith('src/agent/'), - ); - expect(agentFiles).toEqual([ - 'src/agent/default-binding.ts', - 'src/agent/index.ts', - 'src/agent/progress.ts', - 'src/agent/runner/shared/types.ts', - 'src/agent/runner/switchboard/resolve-harness.ts', - 'src/agent/signals.ts', - 'src/agent/tools/tool-names.ts', - ]); - }); -}); diff --git a/src/agent/__tests__/entry-streaming.test.ts b/src/agent/__tests__/entry-streaming.test.ts index 31aefefd8..9ff311d90 100644 --- a/src/agent/__tests__/entry-streaming.test.ts +++ b/src/agent/__tests__/entry-streaming.test.ts @@ -1,6 +1,10 @@ import { rmSync } from 'node:fs'; import { runMcpPromptViaSdk } from '@agent'; import type { AgentChunk } from '@agent/types'; +import { + configureGatewayCredentialsForCI, + resetGatewaySession, +} from '@agent/gateway-session'; import { HostResolution } from '@shared/host-resolution'; const { query } = vi.hoisted(() => ({ @@ -28,15 +32,6 @@ async function consume( }, signal: new AbortController().signal, programId: 'mcp-tutorial', - inferenceAuth: { - resolve: () => - Promise.resolve({ - token: 'test-gateway-token', - teamId: 1, - gatewayUrl: 'https://ai-gateway.us.posthog.com', - refreshAtMs: Infinity, - }), - }, ...overrides, })) { chunks.push(chunk); @@ -54,6 +49,11 @@ describe('public agent prompt stream', () => { ]) { vi.stubEnv(name, process.env[name]); } + configureGatewayCredentialsForCI( + 'test-gateway-token', + 1, + 'https://ai-gateway.us.posthog.com', + ); }); afterEach(() => { @@ -64,6 +64,7 @@ describe('public agent prompt stream', () => { }); } query.mockReset(); + resetGatewaySession(); vi.unstubAllEnvs(); }); @@ -98,11 +99,10 @@ describe('public agent prompt stream', () => { }); it('propagates setup failures instead of silently ending the stream', async () => { - const inferenceAuth = { - resolve: () => Promise.reject(new Error('gateway mint refused')), - }; - await expect(consume({ inferenceAuth })).rejects.toThrow( - 'gateway mint refused', + resetGatewaySession(); + + await expect(consume({ programId: undefined })).rejects.toThrow( + 'this run has no program to attribute its spend to', ); }); }); diff --git a/src/programs/__tests__/gateway-session.test.ts b/src/agent/__tests__/gateway-session.test.ts similarity index 94% rename from src/programs/__tests__/gateway-session.test.ts rename to src/agent/__tests__/gateway-session.test.ts index 4e6e4ee26..cf1eda731 100644 --- a/src/programs/__tests__/gateway-session.test.ts +++ b/src/agent/__tests__/gateway-session.test.ts @@ -2,15 +2,14 @@ import { inspect } from 'node:util'; import { GatewayMintFailed, GatewayMintRefused, - gatewayAuth, - resetGatewaySession, -} from '../gateway-session'; -import { buildWizardPropertiesBlob, + configureGatewayCredentialsForCI, + configureGatewayFromCIEnvironment, + gatewayAuth, isPastRefresh, isTrustedGatewayUrl, -} from '@shared/gateway-auth'; -import { createCiGatewayAuth } from '@shared/ci-gateway-auth'; + resetGatewaySession, +} from '@agent/gateway-session'; import type { HostResolution } from '@shared/host-resolution'; import { ErrorCodes } from '@shared/errors'; import { WizardError } from '@utils/wizard-abort'; @@ -66,6 +65,38 @@ describe('gatewayAuth', () => { vi.unstubAllGlobals(); }); + it('uses the supplied CI bearer across programs and time without minting', async () => { + configureGatewayCredentialsForCI( + ' opaque-ci-token ', + 42, + 'https://ai-gateway.us.posthog.com/', + ); + const auth = { + token: 'opaque-ci-token', + teamId: 42, + gatewayUrl: 'https://ai-gateway.us.posthog.com', + refreshAtMs: Infinity, + }; + const results = await Promise.all( + ['integration', 'audit', undefined].map((program) => + gatewayAuth(host, 'phx_project', program), + ), + ); + expect(results).toEqual([auth, auth, auth]); + const clock = vi + .spyOn(Date, 'now') + .mockReturnValue(Number.MAX_SAFE_INTEGER); + try { + expect(await gatewayAuth(host, 'phx_project', 'integration')).toEqual( + auth, + ); + expect(isPastRefresh(auth)).toBe(false); + } finally { + clock.mockRestore(); + } + expect(fetchMock).not.toHaveBeenCalled(); + }); + it.each([ ['', 42, 'https://ai-gateway.us.posthog.com'], ['token', 0, 'https://ai-gateway.us.posthog.com'], @@ -79,7 +110,9 @@ describe('gatewayAuth', () => { ] as const)( 'rejects invalid CI gateway configuration', (token, projectId, url) => { - expect(() => createCiGatewayAuth(token, projectId, url)).toThrow(); + expect(() => + configureGatewayCredentialsForCI(token, projectId, url), + ).toThrow(); }, ); @@ -87,9 +120,9 @@ describe('gatewayAuth', () => { vi.stubEnv('NODE_ENV', 'production'); vi.resetModules(); try { - const prod = await import('@shared/ci-gateway-auth'); + const prod = await import('@agent/gateway-session'); expect(() => - prod.createCiGatewayAuth( + prod.configureGatewayCredentialsForCI( 'token', 42, 'https://ai-gateway.us.posthog.com', @@ -101,6 +134,17 @@ describe('gatewayAuth', () => { } }); + it('requires an explicit gateway token file for CI', () => { + vi.stubEnv('WIZARD_CI_GATEWAY_TOKEN_FILE', ''); + try { + expect(() => configureGatewayFromCIEnvironment(42, 'us')).toThrow( + 'WIZARD_CI_GATEWAY_TOKEN_FILE is required', + ); + } finally { + vi.unstubAllEnvs(); + } + }); + it('resolves auth from a mint response and caches it', async () => { fetchMock.mockResolvedValue({ ok: true, diff --git a/src/agent/__tests__/harness-capabilities.test.ts b/src/agent/__tests__/harness-capabilities.test.ts deleted file mode 100644 index d491f90d0..000000000 --- a/src/agent/__tests__/harness-capabilities.test.ts +++ /dev/null @@ -1,16 +0,0 @@ -import { Harness } from '@shared/constants'; -import { HARNESS_OPTIONS } from '@agent/runner/switchboard/harness'; -import { HARNESS_RUNS_TASKS } from '@agent/runner/switchboard/resolve-harness'; - -describe('harness capabilities', () => { - it.each(Object.values(Harness))( - 'records whether the %s backend implements runTask', - (harness) => { - const backend = HARNESS_OPTIONS[harness]; - expect(backend).toBeDefined(); - expect(HARNESS_RUNS_TASKS[harness]).toBe( - typeof backend?.runTask === 'function', - ); - }, - ); -}); diff --git a/src/agent/__tests__/public-entry.test.ts b/src/agent/__tests__/public-entry.test.ts deleted file mode 100644 index 37572a63d..000000000 --- a/src/agent/__tests__/public-entry.test.ts +++ /dev/null @@ -1,18 +0,0 @@ -describe('@agent public entry', () => { - it('exports only the agent contract and its data', async () => { - const entry = await import('@agent'); - expect(Object.keys(entry).sort()).toEqual([ - 'AgentSignals', - 'DEFAULT_AGENT_BINDING', - 'OutroKind', - 'RunOutcome', - 'TASK_OUTCOMES_KEY', - 'WIZARD_TOOL_NAMES', - 'downloadSkill', - 'harnessRunsTasks', - 'resolveHarness', - 'runAgent', - 'runMcpPromptViaSdk', - ]); - }); -}); diff --git a/src/agent/__tests__/run-agent-standalone.test.ts b/src/agent/__tests__/run-agent-standalone.test.ts index 4469d8a75..df73af8fe 100644 --- a/src/agent/__tests__/run-agent-standalone.test.ts +++ b/src/agent/__tests__/run-agent-standalone.test.ts @@ -54,6 +54,15 @@ vi.mock('@utils/analytics', () => ({ shutdown: vi.fn().mockResolvedValue(undefined), }, })); +vi.mock('@agent/gateway-session', async (importOriginal) => ({ + ...(await importOriginal()), + gatewayAuth: vi.fn().mockResolvedValue({ + gatewayUrl: 'https://gateway.test', + token: 'phe_run', + teamId: 1, + refreshAtMs: Date.now() + 3_600_000, + }), +})); // The fake harness: reports a little of everything, then returns what the // current test told it to. @@ -162,6 +171,10 @@ vi.mock('@agent/runner/switchboard/harness', () => { harnessState.selected.push(name); return { ...fake, name }; }, + resolveHarness: (ctx: { cliHarness?: Harness }) => ({ + harness: ctx.cliHarness ?? Harness.pi, + model: DEFAULT_AGENT_MODEL, + }), }; }); @@ -200,7 +213,6 @@ import type { RunAgentOptions, RunConfig, RunInput } from '@agent/runner'; import { analytics } from '@utils/analytics'; import { initLogFile } from '@utils/debug'; import { flushScanReport } from '@agent/yara-hooks'; -import { clearCleanup, runCleanups } from '@utils/cleanup-registry'; import { QUEUE_DIR_NAME } from '../runner/sequence/orchestrator/queue'; let tmp: string; @@ -225,6 +237,7 @@ const config = (over: Partial = {}): RunConfig => ({ }, composed: false, binding: { sequence: Sequence.linear, harness: Harness.pi, model: 'm' }, + switchboard: { program: 'test-program', flags: {} }, skillsBaseUrl: 'https://skills.test', wizardFlags: {}, wizardFlagPayloads: {}, @@ -240,15 +253,6 @@ const input = (over: Partial = {}): RunInput => ({ host: HostResolution.fromApiHost('https://us.posthog.com'), projectId: 1, }, - inferenceAuth: { - resolve: () => - Promise.resolve({ - gatewayUrl: 'https://gateway.test', - token: 'phe_run', - teamId: 1, - refreshAtMs: Date.now() + 3_600_000, - }), - }, project: null, apiUser: null, skillId: 'test-integration', @@ -312,6 +316,11 @@ describe('runAgent standalone', () => { const running = runAgent( config({ binding: { harness, sequence, model: DEFAULT_AGENT_MODEL }, + switchboard: { + program: 'test-program', + flags: {}, + cliHarness: harness, + }, }), input(), { interaction: { ask }, onProgress: (event) => events.push(event) }, @@ -378,6 +387,11 @@ describe('runAgent standalone', () => { const result = await runAgent( config({ binding: { harness, sequence, model: DEFAULT_AGENT_MODEL }, + switchboard: { + program: 'test-program', + flags: {}, + cliHarness: harness, + }, }), input(), ); @@ -413,6 +427,11 @@ describe('runAgent standalone', () => { sequence: Sequence.orchestrator, model: DEFAULT_AGENT_MODEL, }, + switchboard: { + program: 'test-program', + flags: {}, + cliHarness: Harness.anthropic, + }, }), input(), ); @@ -435,6 +454,11 @@ describe('runAgent standalone', () => { sequence: Sequence.orchestrator, model: DEFAULT_AGENT_MODEL, }, + switchboard: { + program: 'test-program', + flags: {}, + cliHarness: Harness.anthropic, + }, }), input(), ); @@ -458,6 +482,11 @@ describe('runAgent standalone', () => { sequence: Sequence.orchestrator, model: DEFAULT_AGENT_MODEL, }, + switchboard: { + program: 'test-program', + flags: {}, + cliHarness: Harness.pi, + }, }), input(), ); @@ -503,6 +532,11 @@ describe('runAgent standalone', () => { sequence: Sequence.orchestrator, model: DEFAULT_AGENT_MODEL, }, + switchboard: { + program: 'test-program', + flags: {}, + cliHarness: Harness.pi, + }, }), input(), { @@ -740,41 +774,6 @@ describe('runAgent standalone', () => { expect(analytics.shutdown).not.toHaveBeenCalled(); }); - it('a process drain during a run writes the scan report once, through progress', async () => { - clearCleanup(); - harnessState.askQuestions = [{ id: 'q1', prompt: 'Go?', kind: 'text' }]; - const ask = vi.fn(() => new Promise(() => undefined)); - vi.mocked(flushScanReport).mockReturnValueOnce( - 'YARA scan report: /tmp/scan.json', - ); - const controller = new AbortController(); - const events: AgentProgress[] = []; - const running = runAgent( - config(), - input({ flags: { ...input().flags, yaraReport: true } }), - { - signal: controller.signal, - onProgress: (event) => events.push(event), - interaction: { ask }, - }, - ); - await vi.waitFor(() => expect(ask).toHaveBeenCalled()); - - runCleanups(); - - expect(flushScanReport).toHaveBeenCalledExactlyOnceWith({ - yaraReport: true, - }); - expect(events).toContainEqual({ - kind: 'log', - level: 'info', - message: 'YARA scan report: /tmp/scan.json', - }); - controller.abort(); - expect((await running).outcome).toBe(RunOutcome.Aborted); - expect(flushScanReport).toHaveBeenCalledTimes(1); - }); - it('runs to a complete result with no options at all', async () => { const result = await runAgent(config(), input()); @@ -873,35 +872,6 @@ describe('runAgent standalone', () => { expect(result.failure).toBe(failure); }); - it.each([ - [RunOutcome.Failed, ['preexisting', 'user-owned']], - [RunOutcome.Success, ['installed', 'preexisting', 'user-owned']], - ] as const)('a %s run leaves %j in .claude/skills', async (outcome, left) => { - const skillsDir = path.join(tmp, '.claude', 'skills'); - const makeSkill = (id: string, marked: boolean) => { - fs.mkdirSync(path.join(skillsDir, id), { recursive: true }); - if (marked) - fs.writeFileSync(path.join(skillsDir, id, '.posthog-wizard'), ''); - }; - makeSkill('preexisting', true); - if (outcome === RunOutcome.Failed) - harnessState.result = { - kind: 'failure', - classification: AgentErrorType.NO_PROGRESS, - }; - - const result = await runAgent(config(), input(), { - onProgress: (event) => { - if (event.kind !== 'status') return; - makeSkill('installed', true); - makeSkill('user-owned', false); - }, - }); - - expect(result.outcome).toBe(outcome); - expect(fs.readdirSync(skillsDir).sort()).toEqual(left); - }); - it('skips the terminal outro for a composed sub-run', async () => { const events: AgentProgress[] = []; @@ -918,81 +888,6 @@ describe('runAgent standalone', () => { expect(analytics.shutdown).not.toHaveBeenCalled(); }); - it('collectTranscript fills snapshot.transcriptTail; run.prompt replaces the assembled prompt', async () => { - const events: AgentProgress[] = []; - const long = 'x'.repeat(150); - harnessState.run = ({ middleware }) => { - middleware?.onMessage({ - type: 'assistant', - message: { - content: [ - { type: 'text', text: ' Globbing every manifest. ' }, - { - type: 'tool_use', - name: 'Read', - input: { file_path: 'apps/web/package.json' }, - }, - { type: 'text', text: long }, - ], - }, - }); - middleware?.onMessage({ type: 'result', result: '{"path":"."}' }); - return Promise.resolve({ kind: 'success' }); - }; - - const result = await runAgent( - config({ - run: { - ...config().run, - prompt: () => 'Scan the repo only.', - collectTranscript: true, - }, - }), - input(), - { onProgress: (event) => events.push(event) }, - ); - - expect(result.outcome).toBe(RunOutcome.Success); - expect((harnessState.lastInputs as BackendRunInputs).prompt).toBe( - 'Scan the repo only.', - ); - expect(result.snapshot.transcriptTail).toBe( - ` Globbing every manifest. \n${long}\n{"path":"."}`, - ); - expect(events.filter((event) => event.kind === 'activity')).toEqual([ - { kind: 'activity', line: 'Globbing every manifest.' }, - { kind: 'activity', line: 'Read apps/web/package.json' }, - { kind: 'activity', line: `${'x'.repeat(100)}…` }, - ]); - }); - - it('keeps only the newest 256 KiB of a collected transcript', async () => { - const chunk = 'y'.repeat(100 * 1024); - harnessState.run = ({ middleware }) => { - for (const text of ['oldest', chunk, chunk, chunk]) { - middleware?.onMessage({ - type: 'assistant', - message: { content: [{ type: 'text', text }] }, - }); - } - return Promise.resolve({ kind: 'success' }); - }; - - const result = await runAgent( - config({ run: { ...config().run, collectTranscript: true } }), - input(), - ); - - expect(result.snapshot.transcriptTail).toBe(`${chunk}\n${chunk}\n`); - }); - - it('leaves a deferred scan report to the host run', async () => { - const result = await runAgent(config({ scanReport: 'defer' }), input()); - - expect(result.outcome).toBe(RunOutcome.Success); - expect(flushScanReport).not.toHaveBeenCalled(); - }); - it('finishes when the observer throws on every event', async () => { const result = await runAgent(config(), input(), { onProgress: () => { @@ -1020,8 +915,6 @@ describe('runAgent standalone', () => { ]; const host = new AbortController(); const signals: AbortSignal[] = []; - const postRun = vi.fn(); - const events: AgentProgress[] = []; const running = runAgent( config({ binding: { @@ -1029,12 +922,15 @@ describe('runAgent standalone', () => { sequence, model: DEFAULT_AGENT_MODEL, }, - hooks: { postRun }, + switchboard: { + program: 'test-program', + flags: {}, + cliHarness: Harness.pi, + }, }), input(), { signal: host.signal, - onProgress: (event) => events.push(event), interaction: { ask: (_question, { signal }) => { signals.push(signal); @@ -1050,9 +946,6 @@ describe('runAgent standalone', () => { expect(result.outcome).toBe(RunOutcome.Aborted); // The host's abort reached the open question as its own abort. expect(signals[0].aborted).toBe(true); - expect(postRun).not.toHaveBeenCalled(); - expect(events.some((event) => event.kind === 'completion')).toBe(false); - expect(fs.existsSync(path.join(tmp, QUEUE_DIR_NAME))).toBe(false); }, ); @@ -1103,18 +996,6 @@ describe('runAgent standalone', () => { expect(result.snapshot.statusMessages).toContain('Installing the SDK'); }); - it('ends the run before any harness when the inference provider refuses', async () => { - const refusal = new Error('gateway mint refused'); - const result = await runAgent( - config(), - input({ inferenceAuth: { resolve: () => Promise.reject(refusal) } }), - ); - - expect(result.outcome).toBe(RunOutcome.Crashed); - expect(result.failure?.error).toBe(refusal); - expect(harnessState.lastInputs).toBeUndefined(); - }); - it('sends benchmark output to onProgress when benchmarking', async () => { const benchmarkPath = path.join(tmp, 'benchmark.json'); const configPath = path.join(tmp, '.benchmark-config.json'); diff --git a/src/shared/__tests__/run-tags.test.ts b/src/agent/__tests__/run-tags.test.ts similarity index 97% rename from src/shared/__tests__/run-tags.test.ts rename to src/agent/__tests__/run-tags.test.ts index 9a01156bb..8211b96e8 100644 --- a/src/shared/__tests__/run-tags.test.ts +++ b/src/agent/__tests__/run-tags.test.ts @@ -1,4 +1,4 @@ -import { buildRunTags } from '@shared/run-tags'; +import { buildRunTags } from '@agent/agent-interface'; import { CallType } from '@shared/constants'; describe('buildRunTags', () => { diff --git a/src/programs/__tests__/variant-gating.test.ts b/src/agent/__tests__/variant-gating.test.ts similarity index 89% rename from src/programs/__tests__/variant-gating.test.ts rename to src/agent/__tests__/variant-gating.test.ts index b0d6ed00e..7efd2ca49 100644 --- a/src/programs/__tests__/variant-gating.test.ts +++ b/src/agent/__tests__/variant-gating.test.ts @@ -1,4 +1,4 @@ -import { isOrchestratorEnabled } from '@programs/experiments'; +import { isOrchestratorEnabled } from '@agent/runner/switchboard'; describe('isOrchestratorEnabled', () => { it('is true only when the wizard-orchestrator flag is true', () => { diff --git a/src/agent/agent-interface.ts b/src/agent/agent-interface.ts index 2397afdb3..005315236 100644 --- a/src/agent/agent-interface.ts +++ b/src/agent/agent-interface.ts @@ -28,18 +28,16 @@ import { type AdditionalFeature, ADDITIONAL_FEATURE_PROMPTS, } from '@shared/constants'; -import type { - AgentFailure, - InferenceAuthProvider, -} from './runner/shared/types'; +import type { AgentFailure } from './runner/shared/types'; import type { AgentResult } from './runner/harness/types'; import { createCustomHeaders } from '@utils/custom-headers'; import type { HostResolution } from '@shared/host-resolution'; import { buildWizardPropertiesBlob, + gatewayAuth, isPastRefresh, type GatewayAuth, -} from '@shared/gateway-auth'; +} from '@agent/gateway-session'; import { evaluateBashCommand } from './bash-fence'; import { createWizardToolsServer, WIZARD_TOOL_NAMES } from '@agent/tools'; import { @@ -213,10 +211,6 @@ export type AgentConfig = { * another program's budget, so neither should depend on an optional string bag. */ programId: string; - /** Program-owned inference auth, refreshed at each model call. */ - inferenceAuth: InferenceAuthProvider; - /** Program-owned guidance supplied as data, never looked up here. */ - programCommandments?: readonly string[]; /** Program identifier — selects the model for that program. */ integrationLabel?: string; /** @@ -361,8 +355,8 @@ type AgentRunConfig = { * bearer. */ refreshGatewayAuth?: () => Promise; - /** Program-owned guidance supplied as data. */ - programCommandments?: readonly string[]; + /** Program id, for the program-axis commandments. */ + program?: string; /** Resolved sequence, for the sequence-axis commandments. */ sequence: Sequence; /** Where the run reports. A no-op when the caller passed none. */ @@ -371,6 +365,31 @@ type AgentRunConfig = { const NO_PROGRESS: ProgressEmitter = () => undefined; +/** + * Global identifiers attached to every LLM gateway trace for a run. They ride on + * each `$ai_generation` the gateway emits (in the `X-PostHog-Properties` blob + * `buildAgentEnv` builds), so traces are filterable by program, framework, run, + * and build type for cost attribution and dashboards. `skill_id` is omitted when + * the run has none. + */ +export function buildRunTags(args: { + programId: string; + integration: string; + runId: string; + build: string; + skillId?: string; +}): Record { + return { + program_id: args.programId, + integration: args.integration, + run_id: args.runId, + build: args.build, + // Triage and detection spread these tags and override this one. + call_type: CallType.agent, + ...(args.skillId ? { skill_id: args.skillId } : {}), + }; +} + /** * Whether Warlock/YARA scanning is disabled for this run. Off by default: * scanning is disabled only by the local POSTHOG_WIZARD_WARLOCK_DISABLED env @@ -524,10 +543,13 @@ export async function initializeAgent( const emit = config.emit ?? NO_PROGRESS; try { - // Configure model routing with the program-supplied gateway bearer. + // Configure model routing (inherited by the SDK subprocess). All model + // calls route through the PostHog AI gateway with the scoped token + // gatewayAuth mints for this run. // Disable experimental betas (like input_examples) the gateway doesn't support. process.env.CLAUDE_CODE_DISABLE_EXPERIMENTAL_BETAS = 'true'; - const currentGatewayAuth = () => config.inferenceAuth.resolve(); + const currentGatewayAuth = () => + gatewayAuth(config.host, config.posthogApiKey, config.programId); const auth = await currentGatewayAuth(); const gatewayUrl = auth.gatewayUrl; process.env.ANTHROPIC_BASE_URL = gatewayUrl; @@ -647,7 +669,7 @@ export async function initializeAgent( triageProvider, gatewayAuth: auth, refreshGatewayAuth: currentGatewayAuth, - programCommandments: config.programCommandments, + program: config.integrationLabel, // A queue context is present only on a task run; that is the sequence. sequence: config.orchestrator ? Sequence.orchestrator : Sequence.linear, emit, @@ -1147,7 +1169,7 @@ export async function runAgent( // we keep default Claude Code behaviors. An orchestrator context is // present only on a task run — that is what picks the sequence. append: assembleCommandments({ - programCommandments: agentConfig.programCommandments, + program: agentConfig.program, sequence: agentConfig.sequence, harness: Harness.anthropic, }), diff --git a/src/agent/agent-runner.ts b/src/agent/agent-runner.ts index 002c37a5f..06a907fab 100644 --- a/src/agent/agent-runner.ts +++ b/src/agent/agent-runner.ts @@ -1,10 +1,13 @@ /** * Re-export shim. The runner has been split into agent/runner/. * Import from there directly; this shim keeps existing importers working. + * The session-driven `runProgramAgent(programConfig, session)` lives in + * `src/programs/run-agent-legacy.ts`. */ export { runAgent, + shouldDisableAsk, type AgentRunDefinition, type BootstrapResult, type AbortCase, diff --git a/src/agent/default-binding.ts b/src/agent/default-binding.ts deleted file mode 100644 index 7bad35666..000000000 --- a/src/agent/default-binding.ts +++ /dev/null @@ -1,10 +0,0 @@ -import { GPT5_6_SOL_MODEL, Harness, Sequence } from '@shared/constants'; -import type { ResolvedBinding } from './runner/shared/types'; - -/** Standalone callers can supply this resolved route without a program registry. */ -export const DEFAULT_AGENT_BINDING: ResolvedBinding = { - sequence: Sequence.linear, - harness: Harness.pi, - model: GPT5_6_SOL_MODEL, - thinkingLevel: 'medium', -}; diff --git a/src/programs/gateway-session.ts b/src/agent/gateway-session.ts similarity index 69% rename from src/programs/gateway-session.ts rename to src/agent/gateway-session.ts index 8dc778407..e8a1ef1b2 100644 --- a/src/programs/gateway-session.ts +++ b/src/agent/gateway-session.ts @@ -6,13 +6,30 @@ * unattributed money to hide an outage. */ +import { readFileSync } from 'node:fs'; import { logToFile } from '@utils/debug'; import { analytics } from '@utils/analytics'; import { ErrorCodes, WizardError } from '@shared/errors'; import type { HostResolution } from '@shared/host-resolution'; import { checkLlmGatewayHealth } from '@shared/health-checks/endpoints'; import { ServiceHealthStatus } from '@shared/health-checks/types'; -import { isTrustedGatewayUrl, type GatewayAuth } from '@shared/gateway-auth'; +import { IS_PRODUCTION_BUILD, runtimeEnv } from '@env'; +import type { CloudRegion } from '@utils/types'; + +export interface GatewayAuth { + /** Base URL for model calls (no `/v1`; transports append their route). */ + gatewayUrl: string; + /** Gateway bearer, minted normally or supplied directly by CI. */ + token: string; + /** Team verified by the mint, or explicitly supplied for CI attribution. */ + teamId?: number; + /** + * Instant past which a 401 on this bearer is age rather than a bad + * credential: the cache re-mints past it, and a session still holding the + * old bearer may re-mint once. Before it the mint has to be trusted. + */ + refreshAtMs: number; +} interface CachedAuth { key: string; @@ -27,6 +44,54 @@ let cached: CachedAuth | null = null; * task at once, and each would otherwise take its own token and its own cap. */ let inFlight: { key: string; promise: Promise } | null = null; +let ciAuth: GatewayAuth | null = null; + +// Snapshot CI supplies a gateway bearer without minting or re-minting. +export function configureGatewayCredentialsForCI( + token: string, + projectId: number, + gatewayUrl: string, +): void { + if (IS_PRODUCTION_BUILD) + throw new Error('CI gateway auth requires a non-production build'); + if (!token.trim() || !Number.isSafeInteger(projectId) || projectId <= 0) { + throw new Error('CI gateway auth requires a token and valid project ID'); + } + if ( + !/^https?:\/\//.test(gatewayUrl) || + !isTrustedGatewayUrl(gatewayUrl, '') + ) { + throw new Error('CI gateway auth requires a trusted gateway origin'); + } + resetGatewaySession(); + ciAuth = { + token: token.trim(), + teamId: projectId, + gatewayUrl: gatewayUrl.replace(/\/+$/, ''), + refreshAtMs: Infinity, + }; +} + +// TODO(B2): CI credential loading belongs to the headless provider, not the +// agent. Leaves with the rest of this module once RunInput carries resolved +// inference auth. +export function configureGatewayFromCIEnvironment( + projectId: number, + region: CloudRegion, +): void { + if (IS_PRODUCTION_BUILD) + throw new Error('CI gateway auth requires a non-production build'); + const path = runtimeEnv('WIZARD_CI_GATEWAY_TOKEN_FILE'); + if (!path) throw new Error('WIZARD_CI_GATEWAY_TOKEN_FILE is required for CI'); + const token = readFileSync(path, 'utf8'); + delete process.env.WIZARD_CI_GATEWAY_TOKEN_FILE; + configureGatewayCredentialsForCI( + token, + projectId, + runtimeEnv('WIZARD_CI_GATEWAY_URL') || + `https://ai-gateway.${region}.posthog.com`, + ); +} /** * Adoption floor. The anthropic subprocess holds its credential until a 401 @@ -49,6 +114,7 @@ export async function gatewayAuth( accessToken: string, program: string | undefined, ): Promise { + if (ciAuth) return ciAuth; // Keyed by program: a token pins `wizard:`, so reusing one across // programs bills the wrong budget. const key = `${host.apiHost}\n${accessToken}\n${program ?? ''}`; @@ -122,6 +188,55 @@ async function resolveGatewayAuth( export function resetGatewaySession(): void { cached = null; inFlight = null; + ciAuth = null; +} + +/** Whether a 401 on this bearer may be age (past its refresh instant) rather than a bad credential. */ +export function isPastRefresh(auth: GatewayAuth, now = Date.now()): boolean { + return now >= auth.refreshAtMs; +} + +/** + * Whether a server-supplied origin may receive a bearer and prompt content: + * https (loopback excepted), and either a current cloud gateway or the host the run + * authenticated against. + */ +export function isTrustedGatewayUrl(value: string, apiHost: string): boolean { + let url: URL; + try { + url = new URL(value); + } catch { + return false; + } + // Consumers append routes to this value, so anything beyond an origin + // (path, query, fragment, userinfo) would build a malformed endpoint. + if ( + url.pathname !== '/' || + url.search || + url.hash || + url.username || + url.password + ) { + return false; + } + const localhost = + url.hostname === 'localhost' || + url.hostname === '127.0.0.1' || + url.hostname === 'host.docker.internal'; + // Loopback is the dev gateway, and is the one case allowed over http. + if (localhost) return true; + if (url.protocol !== 'https:') return false; + if (url.hostname.endsWith('.posthog.com')) { + return ( + url.origin === 'https://ai-gateway.us.posthog.com' || + url.origin === 'https://ai-gateway.eu.posthog.com' + ); + } + try { + return url.hostname === new URL(apiHost).hostname; + } catch { + return false; + } } interface MintedToken { @@ -335,3 +450,38 @@ async function mintGatewayToken( ); } } + +/** + * The v2 run-metadata carrier: one JSON blob for the `X-PostHog-Properties` + * header. Plain keys only, since the gateway strips `$`-prefixed keys as reserved, + * so feature-flag variants land as `wizard_flag_` instead of the legacy + * `$feature/` (dashboards keying on `$feature/wizard-*` read the new key + * post-cutover). + */ +export function buildWizardPropertiesBlob( + wizardMetadata: Record, + wizardFlags: Record, + teamId?: number, +): string { + // The gateway pins `$ai_product` to `wizard:`, and rejects a legacy + // product override on a scoped token, so the unprefixed key every cost and + // error consumer reads is only present if this blob declares it. + const props: Record = { ai_product: 'wizard' }; + if (teamId !== undefined) props.team_id = teamId; + for (const [key, value] of Object.entries(wizardMetadata)) { + props[stripPropertyPrefix(key)] = value; + } + for (const [flagKey, variant] of Object.entries(wizardFlags)) { + if (!flagKey.toLowerCase().startsWith('wizard')) continue; + props[`wizard_flag_${flagKey.toLowerCase()}`] = variant; + } + return JSON.stringify(props); +} + +const LEGACY_PROPERTY_PREFIX = 'X-POSTHOG-PROPERTY-'; + +function stripPropertyPrefix(key: string): string { + return key.toUpperCase().startsWith(LEGACY_PROPERTY_PREFIX) + ? key.slice(LEGACY_PROPERTY_PREFIX.length).toLowerCase() + : key; +} diff --git a/src/agent/index.ts b/src/agent/index.ts index 232f8384f..dbafd6b3b 100644 --- a/src/agent/index.ts +++ b/src/agent/index.ts @@ -8,45 +8,41 @@ * Grouped by fate, per the stack plan (sections 4.1 to 4.5 and 7). */ -import type { RunAgentOptions, RunConfig, RunInput, RunResult } from './types'; - /** * Stays. The agent's contract: the one way to run it, the marker strings - * program prompts embed, the tool ids programs put in allowedTools and - * disallowedTools, and what programs resolve a binding with: the default - * binding, the harness axis and each harness's task capability. + * program prompts embed, and the tool ids programs put in allowedTools and + * disallowedTools. */ export type * from './types'; -export { RunOutcome } from './runner/shared/types'; -export { AgentSignals } from './signals'; -export { OutroKind } from './progress'; -export { WIZARD_TOOL_NAMES } from './tools/tool-names'; -export { DEFAULT_AGENT_BINDING } from './default-binding'; -export { - harnessRunsTasks, - resolveHarness, -} from './runner/switchboard/resolve-harness'; - -/** Runs one agent pipeline; the runner loads on the first call. */ -export async function runAgent( - config: RunConfig, - input: RunInput, - options?: RunAgentOptions, -): Promise { - const runner = await import('./runner'); - return runner.runAgent(config, input, options); -} +export { runAgent, RunOutcome } from './runner'; +export { AgentSignals } from './agent-interface'; +export { WIZARD_TOOL_NAMES } from './tools'; -/** Leaves in C3 (M16, then D12), once skill install becomes shared. The installer loads on the first call. */ -export async function downloadSkill( - ...args: Parameters -): ReturnType { - const tools = await import('./tools/tools'); - return tools.downloadSkill(...args); -} +/** + * Leaves in B2 (B1 deferred the bindings). Bindings and program data move to + * programs: resolveBinding + * is keyed by PROGRAM_BINDINGS and the agent keeps only "run from an + * already-resolved binding"; shouldDisableAsk is a flags policy programs + * decide and pass in; LONGER_ASK_TIMEOUT_MS is a tuning number programs own + * as askTimeoutMs. + */ +export { resolveBinding, shouldDisableAsk } from './runner'; +export { LONGER_ASK_TIMEOUT_MS } from './wizard-ask-bridge'; +/** + * Leaves in B2. Programs own credentials and the legacy adapter dies. + * buildRunTags builds the trace tags runProgram and agentic detection send. + * configureGatewayFromCIEnvironment is CI inference auth the headless provider + * owns. flushScanReport becomes a progress event rather than a call. + * downloadSkill leaves once the skill scan runs at load and skill install + * becomes shared. + */ +export { buildRunTags } from './agent-interface'; +export { configureGatewayFromCIEnvironment } from './gateway-session'; +export { flushScanReport } from './yara-hooks'; +export { downloadSkill } from './tools'; /** The frameworkContext slot the legacy adapter fills for the e2e harness. */ -export { TASK_OUTCOMES_KEY } from './runner/shared/types'; +export { TASK_OUTCOMES_KEY } from './runner'; /** * Leaves in C2. The TUI receives agent data through program state. Until diff --git a/src/agent/mcp-prompt-streaming.ts b/src/agent/mcp-prompt-streaming.ts index b4efae653..2c1586550 100644 --- a/src/agent/mcp-prompt-streaming.ts +++ b/src/agent/mcp-prompt-streaming.ts @@ -13,11 +13,10 @@ */ import type { Credentials } from '@shared/api'; -import type { InferenceAuthProvider } from '@agent/types'; import { DEFAULT_AGENT_MODEL, WIZARD_USER_AGENT } from '@shared/constants'; import { logToFile } from '@utils/debug'; -import { buildAgentEnv } from '@agent/agent-interface'; -import { buildRunTags } from '@shared/run-tags'; +import { gatewayAuth } from '@agent/gateway-session'; +import { buildAgentEnv, buildRunTags } from '@agent/agent-interface'; import { sanitizeAgentSubprocessEnv } from '@shared/agent-env-isolation'; import { createIsolatedAgentConfigDir } from '@agent/stored-login'; import { analytics } from '@utils/analytics'; @@ -215,7 +214,6 @@ export function buildTutorialRunTags(args: { export async function* runMcpPromptViaSdk(args: { prompt: string; credentials: Credentials; - inferenceAuth: InferenceAuthProvider; signal: AbortSignal; /** When set, the SDK loads the named session's prior turns as * context so the follow-up prompt can reference what the agent @@ -244,7 +242,11 @@ export async function* runMcpPromptViaSdk(args: { // The url and the bearer are one unit: a run must take both from the same // mint. - const auth = await args.inferenceAuth.resolve(); + const auth = await gatewayAuth( + credentials.host, + credentials.accessToken, + args.programId, + ); const gatewayUrl = auth.gatewayUrl; process.env.ANTHROPIC_BASE_URL = gatewayUrl; process.env.ANTHROPIC_AUTH_TOKEN = auth.token; diff --git a/src/agent/progress.ts b/src/agent/progress.ts index b0d4627ca..24f6a7559 100644 --- a/src/agent/progress.ts +++ b/src/agent/progress.ts @@ -4,7 +4,7 @@ * `runAgent` reports through one optional callback and asks through one * optional set of capabilities. Neither reaches into a UI singleton, a store, * or a session: every payload is copied data, every question is awaited on an - * injected answerer. The legacy adapter in `src/lib/runners/run-program-agent.ts` + * injected answerer. The legacy adapter in `src/programs/run-agent-legacy.ts` * maps these back onto `WizardUI` one call per event, so the terminal output of * every existing runner is unchanged. */ diff --git a/src/programs/__tests__/switchboard.test.ts b/src/agent/runner/__tests__/switchboard.test.ts similarity index 97% rename from src/programs/__tests__/switchboard.test.ts rename to src/agent/runner/__tests__/switchboard.test.ts index d6efea02d..0a16a2048 100644 --- a/src/programs/__tests__/switchboard.test.ts +++ b/src/agent/runner/__tests__/switchboard.test.ts @@ -2,7 +2,7 @@ * Switchboard machinery tests: binding registry lockstep, precedence chains, * trace stamping, model capabilities, and structural clamps. Per-experiment * flag behavior and cross-program isolation live in one file per experiment - * under `experiments/__tests__/`. + * under `switchboard/flags/__tests__/`. * * Every resolution test is a BindingCase: (SwitchboardCtx in) → (full * four-axis resolveBinding out), optionally pinning the trace. @@ -24,10 +24,10 @@ import { } from '@shared/constants'; import { PROGRAM_BINDINGS, - resolveProgramBinding as resolveBinding, -} from '../binding'; -import type { ProgramSwitchboardCtx as SwitchboardCtx } from '@programs/types'; -import { DEFAULT_AGENT_BINDING as DEFAULT_BINDING } from '@agent'; + DEFAULT_BINDING, + resolveBinding, + type SwitchboardCtx, +} from '@agent/runner/switchboard'; import { modelCapabilities, MINT_ALLOWED_EFFORTS, @@ -36,7 +36,7 @@ import { TRIAGE_MODELS, VALID_MODELS, } from '@agent/runner/switchboard/models'; -import { runBindingCases } from '@programs/experiments/__tests__/binding-cases'; +import { runBindingCases } from '@agent/runner/switchboard/flags/__tests__/binding-cases'; const PROGRAM_IDS = PROGRAM_REGISTRY.map((c) => c.id); const DEFAULT_RESOLVED = { diff --git a/src/agent/runner/harness/anthropic/__tests__/pending-question.test.ts b/src/agent/runner/harness/anthropic/__tests__/pending-question.test.ts index 1c0608775..86ff59c3f 100644 --- a/src/agent/runner/harness/anthropic/__tests__/pending-question.test.ts +++ b/src/agent/runner/harness/anthropic/__tests__/pending-question.test.ts @@ -44,7 +44,7 @@ async function initializeHarness( sequence: Sequence.linear, model: 'test', }, - programCommandments: ['Follow the program rule'], + switchboard: { program: 'test', flags: {} }, skillsBaseUrl: 'https://skills.test', wizardFlags: {}, wizardFlagPayloads: {}, @@ -64,14 +64,6 @@ async function initializeHarness( }, host: {}, credentials, - inferenceAuth: { - resolve: () => - Promise.resolve({ - gatewayUrl: 'https://ai-gateway.us.posthog.com', - token: 'phe_test', - refreshAtMs: Infinity, - }), - }, project: null, apiUser: null, }, @@ -121,14 +113,6 @@ async function initializeHarness( }; } -it('forwards the supplied program commandments on both entry points', async () => { - for (const mode of ['linear', 'task'] as const) { - await initializeHarness(mode, undefined); - const [config] = vi.mocked(initializeAgent).mock.calls.at(-1)!; - expect(config.programCommandments).toEqual(['Follow the program rule']); - } -}); - afterEach(() => { vi.useRealTimers(); vi.clearAllMocks(); diff --git a/src/agent/runner/harness/anthropic/index.ts b/src/agent/runner/harness/anthropic/index.ts index af57af45b..72ca937f2 100644 --- a/src/agent/runner/harness/anthropic/index.ts +++ b/src/agent/runner/harness/anthropic/index.ts @@ -58,8 +58,6 @@ export const anthropicBackend: AgentHarness = { wizardFlags, wizardMetadata, programId: boot.programId, - inferenceAuth: input.inferenceAuth, - programCommandments: runConfig.programCommandments, integrationLabel: config.integrationLabel, askBridge, getPendingQuestion: askBridge?.getPendingQuestion, @@ -139,8 +137,6 @@ export const anthropicBackend: AgentHarness = { detectPackageManager: detectNodePackageManagers, skillsBaseUrl: boot.skillsBaseUrl, programId: boot.programId, - inferenceAuth: input.inferenceAuth, - programCommandments: config.programCommandments, wizardFlags: boot.wizardFlags, wizardMetadata: boot.wizardMetadata, integrationLabel: config.programId, diff --git a/src/agent/runner/harness/pi/__tests__/gateway.test.ts b/src/agent/runner/harness/pi/__tests__/gateway.test.ts index 6f896c293..548e77bb6 100644 --- a/src/agent/runner/harness/pi/__tests__/gateway.test.ts +++ b/src/agent/runner/harness/pi/__tests__/gateway.test.ts @@ -5,7 +5,7 @@ import { withGatewayRemint, GATEWAY_PROVIDER, } from '../gateway'; -import type { GatewayAuth } from '@shared/gateway-auth'; +import type { GatewayAuth } from '@agent/gateway-session'; describe('buildGatewayProvider effort', () => { const base = { diff --git a/src/agent/runner/harness/pi/gateway.ts b/src/agent/runner/harness/pi/gateway.ts index a576f7812..bb5b7e20a 100644 --- a/src/agent/runner/harness/pi/gateway.ts +++ b/src/agent/runner/harness/pi/gateway.ts @@ -10,7 +10,7 @@ import { buildWizardPropertiesBlob, isPastRefresh, type GatewayAuth, -} from '@shared/gateway-auth'; +} from '@agent/gateway-session'; import { modelCapabilities, type ThinkingLevel, diff --git a/src/agent/runner/harness/pi/index.ts b/src/agent/runner/harness/pi/index.ts index f16d2abdc..90ea3ac54 100644 --- a/src/agent/runner/harness/pi/index.ts +++ b/src/agent/runner/harness/pi/index.ts @@ -26,7 +26,7 @@ import { AgentErrorType } from '@agent/agent-interface'; import { AgentSignals, REMARK_INSTRUCTION } from '@agent/signals'; import { AgentOutputSignals } from '@agent/output-signals'; import { assembleCommandments } from '../../switchboard/commandments'; -import type { GatewayAuth } from '@shared/gateway-auth'; +import { gatewayAuth, type GatewayAuth } from '@agent/gateway-session'; import { buildGatewayProvider, GATEWAY_PROVIDER, @@ -285,9 +285,14 @@ export const piBackend: AgentHarness = { } = await import('@earendil-works/pi-coding-agent'); // the claude-agent-sdk path. The provider spec is shared with the - // orchestrator's per-task sessions (gateway.ts). Programs supply the - // run's inference auth provider. - const refreshAuth = () => input.inferenceAuth.resolve(); + // orchestrator's per-task sessions (gateway.ts). gatewayAuth mints the + // run's scoped token. + const refreshAuth = () => + gatewayAuth( + boot.credentials.host, + boot.credentials.accessToken, + boot.programId, + ); const auth = await refreshAuth(); const providerInputs = (current: GatewayAuth) => ({ gatewayUrl: current.gatewayUrl, @@ -389,7 +394,7 @@ export const piBackend: AgentHarness = { agentDir: getAgentDir(), systemPrompt: assembleCommandments({ - programCommandments: runConfig.programCommandments, + program: runConfig.programId, sequence: Sequence.linear, harness: Harness.pi, caps: { bash: true, posthogMcp }, diff --git a/src/agent/runner/harness/pi/task.ts b/src/agent/runner/harness/pi/task.ts index be53e7460..290950f40 100644 --- a/src/agent/runner/harness/pi/task.ts +++ b/src/agent/runner/harness/pi/task.ts @@ -35,7 +35,7 @@ import { AgentOutputSignals } from '@agent/output-signals'; import { TaskStatus } from '../../sequence/orchestrator/queue'; import type { OrchestratorToolsContext } from '../../sequence/orchestrator/queue-tools'; import type { AgentResult, TaskRunInputs } from '../types'; -import type { GatewayAuth } from '@shared/gateway-auth'; +import { gatewayAuth, type GatewayAuth } from '@agent/gateway-session'; import { buildGatewayProvider, GATEWAY_PROVIDER, @@ -250,7 +250,12 @@ export async function runPiTask(inputs: TaskRunInputs): Promise { createWriteToolDefinition, } = sdk; - const refreshAuth = () => input.inferenceAuth.resolve(); + const refreshAuth = () => + gatewayAuth( + boot.credentials.host, + boot.credentials.accessToken, + boot.programId, + ); const auth = await refreshAuth(); const providerInputs = (current: GatewayAuth) => ({ gatewayUrl: current.gatewayUrl, @@ -335,7 +340,7 @@ export async function runPiTask(inputs: TaskRunInputs): Promise { cwd: input.installDir, agentDir: getAgentDir(), systemPrompt: assembleCommandments({ - programCommandments: config.programCommandments, + program: config.programId, sequence: Sequence.orchestrator, harness: Harness.pi, caps: { bash: codingTools.has('bash'), posthogMcp }, diff --git a/src/agent/runner/index.ts b/src/agent/runner/index.ts index df43d548c..71819588a 100644 --- a/src/agent/runner/index.ts +++ b/src/agent/runner/index.ts @@ -7,18 +7,20 @@ * snapshot (directory, credentials, project); `options` carries an optional * progress observer, an optional answerer. * - * The pipeline prepares the run (logging, supplied inference auth, scan triage), + * The pipeline prepares the run (logging targets, gateway mint, scan triage), * then forks. The `orchestrator` variant routes to the task-queue runner. * Every other variant runs the fixed linear pipeline: * [skill install] → agent init → prompt → run → errors → [postRun] → outro * - * The agent reports and asks; it never renders, reads a session, exits, or - * sends the process's terminal analytics. - * Coded errors return Failed, uncoded throws return Crashed, and both retain - * the caught error. Scan-report flushing is best effort after the result is - * decided. A program run reaches this call through programs' `runProgram`, - * which builds the config and input. Agentic detection and a standalone caller - * build them and call it directly. + * The agent reports and asks, it never renders, never reads a session, never + * exits the process, never sends the process's terminal analytics and never + * rejects. A decided failure comes back in + * `RunResult.failure` with the same fields `wizardAbort` takes; an error the + * agent did not decide (a refused mint, an SDK crash) comes back as + * `outcome: RunOutcome.Crashed` with the original error attached, so a caller can keep + * handling it the way it always did. The legacy adapter in + * `src/programs/run-agent-legacy.ts` rebuilds today's session-driven + * behavior on top of this call for every existing caller. */ import { Sequence } from '@shared/constants'; @@ -40,17 +42,12 @@ import { } from './shared/transcript-tail'; import { getSequence } from './switchboard'; import { flushScanReport } from '@agent/yara-hooks'; -import type { ProgressEmitter } from '@agent/progress'; -import { captureRunSkillCleanup } from '@shared/skill-run-cleanup'; -import { registerCleanup } from '@utils/cleanup-registry'; -import { hostAborted } from './shared/errors'; export type { AbortCase, AgentFailure, BootstrapResult, Credentials, - InferenceAuthProvider, AgentRunDefinition, PromptContext, ResolvedBinding, @@ -68,6 +65,8 @@ export type { AgentProgress, ProgressEmitter, } from '@agent/progress'; +export { shouldDisableAsk } from './shared/bootstrap'; +export { resolveBinding } from './switchboard'; export type { ProgramBinding, SwitchboardCtx } from './switchboard'; export { TASK_OUTCOMES_KEY } from './sequence/orchestrator/queue'; export type { TaskOutcome } from './sequence/orchestrator/queue'; @@ -85,20 +84,7 @@ export async function runAgent( options: RunAgentOptions = {}, ): Promise { let collector: ReturnType | undefined; - let scanReport: { flush(): void } | undefined; let transcript: TranscriptTail | undefined; - let cleanupInstalledSkills: (() => void) | undefined; - const cleanFailedRun = () => { - try { - cleanupInstalledSkills?.(); - } catch (error) { - try { - logToFile('[agent-runner] failed-run skill cleanup error:', error); - } catch { - // Cleanup diagnostics must not replace the run result. - } - } - }; const snapshot = (): RunResult['snapshot'] => { try { if (collector) { @@ -121,37 +107,35 @@ export async function runAgent( }, }; }; - const settle = (result: RunResult): RunResult => { - if (result.outcome !== RunOutcome.Success) cleanFailedRun(); + const flushReport = (): void => { + // A deferred report keeps counting this run's scans toward the host run's. + if (config.scanReport === 'defer') return; try { - // A deferred report keeps counting this run's scans toward the host run's. - scanReport?.flush(); + const report = flushScanReport({ yaraReport: input.flags.yaraReport }); + if (report) + collector?.emit({ kind: 'log', level: 'info', message: report }); } catch { // Scan reporting is best effort after the run outcome is decided. } - return result; }; let result: RunResult; try { - // The report line reaches the collector once it exists; a drain cannot run before that. - if (config.scanReport !== 'defer') { - scanReport = armScanReportFlush(input.flags.yaraReport, (event) => - collector?.emit(event), - ); - } - // Capture before preparation so pre-harness failures clean new skills too. - cleanupInstalledSkills = captureRunSkillCleanup(input.installDir); collector = createProgressCollector(options.onProgress); const { emit } = collector; if (config.run.collectTranscript) transcript = createTranscriptTail(emit); const log = (message: string) => emit({ kind: 'log', level: 'info', message }); if (options.signal?.aborted) { - return settle({ - ...hostAborted(), + flushReport(); + return { + outcome: RunOutcome.Aborted, skillId: input.skillId, + failure: { + code: ErrorCodes.AgentAbort, + message: 'Agent run cancelled', + }, snapshot: snapshot(), - }); + }; } const boot = await prepareRun(config, input); if (config.binding.sequence === Sequence.orchestrator) { @@ -175,7 +159,16 @@ export async function runAgent( transcript, }); result = { - ...(options.signal?.aborted ? hostAborted() : sequenceResult), + ...(options.signal?.aborted && + sequenceResult.outcome === RunOutcome.Success + ? { + outcome: RunOutcome.Aborted as const, + failure: { + code: ErrorCodes.AgentAbort, + message: 'Agent run cancelled', + }, + } + : sequenceResult), skillId: input.skillId, snapshot: snapshot(), }; @@ -197,13 +190,13 @@ export async function runAgent( } catch { // Logging is best effort. } - if (options.signal?.aborted) { - result = { - ...hostAborted(), - skillId: input.skillId, - snapshot: snapshot(), - }; - } else if (failure.coded) { + let abortError = false; + try { + abortError = original.name === 'AbortError'; + } catch { + /* Hostile Error getter. */ + } + if (failure.coded) { result = { outcome: RunOutcome.Failed, skillId: input?.skillId, @@ -214,6 +207,16 @@ export async function runAgent( }, snapshot: snapshot(), }; + } else if (options.signal?.aborted && abortError) { + result = { + outcome: RunOutcome.Aborted, + skillId: input?.skillId, + failure: { + code: ErrorCodes.AgentAbort, + message: 'Agent run cancelled', + }, + snapshot: snapshot(), + }; } else { result = { outcome: RunOutcome.Crashed, @@ -227,26 +230,8 @@ export async function runAgent( }; } } - return settle(result); -} - -/** - * Write the scan report and emit its line once: from the run's own tail, or - * earlier when a process drain (wizardAbort, a signal handler) runs cleanups. - */ -function armScanReportFlush( - yaraReport: boolean, - emit: ProgressEmitter, -): { flush(): void } { - let flushed = false; - const flush = () => { - if (flushed) return; - flushed = true; - const report = flushScanReport({ yaraReport }); - if (report) emit({ kind: 'log', level: 'info', message: report }); - }; - registerCleanup(flush); - return { flush }; + flushReport(); + return result; } function safeErrorMessage(error: unknown): string { diff --git a/src/agent/runner/sequence/linear.ts b/src/agent/runner/sequence/linear.ts index 599c1516f..4a7e1317b 100644 --- a/src/agent/runner/sequence/linear.ts +++ b/src/agent/runner/sequence/linear.ts @@ -21,10 +21,9 @@ import { formatYaraAbortMessage } from '@agent/yara-hooks'; import { installSkillById } from '@agent/tools'; import { assemblePrompt, type PromptContext } from '../../agent-prompt'; import type { SequenceResult, SequenceContext } from '../shared/types'; -import { failed, hostAborted, installFailure } from '../shared/errors'; +import { failed, installFailure } from '../shared/errors'; import { RunOutcome } from '../shared/types'; -import { runOptions } from '../shared/bootstrap'; -import { isAskDisabled } from '@shared/ask-policy'; +import { shouldDisableAsk, runOptions } from '../shared/bootstrap'; import { createEmitSpinner } from '../shared/progress-collector'; import { createAskBridge } from '../shared/ask'; import { withTranscript } from '../shared/transcript-tail'; @@ -62,7 +61,11 @@ async function executeLinear( const { run, composed } = config; const { skillsBaseUrl, credentials, project } = boot; const { projectApiKey, host, projectId } = credentials; - if (signal?.aborted) return hostAborted(); + const aborted = (): SequenceResult => ({ + outcome: RunOutcome.Aborted, + failure: { code: ErrorCodes.AgentAbort, message: 'Agent run cancelled' }, + }); + if (signal?.aborted) return aborted(); // 5. Skill install (if skillId provided) let skillPath: string | undefined; @@ -74,7 +77,7 @@ async function executeLinear( skillsBaseUrl, { triage: boot.triageProvider }, ); - if (signal?.aborted) return hostAborted(); + if (signal?.aborted) return aborted(); if (installResult.kind !== 'ok') { return failed(installFailure(run.integrationLabel, installResult)); } @@ -92,7 +95,7 @@ async function executeLinear( // CI/signup with neither has no answerer, so we omit the bridge and the tool // returns an actionable error rather than hanging on a never-resolving prompt. const askDisabled = - isAskDisabled(input.flags) && process.env.WIZARD_ASK_AUTODRIVE !== '1'; + shouldDisableAsk(input.flags) && process.env.WIZARD_ASK_AUTODRIVE !== '1'; const ask = askDisabled ? undefined : createAskBridge(interaction, { @@ -129,7 +132,7 @@ async function executeLinear( ? run.prompt(promptContext) : assemblePrompt(run, promptContext); logToFile(`[agent-runner] prompt assembled (${prompt.length} chars)`); - if (signal?.aborted) return hostAborted(); + if (signal?.aborted) return aborted(); // 8. Run the agent through the run-level harness. The harness owns the agent // loop + model transport; everything around it (skill install, prompt, ask @@ -149,7 +152,7 @@ async function executeLinear( thinkingLevel, signal: runSignal, }); - if (signal?.aborted) return hostAborted(); + if (signal?.aborted && agentResult.kind === 'success') return aborted(); // 9. Error handling (full set from both harnesses) if (agentResult.kind === 'decided_failure') { @@ -295,7 +298,7 @@ async function executeLinear( // 10. Post-run hooks if (config.hooks?.postRun) { await config.hooks.postRun(credentials); - if (signal?.aborted) return hostAborted(); + if (signal?.aborted) return aborted(); } // A composed sub-run leaves the terminal outro to its host. diff --git a/src/agent/runner/sequence/orchestrator/orchestrator-runner.ts b/src/agent/runner/sequence/orchestrator/orchestrator-runner.ts index d3d6f6fc4..a5d42f8ab 100644 --- a/src/agent/runner/sequence/orchestrator/orchestrator-runner.ts +++ b/src/agent/runner/sequence/orchestrator/orchestrator-runner.ts @@ -10,7 +10,7 @@ * task resolve to a prompt fetched at startup into the registry. The wizard side * stays product-ignorant: it is the queue, the executor, and the loader. */ -import { failed, hostAborted } from '../../shared/errors'; +import { failed } from '../../shared/errors'; import { RunOutcome } from '../../shared/types'; import { randomUUID } from 'crypto'; import { @@ -43,8 +43,10 @@ import type { import { createEmitSpinner } from '../../shared/progress-collector'; import { createAskBridge } from '../../shared/ask'; import { + areSeededTasksEnabled, getHarness, - resolveRoleHarness, + resolveHarness, + resolveStageOverrides, type HarnessPick, } from '../../switchboard'; import { isValidModel, requireKnownModel } from '../../switchboard/models'; @@ -66,7 +68,8 @@ import { import { RunMetrics } from './run-metrics'; import { dependencyClosure, uncoveredBySink } from './queue-tools'; import { deferSeededTasks } from './seeded-deps'; -import { isAskDisabled, LONGER_ASK_TIMEOUT_MS } from '@shared/ask-policy'; +import { LONGER_ASK_TIMEOUT_MS } from '@agent/wizard-ask-bridge'; +import { shouldDisableAsk } from '../../shared/bootstrap'; import { agentRunTools, assembleSeedPrompt, @@ -186,6 +189,13 @@ function terminalResult( } } +function cancelledRun(): SequenceResult { + return { + outcome: RunOutcome.Aborted, + failure: { code: ErrorCodes.AgentAbort, message: 'Agent run cancelled' }, + }; +} + /** Every skill entry the menu knows, across categories. */ async function fetchSkillMenuEntries( skillsBaseUrl: string, @@ -593,11 +603,16 @@ async function executeOrchestrator( cleanupQueue: () => void, controller: AbortController, ): Promise { - if (signal?.aborted) return hostAborted(); + if (signal?.aborted) return cancelledRun(); const runId = randomUUID(); const { run } = config; const programId = config.programId; + // Switchboard context — reused for every per-role harness resolution below. + // The caller resolved the run-level binding from it; per-task roles overlay + // `binding.contextMillOverride[role]` on the same inputs. + const switchboardCtx = { ...config.switchboard, trace: undefined }; + // The WHAT (agent prompts) is served from context-mill. Fetch the registry // once up front: its types drive enqueue validation, and resolving a task to // its run config is then synchronous, with no mid-drain network latency. @@ -605,9 +620,13 @@ async function executeOrchestrator( const registry = await loadAgentRegistry(boot.skillsBaseUrl, flow, { exclude: effectiveExcludedTaskTypes(config, boot.wizardFlags), // Baked into the prompts at load, so enqueue, dispatch, and telemetry all read one effective spec. - overrides: config.stageOverrides, + overrides: resolveStageOverrides( + programId, + boot.wizardFlags, + boot.wizardFlagPayloads, + ), }); - if (signal?.aborted) return hostAborted(); + if (signal?.aborted) return cancelledRun(); const seedPrompt = registry.seed; if (!seedPrompt) { throw new Error( @@ -619,7 +638,7 @@ async function executeOrchestrator( const taskModels = Object.fromEntries( ['seed', ...registry.types].map((type) => { const prompt = type === 'seed' ? seedPrompt : registry.get(type); - const pick = resolveRoleHarness(config.binding, type); + const pick = resolveHarness(switchboardCtx, type); const specModel = prompt && promptModelFor(prompt, pick.harness).model; return [type, isValidModel(specModel) ? specModel : pick.model]; }), @@ -644,7 +663,7 @@ async function executeOrchestrator( const store = new QueueStore(input.installDir, runId, { onTransition: (event, task) => { - const pick = resolveRoleHarness(config.binding, task.type); + const pick = resolveHarness(switchboardCtx, task.type); // Mirror dispatch's allow-list fallback so attribution names the model that runs. const specModel = taskModelSpec(registry, task, pick.harness).model; const base = { @@ -715,7 +734,7 @@ async function executeOrchestrator( let commandmentsPath: string | undefined; let referenceInstallPath: string | undefined; const menuSkillEntries = await fetchSkillMenuEntries(boot.skillsBaseUrl); - if (signal?.aborted) return hostAborted(); + if (signal?.aborted) return cancelledRun(); // The framework key for reference + variant resolution. `input.integration` // is the detected framework and always wins; `input.skillId` is the // fallback for the basic-integration path, where the caller sets it to the @@ -736,7 +755,7 @@ async function executeOrchestrator( triage: boot.triageProvider, }, ); - if (signal?.aborted) return hostAborted(); + if (signal?.aborted) return cancelledRun(); if (ref.kind === 'ok') { referenceInstallPath = ref.path; const example = path.join(ref.path, 'references', 'EXAMPLE.md'); @@ -861,7 +880,7 @@ async function executeOrchestrator( // depend on them, and no prompt has to remember they are there. // Kill switch: off (or unset), the wizard queues nothing itself and the run // is byte-identical to a project with no detected sources. - const seedEntries = config.seededTasksEnabled + const seedEntries = areSeededTasksEnabled(boot.wizardFlags) ? config.seedTasks?.() ?? [] : []; const seededTypes: string[] = []; @@ -916,7 +935,7 @@ async function executeOrchestrator( signal, }), ); - if (signal?.aborted) return hostAborted(); + if (signal?.aborted) return cancelledRun(); } logToFile(`[orchestrator] runner-seeded task ${seeded.type}`); } @@ -947,7 +966,7 @@ async function executeOrchestrator( // One bridge for the run, handed only to a task whose prompt allows asking. // Absent in CI and signup, where nobody can answer. - const askBridge = isAskDisabled(input.flags) + const askBridge = shouldDisableAsk(input.flags) ? undefined : createAskBridge(interaction, { signal, @@ -978,7 +997,7 @@ async function executeOrchestrator( // Prompt-frontmatter model wins over the switchboard pick (§3.6 of the // switchboard plan) — the switchboard's model is the fallback when the // prompt is silent. - const seedPick = resolveRoleHarness(config.binding, 'seed'); + const seedPick = resolveHarness(switchboardCtx, 'seed'); const seedHarness = requireTaskHarness(seedPick); const seedModel = promptModelFor(seedPrompt, seedPick.harness); const seedResult = await seedHarness.runTask({ @@ -1007,7 +1026,7 @@ async function executeOrchestrator( requestRemark: false, analyticsProperties: { task_type: 'seed', harness: seedPick.harness }, }); - if (signal?.aborted) return hostAborted(); + if (signal?.aborted) return cancelledRun(); const seedTerminal = terminalResult(seedResult); if (seedTerminal) return seedTerminal; analytics.wizardCapture('orchestrator seeded', { @@ -1188,9 +1207,10 @@ async function executeOrchestrator( // panel shows progress); errors still surface — the harness stops the // spinner with its own error text. // - // Per-task role = task.type. Programs resolved any role override before - // invocation; prompt-frontmatter model still wins (§3.6). - const taskPick = resolveRoleHarness(config.binding, task.type); + // Per-task role = task.type — the switchboard consults + // PROGRAM_BINDINGS[id].contextMillOverride?.[task.type] for wizard-side + // per-agent overrides. Prompt-frontmatter model still wins (§3.6). + const taskPick = resolveHarness(switchboardCtx, task.type); const taskHarness = requireTaskHarness(taskPick); const taskModel = taskModelSpec(registry, task, taskPick.harness); let taskResult: AgentResult; @@ -1341,7 +1361,7 @@ async function executeOrchestrator( reportBlockedTasks(store.list(), stoppedBy); return { outcome: fatal.outcome, failure: fatal.failure }; } - if (signal?.aborted) return hostAborted(); + if (signal?.aborted) return cancelledRun(); renderQueue(); diff --git a/src/agent/runner/sequence/orchestrator/queue.ts b/src/agent/runner/sequence/orchestrator/queue.ts index 236e6040b..08c66ccde 100644 --- a/src/agent/runner/sequence/orchestrator/queue.ts +++ b/src/agent/runner/sequence/orchestrator/queue.ts @@ -132,7 +132,9 @@ export interface QueueFile { tasks: QueuedTask[]; } -export { TASK_OUTCOMES_KEY } from '../../shared/types'; +/** Session frameworkContext key holding the drained queue's final outcomes — + * written by the runner before the cache wipe, read by the e2e harness. */ +export const TASK_OUTCOMES_KEY = 'orchestrator-task-outcomes'; export interface TaskOutcome { type: string; diff --git a/src/agent/runner/shared/bootstrap.ts b/src/agent/runner/shared/bootstrap.ts index 799786341..1ed76bae0 100644 --- a/src/agent/runner/shared/bootstrap.ts +++ b/src/agent/runner/shared/bootstrap.ts @@ -2,22 +2,45 @@ * Shared preparation for the runner pipeline. * * Runs before the fork into the linear or orchestrator arm: logging targets, - * caller-owned gateway auth and the scan-triage classifier built on it. Everything the + * the gateway mint and the scan-triage classifier built on it. Everything the * caller must decide first — health gates, settings conflicts, authentication, * the AI opt-in gate, post-auth gates, feature flags, run tags, token refresh — * arrives already resolved in `RunConfig` and `RunInput`. */ import { createTriageLLMProvider } from '@agent/triage-provider'; +import { gatewayAuth } from '@agent/gateway-session'; import { logToFile } from '@utils/debug'; import { CallType, IS_DEV } from '@shared/constants'; import { VERSION } from '@shared/version'; import { mcpUrlFor } from '@shared/host-resolution'; import type { WizardRunOptions } from '@utils/types'; -import type { BootstrapResult, RunConfig, RunInput } from './types'; +import type { BootstrapResult, RunConfig, RunFlags, RunInput } from './types'; // ── Helpers ────────────────────────────────────────────────────────── +/** + * Decide whether the `wizard_ask` overlay should be wired for this run. + * Disabled in non-interactive modes (CI, signup) — there's no human to + * answer. Per-program disabling is done by adding WIZARD_ASK_TOOL_NAME to + * the program's `disallowedTools` so the SDK rejects calls outright. + * Extracted so the policy can be unit-tested directly. + * + * `e2eAsk` is the one escape hatch. The e2e harness runs a `ci` + * session, but it does have an answerer — the driver loop answers each + * `wizard_ask` batch from the program's e2e profile. Without the flag the + * agent-in-the-loop layer (the ask bridge in both sequence arms, and the + * orchestrator's seeded warehouse task) stays unreachable from a test. + * + * Only the e2e TUI host sets the flag, from the `E2E_ASK` env var. No CLI flag + * populates it, so plain `--ci` and `--signup` runs behave exactly as before. + */ +export function shouldDisableAsk( + flags: Pick, +): boolean { + return (flags.ci || flags.signup) && !flags.e2eAsk; +} + /** The option bag the agent interface and the middleware read. */ export function runOptions(input: RunInput): WizardRunOptions { return { @@ -35,8 +58,8 @@ export function runOptions(input: RunInput): WizardRunOptions { // ── Prepare ─────────────────────────────────────────────────────────── /** - * Shared setup for both arms: logging targets, then the supplied gateway auth and - * triage classifier. Throws when auth is refused, so the run fails before + * Shared setup for both arms: logging targets, then the gateway mint and the + * triage classifier. Throws when the mint is refused, so the run fails before * any agent starts — the caller maps that the way it maps any unexpected error. */ export async function prepareRun( @@ -57,9 +80,17 @@ export async function prepareRun( const { credentials } = input; const { wizardFlags, wizardFlagPayloads, wizardMetadata, programId } = config; - // Resolve before starting either sequence, so a refusal stops the run. - const { inferenceAuth } = input; - await inferenceAuth.resolve(); + // Mint now so a refusal fails the boot before any agent starts. Later + // readers re-resolve through the cache, which re-mints past the refresh + // point. + const currentGatewayAuth = () => + // TODO(B2): the agent must not mint inference auth. It receives the + // PostHog token here and derives a gateway token from it, re-minting near + // expiry. Programs own credentials (stack plan 4.5): pass a resolved + // inference-auth provider on RunInput.credentials and move + // gateway-session.ts out of src/agent with it. + gatewayAuth(credentials.host, credentials.accessToken, programId); + await currentGatewayAuth(); return { skillsBaseUrl, @@ -74,7 +105,7 @@ export async function prepareRun( // Resolved once, here: the only place holding both the run-level harness // and the gateway auth. Every skill install downstream reads it off boot. triageProvider: createTriageLLMProvider(async () => { - const auth = await inferenceAuth.resolve(); + const auth = await currentGatewayAuth(); return { baseURL: auth.gatewayUrl, authToken: auth.token, diff --git a/src/agent/runner/shared/errors.ts b/src/agent/runner/shared/errors.ts index 9fd363d99..d227a7f0f 100644 --- a/src/agent/runner/shared/errors.ts +++ b/src/agent/runner/shared/errors.ts @@ -11,12 +11,6 @@ export const failed = (failure: AgentFailure): SequenceResult => ({ failure, }); -/** A host cancellation is a decided abort, regardless of SDK error wording. */ -export const hostAborted = (): SequenceResult => ({ - outcome: RunOutcome.Aborted, - failure: { code: ErrorCodes.AgentAbort, message: 'Run cancelled by host.' }, -}); - /** The failure a skill install error decides. The caller reports and exits. */ export function installFailure( integrationLabel: string, diff --git a/src/agent/runner/shared/types.ts b/src/agent/runner/shared/types.ts index 9842e8a11..f447ba6d7 100644 --- a/src/agent/runner/shared/types.ts +++ b/src/agent/runner/shared/types.ts @@ -5,8 +5,8 @@ * invocation snapshot, reports through `options.onProgress`, asks through * `options.interaction`, and returns a `RunResult`. Nothing here names a UI, * a store, a session or a program registry: the caller resolves those and - * hands over plain data. Programs' `runProgram` is the caller that builds it - * for every host. + * hands over plain data. `src/programs/run-agent-legacy.ts` is the caller + * that rebuilds today's session-driven behavior on top of this contract. */ import type { AdditionalFeature } from '@shared/constants'; @@ -21,16 +21,11 @@ import type { ErrorCode } from '@shared/errors'; import type { LLMProvider } from '@posthog/warlock'; import type { AgentInteraction, ProgressEmitter } from '@agent/progress'; import type { EffortLevel } from '../switchboard/models'; -import type { GatewayAuth } from '@shared/gateway-auth'; +import type { SwitchboardCtx } from '../switchboard'; import type { TranscriptTail } from './transcript-tail'; export type { PromptContext, Credentials }; -/** Agent-facing capability; programs decide where inference auth comes from. */ -export type InferenceAuthProvider = { - resolve(): Promise; -}; - /** * A known `[ABORT] ` case. First matching entry is rendered on * the error outro; unmatched aborts use a generic fallback. @@ -151,11 +146,6 @@ export interface ResolvedBinding { model: string; /** Reasoning-effort override. Absent → the model's table default. */ thinkingLevel?: EffortLevel; - /** Role-specific routes resolved by the caller before the agent starts. */ - roleBindings?: Record< - string, - { harness: Harness; model: string; thinkingLevel?: EffortLevel } - >; } /** @@ -164,7 +154,7 @@ export interface ResolvedBinding { * treats every label as opaque. */ export interface RunConfig { - /** Opaque program label for gateway spend pin and analytics. */ + /** Program id: gateway spend pin, analytics label, commandments axis. */ programId: string; /** The run definition. A program's session-taking hooks are the caller's, see `hooks`. */ run: AgentRunDefinition; @@ -172,12 +162,11 @@ export interface RunConfig { composed: boolean; /** Run-level sequence, harness and model. */ binding: ResolvedBinding; - /** Program text selected by the caller; the agent only assembles it. */ - programCommandments?: readonly string[]; - /** Validated stage policy selected by the caller; absent keeps flow frontmatter. */ - stageOverrides?: Record; - /** The caller's resolved seeded-task experiment. */ - seededTasksEnabled?: boolean; + /** + * The inputs the run-level binding was resolved from. The orchestrator + * re-resolves the harness per task role from these; nothing else reads them. + */ + switchboard: SwitchboardCtx; /** Primary skills origin (context-mill dev or GitHub Releases). */ skillsBaseUrl: string; /** Feature flag key → variant, evaluated before the run. */ @@ -223,8 +212,6 @@ export interface RunInput { installDir: string; /** Resolved credentials, including the host family and its MCP url. */ credentials: Credentials; - /** Caller-owned gateway auth, including refresh policy. */ - inferenceAuth: InferenceAuthProvider; /** Project payload resolved at authentication, for prompt context. */ project: ApiProject | null; /** User payload resolved at authentication, for the AI opt-in prompt line. */ @@ -284,9 +271,6 @@ export interface AgentFailure { authErrorDetail?: AuthErrorDetail; } -/** frameworkContext key for the drained queue's final outcomes, read by the e2e harness. */ -export const TASK_OUTCOMES_KEY = 'orchestrator-task-outcomes'; - export enum RunOutcome { Success = 'success', Aborted = 'aborted', diff --git a/src/agent/runner/switchboard/commandments.ts b/src/agent/runner/switchboard/commandments.ts index a170b271a..5f709b1fb 100644 --- a/src/agent/runner/switchboard/commandments.ts +++ b/src/agent/runner/switchboard/commandments.ts @@ -1,8 +1,11 @@ /** * System-prompt commandments, keyed by the axes the switchboard resolves. * - * Programs select their guidance before the run; the agent assembles that - * supplied text with global, sequence, model and harness guidance. + * A run is a resolved (program, sequence, harness, model). Guidance belongs to + * whichever axis makes it true, declared here beside the tables that resolve + * them, and assembled once by `assembleCommandments`. A rule that is true for + * every run stays in `@lib/agent/commandments`; anything narrower lives here so + * the call sites never re-derive it. * * Leaf module by design — it imports the axis enums and the per-axis text, never * a runner or a harness backend, so the harnesses can call the assembler without @@ -37,6 +40,20 @@ const SEQUENCE_COMMANDMENTS: Record = { [Sequence.orchestrator]: [], }; +// ── Program axis ──────────────────────────────────────────────────────── + +const SELF_DRIVING = [ + 'ALWAYS surface a custom-scout proposal in step 6b: bring the user your one or two strongest candidate scouts even when the built-in troop looks sufficient. The proposal ask leads with a "None — keep the built-in troop" option, so declining costs the user one keystroke — but a proposal you silently skip is coverage they never got to see or judge. Where the skill says to skip the ask when the gap analysis finds no candidate, do NOT skip: pick your best candidates anyway and let the user decide.', + + 'Rank candidates at the discriminator level, not the category level. "Covered" only means an enabled scout would actually FIRE for that failure mode: a conversion-rate watcher does not catch entry volume collapsing; a Stripe-transaction watcher does not catch a lead form going silent. A surface whose failure mode has no firing condition among the enabled scouts is your strongest candidate.', + + 'Be honest in the option descriptions: if a candidate overlaps something an enabled scout partially watches, say so in its description rather than dropping the candidate. The user chooses with full information; you do not gatekeep on their behalf.', +]; + +const PROGRAM_COMMANDMENTS: Record = { + 'self-driving': SELF_DRIVING, +}; + // ── Harness axis ──────────────────────────────────────────────────────── /** @@ -58,8 +75,8 @@ const MODEL_COMMANDMENTS: Record = {}; // ── Assembly ──────────────────────────────────────────────────────────── export interface CommandmentAxes { - /** Selected by programs; the agent only assembles supplied text. */ - programCommandments?: readonly string[]; + /** Program id, as resolved into `PROGRAM_BINDINGS`. */ + program?: string; sequence: Sequence; harness: Harness; /** Gateway model id. */ @@ -70,14 +87,14 @@ export interface CommandmentAxes { /** Every commandment this run's axes call for, broad to narrow. */ export function assembleCommandments(axes: CommandmentAxes): string { - const { programCommandments, sequence, harness, model, caps } = axes; + const { program, sequence, harness, model, caps } = axes; const harnessNotes = HARNESS_NOTES[harness]?.( sequence, caps ?? { bash: true, posthogMcp: true }, ); return [ ...WIZARD_COMMANDMENTS, - ...(programCommandments ?? []), + ...(program ? PROGRAM_COMMANDMENTS[program] ?? [] : []), ...SEQUENCE_COMMANDMENTS[sequence], ...(model ? MODEL_COMMANDMENTS[model] ?? [] : []), // Blank line first: the notes open their own `## This runtime` section. diff --git a/src/programs/experiments/__tests__/binding-cases.ts b/src/agent/runner/switchboard/flags/__tests__/binding-cases.ts similarity index 91% rename from src/programs/experiments/__tests__/binding-cases.ts rename to src/agent/runner/switchboard/flags/__tests__/binding-cases.ts index da7a31439..341a5d9a8 100644 --- a/src/programs/experiments/__tests__/binding-cases.ts +++ b/src/agent/runner/switchboard/flags/__tests__/binding-cases.ts @@ -5,8 +5,11 @@ */ import { describe, it, expect } from 'vitest'; import { GPT5_6_SOL_MODEL, Harness, Sequence } from '@shared/constants'; -import { resolveProgramBinding as resolveBinding } from '../../binding'; -import type { ProgramSwitchboardCtx as SwitchboardCtx } from '@programs/types'; +import { + resolveBinding, + type SwitchboardCtx, + type SwitchboardTrace, +} from '@agent/runner/switchboard'; import type { EffortLevel } from '@agent/runner/switchboard/models'; /** The complete resolved binding — every axis stated, nothing implicit. */ @@ -24,7 +27,7 @@ export interface BindingCase { ctx: Omit; binding: ExpectedBinding; /** Also pin which precedence rung decided each axis. */ - trace?: SwitchboardCtx['trace']; + trace?: SwitchboardTrace; } export function runBindingCases( diff --git a/src/programs/experiments/__tests__/flags.test.ts b/src/agent/runner/switchboard/flags/__tests__/flags.test.ts similarity index 93% rename from src/programs/experiments/__tests__/flags.test.ts rename to src/agent/runner/switchboard/flags/__tests__/flags.test.ts index e34adf86c..8c04f0983 100644 --- a/src/programs/experiments/__tests__/flags.test.ts +++ b/src/agent/runner/switchboard/flags/__tests__/flags.test.ts @@ -15,14 +15,17 @@ import { WIZARD_ORCHESTRATOR_OVERRIDE_FLAG_KEY, WIZARD_SELF_DRIVING_USE_PI_HARNESS_FLAG_KEY, } from '@shared/constants'; -import { resolveProgramBinding as resolveBinding } from '../../binding'; -import type { ProgramSwitchboardCtx as SwitchboardCtx } from '@programs/types'; -import { areSeededTasksEnabled, resolveStageOverrides } from '..'; +import { + areSeededTasksEnabled, + resolveBinding, + resolveStageOverrides, + type SwitchboardCtx, +} from '@agent/runner/switchboard'; import { ORCHESTRATOR_SEQUENCE_ROUTE, ORCHESTRATOR_HARNESS_ROUTE, -} from '@programs/experiments/orchestrator'; -import { SELF_DRIVING_EXPERIMENT } from '@programs/experiments/self-driving'; +} from '@agent/runner/switchboard/flags/orchestrator'; +import { SELF_DRIVING_EXPERIMENT } from '@agent/runner/switchboard/flags/self-driving'; import { runBindingCases } from './binding-cases'; const envState = vi.hoisted(() => ({ @@ -329,17 +332,22 @@ describe('isolation — everything on at once', () => { }); }); -describe('seam scan — routing reads live only in experiments/', () => { - const programsDir = join(dirname(fileURLToPath(import.meta.url)), '..', '..'); +describe('seam scan — routing reads live only in flags/', () => { + const switchboardDir = join( + dirname(fileURLToPath(import.meta.url)), + '..', + '..', + ); + // orchestrator-runner consumes a flags/ resolver; it may pass the snapshot through, never index it. for (const file of [ - '../agent/runner/switchboard/harness.ts', - '../agent/runner/switchboard/sequence.ts', - '../agent/runner/switchboard/models.ts', - '../agent/runner/switchboard/index.ts', - '../agent/runner/sequence/orchestrator/orchestrator-runner.ts', + 'harness.ts', + 'sequence.ts', + 'models.ts', + 'index.ts', + '../sequence/orchestrator/orchestrator-runner.ts', ]) { it(`${file} contains no direct flag reads or flag-key imports`, () => { - const src = readFileSync(join(programsDir, file), 'utf8'); + const src = readFileSync(join(switchboardDir, file), 'utf8'); expect(src).not.toMatch(/ctx\.flags\[/); expect(src).not.toMatch(/flags\[['"`]/); expect(src).not.toMatch(/WIZARD_\w+_FLAG_KEY/); diff --git a/src/programs/experiments/index.ts b/src/agent/runner/switchboard/flags/index.ts similarity index 100% rename from src/programs/experiments/index.ts rename to src/agent/runner/switchboard/flags/index.ts diff --git a/src/programs/experiments/orchestrator.ts b/src/agent/runner/switchboard/flags/orchestrator.ts similarity index 100% rename from src/programs/experiments/orchestrator.ts rename to src/agent/runner/switchboard/flags/orchestrator.ts diff --git a/src/programs/experiments/schemes.ts b/src/agent/runner/switchboard/flags/schemes.ts similarity index 99% rename from src/programs/experiments/schemes.ts rename to src/agent/runner/switchboard/flags/schemes.ts index 1fd30ce34..99f9c4a97 100644 --- a/src/programs/experiments/schemes.ts +++ b/src/agent/runner/switchboard/flags/schemes.ts @@ -15,7 +15,7 @@ import { } from '@shared/constants'; import type { ProgramId } from '@programs/types'; import { logToFile } from '@utils/debug'; -import type { EffortLevel } from '@agent/types'; +import type { EffortLevel } from '../models'; // ── Shared vocabulary ───────────────────────────────────────────────────── diff --git a/src/programs/experiments/self-driving.ts b/src/agent/runner/switchboard/flags/self-driving.ts similarity index 100% rename from src/programs/experiments/self-driving.ts rename to src/agent/runner/switchboard/flags/self-driving.ts diff --git a/src/agent/runner/switchboard/harness.ts b/src/agent/runner/switchboard/harness.ts index b9078ac34..0f5673ed2 100644 --- a/src/agent/runner/switchboard/harness.ts +++ b/src/agent/runner/switchboard/harness.ts @@ -1,12 +1,22 @@ /** - * Harness axis: the registry. Mirrors `sequence.ts`; the resolver is - * `resolve-harness.ts`. + * Harness axis: registry, middleware, resolver. Mirrors `sequence.ts`. */ +import { IS_PRODUCTION_BUILD } from '@env'; import { Harness } from '@shared/constants'; +import { logToFile } from '@utils/debug'; import { anthropicBackend } from '../harness/anthropic'; import { piBackend } from '../harness/pi'; import type { AgentHarness } from '../harness/types'; +import { resolveFlagRoute } from './flags'; +import { + DEFAULT_BINDING, + PROGRAM_BINDINGS, + runChain, + type HarnessPick, + type Middleware, + type SwitchboardCtx, +} from '.'; export const HARNESS_OPTIONS: Partial> = { [Harness.anthropic]: anthropicBackend, @@ -20,3 +30,76 @@ export function getHarness(name: Harness): AgentHarness { } return harness; } + +/** + * PostHog-flag routing to pi (see `./flags`). No valid route — flag off, no + * config, or an invalid payload — keeps the non-flagged binding default. + */ +const flagRunnerOverride: Middleware = (ctx, next) => { + const pick = next(); + const route = resolveFlagRoute(ctx.program, ctx.flags, ctx.flagPayloads); + if (!route) return pick; + if (ctx.trace) { + ctx.trace.harness = 'flag'; + // Harness-only routes keep the binding's model — trace it truthfully so + // analytics never attributes the fallback model to the flag. + if (route.model) ctx.trace.model = 'flag'; + } + return { + harness: route.harness ?? Harness.pi, + model: route.model ?? pick.model, + thinkingLevel: route.thinkingLevel ?? pick.thinkingLevel, + }; +}; + +/** `--harness` override. Dev/test only — the option is gated out of published builds. */ +const cliHarnessOverride: Middleware = (ctx, next) => { + const pick = next(); + if (!ctx.cliHarness) return pick; + if (ctx.trace) ctx.trace.harness = 'cli'; + return { ...pick, harness: ctx.cliHarness }; +}; + +/** `--model` override. Dev/test only — the option is gated out of published builds. */ +const cliModelOverride: Middleware = (ctx, next) => { + const pick = next(); + if (!ctx.cliModel) return pick; + if (ctx.trace) ctx.trace.model = 'cli'; + return { ...pick, model: ctx.cliModel }; +}; + +// Order = precedence: CLI > flag > binding default. The prod spread collapses +// to [], dropping the CLI overrides from the chain. +const HARNESS_MIDDLEWARE: Middleware[] = [ + ...(IS_PRODUCTION_BUILD ? [] : [cliHarnessOverride, cliModelOverride]), + flagRunnerOverride, +]; + +/** + * Resolve the harness for a role. Linear callers omit `role`; orchestrator + * callers pass `'seed'` or `task.type`. `contextMillOverride[role]` overlays. + */ +export function resolveHarness( + ctx: SwitchboardCtx, + role = 'default', +): HarnessPick { + const pick = runChain(HARNESS_MIDDLEWARE, ctx, () => { + if (ctx.trace) + Object.assign(ctx.trace, { harness: 'binding', model: 'binding' }); + const binding = PROGRAM_BINDINGS[ctx.program] ?? DEFAULT_BINDING; + return { + harness: binding.harness, + model: binding.model, + thinkingLevel: binding.thinkingLevel, + ...binding.contextMillOverride?.[role], + }; + }); + logToFile( + `[switchboard] resolved: program=${ctx.program} harness=${pick.harness}` + + `${ctx.trace?.harness ? ` (${ctx.trace.harness})` : ''} model=${ + pick.model + }` + + `${ctx.trace?.model ? ` (${ctx.trace.model})` : ''}`, + ); + return pick; +} diff --git a/src/agent/runner/switchboard/index.ts b/src/agent/runner/switchboard/index.ts index 7b6203f81..b2d84add2 100644 --- a/src/agent/runner/switchboard/index.ts +++ b/src/agent/runner/switchboard/index.ts @@ -1,11 +1,20 @@ // Resolves routing; model additions also require mint allowlists and gateway prompt/transport support. -import { Harness, Sequence } from '@shared/constants'; +import { + DEFAULT_AGENT_MODEL, + GPT5_6_SOL_MODEL, + GPT5_6_TERRA_MODEL, + Harness, + Sequence, +} from '@shared/constants'; +import type { ProgramId } from '@programs/types'; +import { resolveHarness } from './harness'; import type { EffortLevel } from './models'; +import { resolveSequence } from './sequence'; // ── Shared machinery ──────────────────────────────────────────────────── -/** Which precedence rung decided each axis. Stamped by the resolvers as they decide. */ +/** Which precedence rung decided each axis. Stamped by middlewares as they assert. */ export interface SwitchboardTrace { harness?: 'cli' | 'flag' | 'binding'; model?: 'cli' | 'flag' | 'binding'; @@ -18,21 +27,14 @@ export interface SwitchboardTrace { | 'binding'; } -/** Everything a resolver may branch on. Built once per run. */ +/** Everything a resolver middleware may branch on. Built once per run. */ export interface SwitchboardCtx { - /** Opaque log label. Program lookup stays with the caller. */ - program?: string; - /** The caller's selected base binding, before route/CLI overlays. */ - baseBinding?: ProgramBinding; + program: ProgramId; /** Composed sub-run (a dependency inside a parent program). Structurally linear — no override can orchestrate it. */ composed?: boolean; - /** Already validated experiment route; no flag parsing happens in the agent. */ - flagRoute?: { - harness?: Harness; - model?: string; - thinkingLevel?: EffortLevel; - sequence?: Sequence; - }; + flags: Record; + /** Flag payloads from the same snapshot (payload-carrying flags, e.g. self-driving pi). */ + flagPayloads?: Record; /** CLI override (`--harness`). Wins over `flags`. */ cliHarness?: Harness; /** CLI override (`--sequence`). Wins over `flags`. */ @@ -43,6 +45,37 @@ export interface SwitchboardCtx { trace?: SwitchboardTrace; } +/** A resolver middleware: defer via `next()`, or assert by returning a value. */ +export type Middleware = (ctx: SwitchboardCtx, next: () => D) => D; + +/** + * Run a middleware chain over `ctx`. Each middleware receives `next` (which + * runs the rest of the chain) and can either: + * - defer: call `next()` and optionally modify its result (overlay pattern) + * - short-circuit: return a value without calling `next()` (skip the rest) + * + * **Earlier in the array = higher precedence.** Index 0 runs first and can + * short-circuit the rest; index 1 only runs if index 0 deferred. So + * `[cliSequenceMw, orchestratorFeatureFlagMw]` means CLI takes precedence over the + * flag, not the other way around. + * + * `fallback` runs at the end — reached only when every middleware deferred. + * Typically the map read for the base value. + */ +export function runChain( + chain: Middleware[], + ctx: SwitchboardCtx, + fallback: () => D, +): D { + function step(index: number): D { + if (index >= chain.length) return fallback(); + const middleware = chain[index]; + const next = () => step(index + 1); + return middleware(ctx, next); + } + return step(0); +} + // ── Data model ────────────────────────────────────────────────────────── /** Harness + model for one leaf of agent work. */ @@ -69,11 +102,92 @@ export interface ProgramBinding { contextMillOverride?: Record>; } +// Legacy fallback; new programs should explicitly choose Pi and prefer orchestration. +export const DEFAULT_BINDING: ProgramBinding = { + sequence: Sequence.linear, + harness: Harness.pi, + model: GPT5_6_SOL_MODEL, + thinkingLevel: 'medium', +}; + +/** + * Per-program routing. Kept in lockstep with `PROGRAM_REGISTRY` by the + * switchboard test. Anything absent falls back to `DEFAULT_BINDING`. + */ +export const PROGRAM_BINDINGS: Partial> = { + 'posthog-integration': DEFAULT_BINDING, + 'revenue-analytics-setup': DEFAULT_BINDING, + 'warehouse-source': DEFAULT_BINDING, + 'error-tracking-upload-source-maps': { + sequence: Sequence.linear, + harness: Harness.pi, + model: GPT5_6_SOL_MODEL, + thinkingLevel: 'medium', + }, + audit: DEFAULT_BINDING, + 'events-audit': DEFAULT_BINDING, + 'posthog-doctor': DEFAULT_BINDING, + 'web-analytics-doctor': DEFAULT_BINDING, + migration: DEFAULT_BINDING, + 'self-driving': DEFAULT_BINDING, + 'agent-skill': DEFAULT_BINDING, + 'mcp-add': DEFAULT_BINDING, + 'mcp-remove': DEFAULT_BINDING, + 'mcp-tutorial': DEFAULT_BINDING, + 'mcp-analytics': DEFAULT_BINDING, + // Orchestrator on pi. The binding routes only; every stage's model and + // effort are pinned context-mill side in the flow's frontmatter + // (`model_pi`/`effort_pi`: terra seed, sol tasks, luna report). + metrics: { + sequence: Sequence.orchestrator, + harness: Harness.pi, + model: DEFAULT_AGENT_MODEL, + }, + 'replay-vision': { + sequence: Sequence.orchestrator, + harness: Harness.anthropic, + model: DEFAULT_AGENT_MODEL, + }, + // Orchestrator on pi, like metrics. The binding routes only; every stage's + // model and effort are pinned context-mill side in the flow's frontmatter + // (`model_pi`/`effort_pi`: terra seed, install and init, sol tasks, luna report). + 'error-tracking': { + sequence: Sequence.orchestrator, + harness: Harness.pi, + model: DEFAULT_AGENT_MODEL, + }, + 'ai-observability': { + sequence: Sequence.linear, + harness: Harness.pi, + model: GPT5_6_TERRA_MODEL, + thinkingLevel: 'high', + }, + slack: DEFAULT_BINDING, +}; + +// ── Unified resolver ──────────────────────────────────────────────────── + +/** Compose both axes. Callers needing only one axis use the per-axis resolver. */ +export function resolveBinding( + ctx: SwitchboardCtx, + role = 'default', +): ProgramBinding { + ctx.trace ??= {}; + const sequence = resolveSequence(ctx); + const { harness, model, thinkingLevel } = resolveHarness(ctx, role); + return { sequence, harness, model, thinkingLevel }; +} + // ── Unified re-export surface ─────────────────────────────────────────── -export { HARNESS_OPTIONS, getHarness } from './harness'; +export { HARNESS_OPTIONS, getHarness, resolveHarness } from './harness'; +export { + SEQUENCE_OPTIONS, + getSequence, + resolveSequence, + type SequenceRunner, +} from './sequence'; export { - harnessRunsTasks, - resolveHarness, - resolveRoleHarness, -} from './resolve-harness'; -export { SEQUENCE_OPTIONS, getSequence, type SequenceRunner } from './sequence'; + isOrchestratorEnabled, + areSeededTasksEnabled, + resolveStageOverrides, +} from './flags'; diff --git a/src/agent/runner/switchboard/resolve-harness.ts b/src/agent/runner/switchboard/resolve-harness.ts deleted file mode 100644 index af47f4fa5..000000000 --- a/src/agent/runner/switchboard/resolve-harness.ts +++ /dev/null @@ -1,90 +0,0 @@ -/** - * Harness axis: the resolver that picks a harness and model, and which - * harnesses the orchestrator can drive. Data and pure functions only, so the - * agent entry loads this at startup; the registry is `harness.ts`. - */ - -import { IS_PRODUCTION_BUILD } from '@env'; -import { Harness } from '@shared/constants'; -import { logToFile } from '@utils/debug'; -import { DEFAULT_AGENT_BINDING } from '@agent/default-binding'; -import type { ResolvedBinding } from '../shared/types'; -import type { HarnessPick, ProgramBinding, SwitchboardCtx } from '.'; - -/** Which backends implement `runTask`; a registry test keeps this in step with HARNESS_OPTIONS. */ -export const HARNESS_RUNS_TASKS: Record = { - [Harness.anthropic]: true, - [Harness.pi]: true, -}; - -/** Whether the orchestrator can drive this harness. */ -export function harnessRunsTasks(name: Harness): boolean { - return HARNESS_RUNS_TASKS[name] === true; -} - -/** - * Resolve the harness for a role. Linear callers omit `role`; orchestrator - * callers pass `'seed'` or `task.type`. `contextMillOverride[role]` overlays. - * Precedence: CLI > flag route > binding default; the CLI overrides are - * dev/test only and gated out of published builds. - */ -export function resolveHarness( - ctx: SwitchboardCtx, - role = 'default', -): HarnessPick { - if (ctx.trace) - Object.assign(ctx.trace, { harness: 'binding', model: 'binding' }); - const binding: ProgramBinding = ctx.baseBinding ?? DEFAULT_AGENT_BINDING; - let pick: HarnessPick = { - harness: binding.harness, - model: binding.model, - thinkingLevel: binding.thinkingLevel, - ...binding.contextMillOverride?.[role], - }; - const route = ctx.flagRoute; - if (route) { - if (ctx.trace) { - ctx.trace.harness = 'flag'; - // Harness-only routes keep the binding's model — trace it truthfully so - // analytics never attributes the fallback model to the flag. - if (route.model) ctx.trace.model = 'flag'; - } - pick = { - harness: route.harness ?? Harness.pi, - model: route.model ?? pick.model, - thinkingLevel: route.thinkingLevel ?? pick.thinkingLevel, - }; - } - if (!IS_PRODUCTION_BUILD && ctx.cliModel) { - if (ctx.trace) ctx.trace.model = 'cli'; - pick = { ...pick, model: ctx.cliModel }; - } - if (!IS_PRODUCTION_BUILD && ctx.cliHarness) { - if (ctx.trace) ctx.trace.harness = 'cli'; - pick = { ...pick, harness: ctx.cliHarness }; - } - logToFile( - `[switchboard] resolved: program=${ctx.program ?? '?'} harness=${ - pick.harness - }` + - `${ctx.trace?.harness ? ` (${ctx.trace.harness})` : ''} model=${ - pick.model - }` + - `${ctx.trace?.model ? ` (${ctx.trace.model})` : ''}`, - ); - return pick; -} - -/** The agent resolves a task role only from data the caller already supplied. */ -export function resolveRoleHarness( - binding: ResolvedBinding, - role: string, -): HarnessPick { - return ( - binding.roleBindings?.[role] ?? { - harness: binding.harness, - model: binding.model, - thinkingLevel: binding.thinkingLevel, - } - ); -} diff --git a/src/agent/runner/switchboard/sequence.ts b/src/agent/runner/switchboard/sequence.ts index 43881f268..b6d241720 100644 --- a/src/agent/runner/switchboard/sequence.ts +++ b/src/agent/runner/switchboard/sequence.ts @@ -1,11 +1,27 @@ /** - * Sequence axis: the registry. Programs resolve which sequence a run uses. + * Sequence axis: gate helpers, registry, middleware, resolver. + * Percentage rollouts are PostHog-side — the gate just reads the resolved bool. */ +import { IS_PRODUCTION_BUILD } from '@env'; import { Sequence } from '@shared/constants'; +import { logToFile } from '@utils/debug'; +import { + isOrchestratorEnabled, + resolveFlagRoute, + resolveFlagSequence, +} from './flags'; +import { getHarness, resolveHarness } from './harness'; import type { SequenceResult, SequenceContext } from '../shared/types'; import { runLinearProgram } from '../sequence/linear'; import { runOrchestrator } from '../sequence/orchestrator/orchestrator-runner'; +import { + DEFAULT_BINDING, + PROGRAM_BINDINGS, + runChain, + type Middleware, + type SwitchboardCtx, +} from '.'; // ── Registry ──────────────────────────────────────────────────────────── @@ -33,3 +49,82 @@ export function getSequence(name: Sequence): SequenceRunner { } return sequence; } + +// ── Middleware + resolver ─────────────────────────────────────────────── + +/** + * A composed sub-run (integration inside self-driving) is structurally + * linear: the orchestrator owns the full run lifecycle (queue, outro) and + * cannot nest. Sits above every override, including CLI. + */ +const composedClampMw: Middleware = (ctx, next) => { + if (!ctx.composed) return next(); + if (ctx.trace) ctx.trace.sequence = 'composed'; + return Sequence.linear; +}; + +/** `--sequence` override. Dev/test only — the option is gated out of published builds. */ +const cliSequenceMw: Middleware = (ctx, next) => { + if (!ctx.cliSequence) return next(); + if (ctx.trace) ctx.trace.sequence = 'cli'; + return ctx.cliSequence; +}; + +/** A program's own flag route may pin the sequence; wins over the global orchestrator flag. Traced as 'payload' to stay distinguishable from sequence experiments ('flag'). */ +const flagRouteSequenceMw: Middleware = (ctx, next) => { + const route = resolveFlagRoute(ctx.program, ctx.flags, ctx.flagPayloads); + if (!route?.sequence) return next(); + if (ctx.trace) ctx.trace.sequence = 'payload'; + return route.sequence; +}; + +/** Sequence experiments (e.g. wizard-orchestrator), each inert outside its declared programs. */ +const sequenceExperimentMw: Middleware = (ctx, next) => { + const sequence = resolveFlagSequence(ctx.program, ctx.flags); + if (!sequence) return next(); + if (ctx.trace) ctx.trace.sequence = 'flag'; + return sequence; +}; + +/** + * The orchestrator drives harnesses through `runTask`; a harness that has not + * implemented it clamps the run to linear. A capability check, not a harness + * identity check — a harness gains orchestrator support by implementing the + * method, with no switchboard change. Sits below the CLI override so + * `--sequence orchestrator` still reproduces the hard error in dev builds. + */ +const runTaskCapabilityClampMw: Middleware = (ctx, next) => { + const pick = resolveHarness(ctx); + if (getHarness(pick.harness).runTask) return next(); + if (isOrchestratorEnabled(ctx.flags)) { + logToFile( + `[switchboard] wizard-orchestrator ignored: ${pick.harness} has no runTask, clamping to linear`, + ); + } + if (ctx.trace) ctx.trace.sequence = 'runtask-clamp'; + return Sequence.linear; +}; + +// Order = precedence: CLI > capability clamp > flag > binding default. The +// prod spread collapses to [], dropping cliSequenceMw from the chain. +const SEQUENCE_MIDDLEWARE: Middleware[] = [ + composedClampMw, + ...(IS_PRODUCTION_BUILD ? [] : [cliSequenceMw]), + runTaskCapabilityClampMw, + flagRouteSequenceMw, + sequenceExperimentMw, +]; + +/** CLI wins over `wizard-orchestrator` flag wins over binding default. */ +export function resolveSequence(ctx: SwitchboardCtx): Sequence { + const sequence = runChain(SEQUENCE_MIDDLEWARE, ctx, () => { + if (ctx.trace) ctx.trace.sequence = 'binding'; + const binding = PROGRAM_BINDINGS[ctx.program] ?? DEFAULT_BINDING; + return binding.sequence; + }); + logToFile( + `[switchboard] resolved: program=${ctx.program} sequence=${sequence}` + + `${ctx.trace?.sequence ? ` (${ctx.trace.sequence})` : ''}`, + ); + return sequence; +} diff --git a/src/agent/tools/tool-names.ts b/src/agent/tools/tool-names.ts deleted file mode 100644 index 6d4d9d067..000000000 --- a/src/agent/tools/tool-names.ts +++ /dev/null @@ -1,24 +0,0 @@ -/** Tool ids programs name in allowedTools and disallowedTools. Data only: the agent entry loads it at startup. */ - -export const SERVER_NAME = 'wizard-tools'; - -/** Tool names exposed by the wizard-tools server, keyed for selective use. */ -// SDK expects MCP tool names in allowedTools/disallowedTools to be the -// fully-qualified `mcp____` form (sdk.d.ts: "Fully-qualified -// MCP tool name, e.g. mcp__server__tool_name."). The colon form silently -// fails to match, which made every program's `disallowedTools` entry a no-op. -export const WIZARD_TOOL_NAMES = { - checkEnvKeys: `mcp__${SERVER_NAME}__check_env_keys`, - setEnvValues: `mcp__${SERVER_NAME}__set_env_values`, - detectPackageManager: `mcp__${SERVER_NAME}__detect_package_manager`, - loadSkillMenu: `mcp__${SERVER_NAME}__load_skill_menu`, - installSkill: `mcp__${SERVER_NAME}__install_skill`, - auditSeedChecks: `mcp__${SERVER_NAME}__audit_seed_checks`, - auditAddChecks: `mcp__${SERVER_NAME}__audit_add_checks`, - auditResolveChecks: `mcp__${SERVER_NAME}__audit_resolve_checks`, - wizardAsk: `mcp__${SERVER_NAME}__wizard_ask`, - publishHandoff: `mcp__${SERVER_NAME}__publish_handoff`, - enqueueTask: `mcp__${SERVER_NAME}__enqueue_task`, - completeTask: `mcp__${SERVER_NAME}__complete_task`, - readHandoffs: `mcp__${SERVER_NAME}__read_handoffs`, -} as const; diff --git a/src/agent/tools/tools.ts b/src/agent/tools/tools.ts index 0ab9a6283..6751232a7 100644 --- a/src/agent/tools/tools.ts +++ b/src/agent/tools/tools.ts @@ -1156,7 +1156,28 @@ export function appendAuditChecksToLedger( return { ok: true, added: additions.length }; } -export { SERVER_NAME, WIZARD_TOOL_NAMES } from './tool-names'; +export const SERVER_NAME = 'wizard-tools'; + +/** Tool names exposed by the wizard-tools server, keyed for selective use. */ +// SDK expects MCP tool names in allowedTools/disallowedTools to be the +// fully-qualified `mcp____` form (sdk.d.ts: "Fully-qualified +// MCP tool name, e.g. mcp__server__tool_name."). The colon form silently +// fails to match, which made every program's `disallowedTools` entry a no-op. +export const WIZARD_TOOL_NAMES = { + checkEnvKeys: `mcp__${SERVER_NAME}__check_env_keys`, + setEnvValues: `mcp__${SERVER_NAME}__set_env_values`, + detectPackageManager: `mcp__${SERVER_NAME}__detect_package_manager`, + loadSkillMenu: `mcp__${SERVER_NAME}__load_skill_menu`, + installSkill: `mcp__${SERVER_NAME}__install_skill`, + auditSeedChecks: `mcp__${SERVER_NAME}__audit_seed_checks`, + auditAddChecks: `mcp__${SERVER_NAME}__audit_add_checks`, + auditResolveChecks: `mcp__${SERVER_NAME}__audit_resolve_checks`, + wizardAsk: `mcp__${SERVER_NAME}__wizard_ask`, + publishHandoff: `mcp__${SERVER_NAME}__publish_handoff`, + enqueueTask: `mcp__${SERVER_NAME}__enqueue_task`, + completeTask: `mcp__${SERVER_NAME}__complete_task`, + readHandoffs: `mcp__${SERVER_NAME}__read_handoffs`, +} as const; // --------------------------------------------------------------------------- // Test-only exports diff --git a/src/agent/types.ts b/src/agent/types.ts index 2307a4572..6a1118e4e 100644 --- a/src/agent/types.ts +++ b/src/agent/types.ts @@ -11,15 +11,11 @@ export type { AgentFailure, AgentRunDefinition, PromptContext, - InferenceAuthProvider, - RunAgentOptions, RunConfig, ResolvedBinding, - RunFlags, RunHooks, RunInput, RunResult, - SeedTaskEntry, } from './runner'; export type { AgentInteraction, @@ -34,11 +30,10 @@ export type { TokenUsageDelta, } from './progress'; -/** Input types of the exported resolveHarness. */ +/** Leaves in B2 with the bindings table; B1 deferred PROGRAM_BINDINGS. */ export type { ProgramBinding, SwitchboardCtx } from './runner'; -export type { EffortLevel } from './runner/switchboard/models'; -/** Leaves in C3 with downloadSkill. */ +/** Leaves in B2 with downloadSkill. */ export type { InstallSkillResult } from './tools'; /** Leaves in B2 with the legacy adapter that records it. */ diff --git a/src/agent/wizard-ask-bridge.ts b/src/agent/wizard-ask-bridge.ts index 32457933b..5c57a1b79 100644 --- a/src/agent/wizard-ask-bridge.ts +++ b/src/agent/wizard-ask-bridge.ts @@ -92,6 +92,14 @@ export const CANCELLED_SENTINEL = '__cancelled__'; /** Default per-question timeout (5 minutes). */ export const DEFAULT_ASK_TIMEOUT_MS = 5 * 60 * 1000; +/** + * The longer per-question timeout, for asks that send the user on an errand — + * open a database console, mint a restricted API key. The default above is + * sized for a question answerable from memory and expires long before an + * errand is done. + */ +export const LONGER_ASK_TIMEOUT_MS = 20 * 60 * 1000; + function buildCancelledAnswers(questions: AskQuestion[]): AskAnswers { const out: AskAnswers = {}; for (const q of questions) { diff --git a/src/commands/ai-observability.ts b/src/commands/ai-observability.ts index a67b8c17f..c51fe2b0a 100644 --- a/src/commands/ai-observability.ts +++ b/src/commands/ai-observability.ts @@ -1,5 +1,4 @@ import { aiObservabilityConfig } from '@programs/ai-observability/index'; -import { headlessOption, regionOption } from '@lib/headless-mode'; import type { Command } from './command'; import { nativeCommandFactory } from './factories/native-command-factory'; @@ -17,5 +16,4 @@ import { nativeCommandFactory } from './factories/native-command-factory'; */ export const aiObservabilityCommand: Command = nativeCommandFactory( aiObservabilityConfig, - { cliOptions: { ...headlessOption, ...regionOption } }, ); diff --git a/src/commands/audit.ts b/src/commands/audit.ts index 9b0bdb3cd..39ed26dfe 100644 --- a/src/commands/audit.ts +++ b/src/commands/audit.ts @@ -1,5 +1,4 @@ import { auditConfig } from '@programs/audit/index'; -import { headlessOption, regionOption } from '@lib/headless-mode'; import type { Command } from './command'; import { familyCommandFactory } from './factories/family-command-factory'; @@ -20,5 +19,4 @@ export const auditCommand: Command = familyCommandFactory({ family: 'audit', description: auditConfig.description, optionsFrom: auditConfig, - cliOptions: { ...headlessOption, ...regionOption }, }); diff --git a/src/commands/factories/family-command-factory.ts b/src/commands/factories/family-command-factory.ts index 62a056f71..e26a38f55 100644 --- a/src/commands/factories/family-command-factory.ts +++ b/src/commands/factories/family-command-factory.ts @@ -1,11 +1,11 @@ -import type { Arguments, Options } from 'yargs'; +import type { Arguments } from 'yargs'; import type { ProgramConfig } from '@programs/types'; import { buildFamilyPickerChildren, dispatchFamily, pickerChildrenToShow, -} from '../dispatch-family'; +} from '@programs/dispatch-family'; import { getSkillsBaseUrl } from '@shared/constants'; import { fetchSkillMenu } from '@shared/skill-menu'; @@ -24,8 +24,6 @@ export interface FamilyCommandFactoryOpts { * generic agent-skill config. */ optionsFrom: ProgramConfig; - /** Options supplied by the CLI rather than the program. */ - cliOptions?: Record; } /** @@ -48,7 +46,6 @@ export function familyCommandFactory({ family, description, optionsFrom, - cliOptions, }: FamilyCommandFactoryOpts): Command { // Bare `wizard ` in an interactive terminal. With a single option // today (e.g. `audit events`), skip the picker and run it directly so the @@ -71,7 +68,7 @@ export function familyCommandFactory({ return { name: `${family} [skill]`, description, - options: mergeCommandOptions(optionsFrom, cliOptions), + options: mergeCommandOptions(optionsFrom), positionals: { skill: { type: 'string', diff --git a/src/commands/factories/native-command-factory.ts b/src/commands/factories/native-command-factory.ts index da7349e1a..1cb4ae068 100644 --- a/src/commands/factories/native-command-factory.ts +++ b/src/commands/factories/native-command-factory.ts @@ -1,4 +1,3 @@ -import type { Options } from 'yargs'; import type { ProgramConfig } from '@programs/types'; import type { Command } from '../command'; @@ -8,8 +7,6 @@ import { dispatchProgram, mergeCommandOptions } from './shared'; export interface NativeCommandFactoryOpts { /** Subcommands nested under this command. */ children?: readonly Command[]; - /** Options supplied by the CLI rather than the program. */ - cliOptions?: Record; } /** @@ -31,7 +28,7 @@ export function nativeCommandFactory( return { name: config.command, description: config.description, - options: mergeCommandOptions(config, opts.cliOptions), + options: mergeCommandOptions(config), children: opts.children, handler: (argv) => dispatchProgram(config, argv), }; diff --git a/src/commands/factories/shared.ts b/src/commands/factories/shared.ts index 0802faa56..e38e96717 100644 --- a/src/commands/factories/shared.ts +++ b/src/commands/factories/shared.ts @@ -51,17 +51,17 @@ export function dispatchProgram(config: ProgramConfig, argv: Arguments): void { } /** - * Merge standard flags with program options and CLI-owned command options. + * Merge the standard skill-program flags (`--debug`, `--install-dir`, etc.) + * with any program-specific options declared on `cliOptions`. * - * Later entries override earlier ones so command-specific flags win. + * Program-specific options shadow the standard ones — that's intentional, so + * a program can override a default flag if it ever needs to. */ export function mergeCommandOptions( config: ProgramConfig, - cliOptions: Record = {}, ): Record { return { ...skillProgramOptions, ...((config.cliOptions ?? {}) as Record), - ...cliOptions, }; } diff --git a/src/shared/file-watcher.ts b/src/lib/file-watcher.ts similarity index 100% rename from src/shared/file-watcher.ts rename to src/lib/file-watcher.ts diff --git a/src/lib/runners/__tests__/ci-inference-auth.test.ts b/src/lib/runners/__tests__/ci-inference-auth.test.ts deleted file mode 100644 index 67db2d866..000000000 --- a/src/lib/runners/__tests__/ci-inference-auth.test.ts +++ /dev/null @@ -1,43 +0,0 @@ -import { mkdtempSync, rmSync, writeFileSync } from 'node:fs'; -import { tmpdir } from 'node:os'; -import { join } from 'node:path'; -import { loadCiInferenceAuthProvider } from '../ci-inference-auth'; - -describe('CI inference credentials', () => { - let directory: string; - - beforeEach(() => { - directory = mkdtempSync(join(tmpdir(), 'wizard-ci-inference-')); - }); - - afterEach(() => { - vi.unstubAllEnvs(); - rmSync(directory, { recursive: true, force: true }); - }); - - it('reads the required bearer file once and returns a fixed provider', async () => { - const tokenFile = join(directory, 'gateway-token'); - writeFileSync(tokenFile, ' opaque-fixture-bearer \n'); - vi.stubEnv('WIZARD_CI_GATEWAY_TOKEN_FILE', tokenFile); - - const provider = loadCiInferenceAuthProvider(42, 'us'); - expect(process.env.WIZARD_CI_GATEWAY_TOKEN_FILE).toBeUndefined(); - rmSync(tokenFile); - - const auth = { - token: 'opaque-fixture-bearer', - teamId: 42, - gatewayUrl: 'https://ai-gateway.us.posthog.com', - refreshAtMs: Infinity, - }; - expect(await provider.resolve()).toEqual(auth); - expect(await provider.resolve()).toEqual(auth); - }); - - it('requires the token file', () => { - vi.stubEnv('WIZARD_CI_GATEWAY_TOKEN_FILE', ''); - expect(() => loadCiInferenceAuthProvider(42, 'us')).toThrow( - 'WIZARD_CI_GATEWAY_TOKEN_FILE is required', - ); - }); -}); diff --git a/src/lib/runners/__tests__/composed-program-step.test.ts b/src/lib/runners/__tests__/composed-program-step.test.ts deleted file mode 100644 index 14b1ab3ec..000000000 --- a/src/lib/runners/__tests__/composed-program-step.test.ts +++ /dev/null @@ -1,40 +0,0 @@ -import { advanceStep } from '../run-wizard'; -import { runProgramAgent } from '../run-program-agent'; -import { posthogIntegrationConfig } from '@programs/posthog-integration/index'; -import { selfDrivingConfig } from '@programs/self-driving/index'; -import { SELF_DRIVING_INTEGRATE_PATH_KEY } from '@programs/self-driving/detect'; -import { buildSession, RunPhase } from '@lib/wizard-session'; -import { WizardStore } from '@ui/tui/store'; - -vi.mock('../run-program-agent', () => ({ - runProgramAgent: vi.fn().mockResolvedValue(undefined), -})); - -it('dispatches the declared integration child with a scoped session and records completion', async () => { - const step = selfDrivingConfig.steps.find( - (entry) => entry.id === 'integrate-run', - ); - expect(step).toBeDefined(); - if (!step) throw new Error('missing integration run step'); - expect(step).toHaveProperty('runProgramId', 'posthog-integration'); - expect(step).not.toHaveProperty('run'); - - const store = new WizardStore('self-driving'); - const session = buildSession({ installDir: '/repo' }); - session.frameworkContext[SELF_DRIVING_INTEGRATE_PATH_KEY] = 'apps/web'; - store.session = session; - - await advanceStep( - { ...step, onRunPrep: () => Promise.resolve() }, - store, - selfDrivingConfig, - ); - - expect(runProgramAgent).toHaveBeenCalledWith( - posthogIntegrationConfig, - expect.objectContaining({ installDir: '/repo/apps/web' }), - { composed: true }, - ); - expect(store.session.completedRuns).toContain('integrate-run'); - expect(store.session.runPhase).toBe(RunPhase.Idle); -}); diff --git a/src/lib/runners/__tests__/mint-recovery.test.ts b/src/lib/runners/__tests__/mint-recovery.test.ts index 76c1329ff..ada5cf6c2 100644 --- a/src/lib/runners/__tests__/mint-recovery.test.ts +++ b/src/lib/runners/__tests__/mint-recovery.test.ts @@ -1,6 +1,6 @@ import { vi, it, expect, afterEach } from 'vitest'; import { runWizard } from '../run-wizard'; -import { runProgramAgent } from '../run-program-agent'; +import { runProgramAgent } from '@programs/run-agent-legacy'; import { startTUI } from '@ui/tui/start-tui'; import { WizardStore } from '@ui/tui/store'; import { InkUI } from '@ui/tui/ink-ui'; @@ -9,12 +9,8 @@ import { posthogIntegrationConfig } from '@programs/posthog-integration'; import { ScreenId } from '@ui/tui/router'; import { HostResolution } from '@shared/host-resolution'; import { analytics } from '@utils/analytics'; -import { clearCleanup, runCleanups } from '@utils/wizard-abort'; -import * as fs from 'node:fs'; -import * as os from 'node:os'; -import * as path from 'node:path'; -vi.mock('../run-program-agent', () => ({ runProgramAgent: vi.fn() })); +vi.mock('@programs/run-agent-legacy', () => ({ runProgramAgent: vi.fn() })); vi.mock('@ui/tui/start-tui', () => ({ startTUI: vi.fn() })); vi.mock('@shared/local-dev', async (original) => ({ ...(await original()), @@ -44,70 +40,10 @@ vi.mock('@programs/task-stream/destinations/posthog', () => ({ })); afterEach(() => { - clearCleanup(); vi.restoreAllMocks(); vi.clearAllMocks(); }); -it.each([ - ['TUI setup fails before the agent starts', false, 1, 'setup'], - ['SIGTERM arrives before the completion screen exits', false, 130, 'signal'], - ['the completion wait fails after the agent succeeds', false, 1, 'wait'], - ['the completion screen exits after a successful run', true, 0, 'success'], -] as const)( - 'when %s, the run-installed skill is kept: %s', - async (_case, kept, exitCode, ending) => { - const installDir = fs.mkdtempSync(path.join(os.tmpdir(), 'wizard-skills-')); - const skillDir = path.join(installDir, '.claude', 'skills', 'installed'); - const install = () => { - fs.mkdirSync(skillDir, { recursive: true }); - fs.writeFileSync(path.join(skillDir, '.posthog-wizard'), ''); - }; - const store = new WizardStore(); - setUI(new InkUI(store)); - vi.spyOn(store, 'runReadyHooks').mockImplementation(() => { - if (ending !== 'setup') return Promise.resolve(); - install(); - return Promise.reject(new Error('TUI setup failed')); - }); - vi.spyOn(store, 'getGate').mockResolvedValue(undefined); - let dismiss!: () => void; - const completion = vi.spyOn(store, 'waitUntil').mockImplementation(() => { - if (ending === 'wait') return Promise.reject(new Error('wait failed')); - if (ending !== 'signal') return Promise.resolve(); - return new Promise((resolve) => (dismiss = resolve)); - }); - vi.mocked(startTUI).mockReturnValue({ - store, - unmount: vi.fn(), - waitForSetup: () => Promise.resolve(), - }); - vi.mocked(runProgramAgent).mockImplementation(() => { - install(); - return Promise.resolve(); - }); - const exit = vi - .spyOn(process, 'exit') - .mockImplementation(() => undefined as never); - try { - runWizard(posthogIntegrationConfig, { installDir, telemetry: false }); - if (ending === 'signal') { - await vi.waitFor(() => expect(completion).toHaveBeenCalled()); - process.emit('SIGTERM'); - } - await vi.waitFor(() => expect(exit).toHaveBeenCalledWith(exitCode)); - if (ending === 'signal') dismiss(); - await new Promise((resolve) => setImmediate(resolve)); - runCleanups(); - - expect(exit).toHaveBeenCalledOnce(); - expect(fs.existsSync(skillDir)).toBe(kept); - } finally { - fs.rmSync(installDir, { recursive: true, force: true }); - } - }, -); - it.each(['continue', 'exit'] as const)( 'catches a failed run, shows the handoff screen, and exits 1 after %s', async (action) => { diff --git a/src/lib/runners/ci-inference-auth.ts b/src/lib/runners/ci-inference-auth.ts deleted file mode 100644 index 28e001575..000000000 --- a/src/lib/runners/ci-inference-auth.ts +++ /dev/null @@ -1,25 +0,0 @@ -/** CI owns the token-file input and hands a fixed bearer to the program. */ - -import { readFileSync } from 'node:fs'; -import { createCiGatewayAuth } from '@shared/ci-gateway-auth'; -import type { InferenceAuthProvider } from '@agent/types'; -import { IS_PRODUCTION_BUILD, runtimeEnv } from '@env'; -import type { CloudRegion } from '@utils/types'; - -export function loadCiInferenceAuthProvider( - projectId: number, - region: CloudRegion, -): InferenceAuthProvider { - if (IS_PRODUCTION_BUILD) - throw new Error('CI gateway auth requires a non-production build'); - const tokenFile = runtimeEnv('WIZARD_CI_GATEWAY_TOKEN_FILE'); - if (!tokenFile) - throw new Error('WIZARD_CI_GATEWAY_TOKEN_FILE is required for CI'); - const token = readFileSync(tokenFile, 'utf8'); - delete process.env.WIZARD_CI_GATEWAY_TOKEN_FILE; - const gatewayUrl = - runtimeEnv('WIZARD_CI_GATEWAY_URL') || - `https://ai-gateway.${region}.posthog.com`; - const auth = createCiGatewayAuth(token, projectId, gatewayUrl); - return { resolve: () => Promise.resolve(auth) }; -} diff --git a/src/lib/runners/run-non-interactive.ts b/src/lib/runners/run-non-interactive.ts index 16e5c7e1c..267d4bfa0 100644 --- a/src/lib/runners/run-non-interactive.ts +++ b/src/lib/runners/run-non-interactive.ts @@ -9,7 +9,7 @@ import { POSTHOG_LOCAL_URL, } from '@shared/local-dev'; import type { CloudRegion } from '@utils/types'; -import { createUiReducer, getUI, setUI } from '@ui'; +import { getUI, setUI } from '@ui'; import { LoggingUI } from '@ui/logging-ui'; import type { ProgramConfig } from '@programs/types'; import { getAuditChecks } from '@programs/audit/types'; @@ -17,6 +17,7 @@ import { analytics } from '@utils/analytics'; import { resolveNoTelemetry } from './resolve-no-telemetry'; import type { WizardStore } from '@ui/tui/store'; import type { TaskStreamPush } from '@programs/task-stream/task-stream-push'; +import { join } from 'node:path'; import { ErrorCodes, classifyRunFailure, @@ -24,10 +25,6 @@ import { } from '@shared/errors'; import { detectErrorCode } from '@programs/detect-map'; import type { OutroData, RunPhase as RunPhaseT } from '@lib/wizard-session'; -import { - commitRegisteredRunSkillCleanups, - registerRunSkillCleanup, -} from '@shared/skill-run-cleanup'; /** * The two non-interactive run modes. Both drive the same pipeline today; the @@ -107,8 +104,6 @@ export function runNonInteractive( // (cloud / CI/CD) tags 'headless'; a dev/test `--ci` run upgrades 'dev' to // 'ci'. The mode string is the tag value. analytics.setTag('build', mode); - let detachSignalHandlers: () => void = () => undefined; - let runRegisteredCleanups: () => void = () => undefined; void (async () => { const path = await import('path'); @@ -120,10 +115,7 @@ export function runNonInteractive( const { configureLogFileFromEnvironment, logToFile } = await import( '@utils/debug' ); - const { runCleanups, wizardAbort, WizardError } = await import( - '@utils/wizard-abort' - ); - runRegisteredCleanups = runCleanups; + const { wizardAbort, WizardError } = await import('@utils/wizard-abort'); configureLogFileFromEnvironment(); @@ -134,23 +126,6 @@ export function runNonInteractive( ? (options.installDir as string) : path.join(process.cwd(), options.installDir as string); - // Armed until the run completes, so every failed or interrupted exit removes new skills. - registerRunSkillCleanup(installDir); - const onSigint = () => { - runCleanups(); - process.exit(130); - }; - const onSigterm = () => { - runCleanups(); - process.exit(143); - }; - process.once('SIGINT', onSigint); - process.once('SIGTERM', onSigterm); - detachSignalHandlers = () => { - process.off('SIGINT', onSigint); - process.off('SIGTERM', onSigterm); - }; - const session = buildSession({ debug: options.debug as boolean | undefined, installDir, @@ -238,6 +213,9 @@ export function runNonInteractive( store: headlessStore, programId: config.streamWorkflowId ?? config.id, destinations, + eventPlanPath: config.eventPlanFile + ? join(session.installDir, config.eventPlanFile) + : undefined, auditChecks: config.auditLedgerFile ? () => getAuditChecks(headlessStore.session) : undefined, @@ -264,20 +242,18 @@ export function runNonInteractive( try { if (mode === 'ci') { - const { loadCiInferenceAuthProvider } = await import( - './ci-inference-auth' - ); - session.inferenceAuth = loadCiInferenceAuthProvider( + const { configureGatewayFromCIEnvironment } = await import('@agent'); + configureGatewayFromCIEnvironment( Number(session.projectId), session.region ?? 'us', ); } if (config.ciPreRun) { - const ui = getUI(); await config.ciPreRun(session, { - auth: ui, - log: ui.log, - onProgress: createUiReducer(ui), + log: { + info: (message) => getUI().log.info(message), + warn: (message) => getUI().log.warn(message), + }, }); } else { const readyCtx = { @@ -286,9 +262,7 @@ export function runNonInteractive( session.frameworkContext[key] = value; }, setFrameworkConfig: () => undefined, - setDetectedFramework: (label: string) => { - session.detectedFrameworkLabel = label; - }, + setDetectedFramework: () => undefined, // Non-interactive session is a plain object (no nanostore // copy-on-write), so direct assignment is safe here. setSkillId: (skillId: string | null) => { @@ -368,14 +342,9 @@ export function runNonInteractive( } } - if (session.detectedFrameworkLabel) { - getUI().setDetectedFramework(session.detectedFrameworkLabel); - } - - const { runProgramAgent } = await import('./run-program-agent'); + const { runProgramAgent } = await import('@programs/run-agent-legacy'); await runProgramAgent(config, session); await settleStream(RunPhase.Completed); - commitRegisteredRunSkillCleanups(); } catch (error) { const errorMessage = error instanceof Error ? error.message : String(error); @@ -406,14 +375,11 @@ export function runNonInteractive( error: error as Error, }); } - })() - .catch((error: unknown) => { - runRegisteredCleanups(); - emitWizardError({ - code: ErrorCodes.InternalUnhandled, - message: error instanceof Error ? error.message : String(error), - }); - process.exit(1); - }) - .finally(() => detachSignalHandlers()); + })().catch((error: unknown) => { + emitWizardError({ + code: ErrorCodes.InternalUnhandled, + message: error instanceof Error ? error.message : String(error), + }); + process.exit(1); + }); } diff --git a/src/lib/runners/run-wizard.ts b/src/lib/runners/run-wizard.ts index 20b83155d..2ef2ff9cf 100644 --- a/src/lib/runners/run-wizard.ts +++ b/src/lib/runners/run-wizard.ts @@ -1,6 +1,6 @@ import { VERSION } from '@shared/version'; import { logToFile, getLogFilePath } from '@utils/debug'; -import { runProgramAgent } from './run-program-agent'; +import { runProgramAgent } from '@programs/run-agent-legacy'; import { authenticate } from '@programs/authenticate'; import { getProgramConfig } from '@programs'; import { getAuditChecks } from '@programs/audit/types'; @@ -14,14 +14,11 @@ import type { TaskStreamPush as TaskStreamPushClass } from '@programs/task-strea import { resolveNoTelemetry } from './resolve-no-telemetry'; import { checkLocalServices, getLocalDev } from '@shared/local-dev'; import { runCleanups } from '@utils/wizard-abort'; -import { - commitRegisteredRunSkillCleanups, - registerRunSkillCleanup, -} from '@shared/skill-run-cleanup'; import { classifyRunFailure, emitWizardError } from '@shared/errors'; import { isRunFailure } from '@ui/mint-failure'; import { getUI } from '@ui'; import { analytics } from '@utils/analytics'; +import { join } from 'node:path'; const WIZARD_VERSION = VERSION; @@ -33,10 +30,8 @@ type Step = ProgramConfig['steps'][number]; * The frameworkContext copy is shallow and unfiltered — name keys per owning program. */ async function prepareRunSession( step: Step, - store: WizardStore, + live: WizardSession, ): Promise { - const live = store.session; - const previousLabel = live.detectedFrameworkLabel; const session = step.targetDir ? { ...live, @@ -45,37 +40,27 @@ async function prepareRunSession( } : live; if (step.onRunPrep) await step.onRunPrep(session); - if ( - session.detectedFrameworkLabel && - session.detectedFrameworkLabel !== previousLabel - ) { - store.setDetectedFramework(session.detectedFrameworkLabel); - } return session; } /** Advance one step of a composed run to completion: the auth screen - * authenticates (every later run reuses it); a step naming a child program - * runs that agent in its dir and is recorded in `completedRuns`; the + * authenticates (every later run reuses it); a step carrying its own `run` + * thunk runs that agent in its dir and is recorded in `completedRuns`; the * host program's own run screen runs `config.run`; any other screen waits for * the user to satisfy `isComplete`. */ -export async function advanceStep( +async function advanceStep( step: Step, store: WizardStore, config: ProgramConfig, ): Promise { if (step.screenId === 'auth') { - await authenticate(store.session, config.id, getUI()); + await authenticate(store.session, config.id); maybeStampAiSdkDetected(store.session); - } else if (step.runProgramId) { - await runProgramAgent( - getProgramConfig(step.runProgramId), - await prepareRunSession(step, store), - { composed: true }, - ); + } else if (step.run) { + await step.run(await prepareRunSession(step, store.session)); store.completeRunStep(step.id); } else if (step.screenId === 'run') { - await runProgramAgent(config, await prepareRunSession(step, store)); + await runProgramAgent(config, await prepareRunSession(step, store.session)); } else if (step.isComplete) { await store.waitUntil(step.isComplete); } @@ -98,8 +83,6 @@ export function runWizard( void (async () => { try { const installDir = (options.installDir as string) || process.cwd(); - // Armed until a successful exit, so every failed or interrupted exit removes new skills. - registerRunSkillCleanup(installDir); const { startTUI } = await import('@ui/tui/start-tui'); const { buildSession, RunPhase } = await import('@lib/wizard-session'); @@ -210,10 +193,11 @@ export function runWizard( config = getProgramConfig(active); } - // After the switch loop, not before: the stream bakes its program id and - // session id in at construction, so a stream built for the launch program - // would report the whole run under a program the user left on the intro - // screen. Nothing before this point produces a task to push. + // After the switch loop, not before: the stream bakes its program id, + // session id, and event-plan path in at construction, so a stream built + // for the launch program would report the whole run under a program the + // user left on the intro screen. Nothing before this point produces a + // task to push. // Consent gates the push, not the dump: `--no-telemetry` still logs. const fileDestination = createFileDestination(options.taskStreamLog); const destinations = [ @@ -232,6 +216,9 @@ export function runWizard( store: activeTui.store, programId: config.streamWorkflowId ?? config.id, destinations, + eventPlanPath: config.eventPlanFile + ? join(session.installDir, config.eventPlanFile) + : undefined, auditChecks: config.auditLedgerFile ? () => getAuditChecks(activeTui.store.session) : undefined, @@ -247,9 +234,9 @@ export function runWizard( const shown = (s: ProgramConfig['steps'][number]) => !s.show || s.show(activeTui.store.session); - if (config.steps.some((s) => s.runProgramId || s.targetDir)) { - // A composed program: its step list includes a child program run - // (self-driving runs the integration before its own + if (config.steps.some((s) => s.run || s.targetDir)) { + // A composed program: its step list splices in run steps that carry + // their own agent (self-driving runs the integration before its own // run), or scopes its own run to a picked project (error-tracking). // Walk the list once, advancing each step to completion. for (const step of config.steps) { @@ -300,18 +287,13 @@ export function runWizard( if (skipAgent && !runFailed) return s.outroDismissed; return s.skillsComplete; }); - if (signalled) return; - await activeStream.shutdown(2000); - if (signalled) return; - // Handlers stay attached, so a late signal cannot end the process before drain or commit. exitInProgress = true; - if (runFailed) { - runCleanups(); - await analytics.shutdown('error'); - } + await activeStream.shutdown(2000); + process.off('SIGINT', onSignal); + process.off('SIGTERM', onSignal); + if (runFailed) await analytics.shutdown('error'); activeTui.unmount(); - if (!runFailed) commitRegisteredRunSkillCleanups(); process.exit(runFailed ? 1 : 0); } catch (err) { // File-log first — the cleanup below can throw or exit. diff --git a/src/lib/wizard-session.ts b/src/lib/wizard-session.ts index b30035059..4739c3aa6 100644 --- a/src/lib/wizard-session.ts +++ b/src/lib/wizard-session.ts @@ -11,7 +11,6 @@ */ import { POSTHOG_LOCAL_URL, resolveLocalDev } from '@shared/local-dev'; -import { DiscoveredFeature, ScanConsent } from '@shared/scan-consent'; import { AdditionalFeature, ADDITIONAL_FEATURE_LABELS, @@ -24,7 +23,6 @@ import type { FrameworkConfig } from '@programs/types'; import type { WizardReadinessResult } from '@shared/health-checks/readiness'; import type { SettingsConflict } from '@shared/claude-settings'; import type { ApiUser, ApiProject, Credentials } from '@shared/api'; -import type { InferenceAuthProvider } from '@agent/types'; import type { CloudRegion } from '@utils/types'; import type { AskAnswers, @@ -37,7 +35,6 @@ import type { // entry would form a module cycle here. // eslint-disable-next-line @typescript-eslint/no-restricted-imports -- B2: the session becomes a TUI projection import { OutroKind } from '@agent/progress'; -import { McpOutcome, RunPhase } from '@shared/run-state'; // These shapes moved to their owners; re-exported so every session reader // keeps its import path. `Credentials` sits with the API types, the @@ -50,7 +47,6 @@ export { ADDITIONAL_FEATURE_PROMPTS, }; export { OutroKind }; -export { McpOutcome, RunPhase, ScanConsent }; export type { AskAnswers, AskQuestion, OutroData, PendingQuestion, TaskNotice }; function parseProjectIdArg(value: string | undefined): number | undefined { @@ -59,8 +55,38 @@ function parseProjectIdArg(value: string | undefined): number | undefined { return Number.isInteger(n) && n > 0 ? n : undefined; } -/** Compatibility export for session readers; detection owns the shared value. */ -export { DiscoveredFeature }; +/** Lifecycle phase of the main work (agent run, MCP install, etc.) */ +export enum RunPhase { + /** Still gathering input (intro, setup screens) */ + Idle = 'idle', + /** Main work is in progress */ + Running = 'running', + /** Main work finished successfully */ + Completed = 'completed', + /** Main work finished with an error */ + Error = 'error', +} + +/** Features discovered by the feature-discovery subagent */ +export enum DiscoveredFeature { + Stripe = 'stripe', + LLM = 'llm', +} + +/** Consent to report what local detection found (see `scanConsent` below). */ +export enum ScanConsent { + Undecided = 'undecided', + Granted = 'granted', + Declined = 'declined', +} + +/** Outcome of the MCP server installation step */ +export enum McpOutcome { + NoClients = 'no_clients', + Skipped = 'skipped', + Installed = 'installed', + Failed = 'failed', +} /** * PostHog dashboard URL emitted by the agent during a program run. @@ -82,7 +108,7 @@ export interface WizardSession { * * Only the e2e TUI host sets it, from the `E2E_ASK` env var. There is no CLI * flag, `bin.ts` never populates it, and nothing in a published build reads - * the env var — so a normal `--ci` run is unchanged. See `isAskDisabled`. + * the env var — so a normal `--ci` run is unchanged. See `shouldDisableAsk`. * * Guarding `E2E_ASK` is not enough on its own: the CI runner spreads the * whole `POSTHOG_WIZARD_*` bag into `buildSession`, which would let @@ -139,9 +165,9 @@ export interface WizardSession { /** Guards against reporting twice; consent resolves from two paths. */ warehouseSourcesReported: boolean; /** - * Latched once the organization's AI SDK stamp was considered for this login: - * by run-wizard.ts's auth step (`maybeStampAiSdkDetected`), or by runProgram, - * whose latch the legacy adapter mirrors back, whichever logs in first. + * Guards `maybeStampAiSdkDetected` against running twice: it is called from + * both run-wizard.ts's auth step and bootstrap.ts, since either can be the + * first real `authenticate()` to complete depending on the program. */ aiSdkStampReported: boolean; integration: Integration | null; @@ -166,8 +192,6 @@ export interface WizardSession { // From OAuth credentials: Credentials | null; - /** Host-supplied inference auth for legacy steps that run before the callable host. */ - inferenceAuth?: InferenceAuthProvider; /** * `role_at_organization` from `/api/users/@me/`. Null when the upstream @@ -433,3 +457,22 @@ export function buildSession(args: { pendingQuestion: null, }; } + +/** One place to ask, so a new consent state does not need three edits. */ +export function mayReportScanResults(session: WizardSession): boolean { + return session.scanConsent === ScanConsent.Granted; +} + +/** Lives here so analytics infrastructure never learns what consent means. */ +export function reportableDiscoveredFeatures( + session: WizardSession, +): DiscoveredFeature[] | undefined { + return mayReportScanResults(session) ? session.discoveredFeatures : undefined; +} + +/** Also a scan result, so it travels under the same consent as the rest. */ +export function reportablePosthogSdkDetected( + session: WizardSession, +): boolean | undefined { + return mayReportScanResults(session) ? session.posthogSdkDetected : undefined; +} diff --git a/src/programs/__tests__/authenticate.test.ts b/src/programs/__tests__/authenticate.test.ts deleted file mode 100644 index 7cf7d5b5f..000000000 --- a/src/programs/__tests__/authenticate.test.ts +++ /dev/null @@ -1,56 +0,0 @@ -import { authenticate, type AuthSession } from '../authenticate'; -import { getOrAskForProjectData } from '@utils/setup-utils'; -import { HostResolution } from '@shared/host-resolution'; -import type { ApiUser } from '@shared/api'; - -vi.mock('@utils/setup-utils', () => ({ getOrAskForProjectData: vi.fn() })); -vi.mock('@utils/debug', () => ({ logToFile: vi.fn() })); -vi.mock('@utils/analytics', () => ({ - analytics: { identifyUser: vi.fn(), setGroups: vi.fn() }, - groupsFromUser: vi.fn().mockReturnValue({}), -})); -vi.mock('@ui', () => ({ - getUI: () => { - throw new Error('authentication must use the supplied projection'); - }, -})); - -it('publishes the first login through its projection and reuses it', async () => { - const user = { distinct_id: 'user-1' } as ApiUser; - vi.mocked(getOrAskForProjectData).mockResolvedValue({ - accessToken: 'pha_test', - projectApiKey: 'phc_test', - host: HostResolution.fromApiHost('https://us.posthog.com'), - projectId: 42, - roleAtOrganization: 'admin', - user, - project: null, - missingScopes: [], - }); - const session: AuthSession = { - credentials: null, - ci: false, - signup: false, - localMcp: false, - apiProject: null, - roleAtOrganization: null, - apiUser: null, - }; - const projection = { - setCredentials: vi.fn(), - setRoleAtOrganization: vi.fn(), - setApiUser: vi.fn(), - }; - - await authenticate(session, 'metrics', projection); - await authenticate(session, 'metrics', projection); - - expect(getOrAskForProjectData).toHaveBeenCalledOnce(); - expect(projection.setCredentials).toHaveBeenCalledExactlyOnceWith( - session.credentials, - ); - expect(projection.setRoleAtOrganization).toHaveBeenCalledExactlyOnceWith( - 'admin', - ); - expect(projection.setApiUser).toHaveBeenCalledExactlyOnceWith(user); -}); diff --git a/src/programs/__tests__/binding-owner.test.ts b/src/programs/__tests__/binding-owner.test.ts deleted file mode 100644 index 824da3c38..000000000 --- a/src/programs/__tests__/binding-owner.test.ts +++ /dev/null @@ -1,88 +0,0 @@ -import { describe, expect, it } from 'vitest'; -import { - GPT5_6_TERRA_MODEL, - Harness, - Sequence, - WIZARD_ORCHESTRATOR_FLAG_KEY, - WIZARD_SELF_DRIVING_USE_PI_HARNESS_FLAG_KEY, -} from '@shared/constants'; -import { HARNESS_RUNS_TASKS } from '@agent/runner/switchboard/resolve-harness'; -import { PROGRAM_BINDINGS, resolveProgramBinding } from '../binding'; - -describe('program binding owner', () => { - it('resolves and traces a CLI sequence override ahead of an experiment', () => { - const trace = {}; - const binding = resolveProgramBinding({ - program: 'posthog-integration', - flags: { [WIZARD_ORCHESTRATOR_FLAG_KEY]: 'true' }, - cliSequence: Sequence.linear, - trace, - }); - expect(binding.sequence).toBe(Sequence.linear); - expect(trace).toMatchObject({ sequence: 'cli', harness: 'flag' }); - }); - - it('pre-resolves task roles with flag routes above role defaults', () => { - const original = PROGRAM_BINDINGS['posthog-integration']; - PROGRAM_BINDINGS['posthog-integration'] = { - sequence: Sequence.linear, - harness: Harness.anthropic, - model: 'claude-sonnet-4-5', - contextMillOverride: { - seed: { model: GPT5_6_TERRA_MODEL, thinkingLevel: 'high' }, - }, - }; - try { - const binding = resolveProgramBinding({ - program: 'posthog-integration', - flags: { [WIZARD_ORCHESTRATOR_FLAG_KEY]: 'true' }, - }); - expect(binding).toMatchObject({ - harness: Harness.pi, - roleBindings: { - seed: { - harness: Harness.pi, - model: GPT5_6_TERRA_MODEL, - thinkingLevel: 'high', - }, - }, - }); - } finally { - PROGRAM_BINDINGS['posthog-integration'] = original; - } - }); - - it('clamps a flag route without runTask while preserving the dev CLI hard-error route', () => { - HARNESS_RUNS_TASKS[Harness.anthropic] = false; - const input = { - program: 'self-driving', - flags: { [WIZARD_SELF_DRIVING_USE_PI_HARNESS_FLAG_KEY]: 'true' }, - flagPayloads: { - [WIZARD_SELF_DRIVING_USE_PI_HARNESS_FLAG_KEY]: { - model: 'gpt-5-6-sol', - harness: Harness.anthropic, - sequence: Sequence.orchestrator, - }, - }, - }; - try { - const flagTrace = {}; - expect( - resolveProgramBinding({ ...input, trace: flagTrace }).sequence, - ).toBe(Sequence.linear); - expect(flagTrace).toMatchObject({ sequence: 'runtask-clamp' }); - - const cliTrace = {}; - expect( - resolveProgramBinding({ - ...input, - cliSequence: Sequence.orchestrator, - trace: cliTrace, - }).sequence, - ).toBe(Sequence.orchestrator); - expect(cliTrace).toMatchObject({ sequence: 'cli' }); - } finally { - HARNESS_RUNS_TASKS[Harness.anthropic] = true; - } - }); -}); diff --git a/src/programs/__tests__/binding-telemetry.test.ts b/src/programs/__tests__/binding-telemetry.test.ts deleted file mode 100644 index e41b2f512..000000000 --- a/src/programs/__tests__/binding-telemetry.test.ts +++ /dev/null @@ -1,36 +0,0 @@ -import { expect, it, vi } from 'vitest'; -import { Harness, Sequence } from '@shared/constants'; -import { captureSwitchboardDecision } from '../binding-telemetry'; -import { analytics } from '@utils/analytics'; - -vi.mock('@utils/analytics', () => ({ - analytics: { wizardCapture: vi.fn() }, -})); -vi.mock('@utils/debug', async (original) => ({ - ...(await original()), - logToFile: vi.fn(), -})); - -it('records the caller-selected sequence source with the final route', () => { - captureSwitchboardDecision( - { - program: 'posthog-integration', - flags: {}, - cliSequence: Sequence.orchestrator, - trace: { harness: 'binding', model: 'binding', sequence: 'cli' }, - }, - { sequence: Sequence.orchestrator, harness: Harness.pi, model: 'm' }, - ); - - expect(analytics.wizardCapture).toHaveBeenCalledWith( - 'switchboard resolved', - expect.objectContaining({ - program: 'posthog-integration', - sequence_source: 'cli', - sequence: Sequence.orchestrator, - harness: Harness.pi, - model: 'chosen-per-task', - model_source: 'agent-prompts', - }), - ); -}); diff --git a/src/programs/__tests__/credentials.test.ts b/src/programs/__tests__/credentials.test.ts deleted file mode 100644 index 6f7d95b4d..000000000 --- a/src/programs/__tests__/credentials.test.ts +++ /dev/null @@ -1,33 +0,0 @@ -import { HostResolution } from '@shared/host-resolution'; -import type { Credentials } from '@shared/api'; -import { gatewayAuth } from '../gateway-session'; -import { createPosthogInferenceAuthProvider } from '../credentials'; - -vi.mock('../gateway-session', () => ({ gatewayAuth: vi.fn() })); - -const posthog: Credentials = { - accessToken: 'pha_fixture', - projectApiKey: 'phc_fixture', - projectId: 42, - host: HostResolution.fromRegion('us'), -}; - -it('resolves a scoped bearer on each request so the mint cache can refresh it', async () => { - const auth = { - gatewayUrl: 'https://ai-gateway.us.posthog.com', - token: 'phe_fixture', - teamId: 42, - refreshAtMs: 100, - }; - vi.mocked(gatewayAuth).mockResolvedValue(auth); - - const provider = createPosthogInferenceAuthProvider(posthog, 'metrics'); - await provider.resolve(); - expect(await provider.resolve()).toBe(auth); - expect(gatewayAuth).toHaveBeenCalledTimes(2); - expect(gatewayAuth).toHaveBeenLastCalledWith( - posthog.host, - 'pha_fixture', - 'metrics', - ); -}); diff --git a/src/programs/__tests__/error-tracking.test.ts b/src/programs/__tests__/error-tracking.test.ts index 0dde92f92..3032f3723 100644 --- a/src/programs/__tests__/error-tracking.test.ts +++ b/src/programs/__tests__/error-tracking.test.ts @@ -1,11 +1,9 @@ import { beforeEach, describe, expect, test, vi } from 'vitest'; import type { ProgramRun } from '@programs/program-run'; -import type { ProgramRunHost } from '@programs/host-capabilities'; import { Integration } from '@shared/constants'; import type { AgenticDetectionReport } from '@programs/detection/agentic'; import { detectFramework } from '@programs/detection/index'; -import { scopeInstallDirToProject } from '@programs/detection/project-scope'; import { ErrorCodes } from '@shared/errors'; import { ERROR_TRACKING_TIPS } from '@ui/tui/decks/error-tracking/tips'; import { @@ -18,10 +16,15 @@ import { } from '@programs/error-tracking/index'; import { VARIANTS_REQUIRING_POSTHOG_CLI } from '@programs/error-tracking-upload-source-maps/detect'; import { preinstallPostHogCliOnce } from '@programs/shared/posthog-cli-preinstall'; -import { buildSession, type WizardSession } from '@lib/wizard-session'; +import type { WizardSession } from '@lib/wizard-session'; +import type { ProgramRunHost } from '@programs/host-capabilities'; +import { scopeInstallDirToProject } from '@programs/detection/project-scope'; +import { + testProgramCiHost, + testProgramRunHost, +} from '../../../test/program-host'; import { analytics } from '@utils/analytics'; import { wizardAbort } from '@utils/wizard-abort'; -import { testProgramCiHost } from '../../../test/program-host'; vi.mock('@programs/detection/index', async (importOriginal) => ({ ...(await importOriginal()), @@ -43,17 +46,10 @@ vi.mock('@utils/wizard-abort', async (importOriginal) => ({ })); const resolveRun = errorTrackingConfig.run as ( - session: { integration: Integration | null }, + session: WizardSession, host: ProgramRunHost, ) => Promise; -const runHost = (): ProgramRunHost => ({ - getFrameworkContext: vi.fn(), - setFrameworkContext: vi.fn(), - warn: vi.fn(), - uploadEnvironmentVariables: vi.fn().mockResolvedValue([]), -}); - const step = (id: string) => errorTrackingConfig.steps.find((s) => s.id === id); beforeEach(() => { @@ -78,7 +74,10 @@ describe('error-tracking program', () => { // There is no bare `error-tracking` menu entry; a seeded skillId would // send the linear path to a skill-not-found abort and mislead the intro. expect(errorTrackingConfig.skillId).toBeUndefined(); - const run = await resolveRun({ integration: null }, runHost()); + const run = await resolveRun( + { integration: null } as WizardSession, + testProgramRunHost(), + ); expect(run.skillId).toBeUndefined(); }); @@ -177,23 +176,27 @@ describe('error-tracking project picker report', () => { describe('error-tracking ciPreRun', () => { test('stops KMP before it sets the framework', async () => { vi.mocked(detectFramework).mockResolvedValue(Integration.kmp); - const session = buildSession({ installDir: '/tmp/error-tracking-ci' }); - const host = testProgramCiHost(); + const session = { + installDir: '/tmp/error-tracking-ci', + frameworkContext: {}, + } as unknown as WizardSession; + const host = testProgramCiHost(); await errorTrackingConfig.ciPreRun?.(session, host); expect(scopeInstallDirToProject).toHaveBeenCalledWith(session, host); + expect(wizardAbort).toHaveBeenCalledWith( expect.objectContaining({ code: ErrorCodes.DetectUnsupportedPlatform }), ); - expect(session.integration).toBeNull(); + expect(session.integration).toBeUndefined(); }); }); describe('error-tracking run config', () => { test('pre-installs posthog-cli when run resolves, after the project pick', async () => { - const host = runHost(); - await resolveRun({ integration: Integration.swift }, host); + const host = { ...testProgramRunHost(), warn: vi.fn() }; + await resolveRun({ integration: Integration.swift } as WizardSession, host); expect(preinstallPostHogCliOnce).toHaveBeenCalledWith( 'error tracking posthog-cli preinstall failed', @@ -201,13 +204,15 @@ describe('error-tracking run config', () => { expect.any(Function), ); const warn = vi.mocked(preinstallPostHogCliOnce).mock.calls[0]?.[2]; - if (!warn) throw new Error('missing preinstall warning callback'); - warn('install warning'); + warn?.('install warning'); expect(host.warn).toHaveBeenCalledWith('install warning'); }); test('skips the pre-install for platforms without symbol upload', async () => { - await resolveRun({ integration: Integration.nextjs }, runHost()); + await resolveRun( + { integration: Integration.nextjs } as WizardSession, + testProgramRunHost(), + ); expect(preinstallPostHogCliOnce).not.toHaveBeenCalled(); }); diff --git a/src/programs/__tests__/flow-traces.test.ts b/src/programs/__tests__/flow-traces.test.ts index b862ffac6..30f7ec314 100644 --- a/src/programs/__tests__/flow-traces.test.ts +++ b/src/programs/__tests__/flow-traces.test.ts @@ -125,7 +125,7 @@ function advance(store: WizardStore, screen: string): boolean { (!st.show || st.show(s)) && (!st.isComplete || !st.isComplete(s)), ); - if (runStep?.runProgramId) { + if (runStep?.run) { store.completeRunStep(runStep.id); } else { store.setRunPhase(RunPhase.Running); diff --git a/src/programs/__tests__/posthog-cli-preinstall.test.ts b/src/programs/__tests__/posthog-cli-preinstall.test.ts index 9705b2a4c..6870bedce 100644 --- a/src/programs/__tests__/posthog-cli-preinstall.test.ts +++ b/src/programs/__tests__/posthog-cli-preinstall.test.ts @@ -4,10 +4,10 @@ import { preinstallPostHogCliOnce, resetPostHogCliPreinstallForTests, } from '@programs/shared/posthog-cli-preinstall'; -import { installOrUpdatePostHogCli } from '@shared/posthog-cli-install'; +import { installOrUpdatePostHogCli } from '@steps/install-cli-steering'; import { analytics } from '@utils/analytics'; -vi.mock('@shared/posthog-cli-install', () => ({ +vi.mock('@steps/install-cli-steering', () => ({ installOrUpdatePostHogCli: vi.fn(), })); vi.mock('@utils/analytics', () => ({ @@ -26,16 +26,12 @@ describe('preinstallPostHogCliOnce', () => { preinstallPostHogCliOnce( 'source maps posthog-cli preinstall failed', - { - variant: 'ios', - }, + { variant: 'ios' }, warn, ); preinstallPostHogCliOnce( 'error tracking posthog-cli preinstall failed', - { - integration: 'swift', - }, + { integration: 'swift' }, warn, ); @@ -50,9 +46,7 @@ describe('preinstallPostHogCliOnce', () => { preinstallPostHogCliOnce( 'error tracking posthog-cli preinstall failed', - { - integration: 'swift', - }, + { integration: 'swift' }, warn, ); @@ -69,9 +63,7 @@ describe('preinstallPostHogCliOnce', () => { preinstallPostHogCliOnce( 'error tracking posthog-cli preinstall failed', - { - integration: 'swift', - }, + { integration: 'swift' }, warn, ); diff --git a/src/programs/__tests__/program-file-watchers.test.ts b/src/programs/__tests__/program-file-watchers.test.ts deleted file mode 100644 index 2a4a35258..000000000 --- a/src/programs/__tests__/program-file-watchers.test.ts +++ /dev/null @@ -1,143 +0,0 @@ -import { - mkdtempSync, - rmSync, - symlinkSync, - unlinkSync, - writeFileSync, -} from 'node:fs'; -import { tmpdir } from 'node:os'; -import { join } from 'node:path'; -import { watchAuditLedger } from '../audit/watch-ledger'; -import { - ProgramEventPlanWatcher, - normalizeEventPlan, -} from '../posthog-integration/watch-event-plan'; -import { AUDIT_CHECKS_FILE } from '@shared/audit-ledger'; -import { EVENT_PLAN_FILE } from '@shared/constants'; - -describe('program-owned file watchers', () => { - let installDir: string; - - beforeEach(() => { - installDir = mkdtempSync(join(tmpdir(), 'wizard-program-watchers-')); - }); - - afterEach(() => { - rmSync(installDir, { recursive: true, force: true }); - }); - - it('ignores an old audit ledger, then reports this run’s update until stopped', () => { - const path = join(installDir, AUDIT_CHECKS_FILE); - const fresh = [{ id: 'new', area: 'Events', label: 'new', status: 'pass' }]; - writeFileSync(path, JSON.stringify([{ ...fresh[0], status: 'pending' }])); - const onChecks = vi.fn(); - const handle = watchAuditLedger(installDir, AUDIT_CHECKS_FILE, onChecks); - try { - handle.refresh(); - writeFileSync(path, JSON.stringify(fresh)); - handle.refresh(); - handle.stop(); - writeFileSync(path, JSON.stringify([{ ...fresh[0], id: 'later' }])); - handle.refresh(); - - expect(onChecks.mock.calls).toEqual([[fresh]]); - } finally { - handle.stop(); - } - }); - - it('reports only the first non-empty event plan this run writes', () => { - const path = join(installDir, EVENT_PLAN_FILE); - writeFileSync(path, JSON.stringify([{ event_name: 'stale_event' }])); - const onEvents = vi.fn(); - const watcher = new ProgramEventPlanWatcher(path, onEvents); - try { - watcher.start(); - watcher.refresh(); - for (const plan of [ - [], - [{ event_name: 'first' }], - [{ event_name: 'later_event' }], - ]) { - writeFileSync(path, JSON.stringify(plan)); - watcher.refresh(); - } - - expect(onEvents.mock.calls).toEqual([ - [[{ name: 'first', description: '' }]], - ]); - } finally { - watcher.stop(); - } - }); - - it('releases an event-plan watcher even when its consumer throws', () => { - const path = join(installDir, EVENT_PLAN_FILE); - const watcher = new ProgramEventPlanWatcher(path, () => { - throw new Error('store unavailable'); - }); - const stop = vi.spyOn(watcher, 'stop'); - watcher.start(); - writeFileSync(path, JSON.stringify([{ event_name: 'first_event' }])); - watcher.refresh(); - - expect(stop).toHaveBeenCalledOnce(); - }); - - it('rejects oversized and symbolic-link event plan files', () => { - const path = join(installDir, EVENT_PLAN_FILE); - const onEvents = vi.fn(); - const watcher = new ProgramEventPlanWatcher(path, onEvents); - try { - watcher.start(); - writeFileSync( - path, - JSON.stringify([{ event_name: 'x'.repeat(300_000) }]), - ); - watcher.refresh(); - - unlinkSync(path); - const target = join(installDir, 'external-plan.json'); - writeFileSync(target, JSON.stringify([{ event_name: 'linked_event' }])); - symlinkSync(target, path); - watcher.refresh(); - - expect(onEvents).not.toHaveBeenCalled(); - } finally { - watcher.stop(); - } - }); -}); - -describe('normalizeEventPlan', () => { - it('normalizes canonical fields and legacy fallbacks', () => { - expect( - normalizeEventPlan([ - { event_name: 'signed_up', event_description: 'User signs up' }, - { name: 'invited_user', description: 'User sends an invite' }, - { event: 'created_team' }, - { event_name: 42, name: 'valid_fallback' }, - { event_name: 'x'.repeat(401) }, - { event_name: ' ' }, - { description: 'missing name' }, - ]), - ).toEqual([ - { name: 'signed_up', description: 'User signs up' }, - { name: 'invited_user', description: 'User sends an invite' }, - { name: 'created_team', description: '' }, - { name: 'valid_fallback', description: '' }, - ]); - }); - - it('caps event count and description length', () => { - const events = Array.from({ length: 60 }, (_, index) => ({ - event_name: `event_${index}`, - event_description: 'x'.repeat(5000), - })); - - const normalized = normalizeEventPlan(events); - - expect(normalized).toHaveLength(50); - expect(normalized?.[0].description).toHaveLength(4000); - }); -}); diff --git a/src/programs/__tests__/program-store.test.ts b/src/programs/__tests__/program-store.test.ts index a18231b9c..1383114dc 100644 --- a/src/programs/__tests__/program-store.test.ts +++ b/src/programs/__tests__/program-store.test.ts @@ -92,7 +92,6 @@ it('copies invocation data on write, on read and in each emitted snapshot', () = projectId: 42, host: { region: 'us', apiHost: 'https://example.test' }, } as Credentials; - const frameworkValue = { paths: ['apps/web'] }; const binding = { sequence: Sequence.linear, harness: Harness.anthropic, @@ -100,25 +99,18 @@ it('copies invocation data on write, on read and in each emitted snapshot', () = }; store.setAuthenticated({ credentials, apiProject: null, apiUser: null }); - store.setFrameworkContext('selectedProject', frameworkValue); - store.setEventPlan([{ name: 'signup', description: 'Account created' }]); store.setBinding(binding); const written = store.readData(); credentials.accessToken = 'changed input'; - frameworkValue.paths.push('changed input'); binding.model = 'changed input'; - observed[3].data.eventPlan[0].name = 'changed by observer'; - store.readData().eventPlan.push({ name: 'changed', description: 'output' }); + observed[1].data.binding!.model = 'changed by observer'; + store.readData().credentials!.accessToken = 'changed output'; - expect(observed).toHaveLength(4); - expect(observed[0].data.eventPlan).toEqual([]); + expect(observed).toHaveLength(2); + expect(observed[0].data.binding).toBeNull(); expect(store.readData()).toEqual(written); expect(written).toMatchObject({ credentials: { accessToken: 'test-access-token' }, - detection: { - frameworkContext: { selectedProject: { paths: ['apps/web'] } }, - }, - eventPlan: [{ name: 'signup', description: 'Account created' }], binding: { model: 'claude-test' }, }); }); diff --git a/src/programs/__tests__/refresh-access-token-if-needed.test.ts b/src/programs/__tests__/refresh-access-token-if-needed.test.ts new file mode 100644 index 000000000..013d53051 --- /dev/null +++ b/src/programs/__tests__/refresh-access-token-if-needed.test.ts @@ -0,0 +1,134 @@ +import { refreshCredentialsIfNeeded } from '../authenticate'; +import { refreshAccessToken } from '@utils/oauth'; +import { OAuthError } from '@utils/oauth-errors'; +import { + isGrantRevoked, + resetAuthSessionState, +} from '@shared/auth-session-state'; +import type { Credentials } from '@lib/wizard-session'; + +vi.mock('@utils/oauth', () => ({ refreshAccessToken: vi.fn() })); +vi.mock('@utils/debug', () => ({ logToFile: vi.fn() })); +vi.mock('@utils/analytics', () => ({ + analytics: { wizardCapture: vi.fn() }, + groupsFromUser: vi.fn(), +})); + +vi.mock('@ui', () => ({ getUI: vi.fn() })); + +const mockedRefresh = refreshAccessToken as Mock; + +const refresh = (credentials: Partial) => + refreshCredentialsIfNeeded(credentials as Credentials, {}); + +/** Aging enough to be under the 50-minute threshold. */ +const aging = (over: Partial = {}): Partial => ({ + accessToken: 'pha_old', + refreshToken: 'phr_old', + expiresAt: Date.now() + 20 * 60 * 1000, + ...over, +}); + +describe('refreshCredentialsIfNeeded', () => { + beforeEach(() => { + vi.clearAllMocks(); + resetAuthSessionState(); + }); + + it('is a no-op without a refresh token (CI api-key runs, refresh-less grants)', async () => { + await refresh({ accessToken: 'pha_ci_key', expiresAt: 0 }); + expect(mockedRefresh).not.toHaveBeenCalled(); + }); + + it('skips a token that still has most of its lifetime left', async () => { + await refresh(aging({ expiresAt: Date.now() + 59 * 60 * 1000 })); + expect(mockedRefresh).not.toHaveBeenCalled(); + }); + + // `?? 0` would read as "expired" and spend a rotation on every run. + it('skips a credential carrying a refresh token but no expiry', async () => { + await refresh({ accessToken: 'pha_old', refreshToken: 'phr_old' }); + expect(mockedRefresh).not.toHaveBeenCalled(); + }); + + it('refreshes an aging token and stores the rotated refresh token', async () => { + mockedRefresh.mockResolvedValueOnce({ + access_token: 'pha_new', + refresh_token: 'phr_rotated', + expires_in: 3600, + token_type: 'Bearer', + scope: 'project:read', + }); + const refreshed = await refreshCredentialsIfNeeded( + aging({ projectId: 7 }) as Credentials, + { baseUrl: 'https://posthog.example' }, + ); + + expect(mockedRefresh).toHaveBeenCalledWith( + 'phr_old', + 'https://posthog.example', + undefined, + ); + expect(refreshed.accessToken).toBe('pha_new'); + expect(refreshed.refreshToken).toBe('phr_rotated'); + // Unrelated fields survive the swap. + expect(refreshed.projectId).toBe(7); + }); + + it('refreshes under the minting client id when the credential carries one (provisioning signups)', async () => { + mockedRefresh.mockResolvedValueOnce({ + access_token: 'pha_new', + expires_in: 3600, + token_type: 'Bearer', + scope: 'project:read', + }); + + await refresh(aging({ oauthClientId: 'client_us_provisioning' })); + + expect(mockedRefresh).toHaveBeenCalledWith( + 'phr_old', + undefined, + 'client_us_provisioning', + ); + }); + + it('replaces the credentials object rather than mutating it in place', async () => { + mockedRefresh.mockResolvedValueOnce({ + access_token: 'pha_new', + expires_in: 3600, + token_type: 'Bearer', + scope: 'project:read', + }); + const before = aging(); + + const refreshed = await refresh(before); + + expect(refreshed).not.toBe(before); + expect(before.accessToken).toBe('pha_old'); + // No rotation in the response: the old refresh token has to carry over. + expect(refreshed.refreshToken).toBe('phr_old'); + }); + + it('keeps the existing token and does not throw when the refresh fails', async () => { + mockedRefresh.mockRejectedValueOnce(new Error('network down')); + const before = aging(); + + await expect(refresh(before)).resolves.toBe(before); + }); + + it('marks the grant revoked on invalid_grant, so a later 401 can name the cause', async () => { + mockedRefresh.mockRejectedValueOnce(new OAuthError('invalid_grant')); + + await refresh(aging()); + + expect(isGrantRevoked()).toBe(true); + }); + + it('leaves the grant unmarked for a transport failure, which says nothing about the login', async () => { + mockedRefresh.mockRejectedValueOnce(new Error('ETIMEDOUT')); + + await refresh(aging()); + + expect(isGrantRevoked()).toBe(false); + }); +}); diff --git a/src/programs/__tests__/run-agent-legacy.test.ts b/src/programs/__tests__/run-agent-legacy.test.ts index 486899083..1fc0f6b99 100644 --- a/src/programs/__tests__/run-agent-legacy.test.ts +++ b/src/programs/__tests__/run-agent-legacy.test.ts @@ -1,41 +1,23 @@ import { runNonInteractive } from '@lib/runners/run-non-interactive'; import { runWizard } from '@lib/runners/run-wizard'; -import { authenticate } from '@programs/authenticate'; -import { runProgramAgent } from '@lib/runners/run-program-agent'; -import { runAgent, RunOutcome, type RunResult } from '@agent/runner'; -import { Harness, Integration, Sequence } from '@shared/constants'; -import { uploadEnvironmentVariablesStep } from '@steps/upload-environment-variables'; import { - buildSession, - DiscoveredFeature, - OutroKind, - ScanConsent, -} from '@lib/wizard-session'; -import type { ApiUser } from '@shared/api'; + authenticate, + refreshCredentialsIfNeeded, +} from '@programs/authenticate'; +import { runProgramAgent } from '../run-agent-legacy'; +import { runAgent, RunOutcome, type RunResult } from '@agent/runner'; +import { Harness, Sequence } from '@shared/constants'; +import { buildSession, OutroKind } from '@lib/wizard-session'; import { HostResolution } from '@shared/host-resolution'; import { LoggingUI } from '@ui/logging-ui'; import { InkUI } from '@ui/tui/ink-ui'; -import * as ledgerWatch from '../audit/watch-ledger'; -import { auditConfig } from '../audit/index'; -import { AUDIT_SEED_CHECKS } from '../audit/seed'; -import { AUDIT_CHECKS_FILE, AUDIT_CHECKS_KEY } from '../audit/types'; -import { EVENT_PLAN_FILE } from '../posthog-integration/constants'; import { startTUI } from '@ui/tui/start-tui'; import { WizardStore } from '@ui/tui/store'; import { getUI, setUI } from '@ui'; import { analytics } from '@utils/analytics'; import { initLogFile, logToFile } from '@utils/debug'; -import { clearCleanup, runCleanups, wizardAbort } from '@utils/wizard-abort'; -import * as fs from 'node:fs'; -import * as os from 'node:os'; -import * as path from 'node:path'; +import { wizardAbort } from '@utils/wizard-abort'; import { ErrorCodes } from '@shared/errors'; -import { - checkAllSettingsConflicts, - restoreClaudeSettings, -} from '@shared/claude-settings'; -import { refreshAccessToken } from '@utils/oauth-token'; -import { errorTrackingUploadSourceMapsConfig } from '../error-tracking-upload-source-maps/index'; import type { ProgramConfig } from '../program-step'; import type { ProgramRun } from '../program-run'; @@ -52,6 +34,10 @@ vi.mock('@utils/environment', async (original) => ({ ...(await original()), readEnvironment: () => ({}), })); +vi.mock('@agent/gateway-session', async (original) => ({ + ...(await original()), + configureGatewayFromCIEnvironment: vi.fn(), +})); vi.mock('@programs/task-stream/index', () => ({ TaskStreamPush: class { attach = vi.fn(); @@ -85,26 +71,23 @@ vi.mock('@agent/runner', async (original) => ({ })); vi.mock('@programs/authenticate', () => ({ authenticate: vi.fn().mockResolvedValue(undefined), -})); -vi.mock('@steps/upload-environment-variables', async (original) => ({ - ...(await original()), - uploadEnvironmentVariablesStep: vi.fn().mockResolvedValue(['POSTHOG_KEY']), + refreshCredentialsIfNeeded: vi.fn((credentials: unknown) => + Promise.resolve(credentials), + ), })); vi.mock('@shared/claude-settings', () => ({ checkAllSettingsConflicts: vi.fn().mockReturnValue([]), restoreClaudeSettings: vi.fn(), })); -vi.mock('@utils/wizard-abort', async (original) => { - const actual = await original(); - return { - ...actual, - wizardAbort: vi.fn().mockResolvedValue(undefined), - }; -}); +vi.mock('@utils/wizard-abort', async (original) => ({ + ...(await original()), + registerCleanup: vi.fn(), + wizardAbort: vi.fn().mockResolvedValue(undefined), +})); vi.mock('../posthog-integration/detect', () => ({ maybeStampAiSdkDetected: vi.fn(), + stampAiSdkDetected: vi.fn(), })); -vi.mock('@utils/oauth-token', () => ({ refreshAccessToken: vi.fn() })); const program = (id: ProgramConfig['id'] = 'metrics'): ProgramConfig => ({ id, @@ -157,7 +140,6 @@ const finishRun: typeof runAgent = (_config, _input, options) => { }; beforeEach(() => { - clearCleanup(); vi.clearAllMocks(); vi.mocked(authenticate).mockImplementation((sess) => { sess.credentials = session().credentials; @@ -167,10 +149,7 @@ beforeEach(() => { logSpy = vi.spyOn(console, 'log').mockImplementation(() => undefined); vi.mocked(runAgent).mockImplementation(finishRun); }); -afterEach(() => { - clearCleanup(); - logSpy.mockRestore(); -}); +afterEach(() => logSpy.mockRestore()); it.each([ ['metrics', Harness.pi, Sequence.orchestrator], @@ -239,150 +218,25 @@ it('clamps a composed program to linear and keeps host analytics alive', async ( expect(analytics.shutdown).toHaveBeenCalledExactlyOnceWith('success'); }); -it('supplies environment upload through the run host for the requested project', async () => { +it('supplies the live UI as the run host, not the session it was handed', async () => { + const ui = getUI(); + vi.spyOn(ui, 'getFrameworkContext').mockReturnValue('ios'); + const write = vi.spyOn(ui, 'setFrameworkContext'); + const warn = vi.spyOn(ui.log, 'warn'); const config = program(); - config.run = async (_session, host) => { - const uploaded = await host.uploadEnvironmentVariables( - { POSTHOG_KEY: 'phc_test' }, - Integration.nextjs, - '/repo/apps/web', - ); - expect(uploaded).toEqual(['POSTHOG_KEY']); - return program().run as ProgramRun; + let read: unknown; + config.run = (_session, host) => { + read = host.getFrameworkContext('selectedVariant'); + host.setFrameworkContext('sourceMapsCompletedVariant', 'ios'); + host.warn('careful'); + return Promise.resolve(program().run as ProgramRun); }; await runProgramAgent(config, session()); - expect(uploadEnvironmentVariablesStep).toHaveBeenCalledWith( - { POSTHOG_KEY: 'phc_test' }, - { - integration: Integration.nextjs, - session: { installDir: '/repo/apps/web' }, - }, - ); -}); - -it('reads completion data when each hook runs, after late URL updates', async () => { - const currentSession = session(); - const postRun = vi.fn().mockResolvedValue(undefined); - const buildOutroData = vi.fn().mockReturnValue({ - kind: OutroKind.Success, - message: 'Done', - }); - const buildOutroNextSteps = vi.fn().mockReturnValue(undefined); - const config = program(); - config.run = { - ...(config.run as ProgramRun), - postRun, - buildOutroData, - buildOutroNextSteps, - }; - vi.mocked(runAgent).mockImplementationOnce(async (runConfig, input) => { - currentSession.dashboardUrl = 'https://us.posthog.com/dashboard/42'; - await runConfig.hooks?.postRun?.(input.credentials); - currentSession.notebookUrl = 'https://us.posthog.com/notebook/7'; - runConfig.hooks?.buildOutroData?.(input.credentials); - runConfig.hooks?.buildOutroNextSteps?.(input.credentials, ['seeded']); - return { outcome: RunOutcome.Success, snapshot }; - }); - - await runProgramAgent(config, currentSession); - - expect(postRun).toHaveBeenCalledWith( - { - signup: false, - dashboardUrl: 'https://us.posthog.com/dashboard/42', - notebookUrl: null, - }, - currentSession.credentials, - ); - expect(buildOutroData).toHaveBeenCalledWith( - { - signup: false, - dashboardUrl: 'https://us.posthog.com/dashboard/42', - notebookUrl: 'https://us.posthog.com/notebook/7', - }, - currentSession.credentials, - ); - expect(buildOutroNextSteps).toHaveBeenCalledWith( - { - signup: false, - dashboardUrl: 'https://us.posthog.com/dashboard/42', - notebookUrl: 'https://us.posthog.com/notebook/7', - }, - currentSession.credentials, - ['seeded'], - ); -}); - -it('projects an audit ledger update from program data through the legacy runner UI bridge', async () => { - const installDir = fs.mkdtempSync( - path.join(os.tmpdir(), 'wizard-audit-bridge-'), - ); - const currentSession = session(); - currentSession.installDir = installDir; - const config = program('audit'); - config.auditLedgerFile = AUDIT_CHECKS_FILE; - config.auditSeedChecks = AUDIT_SEED_CHECKS; - const ui = new LoggingUI(); - const setFrameworkContext = vi.spyOn(ui, 'setFrameworkContext'); - setUI(ui); - const watchLedger = vi.spyOn(ledgerWatch, 'watchAuditLedger'); - const checksSent = () => - setFrameworkContext.mock.calls.filter(([key]) => key === AUDIT_CHECKS_KEY); - const checks = [{ id: 'new', area: 'Events', label: 'new', status: 'pass' }]; - vi.mocked(runAgent).mockImplementationOnce(async () => { - fs.writeFileSync( - path.join(installDir, AUDIT_CHECKS_FILE), - JSON.stringify(checks), - ); - // The update reaches the UI while the agent still runs. - await vi.waitFor( - () => - expect(setFrameworkContext).toHaveBeenCalledWith( - AUDIT_CHECKS_KEY, - checks, - ), - { timeout: 7000 }, - ); - return { outcome: RunOutcome.Success, snapshot }; - }); - try { - await runProgramAgent(config, currentSession); - // runProgram's watcher is the only one; the projection forwards each value once. - expect(watchLedger).toHaveBeenCalledOnce(); - expect(checksSent()).toEqual([ - [AUDIT_CHECKS_KEY, AUDIT_SEED_CHECKS], - [AUDIT_CHECKS_KEY, checks], - ]); - } finally { - watchLedger.mockRestore(); - fs.rmSync(installDir, { recursive: true, force: true }); - } -}, 8000); - -it('hands the CI bearer from the token file to ciPreRun and the agent through the session', async () => { - const installDir = fs.mkdtempSync(path.join(os.tmpdir(), 'wizard-ci-auth-')); - const tokenFile = path.join(installDir, 'gateway-token'); - fs.writeFileSync(tokenFile, 'fixed-ci-bearer'); - vi.stubEnv('WIZARD_CI_GATEWAY_TOKEN_FILE', tokenFile); - try { - const ciPreRun = vi.fn((_session: ReturnType) => - Promise.resolve(), - ); - runNonInteractive( - { ...program(), ciPreRun }, - { apiKey: 'phx_test', projectId: '42', installDir, telemetry: false }, - 'ci', - ); - await vi.waitFor(() => expect(streamShutdown).toHaveBeenCalledOnce()); - const provider = ciPreRun.mock.calls[0]?.[0].inferenceAuth; - expect(provider).toBeDefined(); - expect(vi.mocked(runAgent).mock.calls[0]?.[1].inferenceAuth).toBe(provider); - } finally { - vi.unstubAllEnvs(); - fs.rmSync(installDir, { recursive: true, force: true }); - } + expect(read).toBe('ios'); + expect(write).toHaveBeenCalledWith('sourceMapsCompletedVariant', 'ios'); + expect(warn).toHaveBeenCalledWith('careful'); }); it.each([ @@ -492,91 +346,29 @@ it('rethrows the original crash for the outer runner', async () => { expect(analytics.shutdown).not.toHaveBeenCalled(); }); -it.each(['ci', 'headless'] as const)( - 'removes new Wizard skills when %s stream settlement fails after agent success', - async (mode) => { - const installDir = fs.mkdtempSync( - path.join(os.tmpdir(), `wizard-${mode}-late-failure-`), - ); - const skillDir = path.join(installDir, '.claude', 'skills', 'unfinished'); - const tokenFile = path.join(installDir, 'gateway-token'); - fs.writeFileSync(tokenFile, 'fixed-ci-bearer'); - vi.stubEnv('WIZARD_CI_GATEWAY_TOKEN_FILE', tokenFile); - const settlementError = new Error('task stream failed to flush'); - streamShutdown.mockRejectedValueOnce(settlementError); - vi.mocked(wizardAbort).mockImplementationOnce(() => { - runCleanups(); - return Promise.resolve(undefined as never); - }); - vi.mocked(runAgent).mockImplementationOnce(() => { - fs.mkdirSync(skillDir, { recursive: true }); - fs.writeFileSync(path.join(skillDir, '.posthog-wizard'), ''); - return Promise.resolve({ outcome: RunOutcome.Success, snapshot }); - }); - - try { - runNonInteractive( - program(), - { apiKey: 'phx_test', projectId: '1', installDir, telemetry: false }, - mode, - ); - await vi.waitFor(() => expect(wizardAbort).toHaveBeenCalledOnce()); - expect(wizardAbort).toHaveBeenCalledWith( - expect.objectContaining({ error: settlementError }), - ); - // The agent run succeeded first, so 'success' goes out before wizardAbort's 'error'. - expect(analytics.shutdown).toHaveBeenCalledExactlyOnceWith('success'); - expect( - vi.mocked(analytics.shutdown).mock.invocationCallOrder[0], - ).toBeLessThan(vi.mocked(wizardAbort).mock.invocationCallOrder[0]); - expect(streamShutdown).toHaveBeenCalledTimes(2); - expect(fs.existsSync(skillDir)).toBe(false); - } finally { - vi.unstubAllEnvs(); - fs.rmSync(installDir, { recursive: true, force: true }); - } - }, -); +it('rethrows a login failure for the CLI roots, before the agent starts', async () => { + const error = new Error('OAuth cancelled'); + vi.mocked(authenticate).mockRejectedValueOnce(error); + await expect(runProgramAgent(program(), session())).rejects.toBe(error); + expect(runAgent).not.toHaveBeenCalled(); + expect(wizardAbort).not.toHaveBeenCalled(); +}); -it("cleans a marked install when program setup throws before the functional runner, through the CLI root's drain", async () => { - const installDir = fs.mkdtempSync( - path.join(os.tmpdir(), 'wizard-setup-cleanup-'), +it('projects a refreshed token and the AI SDK stamp back onto the session', async () => { + const current = session(); + const setAccessToken = vi.spyOn(getUI(), 'setAccessToken'); + vi.mocked(refreshCredentialsIfNeeded).mockImplementationOnce((credentials) => + Promise.resolve({ ...credentials, accessToken: 'pha_refreshed' }), ); - const skillDir = path.join(installDir, '.claude', 'skills', 'setup-install'); - const setupFailure = new Error('program setup failed'); - const failingProgram = program(); - failingProgram.run = () => { - fs.mkdirSync(skillDir, { recursive: true }); - fs.writeFileSync(path.join(skillDir, '.posthog-wizard'), ''); - throw setupFailure; - }; - const actual = await vi.importActual( - '@utils/wizard-abort', + await runProgramAgent(program(), current); + expect(vi.mocked(runAgent).mock.calls[0][1].credentials.accessToken).toBe( + 'pha_refreshed', ); - vi.mocked(wizardAbort).mockImplementationOnce(actual.wizardAbort); - const exit = vi - .spyOn(process, 'exit') - .mockImplementation(() => undefined as never); - const stderr = vi - .spyOn(process.stderr, 'write') - .mockImplementation(() => true); - try { - runNonInteractive( - failingProgram, - { apiKey: 'phx_test', projectId: '1', installDir, telemetry: false }, - 'headless', - ); - await vi.waitFor(() => expect(exit).toHaveBeenCalled()); - expect(wizardAbort).toHaveBeenCalledWith( - expect.objectContaining({ error: setupFailure }), - ); - expect(fs.existsSync(skillDir)).toBe(false); - expect(runAgent).not.toHaveBeenCalled(); - } finally { - exit.mockRestore(); - stderr.mockRestore(); - fs.rmSync(installDir, { recursive: true, force: true }); - } + expect(current.credentials?.accessToken).toBe('pha_refreshed'); + // Only the token fields change; the login keeps its host. + expect(current.credentials?.host).toBeInstanceOf(HostResolution); + expect(setAccessToken).toHaveBeenCalledExactlyOnceWith(current.credentials); + expect(current.aiSdkStampReported).toBe(true); }); it.each([ @@ -640,45 +432,29 @@ it('keeps a headless run a success when its terminal analytics flush fails', asy ); }); -it.each([ - ['SIGINT', [[130]], false], - ['SIGTERM', [[143]], false], - ['completion', [], true], -] as const)( - 'a headless run ended by %s exits %j and keeps its new skill: %s', - async (ending, exits, kept) => { - const installDir = fs.mkdtempSync(path.join(os.tmpdir(), 'wizard-ci-')); - const skillDir = path.join(installDir, '.claude', 'skills', 'installed'); - const exit = vi - .spyOn(process, 'exit') - .mockImplementation(() => undefined as never); - const config: ProgramConfig = { - ...program(), - ciPreRun: () => { - fs.mkdirSync(skillDir, { recursive: true }); - fs.writeFileSync(path.join(skillDir, '.posthog-wizard'), ''); - if (ending !== 'completion') process.emit(ending); - return Promise.resolve(); - }, - }; - try { - runNonInteractive( - config, - { apiKey: 'phx_test', projectId: '1', installDir, telemetry: false }, - 'headless', - ); - await vi.waitFor(() => expect(streamShutdown).toHaveBeenCalledOnce()); - await new Promise((resolve) => setImmediate(resolve)); - runCleanups(); +it('supplies the logging UI as the CI host for ciPreRun', async () => { + const config = program(); + config.ciPreRun = (_session, host) => { + host.log.info('Scanning the repo'); + host.log.warn('Scan failed'); + return Promise.resolve(); + }; - expect(exit.mock.calls).toEqual(exits); - expect(fs.existsSync(skillDir)).toBe(kept); - } finally { - exit.mockRestore(); - fs.rmSync(installDir, { recursive: true, force: true }); - } - }, -); + runNonInteractive( + config, + { + apiKey: 'phx_test', + projectId: '1', + installDir: '/tmp/adapter-test', + telemetry: false, + }, + 'headless', + ); + await vi.waitFor(() => expect(streamShutdown).toHaveBeenCalledOnce()); + + expect(logSpy).toHaveBeenCalledWith('│ Scanning the repo'); + expect(logSpy).toHaveBeenCalledWith('▲ Scan failed'); +}); it('keeps a TUI run a success when its terminal analytics flush fails', async () => { const flushError = new Error('flush timed out'); @@ -713,213 +489,3 @@ it('keeps a TUI run a success when its terminal analytics flush fails', async () ); exit.mockRestore(); }); - -describe('host wiring over runProgram', () => { - const approved = { - organization: { id: 'org-1', is_ai_data_processing_approved: true }, - } as ApiUser; - - it('authenticates through the provider after the settings gate, then awaits AI opt-in and the post-auth gate', async () => { - const order: string[] = []; - const ui = getUI(); - const record = - (name: string, value?: T) => - () => { - order.push(name); - return value as T; - }; - vi.mocked(checkAllSettingsConflicts).mockImplementationOnce( - record('settings check', []), - ); - vi.mocked(authenticate).mockImplementationOnce( - record('authenticate', Promise.resolve()), - ); - vi.spyOn(ui, 'waitForAiOptIn').mockImplementation( - record('waitForAiOptIn', Promise.resolve()), - ); - vi.spyOn(ui, 'waitForGate').mockImplementation((id) => - record(`waitForGate:${id}`, Promise.resolve())(), - ); - vi.mocked(analytics.getAllFlagsForWizard).mockImplementationOnce( - record('getAllFlagsForWizard', Promise.resolve({})), - ); - vi.mocked(runAgent).mockImplementationOnce((...args) => { - order.push('runAgent'); - return finishRun(...args); - }); - - await runProgramAgent(errorTrackingUploadSourceMapsConfig, { - ...buildSession({ ci: false, installDir: '/tmp/adapter-test' }), - credentials: session().credentials, - apiUser: { organization: { is_ai_data_processing_approved: false } }, - } as ReturnType); - - expect(order).toEqual([ - 'settings check', - 'authenticate', - 'waitForAiOptIn', - 'waitForGate:detect', - 'getAllFlagsForWizard', - 'runAgent', - ]); - }); - - it('a refreshed token reaches session and UI', async () => { - vi.mocked(refreshAccessToken).mockResolvedValueOnce({ - access_token: 'pha_new', - refresh_token: 'phr_rotated', - expires_in: 3600, - token_type: 'Bearer', - scope: 'project:read', - }); - // authenticate is a no-op for a session that already has its login. - vi.mocked(authenticate).mockImplementationOnce(() => Promise.resolve()); - const aging = { - ...session().credentials, - refreshToken: 'phr_old', - expiresAt: Date.now() + 20 * 60 * 1000, - projectId: 7, - }; - const refreshing = { ...session(), credentials: aging }; - const setAccessToken = vi.spyOn(getUI(), 'setAccessToken'); - - await runProgramAgent(program(), refreshing); - - expect(refreshing.credentials).toMatchObject({ - accessToken: 'pha_new', - refreshToken: 'phr_rotated', - projectId: 7, - }); - // The login's host keeps its class, not a structured copy. - expect(refreshing.credentials.host).toBe(aging.host); - expect(aging.accessToken).toBe('test'); - expect(setAccessToken).toHaveBeenCalledExactlyOnceWith( - refreshing.credentials, - ); - }); - - it.each([ - [false, 1], - [true, 0], - ])( - 'passes the session stamp latch (%s) to runProgram and latches the session', - async (latched, stamps) => { - const stamping = Object.assign(session(), { - apiUser: approved, - scanConsent: ScanConsent.Granted, - discoveredFeatures: [DiscoveredFeature.LLM], - aiSdkStampReported: latched, - }); - - await runProgramAgent(program(), stamping); - - expect(analytics.groupIdentify).toHaveBeenCalledTimes(stamps); - expect(stamping.aiSdkStampReported).toBe(true); - }, - ); - - it('registers the linear settings restore once, before the run can reach the outro', async () => { - const onEnterScreen = vi.spyOn(getUI(), 'onEnterScreen'); - let registeredBeforeRun = false; - vi.mocked(runAgent).mockImplementationOnce((...args) => { - registeredBeforeRun = onEnterScreen.mock.calls.length === 1; - return finishRun(...args); - }); - - // A composed program is clamped to linear. - await runProgramAgent(program(), session(), { composed: true }); - - expect(registeredBeforeRun).toBe(true); - expect(onEnterScreen).toHaveBeenCalledExactlyOnceWith( - 'outro', - expect.any(Function), - ); - onEnterScreen.mock.calls[0][1](); - expect(restoreClaudeSettings).toHaveBeenCalledExactlyOnceWith( - '/tmp/adapter-test', - ); - - onEnterScreen.mockClear(); - await runProgramAgent(program(), session()); - expect(onEnterScreen).not.toHaveBeenCalled(); - }); - - describe('program files', () => { - let installDir: string; - beforeEach(() => { - installDir = fs.mkdtempSync(path.join(os.tmpdir(), 'wizard-files-')); - }); - afterEach(() => { - fs.rmSync(installDir, { recursive: true, force: true }); - }); - - it('sends the host the seeded audit checks before the run, then each update once', async () => { - const setFrameworkContext = vi.spyOn(getUI(), 'setFrameworkContext'); - const sent = () => - setFrameworkContext.mock.calls.filter( - ([key]) => key === AUDIT_CHECKS_KEY, - ); - const resolved = [{ ...AUDIT_SEED_CHECKS[0], status: 'pass' }]; - let sentBeforeRun: unknown[][] = []; - vi.mocked(runAgent).mockImplementationOnce((...args) => { - sentBeforeRun = sent(); - fs.writeFileSync( - path.join(installDir, AUDIT_CHECKS_FILE), - JSON.stringify(resolved), - ); - return finishRun(...args); - }); - - await runProgramAgent(auditConfig, { ...session(), installDir }); - - expect(sentBeforeRun).toEqual([[AUDIT_CHECKS_KEY, AUDIT_SEED_CHECKS]]); - expect(sent()).toEqual([ - [AUDIT_CHECKS_KEY, AUDIT_SEED_CHECKS], - [AUDIT_CHECKS_KEY, resolved], - ]); - }); - - it('sends the host the event plan an integration run wrote', async () => { - const setEventPlan = vi.spyOn(getUI(), 'setEventPlan'); - vi.mocked(runAgent).mockImplementationOnce((...args) => { - fs.writeFileSync( - path.join(installDir, EVENT_PLAN_FILE), - JSON.stringify([{ event_name: 'checkout_started' }]), - ); - return finishRun(...args); - }); - - await runProgramAgent( - { ...program('posthog-integration'), eventPlanFile: EVENT_PLAN_FILE }, - { ...session(), installDir }, - ); - - expect(setEventPlan).toHaveBeenCalledExactlyOnceWith([ - { name: 'checkout_started', description: '' }, - ]); - }); - }); - - const hostFailure = new Error('host capability failed'); - it.each([ - [ - 'a failed login', - () => vi.mocked(authenticate).mockRejectedValueOnce(hostFailure), - ], - [ - 'a malformed flag override', - () => - vi - .mocked(analytics.getAllFlagsForWizard) - .mockRejectedValueOnce(hostFailure), - ], - ])('rethrows %s for the CLI root, as before', async (_name, arrange) => { - arrange(); - - await expect(runProgramAgent(program(), session())).rejects.toBe( - hostFailure, - ); - expect(runAgent).not.toHaveBeenCalled(); - expect(wizardAbort).not.toHaveBeenCalled(); - }); -}); diff --git a/src/programs/__tests__/run-program.test.ts b/src/programs/__tests__/run-program.test.ts index a6472d367..5dca801a7 100644 --- a/src/programs/__tests__/run-program.test.ts +++ b/src/programs/__tests__/run-program.test.ts @@ -1,43 +1,20 @@ import { runAgent, RunOutcome } from '@agent'; -import fs from 'node:fs'; -import os from 'node:os'; -import path from 'node:path'; -import { EVENT_PLAN_FILE, Harness, Sequence } from '@shared/constants'; -import { AUDIT_CHECKS_FILE, type AuditCheck } from '@shared/audit-ledger'; +import { Harness, Sequence } from '@shared/constants'; import { HostResolution } from '@shared/host-resolution'; import type { ApiUser } from '@shared/api'; import type { RunResult } from '@agent/types'; -import type { DetectedSource } from '../warehouse-sources/types'; +import { ErrorCodes } from '@shared/errors'; +import { DiscoveredFeature } from '@lib/wizard-session'; +import { analytics } from '@utils/analytics'; +import { refreshAccessToken } from '@utils/oauth'; import type { ResolvedProgramCredentials } from '../credentials'; import type { ProgramInput, ProgramOptions } from '../run-program'; -import { ErrorCodes } from '@shared/errors'; -import { ProgramEventPlanWatcher } from '../posthog-integration/watch-event-plan'; import { runProgram } from '@programs'; -import { runProgram as runProgramDirect } from '../run-program'; -import { analytics } from '@utils/analytics'; -import { refreshAccessToken } from '@utils/oauth-token'; -import { DiscoveredFeature } from '@shared/scan-consent'; -import { captureSwitchboardDecision } from '../binding-telemetry'; -import { gatewayAuth } from '../gateway-session'; -import { clearCleanup, runCleanups } from '@utils/cleanup-registry'; -import { commitRegisteredRunSkillCleanups } from '@shared/skill-run-cleanup'; vi.mock('@agent', async (importOriginal) => ({ ...(await importOriginal()), runAgent: vi.fn(), - DEFAULT_AGENT_BINDING: { - sequence: 'linear', - harness: 'anthropic', - model: 'claude-test', - }, })); -vi.mock('../binding-telemetry', async (importOriginal) => { - const actual = await importOriginal(); - return { - ...actual, - captureSwitchboardDecision: vi.fn(actual.captureSwitchboardDecision), - }; -}); vi.mock('@utils/analytics', async (importOriginal) => ({ ...(await importOriginal()), analytics: { @@ -51,8 +28,8 @@ vi.mock('@utils/analytics', async (importOriginal) => ({ groupIdentify: vi.fn(), }, })); -vi.mock('@utils/oauth-token', () => ({ refreshAccessToken: vi.fn() })); -vi.mock('../gateway-session', () => ({ gatewayAuth: vi.fn() })); +vi.mock('@utils/oauth', () => ({ refreshAccessToken: vi.fn() })); +vi.mock('@utils/debug'); const run = { integrationLabel: 'metrics', @@ -81,7 +58,6 @@ const credentials: ResolvedProgramCredentials = { projectId: 42, host: HostResolution.fromRegion('us'), }, - inferenceAuth: { resolve: vi.fn() }, project: null, apiUser: { organization: { is_ai_data_processing_approved: true }, @@ -115,25 +91,6 @@ const gated: ProgramInput = { program: { postAuthGates: ['detect'] }, }; -const warehouseSource = (kind: string): DetectedSource => ({ - kind, - label: kind, - mode: 'in-cli', - matchedSignal: `dependency: ${kind.toLowerCase()}`, -}); - -const tempDirs: string[] = []; -const tempDir = () => { - const dir = fs.mkdtempSync(path.join(os.tmpdir(), 'wizard-run-program-')); - tempDirs.push(dir); - return dir; -}; - -const markSkill = (dir: string) => { - fs.mkdirSync(dir, { recursive: true }); - fs.writeFileSync(path.join(dir, '.posthog-wizard'), ''); -}; - describe('runProgram', () => { beforeEach(() => { vi.clearAllMocks(); @@ -142,14 +99,8 @@ describe('runProgram', () => { snapshot, }); }); - afterEach(() => { - vi.restoreAllMocks(); - for (const dir of tempDirs.splice(0)) { - fs.rmSync(dir, { recursive: true, force: true }); - } - }); - it("runs the caller's run definition and program settings, and returns its final results", async () => { + it("runs the caller's run definition, settings and route, and returns its final results", async () => { const excludedTaskTypes = () => ['logs']; const { signal } = new AbortController(); const outcome = await runProgram( @@ -165,6 +116,7 @@ describe('runProgram', () => { excludedTaskTypes, }, credentials, + overrides: { harness: Harness.anthropic, sequence: Sequence.linear }, }, { signal }, ); @@ -176,41 +128,41 @@ describe('runProgram', () => { agentFlow: 'metrics-flow', allowedTools: ['Agent'], disallowedTools: ['wizard_ask'], + binding: { sequence: Sequence.linear, harness: Harness.anthropic }, + switchboard: { + program: 'metrics', + cliHarness: Harness.anthropic, + cliSequence: Sequence.linear, + }, + wizardMetadata: { + program_id: 'metrics', + integration: 'metrics', + run_id: 'analytics-run-id', + build: 'test', + call_type: 'agent', + SEQUENCE: Sequence.linear, + HARNESS: Harness.anthropic, + }, }); expect(config.excludedTaskTypes).toBe(excludedTaskTypes); - expect(input.installDir).toBe('/project'); expect(input.credentials).toBe(credentials.posthog); - expect(input.inferenceAuth).toBe(credentials.inferenceAuth); expect(agentOptions?.signal).toBe(signal); + expect(analytics.setTag).toHaveBeenCalledWith('harness', Harness.anthropic); + expect(analytics.wizardCapture).toHaveBeenCalledWith( + 'switchboard resolved', + expect.objectContaining({ program: 'metrics', cli_harness: 'anthropic' }), + ); expect(outcome).toMatchObject({ programId: 'metrics', outcome: RunOutcome.Success, - data: { credentials: { projectId: 42 } }, + data: { credentials: { projectId: 42 }, binding: config.binding }, settledRuns: [ { runId: 'run-1', result: { outcome: RunOutcome.Success } }, ], diagnostics: [], artifacts: { reportFile: '/project/posthog-metrics-report.md' }, }); - }); - - it('every run carries the standard trace tags and its route', async () => { - await runProgram('metrics', { - installDir: '/project', - run, - credentials, - overrides: { harness: Harness.anthropic, sequence: Sequence.linear }, - }); - - expect(vi.mocked(runAgent).mock.calls[0][0].wizardMetadata).toEqual({ - program_id: 'metrics', - integration: 'metrics', - run_id: 'analytics-run-id', - build: 'test', - call_type: 'agent', - SEQUENCE: Sequence.linear, - HARNESS: Harness.anthropic, - }); + expect(outcome.failure).toBeUndefined(); }); const closed = () => Promise.reject(new Error('host closed')); @@ -221,6 +173,12 @@ describe('runProgram', () => { RunOutcome.Failed, 'Credentials are required to run metrics.', ], + [ + 'a rejecting credential provider', + () => ({ credentials: { resolve: closed } }), + RunOutcome.Failed, + 'host closed', + ], [ 'no org AI approval and no host approval capability', () => ({ credentials: login(null) }), @@ -236,12 +194,6 @@ describe('runProgram', () => { RunOutcome.Aborted, 'AI processing approval declined.', ], - [ - 'a rejecting credential provider', - () => ({ credentials: { resolve: closed } }), - RunOutcome.Failed, - 'host closed', - ], ])( '%s is a decided result before the agent it guards', async (_case, options, outcome, message) => { @@ -270,6 +222,41 @@ describe('runProgram', () => { expect(awaitAiApproval).not.toHaveBeenCalled(); }); + it.each<[RunOutcome, RunResult['failure']]>([ + [ + RunOutcome.Failed, + { code: ErrorCodes.AgentApiError, message: 'API Error' }, + ], + [ + RunOutcome.Aborted, + { code: ErrorCodes.AgentAbort, message: 'Agent run cancelled' }, + ], + [ + RunOutcome.Crashed, + { + code: ErrorCodes.InternalUnhandled, + message: 'mint refused', + error: new Error('mint refused'), + }, + ], + ])( + 'a %s agent run settles with its failure and its run', + async (outcome, failure) => { + const result = { outcome, failure, snapshot } as RunResult; + vi.mocked(runAgent).mockResolvedValueOnce(result); + + const settled = await runProgram('metrics', { + installDir: '/project', + runId: 'run-1', + run, + credentials, + }); + + expect(settled).toMatchObject({ outcome, failure }); + expect(settled.settledRuns).toEqual([{ runId: 'run-1', result }]); + }, + ); + it.each([ 'credential resolution', 'AI approval', @@ -312,50 +299,6 @@ describe('runProgram', () => { }, ); - it('settles a pre-aborted host signal before credentials or agent startup', async () => { - const controller = new AbortController(); - controller.abort(); - const resolve = vi.fn(); - - const result = await runProgram( - 'metrics', - { installDir: '/project', run }, - { credentials: { resolve }, signal: controller.signal }, - ); - - expect(result).toMatchObject({ - outcome: RunOutcome.Aborted, - failure: { code: ErrorCodes.AgentAbort }, - }); - expect(resolve).not.toHaveBeenCalled(); - expect(runAgent).not.toHaveBeenCalled(); - }); - - it('overrides reach the binding and the decision is captured once', async () => { - const result = await runProgram('metrics', { - installDir: '/project', - run, - credentials, - overrides: { harness: Harness.anthropic, sequence: Sequence.linear }, - }); - - const resolved = vi.mocked(runAgent).mock.calls[0][0].binding; - expect(resolved).toMatchObject({ - sequence: Sequence.linear, - harness: Harness.anthropic, - }); - expect(captureSwitchboardDecision).toHaveBeenCalledExactlyOnceWith( - expect.objectContaining({ - program: 'metrics', - cliHarness: Harness.anthropic, - cliSequence: Sequence.linear, - }), - resolved, - ); - expect(analytics.setTag).toHaveBeenCalledWith('harness', Harness.anthropic); - expect(result.data.binding).toEqual(resolved); - }); - it('runs in order: agent started, credentials, approval, post-auth, flags, refresh, route, agent', async () => { const order: string[] = []; const answer = (name: string, value: T) => @@ -428,51 +371,25 @@ describe('runProgram', () => { }); }); - it('prefers the input flags over the loader', async () => { - const featureFlags = vi.fn(); + it('a host mutation after the call does not reach the run', async () => { + const flags = { ci: false }; + const host: NonNullable = { region: 'us' }; - await runProgram( + const pending = runProgram( 'metrics', - { - installDir: '/project', - run, - credentials, - wizardFlags: { 'wizard-test-flag': 'input' }, - }, - { featureFlags }, + { installDir: '/project', run, flags, host }, + { credentials: login() }, ); + flags.ci = true; + host.region = 'eu'; + await pending; - expect(featureFlags).not.toHaveBeenCalled(); - expect(vi.mocked(runAgent).mock.calls[0][0].wizardFlags).toEqual({ - 'wizard-test-flag': 'input', - }); + const [, runInput] = vi.mocked(runAgent).mock.calls[0]; + expect(runInput.flags.ci).toBe(false); + expect(runInput.host.region).toBe('us'); }); - it.each([ - ['the lazy entry', runProgram], - ['run-program', runProgramDirect], - ])( - 'a host mutation after the call does not reach the run, through %s', - async (_entry, callProgram) => { - const flags = { ci: false }; - const host: NonNullable = { region: 'us' }; - - const pending = callProgram( - 'metrics', - { installDir: '/project', run, flags, host }, - { credentials: login() }, - ); - flags.ci = true; - host.region = 'eu'; - await pending; - - const [, runInput] = vi.mocked(runAgent).mock.calls[0]; - expect(runInput.flags.ci).toBe(false); - expect(runInput.host.region).toBe('us'); - }, - ); - - it('a provider is resolved once, then stamped, and refreshed before the agent starts', async () => { + it('a provider is resolved once, then identified and stamped, and refreshed before the agent starts', async () => { vi.mocked(refreshAccessToken).mockResolvedValueOnce(refreshedToken); const apiUser = { distinct_id: 'user-1', @@ -490,7 +407,6 @@ describe('runProgram', () => { host: { baseUrl: 'https://posthog.example' }, mayReportScanResults: true, discoveredFeatures: [DiscoveredFeature.LLM], - warehouseSources: [warehouseSource('Stripe')], }, { credentials: { resolve } }, ); @@ -514,146 +430,5 @@ describe('runProgram', () => { credentials: { refreshToken: 'phr_rotated' }, aiSdkStampReported: true, }); - - await vi.mocked(runAgent).mock.calls[0][1].inferenceAuth.resolve(); - expect(gatewayAuth).toHaveBeenCalledExactlyOnceWith( - credentials.posthog.host, - 'pha_refreshed', - 'metrics', - ); - }); - - it('seeds the audit ledger before the agent and returns the checks this run wrote', async () => { - const installDir = tempDir(); - const ledgerFile = path.join(installDir, AUDIT_CHECKS_FILE); - const seed: AuditCheck[] = [ - { id: 'seed', area: 'Events', label: 'seed', status: 'pending' }, - ]; - const updated = [{ ...seed[0], status: 'pass' }]; - let seededBeforeRun: unknown; - vi.mocked(runAgent).mockImplementation(() => { - seededBeforeRun = JSON.parse(fs.readFileSync(ledgerFile, 'utf8')); - fs.writeFileSync(ledgerFile, JSON.stringify(updated)); - return Promise.resolve({ outcome: RunOutcome.Success, snapshot }); - }); - - const result = await runProgram('audit', { - installDir, - run, - credentials, - program: { auditLedgerFile: AUDIT_CHECKS_FILE, auditSeedChecks: seed }, - }); - - expect(seededBeforeRun).toEqual(seed); - expect(result.data.detection.frameworkContext.auditChecks).toEqual(updated); - }); - - it('returns the event plan this run wrote, and releases its watcher when the agent throws', async () => { - const installDir = tempDir(); - const planFile = path.join(installDir, EVENT_PLAN_FILE); - const stop = vi.spyOn(ProgramEventPlanWatcher.prototype, 'stop'); - vi.mocked(runAgent) - .mockImplementationOnce(() => { - fs.writeFileSync( - planFile, - JSON.stringify([ - { event_name: 'checkout_started', description: 'A' }, - ]), - ); - return Promise.resolve({ outcome: RunOutcome.Success, snapshot }); - }) - .mockRejectedValueOnce(new Error('agent crashed')); - const input: ProgramInput = { - installDir, - credentials, - run, - program: { eventPlanFile: EVENT_PLAN_FILE }, - }; - - const result = await runProgram('posthog-integration', input); - stop.mockClear(); - await expect(runProgram('posthog-integration', input)).rejects.toThrow( - 'agent crashed', - ); - - expect(result.data.eventPlan).toEqual([ - { name: 'checkout_started', description: 'A' }, - ]); - expect(stop).toHaveBeenCalledOnce(); - }); - - describe('skill cleanup', () => { - let installDir: string; - let newSkill: string; - let oldSkill: string; - - beforeEach(() => { - clearCleanup(); - installDir = tempDir(); - const skillRoot = path.join(installDir, '.claude', 'skills'); - oldSkill = path.join(skillRoot, 'before-run'); - newSkill = path.join(skillRoot, 'during-run'); - markSkill(oldSkill); - vi.mocked(runAgent).mockImplementation(() => { - markSkill(newSkill); - return Promise.resolve({ outcome: RunOutcome.Success, snapshot }); - }); - }); - afterEach(() => clearCleanup()); - - it.each<[string, () => RunResult]>([ - [ - 'a failed run', - () => ({ - outcome: RunOutcome.Failed, - failure: { code: ErrorCodes.AgentApiError, message: 'failed' }, - snapshot, - }), - ], - [ - // What wizardAbort and the CLI roots' signal handlers call. - 'a process drain mid-run', - () => { - runCleanups(); - return { outcome: RunOutcome.Success, snapshot }; - }, - ], - ])("%s removes this invocation's new skills", async (_case, settle) => { - vi.mocked(runAgent).mockImplementation(() => { - markSkill(newSkill); - return Promise.resolve(settle()); - }); - - await runProgram('metrics', { installDir, run, credentials }); - - expect(fs.existsSync(newSkill)).toBe(false); - expect(fs.existsSync(oldSkill)).toBe(true); - }); - - it('commits its handle after success, or leaves it to the host with deferSkillCommit', async () => { - await runProgram('metrics', { installDir, run, credentials }); - runCleanups(); - expect(fs.existsSync(newSkill)).toBe(true); - - fs.rmSync(newSkill, { recursive: true }); - await runProgram( - 'metrics', - { installDir, run, credentials }, - { deferSkillCommit: true }, - ); - // A host that fails after the run still drains this invocation's skills. - runCleanups(); - expect(fs.existsSync(newSkill)).toBe(false); - - await runProgram( - 'metrics', - { installDir, run, credentials }, - { deferSkillCommit: true }, - ); - commitRegisteredRunSkillCleanups(); - runCleanups(); - expect(fs.existsSync(newSkill)).toBe(true); - expect(fs.existsSync(oldSkill)).toBe(true); - }); }); }); diff --git a/src/programs/__tests__/self-driving-deck.test.ts b/src/programs/__tests__/self-driving-deck.test.ts index 84c80dfbe..94a39e55c 100644 --- a/src/programs/__tests__/self-driving-deck.test.ts +++ b/src/programs/__tests__/self-driving-deck.test.ts @@ -8,10 +8,6 @@ import type { ReactNode, ReactElement } from 'react'; import { getContentBlocks } from '@ui/tui/decks/self-driving/index'; -import { - getProgramContentBlocks, - getProgramTips, -} from '@ui/tui/decks/registry'; /** paneWidth in LearnCard at 80 cols: (min(120, 80) - 2) / 2 - 2 */ const PANE_WIDTH_80COL = 37; @@ -31,15 +27,6 @@ describe('self-driving learn deck', () => { expect(blocks.length).toBeGreaterThan(0); }); - it('selects the self-driving deck and tips for its program', () => { - const selected = getProgramContentBlocks('self-driving'); - const last = selected[selected.length - 1]; - expect( - typeof last === 'object' && 'content' in last ? last.content : '', - ).toBe('Your product drives itself.'); - expect(getProgramTips('self-driving')?.[0]?.id).toBe('signal-source'); - }); - it('keeps every fixed-layout line within the 80-col pane', () => { const wide: string[] = []; for (const b of blocks) { diff --git a/src/programs/__tests__/self-driving-detect.test.ts b/src/programs/__tests__/self-driving-detect.test.ts index e52f6233a..82fce06a9 100644 --- a/src/programs/__tests__/self-driving-detect.test.ts +++ b/src/programs/__tests__/self-driving-detect.test.ts @@ -25,8 +25,8 @@ import { import { Integration } from '@shared/constants'; import { WIZARD_TOOL_NAMES } from '@agent/tools'; import { buildSession } from '@lib/wizard-session'; -import type { Mock } from 'vitest'; import { testProgramRunHost } from '../../../test/program-host'; +import type { Mock } from 'vitest'; function makeTmpDir(): string { return fs.mkdtempSync(path.join(os.tmpdir(), 'self-driving-detect-')); @@ -189,6 +189,15 @@ describe('selfDrivingConfig', () => { ); }); + it('ships its own Learn deck ending on the self-driving closer', () => { + const blocks = selfDrivingConfig.getContentBlocks?.() ?? []; + expect(blocks.length).toBeGreaterThan(0); + const last = blocks[blocks.length - 1]; + expect( + typeof last === 'object' && 'content' in last ? last.content : '', + ).toBe('Your product drives itself.'); + }); + it('gives wizard_ask a 30-min timeout for the browser-handoff steps', async () => { // `run` is resolved per-session so the prompt can carry the integrate flag. const { run } = selfDrivingConfig; diff --git a/src/programs/__tests__/token-refresh.test.ts b/src/programs/__tests__/token-refresh.test.ts deleted file mode 100644 index 53799d66c..000000000 --- a/src/programs/__tests__/token-refresh.test.ts +++ /dev/null @@ -1,128 +0,0 @@ -import { refreshCredentialsIfNeeded } from '../token-refresh'; -import { refreshAccessToken } from '@utils/oauth-token'; -import { OAuthError } from '@utils/oauth-errors'; -import { analytics } from '@utils/analytics'; -import { - isGrantRevoked, - resetAuthSessionState, -} from '@shared/auth-session-state'; -import type { Credentials } from '@shared/api'; -import { HostResolution } from '@shared/host-resolution'; - -vi.mock('@utils/oauth-token', () => ({ refreshAccessToken: vi.fn() })); -vi.mock('@utils/debug', () => ({ logToFile: vi.fn() })); -vi.mock('@utils/analytics', () => ({ - analytics: { wizardCapture: vi.fn() }, -})); - -const mockedRefresh = vi.mocked(refreshAccessToken); - -function credentialsWith(over: Partial = {}): Credentials { - return { - accessToken: 'pha_old', - projectApiKey: 'phc_test', - projectId: 7, - host: HostResolution.fromRegion('us'), - ...over, - }; -} - -/** Aging enough to be under the 50-minute threshold. */ -const aging = (over: Partial = {}): Credentials => - credentialsWith({ - refreshToken: 'phr_old', - expiresAt: Date.now() + 20 * 60 * 1000, - ...over, - }); - -const token = (over: Record = {}) => ({ - access_token: 'pha_new', - expires_in: 3600, - token_type: 'Bearer', - scope: 'project:read', - ...over, -}); - -describe('refreshCredentialsIfNeeded', () => { - beforeEach(() => { - vi.clearAllMocks(); - resetAuthSessionState(); - }); - - it.each([ - [ - 'without a refresh token (CI api-key runs, refresh-less grants)', - credentialsWith({ accessToken: 'pha_ci_key', expiresAt: 0 }), - ], - [ - 'while most of the lifetime is left', - aging({ expiresAt: Date.now() + 59 * 60 * 1000 }), - ], - // `?? 0` would read as "expired" and spend a rotation on every run. - [ - 'with a refresh token but no expiry', - credentialsWith({ refreshToken: 'phr_old' }), - ], - ])('returns the same credentials %s', async (_case, credentials) => { - await expect(refreshCredentialsIfNeeded(credentials, {})).resolves.toBe( - credentials, - ); - expect(mockedRefresh).not.toHaveBeenCalled(); - }); - - it('refreshes an aging token under the base URL and minting client id, and keeps the rotated refresh token', async () => { - mockedRefresh.mockResolvedValueOnce( - token({ refresh_token: 'phr_rotated' }), - ); - - const refreshed = await refreshCredentialsIfNeeded( - aging({ oauthClientId: 'client_us_provisioning' }), - { baseUrl: 'https://posthog.example' }, - ); - - expect(mockedRefresh).toHaveBeenCalledWith( - 'phr_old', - 'https://posthog.example', - 'client_us_provisioning', - ); - expect(refreshed.accessToken).toBe('pha_new'); - expect(refreshed.refreshToken).toBe('phr_rotated'); - // Unrelated fields survive the swap. - expect(refreshed.projectId).toBe(7); - expect(refreshed.expiresAt).toBeGreaterThan(Date.now() + 59 * 60 * 1000); - }); - - it('returns new credentials rather than mutating the old ones', async () => { - mockedRefresh.mockResolvedValueOnce(token()); - const before = aging(); - - const refreshed = await refreshCredentialsIfNeeded(before, {}); - - expect(refreshed).not.toBe(before); - expect(before.accessToken).toBe('pha_old'); - // No rotation in the response: the old refresh token has to carry over. - expect(refreshed.refreshToken).toBe('phr_old'); - }); - - it('marks the grant revoked on invalid_grant, so a later 401 can name the cause', async () => { - mockedRefresh.mockRejectedValueOnce(new OAuthError('invalid_grant')); - - await refreshCredentialsIfNeeded(aging(), {}); - - expect(isGrantRevoked()).toBe(true); - expect(analytics.wizardCapture).toHaveBeenCalledWith( - 'auth session expired', - { reason: 'invalid_grant' }, - ); - }); - - it('keeps the same credentials and leaves the grant unmarked for a transport failure, which says nothing about the login', async () => { - mockedRefresh.mockRejectedValueOnce(new Error('ETIMEDOUT')); - const before = aging(); - - await expect(refreshCredentialsIfNeeded(before, {})).resolves.toBe(before); - expect(before.accessToken).toBe('pha_old'); - expect(isGrantRevoked()).toBe(false); - expect(analytics.wizardCapture).not.toHaveBeenCalled(); - }); -}); diff --git a/src/programs/__tests__/warehouse-ask-timeout.test.ts b/src/programs/__tests__/warehouse-ask-timeout.test.ts index 2c08f3e8d..591e57aa1 100644 --- a/src/programs/__tests__/warehouse-ask-timeout.test.ts +++ b/src/programs/__tests__/warehouse-ask-timeout.test.ts @@ -20,8 +20,10 @@ vi.mock('@utils/analytics', () => ({ })); import { warehouseSourceConfig } from '@programs/warehouse-source/index'; -import { DEFAULT_ASK_TIMEOUT_MS } from '@agent/wizard-ask-bridge'; -import { LONGER_ASK_TIMEOUT_MS } from '@shared/ask-policy'; +import { + LONGER_ASK_TIMEOUT_MS, + DEFAULT_ASK_TIMEOUT_MS, +} from '@agent/wizard-ask-bridge'; function session(): WizardSession { return { installDir: '/tmp/app', frameworkContext: {} } as WizardSession; diff --git a/src/programs/__tests__/warehouse-suggestion.test.ts b/src/programs/__tests__/warehouse-suggestion.test.ts index 9db9ad9e9..3966f9dc4 100644 --- a/src/programs/__tests__/warehouse-suggestion.test.ts +++ b/src/programs/__tests__/warehouse-suggestion.test.ts @@ -173,9 +173,9 @@ describe('flow shape', () => { }); it('keeps the program single-run, so the outro stays terminal', () => { - // A step declaring a child program would flip run-wizard into the composed + // A step carrying its own `run` would flip run-wizard into the composed // walk, where a second agent run could abort before the outro is pushed. - expect(POSTHOG_INTEGRATION_PROGRAM.some((s) => s.runProgramId)).toBe(false); + expect(POSTHOG_INTEGRATION_PROGRAM.some((s) => s.run)).toBe(false); }); }); diff --git a/src/programs/agent-skill/index.ts b/src/programs/agent-skill/index.ts index 3d75387c3..75752731e 100644 --- a/src/programs/agent-skill/index.ts +++ b/src/programs/agent-skill/index.ts @@ -20,13 +20,36 @@ */ import type { ProgramConfig } from '@programs/program-step'; +import type { AbortCase } from '@agent/types'; +import type { ProgramRun } from '@programs/program-run'; import { AGENT_SKILL_STEPS } from './steps.js'; -import { - skillRunDefinition, - type SkillProgramOptions, -} from './run-definition.js'; +import { getContentBlocks } from '../../ui/tui/decks/agent-skill/index.js'; -export type { SkillProgramOptions } from './run-definition.js'; +export interface SkillProgramOptions { + /** Context-mill skill ID to install */ + skillId: string; + /** CLI subcommand name */ + command: string; + /** Unique flow key — must match a Program enum entry */ + id: string; + /** CLI description shown in --help */ + description: string; + /** Analytics integration label */ + integrationLabel: string; + /** Custom prompt instruction. Appended after default project prompt. */ + customPrompt?: string; + successMessage: string; + reportFile: string; + docsUrl: string; + spinnerMessage: string; + estimatedDurationMinutes: number; + /** Other program ids that must be satisfied first */ + requires?: string[]; + /** Override the default outro. Receives the same args as ProgramRun.buildOutroData. */ + buildOutroData?: ProgramRun['buildOutroData']; + /** Known `[ABORT] ` cases the skill can emit. */ + abortCases?: AbortCase[]; +} export function createSkillProgram(opts: SkillProgramOptions): ProgramConfig { return { @@ -36,7 +59,19 @@ export function createSkillProgram(opts: SkillProgramOptions): ProgramConfig { skillId: opts.skillId, steps: AGENT_SKILL_STEPS, reportFile: opts.reportFile, - run: skillRunDefinition(opts), + getContentBlocks, + run: { + skillId: opts.skillId, + integrationLabel: opts.integrationLabel, + customPrompt: opts.customPrompt ? () => opts.customPrompt! : undefined, + successMessage: opts.successMessage, + reportFile: opts.reportFile, + docsUrl: opts.docsUrl, + spinnerMessage: opts.spinnerMessage, + estimatedDurationMinutes: opts.estimatedDurationMinutes, + buildOutroData: opts.buildOutroData, + abortCases: opts.abortCases, + }, requires: opts.requires, }; } diff --git a/src/programs/agent-skill/run-definition.ts b/src/programs/agent-skill/run-definition.ts deleted file mode 100644 index 6a13a9570..000000000 --- a/src/programs/agent-skill/run-definition.ts +++ /dev/null @@ -1,43 +0,0 @@ -import type { AbortCase } from '@agent/types'; -import type { ProgramRun } from '@programs/program-run'; - -export interface SkillProgramOptions { - /** Context-mill skill ID to install */ - skillId: string; - /** CLI subcommand name */ - command: string; - /** Unique flow key — must match a Program enum entry */ - id: string; - /** CLI description shown in --help */ - description: string; - /** Analytics integration label */ - integrationLabel: string; - /** Custom prompt instruction. Appended after default project prompt. */ - customPrompt?: string; - successMessage: string; - reportFile: string; - docsUrl: string; - spinnerMessage: string; - estimatedDurationMinutes: number; - /** Other program ids that must be satisfied first */ - requires?: string[]; - /** Override the default outro. Receives the same args as ProgramRun.buildOutroData. */ - buildOutroData?: ProgramRun['buildOutroData']; - /** Known `[ABORT] ` cases the skill can emit. */ - abortCases?: AbortCase[]; -} - -export function skillRunDefinition(opts: SkillProgramOptions): ProgramRun { - return { - skillId: opts.skillId, - integrationLabel: opts.integrationLabel, - customPrompt: opts.customPrompt ? () => opts.customPrompt! : undefined, - successMessage: opts.successMessage, - reportFile: opts.reportFile, - docsUrl: opts.docsUrl, - spinnerMessage: opts.spinnerMessage, - estimatedDurationMinutes: opts.estimatedDurationMinutes, - buildOutroData: opts.buildOutroData, - abortCases: opts.abortCases, - }; -} diff --git a/src/programs/agent-skill/steps.ts b/src/programs/agent-skill/steps.ts index 5474ee865..faafee60b 100644 --- a/src/programs/agent-skill/steps.ts +++ b/src/programs/agent-skill/steps.ts @@ -6,7 +6,7 @@ */ import type { ProgramStep } from '@programs/program-step'; -import { RunPhase } from '@shared/run-state'; +import { RunPhase } from '@lib/wizard-session'; import { HEALTH_CHECK_STEP } from '@programs/shared/health-check-step'; export const AGENT_SKILL_STEPS: ProgramStep[] = [ diff --git a/src/programs/ai-observability/index.ts b/src/programs/ai-observability/index.ts index 707966f31..b71697d30 100644 --- a/src/programs/ai-observability/index.ts +++ b/src/programs/ai-observability/index.ts @@ -1,11 +1,14 @@ import type { ProgramConfig, ProgramStep } from '@programs/program-step'; import { AGENT_SKILL_STEPS } from '@programs/agent-skill/index'; -import { AI_OBSERVABILITY_REPORT_FILE, AI_OBSERVABILITY_RUN } from './run.js'; +import { getContentBlocks } from '@ui/tui/decks/agent-skill/index'; +import { headlessOption, regionOption } from '@lib/headless-mode'; const AI_OBSERVABILITY_STEPS: ProgramStep[] = AGENT_SKILL_STEPS.map((step) => step.id === 'intro' ? { ...step, screenId: 'ai-observability-intro' } : step, ); +const AI_OBSERVABILITY_REPORT_FILE = 'posthog-ai-observability-report.md'; + /** * `wizard ai-observability` — wrap the project's LLM client calls so they emit * `$ai_generation` events into AI Observability. @@ -20,7 +23,42 @@ export const aiObservabilityConfig: ProgramConfig = { command: 'ai-observability', description: 'Add PostHog AI Observability to your LLM calls', id: 'ai-observability', + cliOptions: { ...headlessOption, ...regionOption }, steps: AI_OBSERVABILITY_STEPS, reportFile: AI_OBSERVABILITY_REPORT_FILE, - run: AI_OBSERVABILITY_RUN, + getContentBlocks, + run: { + integrationLabel: 'ai-observability', + // No `skillId`: linear.ts skips its pre-install step (see the gate on + // `linear.ts:47`), so the agent must load the menu and install the right + // variant itself. The prompt below tells it how. + customPrompt: + () => `Instrument this project's LLM calls with PostHog AI Observability. + +This flow has no pre-installed skill — you install the right one yourself: + +1. Call \`load_skill_menu\` with \`category: "ai-observability"\`. The menu is + the source of truth: one variant per (LLM provider × language), plus a + \`manual-capture\` variant for projects with no vendor SDK. + +2. Scan the project manifest (\`package.json\`, \`pyproject.toml\`, + \`requirements.txt\`) for a vendor LLM SDK and pick the variant that matches + it — the language follows the manifest (\`package.json\` → Node, Python + tooling → Python). Multiple SDKs (e.g. LangChain wrapping OpenAI) → prefer + the higher abstraction. No vendor SDK → the \`manual-capture\` variant. + Genuinely ambiguous → \`wizard_ask\` with a multi-choice picker. + +3. Call \`install_skill\` with the picked variant id. Then follow that skill's + \`SKILL.md\` and references end-to-end. The skill itself will install + packages, wire OTel, set env vars, and describe verification. + +Make only additive changes — do not touch existing PostHog init, identify +calls, event capture, or dashboards. Those belong to other skills. The +final report is written to ./${AI_OBSERVABILITY_REPORT_FILE}.`, + successMessage: `AI Observability configured! View the report at ./${AI_OBSERVABILITY_REPORT_FILE}`, + reportFile: AI_OBSERVABILITY_REPORT_FILE, + docsUrl: 'https://posthog.com/docs/ai-observability', + spinnerMessage: 'Setting up AI Observability...', + estimatedDurationMinutes: 5, + }, }; diff --git a/src/programs/ai-observability/run.ts b/src/programs/ai-observability/run.ts deleted file mode 100644 index 1e9510e42..000000000 --- a/src/programs/ai-observability/run.ts +++ /dev/null @@ -1,39 +0,0 @@ -import type { ProgramRun } from '@programs/program-run'; - -export const AI_OBSERVABILITY_REPORT_FILE = - 'posthog-ai-observability-report.md'; - -export const AI_OBSERVABILITY_RUN: ProgramRun = { - integrationLabel: 'ai-observability', - // No `skillId`: linear.ts skips its pre-install step (see the gate on - // `linear.ts:47`), so the agent must load the menu and install the right - // variant itself. The prompt below tells it how. - customPrompt: - () => `Instrument this project's LLM calls with PostHog AI Observability. - -This flow has no pre-installed skill — you install the right one yourself: - -1. Call \`load_skill_menu\` with \`category: "ai-observability"\`. The menu is - the source of truth: one variant per (LLM provider × language), plus a - \`manual-capture\` variant for projects with no vendor SDK. - -2. Scan the project manifest (\`package.json\`, \`pyproject.toml\`, - \`requirements.txt\`) for a vendor LLM SDK and pick the variant that matches - it — the language follows the manifest (\`package.json\` → Node, Python - tooling → Python). Multiple SDKs (e.g. LangChain wrapping OpenAI) → prefer - the higher abstraction. No vendor SDK → the \`manual-capture\` variant. - Genuinely ambiguous → \`wizard_ask\` with a multi-choice picker. - -3. Call \`install_skill\` with the picked variant id. Then follow that skill's - \`SKILL.md\` and references end-to-end. The skill itself will install - packages, wire OTel, set env vars, and describe verification. - -Make only additive changes — do not touch existing PostHog init, identify -calls, event capture, or dashboards. Those belong to other skills. The -final report is written to ./${AI_OBSERVABILITY_REPORT_FILE}.`, - successMessage: `AI Observability configured! View the report at ./${AI_OBSERVABILITY_REPORT_FILE}`, - reportFile: AI_OBSERVABILITY_REPORT_FILE, - docsUrl: 'https://posthog.com/docs/ai-observability', - spinnerMessage: 'Setting up AI Observability...', - estimatedDurationMinutes: 5, -}; diff --git a/src/programs/ai-opt-in-gate.ts b/src/programs/ai-opt-in-gate.ts index 9e9c44026..d101a02a3 100644 --- a/src/programs/ai-opt-in-gate.ts +++ b/src/programs/ai-opt-in-gate.ts @@ -27,18 +27,18 @@ * (`WIZARD_PROVISIONING_SCOPES` in constants.ts), so the org's * approval can never be read back — `apiUser` stays null and the gate * could never clear. Creating an account through the wizard to run the - * AI agent is itself the consent, mirroring how `isAskDisabled` + * AI agent is itself the consent, mirroring how `shouldDisableAsk` * already treats `ci || signup` as one non-interactive mode. */ -import type { ApiUser } from '@shared/api'; +import type { WizardSession } from '@lib/wizard-session'; import type { ProgramConfig, ProgramStep } from './program-step.js'; /** Step id — also the ScreenId.AiOptIn enum value in screen-sequences. */ export const AI_OPT_IN_STEP_ID = 'ai-opt-in'; -function aiApproved(user: ApiUser | null): boolean { - return !!user?.organization?.is_ai_data_processing_approved; +function aiApproved(session: WizardSession): boolean { + return !!session.apiUser?.organization?.is_ai_data_processing_approved; } /** @@ -64,11 +64,10 @@ export function withAiOptInGate(config: ProgramConfig): ProgramStep[] { !session.ci && !session.signup && session.apiUser != null && - !aiApproved(session.apiUser), + !aiApproved(session), isComplete: (session) => - session.ci || session.signup || aiApproved(session.apiUser), - gate: (session) => - session.ci || session.signup || aiApproved(session.apiUser), + session.ci || session.signup || aiApproved(session), + gate: (session) => session.ci || session.signup || aiApproved(session), }; return [ diff --git a/src/programs/audit/index.ts b/src/programs/audit/index.ts index aac6f8b95..ee733398c 100644 --- a/src/programs/audit/index.ts +++ b/src/programs/audit/index.ts @@ -1,16 +1,21 @@ import { AGENT_SKILL_STEPS, createSkillProgram, - type SkillProgramOptions, } from '@programs/agent-skill/index'; import type { ProgramStep, ProgramConfig } from '@programs/program-step'; import type { ProgramRun } from '@programs/program-run'; -import { OutroKind } from '@agent'; +import type { ProgramRunHost } from '@programs/host-capabilities'; +import type { WizardSession } from '@lib/wizard-session'; +import { OutroKind } from '@lib/wizard-session'; import { WIZARD_TOOL_NAMES } from '@agent'; -import { skillRunDefinition } from '@programs/agent-skill/run-definition'; +import { headlessOption, regionOption } from '@lib/headless-mode'; import { AUDIT_ABORT_CASES } from './detect.js'; -import { AUDIT_CHECKS_FILE, AUDIT_REPORT_FILE } from './types.js'; -import { AUDIT_SEED_CHECKS } from './seed.js'; +import { + AUDIT_CHECKS_FILE, + AUDIT_CHECKS_KEY, + AUDIT_REPORT_FILE, +} from './types.js'; +import { AUDIT_SEED_CHECKS, seedAuditLedger } from './seed.js'; /** Audit-specific screens for the shared agent-skill pipeline. */ const AUDIT_SCREEN_BY_STEP: Record = { @@ -19,9 +24,9 @@ const AUDIT_SCREEN_BY_STEP: Record = { outro: 'audit-outro', }; -type AuditRunState = { - dashboardUrl: string | null; - notebookUrl: string | null; +const seedBeforeAuditRun = (session: WizardSession): void => { + seedAuditLedger(session.installDir); + session.frameworkContext[AUDIT_CHECKS_KEY] = AUDIT_SEED_CHECKS; }; const withAuditScreens = (steps: ProgramStep[]): ProgramStep[] => @@ -32,7 +37,7 @@ const withAuditScreens = (steps: ProgramStep[]): ProgramStep[] => const auditSteps: ProgramStep[] = withAuditScreens(AGENT_SKILL_STEPS); -const AUDIT_OPTIONS: SkillProgramOptions = { +const baseConfig = createSkillProgram({ skillId: 'audit', command: 'audit', id: 'audit', @@ -48,13 +53,24 @@ const AUDIT_OPTIONS: SkillProgramOptions = { estimatedDurationMinutes: 5, requires: ['posthog-integration'], abortCases: AUDIT_ABORT_CASES, -}; +}); + +const auditRun = async ( + session: WizardSession, + host: ProgramRunHost, +): Promise => { + seedBeforeAuditRun(session); -const baseConfig = createSkillProgram(AUDIT_OPTIONS); -const baseRun = skillRunDefinition(AUDIT_OPTIONS); + if (!baseConfig.run) { + throw new Error('Audit program has no run configuration.'); + } -const auditRun = (session: AuditRunState): Promise => - Promise.resolve({ + const baseRun = + typeof baseConfig.run === 'function' + ? await baseConfig.run(session, host) + : baseConfig.run; + + return { ...baseRun, // Override the default outro so the dashboard + notebook URLs the // agent emits via `[DASHBOARD_URL]` / `[NOTEBOOK_URL]` are surfaced @@ -81,14 +97,14 @@ const auditRun = (session: AuditRunState): Promise => notebookUrl: session.notebookUrl ?? undefined, }; }, - }); + }; +}; export const auditConfig: ProgramConfig = { ...baseConfig, steps: auditSteps, run: auditRun, auditLedgerFile: AUDIT_CHECKS_FILE, - auditSeedChecks: AUDIT_SEED_CHECKS, // Ledger tools are opt-in per program; pi matches on the short name. allowedTools: [ 'Agent', @@ -97,4 +113,8 @@ export const auditConfig: ProgramConfig = { WIZARD_TOOL_NAMES.auditResolveChecks, ], disallowedTools: [WIZARD_TOOL_NAMES.wizardAsk], + // The experimental headless flag — declared on `audit` (and basic + // integration) rather than globally. mergeCommandOptions lands it on the + // `wizard audit` command; dispatchProgram routes it to runWizardHeadless. + cliOptions: { ...headlessOption, ...regionOption }, }; diff --git a/src/programs/audit/watch-ledger.ts b/src/programs/audit/ledger-watcher.ts similarity index 52% rename from src/programs/audit/watch-ledger.ts rename to src/programs/audit/ledger-watcher.ts index 09ee912f1..f7a312543 100644 --- a/src/programs/audit/watch-ledger.ts +++ b/src/programs/audit/ledger-watcher.ts @@ -1,19 +1,24 @@ -import path from 'node:path'; +/** + * Mirrors the agent's `.posthog-audit-checks.json` into the session, so the TUI + * screens and the task stream read one value. `runAgent` owns the lifecycle, so + * every path gets it — including the e2e host, which builds no task stream. + */ + +import path from 'path'; +import { getUI } from '@ui'; import { startFileWatcher, type FileWatcherHandle, type FileWatcherOptions, -} from '@shared/file-watcher'; -import { coerceAuditChecks, type AuditCheck } from '@shared/audit-ledger'; +} from '@lib/file-watcher'; import { logToFile } from '@utils/debug'; +import { AUDIT_CHECKS_KEY, coerceAuditChecks } from './types.js'; const MAX_LEDGER_FILE_BYTES = 256 * 1024; -/** Watch this run's audit ledger and project valid ledger arrays to a caller. */ -export function watchAuditLedger( +export function startAuditLedgerWatcher( installDir: string, file: string, - onChecks: (checks: AuditCheck[]) => void, options: FileWatcherOptions = {}, ): FileWatcherHandle { const target = path.join(installDir, file); @@ -21,7 +26,8 @@ export function watchAuditLedger( return startFileWatcher( target, - (parsed) => onChecks(coerceAuditChecks(parsed)), + (parsed) => + getUI().setFrameworkContext(AUDIT_CHECKS_KEY, coerceAuditChecks(parsed)), { // A ledger an earlier run left behind stays ignored until this run writes. ignoreInitialFile: true, diff --git a/src/programs/audit/types.ts b/src/programs/audit/types.ts index e474357b9..a799d0f3b 100644 --- a/src/programs/audit/types.ts +++ b/src/programs/audit/types.ts @@ -1,3 +1,4 @@ +import type { WizardSession } from '@lib/wizard-session'; import { AUDIT_CHECKS_FILE, AUDIT_REPORT_FILE, @@ -27,9 +28,7 @@ export const AUDIT_SEVERITY_STYLE: Record = { export const AUDIT_CHECKS_KEY = 'auditChecks'; -export function getAuditChecks(session: { - frameworkContext: Record; -}): AuditCheck[] { +export function getAuditChecks(session: WizardSession): AuditCheck[] { const raw = session.frameworkContext[AUDIT_CHECKS_KEY]; return Array.isArray(raw) ? (raw as AuditCheck[]) : []; } diff --git a/src/programs/authenticate.ts b/src/programs/authenticate.ts index 206938bce..b33043ec5 100644 --- a/src/programs/authenticate.ts +++ b/src/programs/authenticate.ts @@ -10,39 +10,19 @@ * back rather than fetching again. */ -import type { ApiProject, ApiUser, Credentials } from '@shared/api'; +import type { Credentials, WizardSession } from '@lib/wizard-session'; import type { ProgramId } from '@programs/program-registry'; -import type { CloudRegion } from '@utils/types'; import { getOrAskForProjectData } from '@utils/setup-utils'; +import { refreshAccessToken } from '@utils/oauth'; +import { OAuthError } from '@utils/oauth-errors'; +import { markGrantRevoked } from '@shared/auth-session-state'; import { analytics, groupsFromUser } from '@utils/analytics'; +import { getUI } from '@ui'; import { logToFile } from '@utils/debug'; -export type AuthProjection = { - setCredentials(credentials: Credentials): void; - setRoleAtOrganization(role: string | null): void; - setApiUser(user: ApiUser | null): void; -}; - -/** Authentication state shared with the CLI host, without TUI session fields. */ -export interface AuthSession { - signup: boolean; - ci: boolean; - apiKey?: string; - projectId?: number; - email?: string; - region?: CloudRegion; - baseUrl?: string; - localMcp: boolean; - credentials: Credentials | null; - apiProject: ApiProject | null; - roleAtOrganization: string | null; - apiUser: ApiUser | null; -} - export async function authenticate( - session: AuthSession, + session: WizardSession, programId: ProgramId, - projection: AuthProjection, ): Promise { if (session.credentials) return; @@ -85,12 +65,67 @@ export async function authenticate( session.roleAtOrganization = roleAtOrganization; session.apiUser = user; - projection.setCredentials(session.credentials); - projection.setRoleAtOrganization(roleAtOrganization); - projection.setApiUser(user); + getUI().setCredentials(session.credentials); + getUI().setRoleAtOrganization(roleAtOrganization); + getUI().setApiUser(user); // Identify the user (email, name) before flags are evaluated, so flags can // target the individual user and not just $app_name. if (user) analytics.identifyUser(user); analytics.setGroups(groupsFromUser(user, host.apiHost)); } + +// Below this remaining lifetime a run risks outliving its token; just-minted and 7-day tokens skip. +// Only a second agent run in one invocation can be this old — see self-driving's chained phases. +const REFRESH_WHEN_REMAINING_MS = 50 * 60 * 1000; + +/** + * Grants the token endpoint refuses permanently. A dead grant means the login + * is gone, not that the network blipped, so only these mark the session. + */ +const DEAD_GRANT_CODES = new Set(['invalid_grant', 'invalid_client']); + +/** Best-effort pre-run refresh; the same object comes back unless the token was refreshed. */ +export async function refreshCredentialsIfNeeded( + credentials: Credentials, + options: { baseUrl?: string }, +): Promise { + if (!credentials.refreshToken) return credentials; + + // No expiry means we cannot tell how much life is left, so leave it alone — + // refreshing every run would spend a rotation for nothing. + if (credentials.expiresAt === undefined) return credentials; + if (credentials.expiresAt - Date.now() >= REFRESH_WHEN_REMAINING_MS) { + return credentials; + } + + try { + const token = await refreshAccessToken( + credentials.refreshToken, + options.baseUrl, + credentials.oauthClientId, + ); + // Replaced, not mutated: readers hold this object, and a new one keeps the + // store and the (possibly shallow-copied) session explicitly in step. + return { + ...credentials, + accessToken: token.access_token, + // Rotation: keep the returned refresh token or the old one stops working. + refreshToken: token.refresh_token ?? credentials.refreshToken, + expiresAt: Date.now() + token.expires_in * 1000, + }; + } catch (error) { + // A dead grant is recorded but not thrown: the current token may still have + // minutes of life, and failing here would break runs that would have worked. + // If a 401 does follow, the auth-error screen can finally name the cause. + if (error instanceof OAuthError && DEAD_GRANT_CODES.has(error.code)) { + markGrantRevoked(); + analytics.wizardCapture('auth session expired', { reason: error.code }); + } + logToFile( + '[oauth] pre-run token refresh failed, continuing with the existing token:', + error instanceof Error ? error.message : error, + ); + return credentials; + } +} diff --git a/src/programs/binding-telemetry.ts b/src/programs/binding-telemetry.ts deleted file mode 100644 index fa6a53bef..000000000 --- a/src/programs/binding-telemetry.ts +++ /dev/null @@ -1,50 +0,0 @@ -import { - Sequence, - WIZARD_ORCHESTRATOR_FLAG_KEY, - WIZARD_SELF_DRIVING_USE_PI_HARNESS_FLAG_KEY, -} from '@shared/constants'; -import type { ResolvedBinding } from '@agent/types'; -import { analytics } from '@utils/analytics'; -import { logToFile } from '@utils/debug'; -import type { ProgramSwitchboardCtx } from './binding'; - -/** Record the selected route and its precedence sources once per program run. */ -export function captureSwitchboardDecision( - ctx: ProgramSwitchboardCtx, - binding: ResolvedBinding, -): void { - const trace = ctx.trace ?? {}; - const perTaskModel = - binding.sequence === Sequence.orchestrator && trace.model === 'binding'; - const model = perTaskModel ? 'chosen-per-task' : binding.model; - const modelSource = perTaskModel ? 'agent-prompts' : trace.model; - analytics.wizardCapture('switchboard resolved', { - program: ctx.program, - flag_self_driving_use_pi_harness: - ctx.flags[WIZARD_SELF_DRIVING_USE_PI_HARNESS_FLAG_KEY], - flag_self_driving_pi_payload: JSON.stringify( - ctx.flagPayloads?.[WIZARD_SELF_DRIVING_USE_PI_HARNESS_FLAG_KEY] ?? null, - ), - flag_orchestrator: ctx.flags[WIZARD_ORCHESTRATOR_FLAG_KEY], - cli_harness: ctx.cliHarness, - cli_sequence: ctx.cliSequence, - cli_model: ctx.cliModel, - harness_source: trace.harness, - model_source: modelSource, - sequence_source: trace.sequence, - harness: binding.harness, - model, - thinking_level: binding.thinkingLevel, - sequence: binding.sequence, - }); - logToFile( - `[switchboard] decision: program=${ctx.program}` + - ` in(orchestrator=${ctx.flags[WIZARD_ORCHESTRATOR_FLAG_KEY] ?? '-'},` + - ` cli=${ctx.cliHarness ?? '-'}/${ctx.cliSequence ?? '-'}/${ - ctx.cliModel ?? '-' - })` + - ` → harness=${binding.harness} (${trace.harness ?? '?'})` + - ` model=${model} (${modelSource ?? '?'})` + - ` sequence=${binding.sequence} (${trace.sequence ?? '?'})`, - ); -} diff --git a/src/programs/binding.ts b/src/programs/binding.ts deleted file mode 100644 index 59c30b348..000000000 --- a/src/programs/binding.ts +++ /dev/null @@ -1,160 +0,0 @@ -import { IS_PRODUCTION_BUILD } from '@env'; -import { - DEFAULT_AGENT_MODEL, - GPT5_6_SOL_MODEL, - GPT5_6_TERRA_MODEL, - Harness, - Sequence, -} from '@shared/constants'; -import { logToFile } from '@utils/debug'; -import { - DEFAULT_AGENT_BINDING, - harnessRunsTasks, - resolveHarness, -} from '@agent'; -import type { - ProgramBinding, - ResolvedBinding, - SwitchboardCtx as HarnessCtx, -} from '@agent/types'; -import type { ProgramId } from './program-registry'; -import { - isOrchestratorEnabled, - resolveFlagRoute, - resolveFlagSequence, -} from './experiments'; - -/** The agent's harness inputs plus the sequence inputs only programs read. */ -type SwitchboardCtx = HarnessCtx & { - flagSequence?: Sequence; - orchestratorFlagOn?: boolean; -}; - -export interface ProgramSwitchboardCtx { - program: ProgramId; - composed?: boolean; - flags: Record; - flagPayloads?: Record; - cliHarness?: Harness; - cliSequence?: Sequence; - cliModel?: string; - trace?: SwitchboardCtx['trace']; -} - -/** Program routes. The registry lockstep contract is tested at this boundary. */ -export const PROGRAM_BINDINGS: Partial> = { - 'posthog-integration': DEFAULT_AGENT_BINDING, - 'revenue-analytics-setup': DEFAULT_AGENT_BINDING, - 'warehouse-source': DEFAULT_AGENT_BINDING, - 'error-tracking-upload-source-maps': { - sequence: Sequence.linear, - harness: Harness.pi, - model: GPT5_6_SOL_MODEL, - thinkingLevel: 'medium', - }, - audit: DEFAULT_AGENT_BINDING, - 'events-audit': DEFAULT_AGENT_BINDING, - 'posthog-doctor': DEFAULT_AGENT_BINDING, - 'web-analytics-doctor': DEFAULT_AGENT_BINDING, - migration: DEFAULT_AGENT_BINDING, - 'self-driving': DEFAULT_AGENT_BINDING, - 'agent-skill': DEFAULT_AGENT_BINDING, - 'mcp-add': DEFAULT_AGENT_BINDING, - 'mcp-remove': DEFAULT_AGENT_BINDING, - 'mcp-tutorial': DEFAULT_AGENT_BINDING, - 'mcp-analytics': DEFAULT_AGENT_BINDING, - metrics: { - sequence: Sequence.orchestrator, - harness: Harness.pi, - model: DEFAULT_AGENT_MODEL, - }, - 'replay-vision': { - sequence: Sequence.orchestrator, - harness: Harness.anthropic, - model: DEFAULT_AGENT_MODEL, - }, - 'error-tracking': { - sequence: Sequence.orchestrator, - harness: Harness.pi, - model: DEFAULT_AGENT_MODEL, - }, - 'ai-observability': { - sequence: Sequence.linear, - harness: Harness.pi, - model: GPT5_6_TERRA_MODEL, - thinkingLevel: 'high', - }, - slack: DEFAULT_AGENT_BINDING, -}; - -/** Resolve product policy once; the agent receives only the resulting route. */ -export function resolveProgramBinding( - ctx: ProgramSwitchboardCtx, -): ResolvedBinding { - ctx.trace ??= {}; - const baseBinding: ProgramBinding = - PROGRAM_BINDINGS[ctx.program] ?? DEFAULT_AGENT_BINDING; - const resolution = { - program: ctx.program, - baseBinding, - composed: ctx.composed, - flagRoute: resolveFlagRoute(ctx.program, ctx.flags, ctx.flagPayloads), - flagSequence: resolveFlagSequence(ctx.program, ctx.flags), - orchestratorFlagOn: isOrchestratorEnabled(ctx.flags), - cliHarness: ctx.cliHarness, - cliSequence: ctx.cliSequence, - cliModel: ctx.cliModel, - trace: ctx.trace, - }; - const sequence = resolveSequence(resolution); - const { harness, model, thinkingLevel } = resolveHarness(resolution); - const binding = { sequence, harness, model, thinkingLevel }; - const roles = Object.keys(baseBinding.contextMillOverride ?? {}); - if (roles.length === 0) return binding; - return { - ...binding, - roleBindings: Object.fromEntries( - roles.map((role) => [ - role, - resolveHarness({ ...resolution, trace: undefined }, role), - ]), - ), - }; -} - -function resolveSequence(ctx: SwitchboardCtx): Sequence { - const [source, sequence] = pickSequence(ctx); - if (ctx.trace) ctx.trace.sequence = source; - logToFile( - `[switchboard] resolved: program=${ - ctx.program ?? '?' - } sequence=${sequence} (${source})`, - ); - return sequence; -} - -/** - * The first rung that decides wins: the composed clamp, the dev-build CLI - * override, the runTask capability clamp, the flag route, the experiment, then - * the base binding. CLI sits above the capability clamp, so `--sequence - * orchestrator` still reaches the orchestrator's hard error in dev builds. - */ -function pickSequence( - ctx: SwitchboardCtx, -): [Required>['sequence'], Sequence] { - // The orchestrator owns the whole run lifecycle and cannot nest. - if (ctx.composed) return ['composed', Sequence.linear]; - if (!IS_PRODUCTION_BUILD && ctx.cliSequence) return ['cli', ctx.cliSequence]; - const { harness } = resolveHarness(ctx); - if (!harnessRunsTasks(harness)) { - if (ctx.orchestratorFlagOn) { - logToFile( - `[switchboard] wizard-orchestrator ignored: ${harness} has no runTask, clamping to linear`, - ); - } - return ['runtask-clamp', Sequence.linear]; - } - if (ctx.flagRoute?.sequence) return ['payload', ctx.flagRoute.sequence]; - if (ctx.flagSequence) return ['flag', ctx.flagSequence]; - return ['binding', (ctx.baseBinding ?? DEFAULT_AGENT_BINDING).sequence]; -} diff --git a/src/programs/commandments.ts b/src/programs/commandments.ts deleted file mode 100644 index d89d267de..000000000 --- a/src/programs/commandments.ts +++ /dev/null @@ -1,18 +0,0 @@ -// ── Program axis ──────────────────────────────────────────────────────── - -const SELF_DRIVING = [ - 'ALWAYS surface a custom-scout proposal in step 6b: bring the user your one or two strongest candidate scouts even when the built-in troop looks sufficient. The proposal ask leads with a "None — keep the built-in troop" option, so declining costs the user one keystroke — but a proposal you silently skip is coverage they never got to see or judge. Where the skill says to skip the ask when the gap analysis finds no candidate, do NOT skip: pick your best candidates anyway and let the user decide.', - - 'Rank candidates at the discriminator level, not the category level. "Covered" only means an enabled scout would actually FIRE for that failure mode: a conversion-rate watcher does not catch entry volume collapsing; a Stripe-transaction watcher does not catch a lead form going silent. A surface whose failure mode has no firing condition among the enabled scouts is your strongest candidate.', - - 'Be honest in the option descriptions: if a candidate overlaps something an enabled scout partially watches, say so in its description rather than dropping the candidate. The user chooses with full information; you do not gatekeep on their behalf.', -]; - -const PROGRAM_COMMANDMENTS: Record = { - 'self-driving': SELF_DRIVING, -}; - -/** Select program-specific guidance before invoking the agent. */ -export function getProgramCommandments(program: string): readonly string[] { - return PROGRAM_COMMANDMENTS[program] ?? []; -} diff --git a/src/programs/credentials.ts b/src/programs/credentials.ts index 94a81ab63..a3c197e1d 100644 --- a/src/programs/credentials.ts +++ b/src/programs/credentials.ts @@ -1,14 +1,9 @@ /** Resolved credentials passed from a program host to one agent run. */ -import { gatewayAuth } from './gateway-session'; -import type { InferenceAuthProvider } from '@agent/types'; -import type { GatewayAuth } from '@shared/gateway-auth'; import type { ApiProject, ApiUser, Credentials } from '@shared/api'; export type ResolvedProgramCredentials = { posthog: Credentials; - /** When absent, runProgram mints first-party gateway auth from the refreshed login. */ - inferenceAuth?: InferenceAuthProvider; project: ApiProject | null; apiUser: ApiUser | null; }; @@ -20,14 +15,3 @@ export type CredentialsProvider = { context: { signal: AbortSignal }, ): Promise; }; - -/** First-party inference auth; each resolve reuses the gateway session's cache and near-expiry refresh. */ -export function createPosthogInferenceAuthProvider( - posthog: Credentials, - programId: string, -): InferenceAuthProvider { - return { - resolve: (): Promise => - gatewayAuth(posthog.host, posthog.accessToken, programId), - }; -} diff --git a/src/programs/detection/__tests__/agentic-progress.test.ts b/src/programs/detection/__tests__/agentic-progress.test.ts index f63f1fbdd..40d7a4c38 100644 --- a/src/programs/detection/__tests__/agentic-progress.test.ts +++ b/src/programs/detection/__tests__/agentic-progress.test.ts @@ -8,16 +8,28 @@ import { buildSession } from '@lib/wizard-session'; import { HostResolution } from '@shared/host-resolution'; import { ErrorCodes } from '@shared/errors'; import { getUI } from '@ui'; -import { createUiReducer } from '@ui/agent-progress'; vi.mock('@utils/debug'); +// Detection runs the real runAgent pipeline: no analytics or gateway mint may leave the process. +vi.mock('@utils/analytics'); +vi.mock('@agent/gateway-session', async (original) => ({ + ...(await original()), + gatewayAuth: vi.fn(() => + Promise.resolve({ + gatewayUrl: 'https://gateway.test', + token: 'phe_test', + refreshAtMs: Infinity, + }), + ), +})); vi.mock('@ui', () => ({ getUI: () => ui })); const ui = vi.hoisted(() => ({ addTokenUsage: vi.fn(), setStage: vi.fn(), pushStatus: vi.fn(), showAuthError: vi.fn(), - log: { error: vi.fn() }, + startRun: vi.fn(), + log: { error: vi.fn(), warn: vi.fn(), info: vi.fn() }, })); vi.mock('@agent/agent-interface', async (original) => ({ ...(await original()), @@ -37,6 +49,7 @@ function detectionSession() { } beforeEach(() => { + vi.clearAllMocks(); vi.mocked(initializeAgent).mockReset(); vi.mocked(executeAgent).mockReset(); }); @@ -78,26 +91,11 @@ it('keeps initialization and execution progress visible during detection', async return Promise.resolve({ kind: 'success' }); }, ); - const session = detectionSession(); - const inferenceAuth = { resolve: vi.fn() }; - session.inferenceAuth = inferenceAuth; - const onProgress = vi.fn(createUiReducer(getUI())); - const report = await detectProjectsWithAgent(session, { + const report = await detectProjectsWithAgent(detectionSession(), { programId: 'posthog-integration', targets: [{ id: 'node', name: 'Node.js' }], - onProgress, }); expect(report.projects[0].targetId).toBe('node'); - expect(onProgress.mock.calls.map(([event]) => event.kind)).toEqual([ - 'log', - 'usage', - 'stage', - 'status', - 'log', - ]); - expect(vi.mocked(initializeAgent).mock.calls[0][0].inferenceAuth).toBe( - inferenceAuth, - ); expect(getUI().addTokenUsage).toHaveBeenCalledWith(delta); expect(ui.setStage).toHaveBeenCalledWith('Scanning'); expect(ui.pushStatus).toHaveBeenCalledWith('Found a project'); @@ -107,105 +105,7 @@ it('keeps initialization and execution progress visible during detection', async ]); }); -import type { AgentProgress } from '@agent/types'; - -// The test below runs the real runAgent pipeline, so no analytics or gateway mint may leave the process. -vi.mock('@utils/analytics'); -vi.mock('@programs/credentials', () => ({ - createPosthogInferenceAuthProvider: vi.fn(() => ({ - resolve: () => - Promise.resolve({ - gatewayUrl: 'https://gateway.test', - token: 'phe_test', - refreshAtMs: Infinity, - }), - })), -})); - -afterEach(() => vi.restoreAllMocks()); - -const cancelled = { - kind: 'abort', - classification: AgentErrorType.ABORT, - message: 'Agent run cancelled', -} as const; - -/** Each attempt's deadline, fired by the test instead of the clock. */ -function fakeDeadlines(): AbortController[] { - const deadlines: AbortController[] = []; - vi.spyOn(AbortSignal, 'timeout').mockImplementation(() => { - const deadline = new AbortController(); - deadlines.push(deadline); - return deadline.signal; - }); - return deadlines; -} - -it('keeps both detection attempts on the host progress sink and the session provider', async () => { - ui.setStage.mockClear(); - ui.pushStatus.mockClear(); - const deadlines = fakeDeadlines(); - vi.mocked(initializeAgent).mockImplementation((config) => { - config.emit?.({ kind: 'status', message: 'Initializing' }); - return Promise.resolve({ emit: config.emit } as Awaited< - ReturnType - >); - }); - vi.mocked(executeAgent) - .mockImplementationOnce((config) => { - config.emit?.({ kind: 'stage', stage: 'First scan' }); - deadlines.at(-1)?.abort(); - return Promise.resolve(cancelled); - }) - .mockImplementationOnce( - (config, _prompt, _options, _spinner, _messages, middleware) => { - config.emit?.({ kind: 'stage', stage: 'Second scan' }); - middleware?.onMessage({ - type: 'result', - result: - '{"projects":[{"path":".","targetId":"node","framework":"Node.js"}]}', - }); - return Promise.resolve({ kind: 'success' }); - }, - ); - const session = detectionSession(); - const inferenceAuth = { resolve: vi.fn() }; - session.inferenceAuth = inferenceAuth; - const onProgress = vi.fn(); - const onEvent = vi.fn(); - - const report = await detectProjectsWithAgent(session, { - programId: 'posthog-integration', - targets: [{ id: 'node', name: 'Node.js' }], - onProgress, - onEvent, - }); - - expect(report.projects[0].targetId).toBe('node'); - expect(onEvent).toHaveBeenCalledWith('Project scan timed out; retrying...'); - const configs = vi - .mocked(initializeAgent) - .mock.calls.map(([config]) => config); - expect(configs).toHaveLength(2); - for (const config of configs) { - expect(config.inferenceAuth).toBe(inferenceAuth); - } - expect( - onProgress.mock.calls - .map(([event]) => event as AgentProgress) - .filter((event) => event.kind === 'status' || event.kind === 'stage'), - ).toEqual([ - { kind: 'status', message: 'Initializing' }, - { kind: 'stage', stage: 'First scan' }, - { kind: 'status', message: 'Initializing' }, - { kind: 'stage', stage: 'Second scan' }, - ]); - // The host owns the sink, so detection itself never reaches for the UI. - expect(ui.setStage).not.toHaveBeenCalled(); - expect(ui.pushStatus).not.toHaveBeenCalled(); -}); - -it('sends each agent step to onEvent and the host only the progress it saw before', async () => { +it('sends each agent step to onEvent and the UI only the progress it saw before runAgent', async () => { vi.mocked(initializeAgent).mockImplementation((config) => Promise.resolve({ emit: config.emit } as Awaited< ReturnType @@ -236,24 +136,21 @@ it('sends each agent step to onEvent and the host only the progress it saw befor return Promise.resolve({ kind: 'success' }); }, ); - const events: AgentProgress[] = []; const lines: string[] = []; const report = await detectProjectsWithAgent(detectionSession(), { programId: 'posthog-integration', targets: [{ id: 'node', name: 'Node.js' }], onEvent: (line) => lines.push(line), - onProgress: (event) => events.push(event), }); expect(report.projects[0].targetId).toBe('node'); expect(lines).toEqual(['Reading the root manifest.', 'Read package.json']); - expect( - events.map((event) => - event.kind === 'log' ? `log:${event.level}` : event.kind, - ), - ).toEqual(['status', 'log:warn', 'activity', 'activity']); - expect(ui.pushStatus).not.toHaveBeenCalledWith('Scanning'); + expect(ui.pushStatus).toHaveBeenCalledWith('Scanning'); + expect(ui.log.warn).toHaveBeenCalledWith('Warn line'); + // The scan's run lifecycle and setup logs never reach the program's UI. + expect(ui.log.info).not.toHaveBeenCalled(); + expect(ui.startRun).not.toHaveBeenCalled(); }); it('stops optional detection on a data-only 401 before parsing partial JSON', async () => { @@ -280,11 +177,7 @@ it('stops optional detection on a data-only 401 before parsing partial JSON', as programId: 'posthog-integration', targets: [{ id: 'node', name: 'Node.js' }], }), - ).rejects.toMatchObject({ - name: 'WizardError', - code: ErrorCodes.AuthInvalidOrExpired, - message: 'Authentication failed (401)', - }); + ).rejects.toThrow('Authentication failed (401)'); expect(ui.showAuthError).not.toHaveBeenCalled(); }); @@ -310,7 +203,7 @@ it('preserves the original error from a decided failure', async () => { ).rejects.toBe(original); }); -it('rejects classified agent failures without retrying', async () => { +it('rejects classified agent failures', async () => { vi.mocked(initializeAgent).mockResolvedValue( {} as Awaited>, ); @@ -326,5 +219,4 @@ it('rejects classified agent failures without retrying', async () => { targets: [{ id: 'node', name: 'Node.js' }], }), ).rejects.toThrow('Agent API unavailable'); - expect(vi.mocked(executeAgent)).toHaveBeenCalledOnce(); }); diff --git a/src/programs/detection/__tests__/agentic-retry.test.ts b/src/programs/detection/__tests__/agentic-retry.test.ts index d9c373cc2..f0454c4eb 100644 --- a/src/programs/detection/__tests__/agentic-retry.test.ts +++ b/src/programs/detection/__tests__/agentic-retry.test.ts @@ -10,6 +10,7 @@ import { } from '@agent/agent-interface'; import { buildSession } from '@lib/wizard-session'; import { HostResolution } from '@shared/host-resolution'; +import { flushScanReport } from '@agent/yara-hooks'; import { Harness, HAIKU_MODEL, Sequence } from '@shared/constants'; vi.mock('@utils/analytics'); @@ -23,15 +24,20 @@ vi.mock('@agent', async (importOriginal) => { const actual = await importOriginal(); return { ...actual, runAgent: vi.fn(actual.runAgent) }; }); -vi.mock('@programs/credentials', () => ({ - createPosthogInferenceAuthProvider: vi.fn(() => ({ - resolve: () => - Promise.resolve({ - gatewayUrl: 'https://gateway.test', - token: 'phe_test', - refreshAtMs: Infinity, - }), - })), +vi.mock('@agent/yara-hooks', async (importOriginal) => ({ + ...(await importOriginal()), + flushScanReport: vi.fn(), +})); +// The runner mints before each attempt; no mint may leave the process. +vi.mock('@agent/gateway-session', async (importOriginal) => ({ + ...(await importOriginal()), + gatewayAuth: vi.fn(() => + Promise.resolve({ + gatewayUrl: 'https://gateway.test', + token: 'phe_test', + refreshAtMs: Infinity, + }), + ), })); const init = vi.mocked(initializeAgent); @@ -98,7 +104,7 @@ describe('agentic detection retry', () => { afterEach(() => vi.restoreAllMocks()); - it('runs both attempts through runAgent with the detection binding, read-only tools, a deferred scan report and one inference provider', async () => { + it('runs both attempts through runAgent on linear Haiku with read-only tools, its own prompt, no remark and a deferred scan report', async () => { timeOut(); emitResult(verdict); @@ -114,10 +120,18 @@ describe('agentic detection retry', () => { }); expect(config.allowedTools).toEqual(['Read', 'Grep', 'Glob']); expect(config.scanReport).toBe('defer'); + expect(config.run).toMatchObject({ + collectTranscript: true, + requestRemark: false, + }); } - const [[, first], [, second]] = calls; - expect(first.inferenceAuth).toBeDefined(); - expect(second.inferenceAuth).toBe(first.inferenceAuth); + // The run definition's prompt replaces the assembled program prompt. + expect(execute.mock.calls[0][1]).toContain( + 'You are scanning a code repository', + ); + expect(execute.mock.calls[0][4]).toMatchObject({ requestRemark: false }); + // The program run's report counts the scan's scans. + expect(flushScanReport).not.toHaveBeenCalled(); }); it('restarts the scan once when the first result has no JSON', async () => { diff --git a/src/programs/detection/__tests__/framework-labels.test.ts b/src/programs/detection/__tests__/framework-labels.test.ts deleted file mode 100644 index 17f6a635e..000000000 --- a/src/programs/detection/__tests__/framework-labels.test.ts +++ /dev/null @@ -1,168 +0,0 @@ -import { ASTRO_AGENT_CONFIG } from '@programs/frameworks/astro/astro-wizard-agent'; -import { AstroRenderingMode } from '@programs/frameworks/astro/utils'; -import { DJANGO_AGENT_CONFIG } from '@programs/frameworks/django/django-wizard-agent'; -import { DjangoProjectType } from '@programs/frameworks/django/utils'; -import { FASTAPI_AGENT_CONFIG } from '@programs/frameworks/fastapi/fastapi-wizard-agent'; -import { FastAPIProjectType } from '@programs/frameworks/fastapi/utils'; -import { FLASK_AGENT_CONFIG } from '@programs/frameworks/flask/flask-wizard-agent'; -import { FlaskProjectType } from '@programs/frameworks/flask/utils'; -import { LARAVEL_AGENT_CONFIG } from '@programs/frameworks/laravel/laravel-wizard-agent'; -import { LaravelProjectType } from '@programs/frameworks/laravel/utils'; -import { NEXTJS_AGENT_CONFIG } from '@programs/frameworks/nextjs/nextjs-wizard-agent'; -import { NextJsRouter } from '@programs/frameworks/nextjs/utils'; -import { RAILS_AGENT_CONFIG } from '@programs/frameworks/rails/rails-wizard-agent'; -import { RailsProjectType } from '@programs/frameworks/rails/utils'; -import { REACT_NATIVE_AGENT_CONFIG } from '@programs/frameworks/react-native/react-native-wizard-agent'; -import { ReactNativeVariant } from '@programs/frameworks/react-native/utils'; -import { REACT_ROUTER_AGENT_CONFIG } from '@programs/frameworks/react-router/react-router-wizard-agent'; -import { ReactRouterMode } from '@programs/frameworks/react-router/utils'; -import { TANSTACK_ROUTER_AGENT_CONFIG } from '@programs/frameworks/tanstack-router/tanstack-router-wizard-agent'; -import { TanStackRouterMode } from '@programs/frameworks/tanstack-router/utils'; - -describe('framework detection labels', () => { - it.each([ - [ - 'Next.js app router', - () => - NEXTJS_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - router: NextJsRouter.APP_ROUTER, - }), - 'Next.js app router 📱', - ], - [ - 'Next.js pages router', - () => - NEXTJS_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - router: NextJsRouter.PAGES_ROUTER, - }), - 'Next.js pages router 📃', - ], - [ - 'Next.js unknown router', - () => NEXTJS_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({}), - undefined, - ], - [ - 'Astro SSR', - () => - ASTRO_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - renderingMode: AstroRenderingMode.SSR, - }), - 'Astro Server (SSR)', - ], - [ - 'React Router data mode', - () => - REACT_ROUTER_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - routerMode: ReactRouterMode.V7_DATA, - }), - 'React Router v7 Data mode', - ], - [ - 'TanStack file routes', - () => - TANSTACK_ROUTER_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - routerMode: TanStackRouterMode.FILE_BASED, - }), - 'TanStack Router File-based routing', - ], - [ - 'Django Wagtail', - () => - DJANGO_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - projectType: DjangoProjectType.WAGTAIL, - }), - 'Django with Wagtail CMS', - ], - [ - 'Django standard', - () => - DJANGO_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - projectType: DjangoProjectType.STANDARD, - }), - 'Django', - ], - [ - 'Flask RESTX', - () => - FLASK_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - projectType: FlaskProjectType.RESTX, - }), - 'Flask-RESTX', - ], - [ - 'Flask standard', - () => - FLASK_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - projectType: FlaskProjectType.STANDARD, - }), - 'Flask', - ], - [ - 'FastAPI fullstack', - () => - FASTAPI_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - projectType: FastAPIProjectType.FULLSTACK, - }), - 'FastAPI fullstack with templates', - ], - [ - 'FastAPI standard', - () => - FASTAPI_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - projectType: FastAPIProjectType.STANDARD, - }), - 'FastAPI', - ], - [ - 'Laravel Inertia', - () => - LARAVEL_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - projectType: LaravelProjectType.INERTIA, - }), - 'Laravel with Inertia.js', - ], - [ - 'Laravel standard', - () => - LARAVEL_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - projectType: LaravelProjectType.STANDARD, - }), - 'Laravel', - ], - [ - 'Rails API', - () => - RAILS_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - projectType: RailsProjectType.API, - }), - 'Rails API-only', - ], - [ - 'Rails standard', - () => - RAILS_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - projectType: RailsProjectType.STANDARD, - }), - 'Rails', - ], - [ - 'React Native Expo', - () => - REACT_NATIVE_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - variant: ReactNativeVariant.EXPO, - }), - 'Expo 📱', - ], - [ - 'React Native bare', - () => - REACT_NATIVE_AGENT_CONFIG.metadata.getDetectedFrameworkLabel?.({ - variant: ReactNativeVariant.REACT_NATIVE, - }), - 'React Native 📱', - ], - ])('keeps the %s label', (_name, getLabel, expected) => { - expect(getLabel()).toBe(expected); - }); -}); diff --git a/src/programs/detection/__tests__/project-scope.test.ts b/src/programs/detection/__tests__/project-scope.test.ts index 70c00cf0a..79f5ad8d8 100644 --- a/src/programs/detection/__tests__/project-scope.test.ts +++ b/src/programs/detection/__tests__/project-scope.test.ts @@ -73,15 +73,7 @@ describe('chooseIntegrationProject', () => { }); describe('scopeInstallDirToProject', () => { - const host: ProgramCiHost = { - auth: { - setCredentials: vi.fn(), - setRoleAtOrganization: vi.fn(), - setApiUser: vi.fn(), - }, - log: { info: vi.fn(), warn: vi.fn() }, - onProgress: vi.fn(), - }; + const host: ProgramCiHost = { log: { info: vi.fn(), warn: vi.fn() } }; const scan = vi.mocked(detectProjectsWithAgent); const FLAG_ON = { [WIZARD_BASIC_INTEGRATION_AGENTIC_DETECTION_FLAG_KEY]: 'true', @@ -127,11 +119,7 @@ describe('scopeInstallDirToProject', () => { const session = buildSession({ installDir: '/repo' }); await scopeInstallDirToProject(session, host); - expect(vi.mocked(authenticate)).toHaveBeenCalledWith( - session, - 'posthog-integration', - host.auth, - ); + expect(vi.mocked(authenticate)).toHaveBeenCalledTimes(1); expect(session.installDir).toBe('/repo'); expect(outcomeEvent()).toMatchObject({ outcome: 'flag-off' }); expect(scan).not.toHaveBeenCalled(); @@ -146,14 +134,8 @@ describe('scopeInstallDirToProject', () => { expect(scan).toHaveBeenCalledWith( expect.anything(), - expect.objectContaining({ - programId: 'posthog-integration', - onProgress: expect.any(Function), - }), + expect.objectContaining({ programId: 'posthog-integration' }), ); - const progress = { kind: 'status' as const, message: 'Scanning' }; - scan.mock.calls[0]?.[1].onProgress?.(progress); - expect(host.onProgress).toHaveBeenCalledWith(progress); }); it('re-points installDir at the recommended project and fires recommended with scan facts', async () => { @@ -227,7 +209,6 @@ describe('scopeInstallDirToProject', () => { expect(host.log.warn).toHaveBeenCalledWith( 'Project scan attempt 2 timed out after 90s; continuing with the install dir as-is.', ); - expect(exceptionSpy).not.toHaveBeenCalled(); }); it('uses a valid retry result after the old 60-second caller deadline', async () => { diff --git a/src/programs/detection/agentic.ts b/src/programs/detection/agentic.ts index c946ad576..56ffd0d15 100644 --- a/src/programs/detection/agentic.ts +++ b/src/programs/detection/agentic.ts @@ -13,11 +13,10 @@ * program uses, which needs credentials. */ -import { AgentSignals, runAgent, RunOutcome } from '@agent'; +import { AgentSignals, buildRunTags, runAgent, RunOutcome } from '@agent'; import type { AgentProgress, AgentRunDefinition, - InferenceAuthProvider, ResolvedBinding, RunConfig, RunInput, @@ -34,12 +33,10 @@ import { POSTHOG_DOCS_URL, Sequence, } from '@shared/constants'; -import type { Credentials } from '@shared/api'; -import { buildRunTags } from '@shared/run-tags'; import { analytics } from '@utils/analytics'; -import type { WizardRunOptions } from '@utils/types'; -import { createPosthogInferenceAuthProvider } from '@programs/credentials'; -import { WizardError } from '@shared/errors'; +import type { WizardSession } from '@lib/wizard-session'; +import { getUI } from '@ui'; +import { createUiReducer } from '@ui/agent-progress'; /** A category the agent classifies each project into (id the agent returns). */ export type DetectTarget = { id: string; name: string }; @@ -153,14 +150,6 @@ export type AgenticDetectOptions = { rerankIds?: readonly string[]; /** Streaming activity callback for the UI. */ onEvent?: DetectEvent; - /** Host-owned sink for the scan's run progress. */ - onProgress?: (event: AgentProgress) => void; -}; - -/** Data the detection agent needs from its host; no UI or session ownership. */ -export type AgenticDetectionContext = WizardRunOptions & { - credentials: Credentials | null; - inferenceAuth?: InferenceAuthProvider; }; function buildPrompt( @@ -346,7 +335,7 @@ function detectionRunDefinition(prompt: string): AgentRunDefinition { }; } -/** What a detect host saw before `runAgent`: no run lifecycle, spinner, outro, or setup logs below warn. */ +/** What the UI saw before `runAgent`: no run lifecycle, spinner, outro, or setup logs below warn. */ function reachesHost(event: AgentProgress): boolean { switch (event.kind) { case 'lifecycle': @@ -362,7 +351,7 @@ function reachesHost(event: AgentProgress): boolean { /** Scan the repo with Haiku through `runAgent`; each attempt is a fresh run with its own deadline. */ export async function detectProjectsWithAgent( - session: AgenticDetectionContext, + session: WizardSession, options: AgenticDetectOptions, ): Promise { if (!session.credentials) { @@ -375,9 +364,7 @@ export async function detectProjectsWithAgent( recommend = false, rerankIds, onEvent, - onProgress, } = options; - const { credentials } = session; // Built here: the scan runs before the program's own run tags exist. const wizardMetadata = { @@ -396,6 +383,13 @@ export async function detectProjectsWithAgent( ), composed: true, binding: AGENTIC_DETECTION_BINDING, + // Only the orchestrator reads it; the scan is linear. + switchboard: { + program: programId, + composed: true, + flags: {}, + flagPayloads: {}, + }, skillsBaseUrl: getSkillsBaseUrl(), wizardFlags: {}, wizardFlagPayloads: {}, @@ -406,11 +400,7 @@ export async function detectProjectsWithAgent( }; const input: RunInput = { installDir: session.installDir, - credentials, - // One provider for both attempts: each resolves its own gateway bearer. - inferenceAuth: - session.inferenceAuth ?? - createPosthogInferenceAuthProvider(credentials, programId), + credentials: session.credentials, project: null, apiUser: null, // No benchmark pipeline and no AIO capture: the scan never had either. @@ -426,9 +416,10 @@ export async function detectProjectsWithAgent( }, host: { projectId: session.projectId, apiKey: session.apiKey }, }; + const reduceUi = createUiReducer(getUI()); const forward = (event: AgentProgress): void => { if (event.kind === 'activity') onEvent?.(event.line); - if (reachesHost(event)) onProgress?.(event); + if (reachesHost(event)) reduceUi(event); }; for (let attempt = 0; attempt < 2; attempt++) { @@ -450,10 +441,7 @@ export async function detectProjectsWithAgent( throw new AgenticDetectionTimeoutError(attempt + 1, timeoutMs); } if (result.outcome !== RunOutcome.Success) { - throw ( - result.failure.error ?? - new WizardError(result.failure.message, undefined, result.failure.code) - ); + throw result.failure.error ?? new Error(result.failure.message); } // Transcript first, final message last — its verdicts win path conflicts. diff --git a/src/programs/detection/context.ts b/src/programs/detection/context.ts index 70f161e09..23ab66832 100644 --- a/src/programs/detection/context.ts +++ b/src/programs/detection/context.ts @@ -11,16 +11,6 @@ import { DETECTION_TIMEOUT_MS } from '@shared/constants'; import type { FrameworkConfig } from '@programs/framework-config'; import type { WizardRunOptions } from '@utils/types'; -/** Host data used when gathering context for a chosen project. */ -export type FrameworkDetectionState = Pick< - WizardRunOptions, - 'installDir' | 'debug' | 'signup' | 'ci' | 'benchmark' | 'yaraReport' -> & { - frameworkConfig: FrameworkConfig | null; - frameworkContext: Record; - detectedFrameworkLabel: string | null; -}; - /** * Run a framework's `gatherContext()` to collect variant-specific * metadata (e.g., router type for Next.js, Expo vs bare for React Native). diff --git a/src/programs/detection/features.ts b/src/programs/detection/features.ts index bd363ef29..2975d5f6c 100644 --- a/src/programs/detection/features.ts +++ b/src/programs/detection/features.ts @@ -8,7 +8,7 @@ import { join } from 'path'; import { readProjectFile } from '@utils/bounded-fs'; -import { DiscoveredFeature } from '@shared/scan-consent'; +import { DiscoveredFeature } from '@lib/wizard-session'; const STRIPE_PACKAGES = new Set(['stripe', '@stripe/stripe-js']); diff --git a/src/programs/detection/project-scope.ts b/src/programs/detection/project-scope.ts index a072a6bc5..cfe55176f 100644 --- a/src/programs/detection/project-scope.ts +++ b/src/programs/detection/project-scope.ts @@ -5,19 +5,18 @@ import { AgenticDetectionTimeoutError, resolveProjectDir, type AgenticDetectionReport, - type AgenticDetectionContext, - type AgenticDetectOptions, type AgenticProject, type DetectEvent, type DetectTarget, } from './agentic.js'; -import { authenticate, type AuthSession } from '@programs/authenticate'; +import { authenticate } from '@programs/authenticate'; import type { ProgramCiHost } from '@programs/host-capabilities'; import { FRAMEWORK_REGISTRY } from '@programs/registry'; import { Integration, WIZARD_BASIC_INTEGRATION_AGENTIC_DETECTION_FLAG_KEY, } from '@shared/constants'; +import type { WizardSession } from '@lib/wizard-session'; import { analytics } from '@utils/analytics'; import { logToFile } from '@utils/debug'; @@ -60,13 +59,12 @@ export function toIntegrationCandidates( /** Run the agentic detector for the wizard's integration frameworks — the single home of targets + purpose. */ export async function detectIntegrationProjects( - session: AgenticDetectionContext, + session: WizardSession, options: { /** Program the scan bills to. Required so no caller can go unattributed. */ programId: string; recommend?: boolean; onEvent?: DetectEvent; - onProgress?: AgenticDetectOptions['onProgress']; }, ): Promise { // Spread first so the targets and purpose this function owns always win. @@ -103,15 +101,13 @@ function captureOutcome( analytics.wizardCapture('agentic detection', { outcome, ...properties }); } -export type ProjectScopeSession = AuthSession & AgenticDetectionContext; - /** Flag-gated non-interactive monorepo phase: scan, auto-choose the recommended project, re-point session.installDir; every failure leaves the session untouched. */ export async function scopeInstallDirToProject( - session: ProjectScopeSession, + session: WizardSession, host: ProgramCiHost, ): Promise { // Idempotent early auth: the detector needs credentials and the flag must evaluate as the logged-in user. - await authenticate(session, 'posthog-integration', host.auth); + await authenticate(session, 'posthog-integration'); const flags = await analytics.getAllFlagsForWizard(); if (flags[WIZARD_BASIC_INTEGRATION_AGENTIC_DETECTION_FLAG_KEY] !== 'true') { // A failed flag fetch surfaces as an empty map, so flag-off also covers "flags unavailable". @@ -129,7 +125,6 @@ export async function scopeInstallDirToProject( programId: 'posthog-integration', recommend: true, onEvent: (line) => logToFile('[agentic detect]', line), - onProgress: (event) => host.onProgress(event), }); } catch (err) { const error = err instanceof Error ? err : new Error(String(err)); diff --git a/src/commands/dispatch-family.ts b/src/programs/dispatch-family.ts similarity index 92% rename from src/commands/dispatch-family.ts rename to src/programs/dispatch-family.ts index c258e80bd..b92a25b6a 100644 --- a/src/commands/dispatch-family.ts +++ b/src/programs/dispatch-family.ts @@ -1,15 +1,17 @@ import type { Arguments } from 'yargs'; -import { AUDIT_CHECKS_FILE } from '@shared/audit-ledger'; +import { auditConfig } from '@programs/audit/index'; +import { AUDIT_CHECKS_FILE } from '@programs/audit/types'; import { WIZARD_TOOL_NAMES } from '@agent'; -import { getProgramConfig } from '@programs'; -import type { ProgramConfig } from '@programs/types'; +import { agentSkillConfig } from '@programs/program-registry'; +import { webAnalyticsDoctorConfig } from '@programs/web-analytics-doctor/index'; +import type { ProgramConfig } from '@programs/program-step'; import { getSkillsBaseUrl } from '@shared/constants'; import { fetchSkillMenu, type CliEntry } from '@shared/skill-menu'; import { analytics } from '@utils/analytics'; -import { dispatchProgram } from './factories/shared'; -import type { Command } from './command'; +import { dispatchProgram } from '../commands/factories/shared'; +import type { Command } from '../commands/command'; import { ErrorCodes } from '@shared/errors'; import { emitWizardError } from '@shared/errors'; @@ -48,7 +50,7 @@ async function exitDispatchError( /** Wizard-native subcommands keyed by family. */ const NATIVE_HANDLERS: Record> = { - audit: { 'web-analytics': getProgramConfig('web-analytics-doctor') }, + audit: { 'web-analytics': webAnalyticsDoctorConfig }, }; /** @@ -62,8 +64,7 @@ const NATIVE_HANDLERS: Record> = { * generic skill program picks up the ledger here rather than for every skill. */ function configForCliEntry(entry: CliEntry, family: string): ProgramConfig { - if (entry.skillId === 'audit') return getProgramConfig('audit'); - const agentSkillConfig = getProgramConfig('agent-skill'); + if (entry.skillId === 'audit') return auditConfig; return { ...agentSkillConfig, skillId: entry.skillId, diff --git a/src/programs/error-tracking-upload-source-maps/detect-agentic.ts b/src/programs/error-tracking-upload-source-maps/detect-agentic.ts index 574fca391..033635b19 100644 --- a/src/programs/error-tracking-upload-source-maps/detect-agentic.ts +++ b/src/programs/error-tracking-upload-source-maps/detect-agentic.ts @@ -16,9 +16,8 @@ import { type DetectTarget, type AgenticDetectionReport, type DetectEvent, - type AgenticDetectionContext, - type AgenticDetectOptions, } from '@programs/detection/agentic'; +import type { WizardSession } from '@lib/wizard-session'; import { VARIANT_DISPLAY_NAME, AUTOMATABLE_VARIANTS, @@ -402,9 +401,8 @@ export function coerceReport( /** Run the Haiku detector over the repo and classify projects for source maps. */ export async function detectSourceMapsProjects( - session: AgenticDetectionContext, + session: WizardSession, onEvent?: DetectEvent, - onProgress?: AgenticDetectOptions['onProgress'], ): Promise { const report = await detectProjectsWithAgent(session, { targets: SOURCE_MAPS_TARGETS, @@ -412,7 +410,6 @@ export async function detectSourceMapsProjects( purpose: 'set up PostHog Error Tracking source-map upload', rerankIds: JS_RERANK_VARIANTS, onEvent, - onProgress, }); return toSourceMapsReport(report, { hasBuildTarget: (path) => projectHasBuildTarget(session.installDir, path), diff --git a/src/programs/error-tracking-upload-source-maps/detect.ts b/src/programs/error-tracking-upload-source-maps/detect.ts index c78888e9f..3574ea1b9 100644 --- a/src/programs/error-tracking-upload-source-maps/detect.ts +++ b/src/programs/error-tracking-upload-source-maps/detect.ts @@ -16,6 +16,7 @@ import { MAX_WALK_FILES, safeReadFile, } from '@utils/bounded-fs'; +import type { WizardSession } from '@lib/wizard-session'; import type { AbortCase } from '@agent/types'; import { ErrorCodes } from '@shared/errors'; @@ -356,7 +357,7 @@ export const SOURCE_MAPS_CONTEXT_KEYS = { * only picks which variant the prompt should ask the agent to load. */ export function detectSourceMapsPrerequisites( - session: { installDir: string }, + session: WizardSession, setFrameworkContext: (key: string, value: unknown) => void, ): void { const fail = (error: SourceMapsDetectError) => diff --git a/src/programs/error-tracking-upload-source-maps/index.ts b/src/programs/error-tracking-upload-source-maps/index.ts index cb17fbcbe..b04956d89 100644 --- a/src/programs/error-tracking-upload-source-maps/index.ts +++ b/src/programs/error-tracking-upload-source-maps/index.ts @@ -1,7 +1,8 @@ import type { ProgramConfig } from '@programs/program-step'; import type { ProgramRun } from '@programs/program-run'; +import type { WizardSession } from '@lib/wizard-session'; +import { OutroKind } from '@lib/wizard-session'; import type { ProgramRunHost } from '@programs/host-capabilities'; -import { OutroKind } from '@agent'; import { ERROR_TRACKING_UPLOAD_SOURCE_MAPS_PROGRAM } from './steps.js'; import { buildSourceMapsUploadPrompt, @@ -13,6 +14,7 @@ import { VARIANTS_REQUIRING_POSTHOG_CLI, type SkillVariant, } from './detect.js'; +import { getContentBlocks } from '../../ui/tui/decks/error-tracking-upload-source-maps/index.js'; import { preinstallPostHogCliOnce } from '@programs/shared/posthog-cli-preinstall'; const REPORT_FILE = 'posthog-source-maps-report.md'; @@ -41,9 +43,10 @@ export const errorTrackingUploadSourceMapsConfig: ProgramConfig = { requiresAi: true, steps: ERROR_TRACKING_UPLOAD_SOURCE_MAPS_PROGRAM, reportFile: REPORT_FILE, + getContentBlocks, requires: ['posthog-integration'], - run: (_session, host: ProgramRunHost): Promise => { + run: (_session: WizardSession, host: ProgramRunHost): Promise => { // Read the picked project LIVE at prompt-build time, not here: the picker // screen runs AFTER this run config is resolved (post-auth), and the store // forks the session reference, so the `session` passed in never sees the diff --git a/src/programs/error-tracking-upload-source-maps/steps.ts b/src/programs/error-tracking-upload-source-maps/steps.ts index bda5bb353..466ae9074 100644 --- a/src/programs/error-tracking-upload-source-maps/steps.ts +++ b/src/programs/error-tracking-upload-source-maps/steps.ts @@ -8,12 +8,11 @@ */ import type { ProgramStep } from '@programs/program-step'; -import { RunPhase } from '@shared/run-state'; +import type { WizardSession } from '@lib/wizard-session'; +import { RunPhase } from '@lib/wizard-session'; import { SOURCE_MAPS_CONTEXT_KEYS } from './detect.js'; -function projectSelected(session: { - frameworkContext: Record; -}): boolean { +function projectSelected(session: WizardSession): boolean { return ( session.frameworkContext[SOURCE_MAPS_CONTEXT_KEYS.selectedVariant] != null ); diff --git a/src/programs/error-tracking/detect-agentic.ts b/src/programs/error-tracking/detect-agentic.ts index 733cb7dfb..654de6c68 100644 --- a/src/programs/error-tracking/detect-agentic.ts +++ b/src/programs/error-tracking/detect-agentic.ts @@ -14,17 +14,15 @@ import { Integration } from '@shared/constants'; import { resolveProjectDir, - type AgenticDetectionContext, - type AgenticDetectOptions, type AgenticDetectionReport, type DetectEvent, } from '@programs/detection/agentic'; import { gatherFrameworkContext } from '@programs/detection/index'; -import type { FrameworkDetectionState } from '@programs/detection/context'; import { detectIntegrationProjects, toIntegrationCandidates, } from '@programs/detection/project-scope'; +import type { WizardSession } from '@lib/wizard-session'; /** frameworkContext key for the picked project's path, relative to the repo root. */ export const ERROR_TRACKING_PROJECT_PATH_KEY = 'errorTrackingProjectPath'; @@ -72,24 +70,19 @@ export function toErrorTrackingReport( /** Scan the repo for projects, billed to error tracking. */ export async function detectErrorTrackingProjects( - session: AgenticDetectionContext, + session: WizardSession, onEvent?: DetectEvent, - onProgress?: AgenticDetectOptions['onProgress'], ): Promise { const report = await detectIntegrationProjects(session, { programId: 'error-tracking', recommend: true, onEvent, - onProgress, }); return toErrorTrackingReport(report); } /** The run's working directory: the picked project, else the repo root. */ -export function errorTrackingProjectDir(session: { - installDir: string; - frameworkContext: Record; -}): string { +export function errorTrackingProjectDir(session: WizardSession): string { return resolveProjectDir( session.installDir, session.frameworkContext[ERROR_TRACKING_PROJECT_PATH_KEY], @@ -98,7 +91,7 @@ export function errorTrackingProjectDir(session: { /** Gather framework context for `session.installDir`, keeping keys already set. */ export async function gatherErrorTrackingContext( - session: FrameworkDetectionState, + session: WizardSession, ): Promise { const frameworkConfig = session.frameworkConfig; if (!frameworkConfig) return; @@ -110,9 +103,6 @@ export async function gatherErrorTrackingContext( benchmark: session.benchmark, yaraReport: session.yaraReport, }); - const detectedLabel = - frameworkConfig.metadata.getDetectedFrameworkLabel?.(context); - if (detectedLabel) session.detectedFrameworkLabel = detectedLabel; for (const [key, value] of Object.entries(context)) { if (!(key in session.frameworkContext)) { session.frameworkContext[key] = value; diff --git a/src/programs/error-tracking/index.ts b/src/programs/error-tracking/index.ts index dd75bd68e..f9a712efc 100644 --- a/src/programs/error-tracking/index.ts +++ b/src/programs/error-tracking/index.ts @@ -1,19 +1,18 @@ import { Integration } from '@shared/constants'; import { detectFramework } from '@programs/detection/index'; -import { - scopeInstallDirToProject, - type ProjectScopeSession, -} from '@programs/detection/project-scope'; -import type { FrameworkDetectionState } from '@programs/detection/context'; +import { scopeInstallDirToProject } from '@programs/detection/project-scope'; import { FRAMEWORK_REGISTRY } from '@programs/registry'; import type { ProgramRun } from '@programs/program-run'; import { AGENT_SKILL_STEPS } from '@programs/agent-skill/steps'; +import { getContentBlocks } from '@ui/tui/decks/error-tracking/index'; +import { getTips } from '@ui/tui/decks/error-tracking/tips'; import { ERROR_TRACKING_UNSUPPORTED, errorTrackingProjectDir, gatherErrorTrackingContext, } from '@programs/error-tracking/detect-agentic'; import type { ProgramConfig, ProgramStep } from '@programs/program-step'; +import type { WizardSession } from '@lib/wizard-session'; import type { ProgramCiHost, ProgramRunHost, @@ -71,19 +70,11 @@ function maybePreinstallPostHogCli( if (!integration || !SYMBOL_UPLOAD_CLI_FRAMEWORKS.has(integration)) return; preinstallPostHogCliOnce( 'error tracking posthog-cli preinstall failed', - { - integration, - }, + { integration }, warn, ); } -type ErrorTrackingCiSession = ProjectScopeSession & - FrameworkDetectionState & { - integration: Integration | null; - skillId: string | null; - }; - /** * After login, the scan lists the repo's projects and the user picks one, as in * the legacy upload-source-maps program. The pick sets the framework preflight @@ -189,11 +180,10 @@ export const errorTrackingConfig: ProgramConfig = { agentFlow: 'error-tracking', steps: ERROR_TRACKING_STEPS, reportFile: ERROR_TRACKING_REPORT_FILE, + getContentBlocks, + getTips, - run: ( - session: { integration: Integration | null }, - host: ProgramRunHost, - ): Promise => { + run: (session: WizardSession, host: ProgramRunHost): Promise => { maybePreinstallPostHogCli(session.integration, (message) => host.warn(message), ); @@ -201,7 +191,7 @@ export const errorTrackingConfig: ProgramConfig = { }, ciPreRun: async ( - session: ErrorTrackingCiSession, + session: WizardSession, host: ProgramCiHost, ): Promise => { await scopeInstallDirToProject(session, host); diff --git a/src/programs/events-audit/index.ts b/src/programs/events-audit/index.ts index a7052b564..c0de7ae5d 100644 --- a/src/programs/events-audit/index.ts +++ b/src/programs/events-audit/index.ts @@ -1,12 +1,13 @@ import type { ProgramConfig } from '@programs/program-step'; import type { ProgramRun } from '@programs/program-run'; -import type { AdditionalFeature } from '@shared/constants'; -import { OutroKind } from '@agent'; +import type { WizardSession } from '@lib/wizard-session'; +import { OutroKind } from '@lib/wizard-session'; import { SPINNER_MESSAGE } from '@programs/framework-config'; import { isUsingTypeScript } from '@utils/setup-utils'; import { WIZARD_TOOL_NAMES } from '@agent'; import { EVENTS_AUDIT_PROGRAM } from './steps.js'; -import { AUDIT_CHECKS_FILE } from '@programs/audit/types'; +import { AUDIT_CHECKS_FILE, AUDIT_CHECKS_KEY } from '@programs/audit/types'; +import { seedAuditLedger } from '@programs/audit/seed'; import { EVENTS_AUDIT_SEED_CHECKS } from './seed.js'; // SETUP_REPORT_FILE is also re-exported for backward compat with existing @@ -16,12 +17,6 @@ import { EVENTS_AUDIT_SEED_CHECKS } from './seed.js'; import { SETUP_REPORT_FILE } from './constants.js'; export { SETUP_REPORT_FILE }; -type EventsAuditRunState = { - installDir: string; - typescript: boolean; - additionalFeatureQueue: AdditionalFeature[]; -}; - const DOCS_URL = 'https://posthog.com/docs/product-analytics/best-practices'; /** @@ -39,9 +34,6 @@ export const eventsAuditConfig: ProgramConfig = { // synchronously without unwrapping the deferred `run` function. reportFile: SETUP_REPORT_FILE, auditLedgerFile: AUDIT_CHECKS_FILE, - // The events-audit ledger is the 6-phase pipeline, not the doctor's 10 - // integrity checks. - auditSeedChecks: EVENTS_AUDIT_SEED_CHECKS, allowedTools: [ 'Agent', WIZARD_TOOL_NAMES.auditSeedChecks, @@ -50,12 +42,18 @@ export const eventsAuditConfig: ProgramConfig = { ], disallowedTools: [WIZARD_TOOL_NAMES.wizardAsk], - run: (session: EventsAuditRunState): Promise => { + run: (session: WizardSession): Promise => { const typeScriptDetected = isUsingTypeScript({ installDir: session.installDir, }); session.typescript = typeScriptDetected; + // Seed the audit ledger so AuditRunScreen has something to render + // before the agent emits its first check update. The events-audit + // ledger is the 6-phase pipeline, not the doctor's 10 integrity checks. + seedAuditLedger(session.installDir, EVENTS_AUDIT_SEED_CHECKS); + session.frameworkContext[AUDIT_CHECKS_KEY] = EVENTS_AUDIT_SEED_CHECKS; + return Promise.resolve({ skillId: 'events-audit', integrationLabel: 'events-audit', diff --git a/src/programs/events-audit/steps.ts b/src/programs/events-audit/steps.ts index a9feef6fb..53b2f59a2 100644 --- a/src/programs/events-audit/steps.ts +++ b/src/programs/events-audit/steps.ts @@ -9,14 +9,11 @@ */ import type { ProgramStep } from '@programs/program-step'; -import type { FrameworkConfig } from '@programs/framework-config'; -import { RunPhase } from '@shared/run-state'; +import type { WizardSession } from '@lib/wizard-session'; +import { RunPhase } from '@lib/wizard-session'; import { HEALTH_CHECK_STEP } from '@programs/shared/health-check-step'; -function needsSetup(session: { - frameworkConfig: FrameworkConfig | null; - frameworkContext: Record; -}): boolean { +function needsSetup(session: WizardSession): boolean { const config = session.frameworkConfig; if (!config?.metadata.setup?.questions) return false; diff --git a/src/programs/framework-config.ts b/src/programs/framework-config.ts index 2c082bc2d..dd33fbe7b 100644 --- a/src/programs/framework-config.ts +++ b/src/programs/framework-config.ts @@ -72,9 +72,6 @@ export interface FrameworkMetadata< */ gatherContext?: (options: WizardRunOptions) => Promise; - /** Label for the gathered variant, if gathering found a more specific name. */ - getDetectedFrameworkLabel?: (context: TContext) => string | undefined; - /** Optional additional MCP servers for this framework (e.g., Svelte MCP). */ additionalMcpServers?: Record; diff --git a/src/programs/frameworks/astro/astro-wizard-agent.ts b/src/programs/frameworks/astro/astro-wizard-agent.ts index d2ff4aee8..a84ad23b3 100644 --- a/src/programs/frameworks/astro/astro-wizard-agent.ts +++ b/src/programs/frameworks/astro/astro-wizard-agent.ts @@ -10,6 +10,7 @@ import { type PackageJson, } from '@utils/package-json'; import { tryGetPackageJson } from '@utils/setup-utils'; +import { getUI } from '@ui'; import { getAstroRenderingMode, getAstroVersionBucket, @@ -28,12 +29,11 @@ export const ASTRO_AGENT_CONFIG: FrameworkConfig = { docsUrl: 'https://posthog.com/docs/libraries/astro', gatherContext: async (options: WizardRunOptions) => { const renderingMode = await getAstroRenderingMode(options); + getUI().setDetectedFramework( + `Astro ${getAstroRenderingModeName(renderingMode)}`, + ); return { renderingMode }; }, - getDetectedFrameworkLabel: (context) => - context.renderingMode - ? `Astro ${getAstroRenderingModeName(context.renderingMode)}` - : undefined, }, detection: { diff --git a/src/programs/frameworks/django/django-wizard-agent.ts b/src/programs/frameworks/django/django-wizard-agent.ts index 91f63f602..6c93bbc4e 100644 --- a/src/programs/frameworks/django/django-wizard-agent.ts +++ b/src/programs/frameworks/django/django-wizard-agent.ts @@ -33,18 +33,6 @@ export const DJANGO_AGENT_CONFIG: FrameworkConfig = { const settingsFile = await findDjangoSettingsFile(options); return { projectType, settingsFile }; }, - getDetectedFrameworkLabel: (context) => { - switch (context.projectType) { - case DjangoProjectType.WAGTAIL: - return 'Django with Wagtail CMS'; - case DjangoProjectType.DRF: - return 'Django REST Framework'; - case DjangoProjectType.CHANNELS: - return 'Django Channels'; - case DjangoProjectType.STANDARD: - return 'Django'; - } - }, }, detection: { diff --git a/src/programs/frameworks/django/utils.ts b/src/programs/frameworks/django/utils.ts index 4c6ef112a..276c2b746 100644 --- a/src/programs/frameworks/django/utils.ts +++ b/src/programs/frameworks/django/utils.ts @@ -1,4 +1,5 @@ import { boundedGlob, readProjectFile } from '@utils/bounded-fs'; +import { getUI } from '@ui'; import type { WizardRunOptions } from '@utils/types'; import { createVersionBucket } from '@utils/semver'; import * as fs from 'node:fs'; @@ -154,20 +155,24 @@ export async function getDjangoProjectType( // Check for Wagtail first (CMS) if (await hasWagtail({ installDir })) { + getUI().setDetectedFramework('Django with Wagtail CMS'); return DjangoProjectType.WAGTAIL; } // Check for Django REST Framework if (await hasDRF({ installDir })) { + getUI().setDetectedFramework('Django REST Framework'); return DjangoProjectType.DRF; } // Check for Django Channels if (await hasChannels({ installDir })) { + getUI().setDetectedFramework('Django Channels'); return DjangoProjectType.CHANNELS; } // Default to standard Django + getUI().setDetectedFramework('Django'); return DjangoProjectType.STANDARD; } diff --git a/src/programs/frameworks/fastapi/fastapi-wizard-agent.ts b/src/programs/frameworks/fastapi/fastapi-wizard-agent.ts index 739d2664f..22326fc49 100644 --- a/src/programs/frameworks/fastapi/fastapi-wizard-agent.ts +++ b/src/programs/frameworks/fastapi/fastapi-wizard-agent.ts @@ -17,16 +17,11 @@ import * as path from 'node:path'; const EXTRA_IGNORE = ['**/env/**', '**/.env/**']; -type FastAPIContext = { - projectType?: FastAPIProjectType; - appFile?: string; -}; - /** * FastAPI framework configuration for the universal agent runner */ -export const FASTAPI_AGENT_CONFIG: FrameworkConfig = { +export const FASTAPI_AGENT_CONFIG: FrameworkConfig = { metadata: { name: 'FastAPI', integration: Integration.fastapi, @@ -37,16 +32,6 @@ export const FASTAPI_AGENT_CONFIG: FrameworkConfig = { const appFile = await findFastAPIAppFile(options); return { projectType, appFile }; }, - getDetectedFrameworkLabel: (context) => { - switch (context.projectType) { - case FastAPIProjectType.FULLSTACK: - return 'FastAPI fullstack with templates'; - case FastAPIProjectType.ROUTER: - return 'FastAPI with APIRouter'; - case FastAPIProjectType.STANDARD: - return 'FastAPI'; - } - }, }, detection: { diff --git a/src/programs/frameworks/fastapi/utils.ts b/src/programs/frameworks/fastapi/utils.ts index 8f5c08417..b208dd72f 100644 --- a/src/programs/frameworks/fastapi/utils.ts +++ b/src/programs/frameworks/fastapi/utils.ts @@ -1,5 +1,6 @@ import { major, minVersion } from 'semver'; import { boundedGlob, readProjectFile } from '@utils/bounded-fs'; +import { getUI } from '@ui'; import type { WizardRunOptions } from '@utils/types'; import * as path from 'node:path'; @@ -150,15 +151,18 @@ export async function getFastAPIProjectType( // Check for fullstack pattern (templates) if (await hasTemplates({ installDir })) { + getUI().setDetectedFramework('FastAPI fullstack with templates'); return FastAPIProjectType.FULLSTACK; } // Check for APIRouter (modular structure) if (await hasAPIRouter({ installDir })) { + getUI().setDetectedFramework('FastAPI with APIRouter'); return FastAPIProjectType.ROUTER; } // Default to standard FastAPI + getUI().setDetectedFramework('FastAPI'); return FastAPIProjectType.STANDARD; } diff --git a/src/programs/frameworks/flask/flask-wizard-agent.ts b/src/programs/frameworks/flask/flask-wizard-agent.ts index 2045d3d31..b781777fb 100644 --- a/src/programs/frameworks/flask/flask-wizard-agent.ts +++ b/src/programs/frameworks/flask/flask-wizard-agent.ts @@ -33,12 +33,6 @@ export const FLASK_AGENT_CONFIG: FrameworkConfig = { const appFile = await findFlaskAppFile(options); return { projectType, appFile }; }, - getDetectedFrameworkLabel: (context) => - context.projectType === FlaskProjectType.STANDARD - ? 'Flask' - : context.projectType - ? getFlaskProjectTypeName(context.projectType) - : undefined, }, detection: { diff --git a/src/programs/frameworks/flask/utils.ts b/src/programs/frameworks/flask/utils.ts index 70442f5b1..66eb46a1c 100644 --- a/src/programs/frameworks/flask/utils.ts +++ b/src/programs/frameworks/flask/utils.ts @@ -1,4 +1,5 @@ import { boundedGlob, readProjectFile } from '@utils/bounded-fs'; +import { getUI } from '@ui'; import type { WizardRunOptions } from '@utils/types'; import { createVersionBucket } from '@utils/semver'; import * as path from 'node:path'; @@ -236,25 +237,30 @@ export async function getFlaskProjectType( // Check for Flask-RESTX first (most specific - includes Swagger) if (await hasFlaskRESTX({ installDir })) { + getUI().setDetectedFramework('Flask-RESTX'); return FlaskProjectType.RESTX; } // Check for flask-smorest (OpenAPI-first) if (await hasFlaskSmorest({ installDir })) { + getUI().setDetectedFramework('flask-smorest'); return FlaskProjectType.SMOREST; } // Check for Flask-RESTful if (await hasFlaskRESTful({ installDir })) { + getUI().setDetectedFramework('Flask-RESTful'); return FlaskProjectType.RESTFUL; } // Check for Blueprints (large app structure) if (await hasBlueprints({ installDir })) { + getUI().setDetectedFramework('Flask with Blueprints'); return FlaskProjectType.BLUEPRINT; } // Default to standard Flask + getUI().setDetectedFramework('Flask'); return FlaskProjectType.STANDARD; } diff --git a/src/programs/frameworks/laravel/laravel-wizard-agent.ts b/src/programs/frameworks/laravel/laravel-wizard-agent.ts index eb07a12ba..cc0849617 100644 --- a/src/programs/frameworks/laravel/laravel-wizard-agent.ts +++ b/src/programs/frameworks/laravel/laravel-wizard-agent.ts @@ -43,12 +43,6 @@ export const LARAVEL_AGENT_CONFIG: FrameworkConfig = { laravelStructure, }; }, - getDetectedFrameworkLabel: (context) => - context.projectType === LaravelProjectType.STANDARD - ? 'Laravel' - : context.projectType - ? getLaravelProjectTypeName(context.projectType) - : undefined, }, detection: { diff --git a/src/programs/frameworks/laravel/utils.ts b/src/programs/frameworks/laravel/utils.ts index 76c491249..08f9e6900 100644 --- a/src/programs/frameworks/laravel/utils.ts +++ b/src/programs/frameworks/laravel/utils.ts @@ -1,4 +1,5 @@ import { boundedGlob, readProjectFile } from '@utils/bounded-fs'; +import { getUI } from '@ui'; import type { WizardRunOptions } from '@utils/types'; import { createVersionBucket } from '@utils/semver'; import * as fs from 'node:fs'; @@ -165,13 +166,16 @@ export async function getLaravelProjectType( ): Promise { // Check for SPA/Reactive frameworks (important to detect - affects SDK needs) if (await hasInertia(options)) { + getUI().setDetectedFramework('Laravel with Inertia.js'); return LaravelProjectType.INERTIA; } if (await hasLivewire(options)) { + getUI().setDetectedFramework('Laravel with Livewire'); return LaravelProjectType.LIVEWIRE; } // Default to standard + getUI().setDetectedFramework('Laravel'); return LaravelProjectType.STANDARD; } diff --git a/src/programs/frameworks/nextjs/nextjs-wizard-agent.ts b/src/programs/frameworks/nextjs/nextjs-wizard-agent.ts index 45f706415..83c180672 100644 --- a/src/programs/frameworks/nextjs/nextjs-wizard-agent.ts +++ b/src/programs/frameworks/nextjs/nextjs-wizard-agent.ts @@ -10,6 +10,7 @@ import { type PackageJson, } from '@utils/package-json'; import { tryGetPackageJson } from '@utils/setup-utils'; +import { getUI } from '@ui'; import { getNextJsRouter, getNextJsVersionBucket, @@ -30,16 +31,15 @@ export const NEXTJS_AGENT_CONFIG: FrameworkConfig = { gatherContext: async (options: WizardRunOptions) => { const router = await getNextJsRouter(options); if (router) { + const emoji = + router === NextJsRouter.APP_ROUTER ? '\u{1F4F1}' : '\u{1F4C3}'; + getUI().setDetectedFramework( + `Next.js ${getNextJsRouterName(router)} ${emoji}`, + ); return { router }; } return {}; }, - getDetectedFrameworkLabel: (context) => { - if (!context.router) return undefined; - const emoji = - context.router === NextJsRouter.APP_ROUTER ? '\u{1F4F1}' : '\u{1F4C3}'; - return `Next.js ${getNextJsRouterName(context.router)} ${emoji}`; - }, setup: { questions: [ { diff --git a/src/programs/frameworks/rails/rails-wizard-agent.ts b/src/programs/frameworks/rails/rails-wizard-agent.ts index d624a6f7c..08c8c39a6 100644 --- a/src/programs/frameworks/rails/rails-wizard-agent.ts +++ b/src/programs/frameworks/rails/rails-wizard-agent.ts @@ -29,12 +29,6 @@ export const RAILS_AGENT_CONFIG: FrameworkConfig = { const initializersDir = findInitializersDir(options); return Promise.resolve({ projectType, initializersDir }); }, - getDetectedFrameworkLabel: (context) => - context.projectType === RailsProjectType.API - ? 'Rails API-only' - : context.projectType === RailsProjectType.STANDARD - ? 'Rails' - : undefined, }, detection: { diff --git a/src/programs/frameworks/rails/utils.ts b/src/programs/frameworks/rails/utils.ts index 5b7820d60..ab7993e4d 100644 --- a/src/programs/frameworks/rails/utils.ts +++ b/src/programs/frameworks/rails/utils.ts @@ -1,4 +1,5 @@ import fg from 'fast-glob'; +import { getUI } from '@ui'; import type { WizardRunOptions } from '@utils/types'; import { createVersionBucket } from '@utils/semver'; import * as fs from 'node:fs'; @@ -95,6 +96,7 @@ export function getRailsProjectType( try { const content = fs.readFileSync(appConfigPath, 'utf-8'); if (content.includes('config.api_only = true')) { + getUI().setDetectedFramework('Rails API-only'); return RailsProjectType.API; } } catch { @@ -102,6 +104,7 @@ export function getRailsProjectType( } } + getUI().setDetectedFramework('Rails'); return RailsProjectType.STANDARD; } diff --git a/src/programs/frameworks/react-native/react-native-wizard-agent.ts b/src/programs/frameworks/react-native/react-native-wizard-agent.ts index bf23d6bb0..2ba1b8f85 100644 --- a/src/programs/frameworks/react-native/react-native-wizard-agent.ts +++ b/src/programs/frameworks/react-native/react-native-wizard-agent.ts @@ -30,10 +30,6 @@ export const REACT_NATIVE_AGENT_CONFIG: FrameworkConfig = { const variant = await detectReactNativeVariant(options); return { variant }; }, - getDetectedFrameworkLabel: (context) => - context.variant - ? `${getReactNativeVariantName(context.variant)} 📱` - : undefined, }, detection: { diff --git a/src/programs/frameworks/react-native/utils.ts b/src/programs/frameworks/react-native/utils.ts index e8617ca71..a3601217f 100644 --- a/src/programs/frameworks/react-native/utils.ts +++ b/src/programs/frameworks/react-native/utils.ts @@ -1,6 +1,7 @@ import { createVersionBucket } from '@utils/semver'; import { tryGetPackageJson } from '@utils/setup-utils'; import { hasDeclaredDependency } from '@utils/package-json'; +import { getUI } from '@ui'; import type { WizardRunOptions } from '@utils/types'; export const getReactNativeVersionBucket = createVersionBucket(); @@ -20,8 +21,14 @@ export async function detectReactNativeVariant( const packageJson = await tryGetPackageJson(options); if (packageJson && hasDeclaredDependency('expo', packageJson)) { + getUI().setDetectedFramework( + `${getReactNativeVariantName(ReactNativeVariant.EXPO)} 📱`, + ); return ReactNativeVariant.EXPO; } + getUI().setDetectedFramework( + `${getReactNativeVariantName(ReactNativeVariant.REACT_NATIVE)} 📱`, + ); return ReactNativeVariant.REACT_NATIVE; } diff --git a/src/programs/frameworks/react-router/react-router-wizard-agent.ts b/src/programs/frameworks/react-router/react-router-wizard-agent.ts index 8bfc99447..b4dd9b38a 100644 --- a/src/programs/frameworks/react-router/react-router-wizard-agent.ts +++ b/src/programs/frameworks/react-router/react-router-wizard-agent.ts @@ -10,6 +10,7 @@ import { type PackageJson, } from '@utils/package-json'; import { tryGetPackageJson } from '@utils/setup-utils'; +import { getUI } from '@ui'; import { getReactRouterMode, getReactRouterModeName, @@ -30,14 +31,13 @@ export const REACT_ROUTER_AGENT_CONFIG: FrameworkConfig = { gatherContext: async (options: WizardRunOptions) => { const routerMode = await getReactRouterMode(options); if (routerMode) { + getUI().setDetectedFramework( + `React Router ${getReactRouterModeName(routerMode)}`, + ); return { routerMode }; } return {}; }, - getDetectedFrameworkLabel: (context) => - context.routerMode - ? `React Router ${getReactRouterModeName(context.routerMode)}` - : undefined, }, detection: { diff --git a/src/programs/frameworks/tanstack-router/tanstack-router-wizard-agent.ts b/src/programs/frameworks/tanstack-router/tanstack-router-wizard-agent.ts index 6c41df8a4..d93e337bf 100644 --- a/src/programs/frameworks/tanstack-router/tanstack-router-wizard-agent.ts +++ b/src/programs/frameworks/tanstack-router/tanstack-router-wizard-agent.ts @@ -10,6 +10,7 @@ import { type PackageJson, } from '@utils/package-json'; import { tryGetPackageJson } from '@utils/setup-utils'; +import { getUI } from '@ui'; import { getTanStackRouterMode, getTanStackRouterModeName, @@ -30,14 +31,13 @@ export const TANSTACK_ROUTER_AGENT_CONFIG: FrameworkConfig { const routerMode = await getTanStackRouterMode(options); if (routerMode) { + getUI().setDetectedFramework( + `TanStack Router ${getTanStackRouterModeName(routerMode)}`, + ); return { routerMode }; } return {}; }, - getDetectedFrameworkLabel: (context) => - context.routerMode - ? `TanStack Router ${getTanStackRouterModeName(context.routerMode)}` - : undefined, }, detection: { diff --git a/src/programs/host-capabilities.ts b/src/programs/host-capabilities.ts index 8896060b0..a6264b512 100644 --- a/src/programs/host-capabilities.ts +++ b/src/programs/host-capabilities.ts @@ -1,25 +1,14 @@ -import type { AgentProgress } from '@agent/types'; -import type { AuthProjection } from '@programs/authenticate'; -import type { Integration } from '@shared/constants'; - /** Effects the non-interactive host supplies while a program scopes its project. */ export type ProgramCiHost = { - auth: AuthProjection; log: { info(message: string): void; warn(message: string): void; }; - onProgress(event: AgentProgress): void; }; -/** Live UI effects a legacy program may need after its run definition resolves. */ +/** Live UI effects a program may need after its run definition resolves. */ export type ProgramRunHost = { getFrameworkContext(key: string): unknown; setFrameworkContext(key: string, value: unknown): void; warn(message: string): void; - uploadEnvironmentVariables( - envVars: Record, - integration: Integration, - installDir: string, - ): Promise; }; diff --git a/src/programs/index.ts b/src/programs/index.ts index 69be7232a..e3f438250 100644 --- a/src/programs/index.ts +++ b/src/programs/index.ts @@ -1,31 +1,6 @@ /** Public runtime entry for the programs surface. */ -import { snapshotProgramInput } from './snapshot-program-input'; export type * from './types'; -/** Load gateway minting only when the caller requests model auth. */ -export function createPosthogInferenceAuthProvider( - posthog: import('@shared/api').Credentials, - programId: string, -): import('@agent/types').InferenceAuthProvider { - return { - resolve: async () => { - const { createPosthogInferenceAuthProvider } = await import( - './credentials' - ); - return createPosthogInferenceAuthProvider(posthog, programId).resolve(); - }, - }; -} -/** Keep agent and execution imports out of CLI startup until a program runs. */ -export async function runProgram( - programId: string, - input: import('./run-program').ProgramInput, - options?: import('./run-program').ProgramOptions, -): Promise { - // Copy before the load, so host writes while it loads cannot reach the run. - const snapshot = snapshotProgramInput(input); - const entry = await import('./run-program'); - return entry.runProgram(programId, snapshot, options); -} +export { runProgram } from './run-program'; export { Program, PROGRAM_REGISTRY, @@ -34,9 +9,3 @@ export { getCommandPath, getLaunchablePrograms, } from './program-registry'; -/** Step-based host helpers for the session adapter. */ -export { postAuthGateSteps } from './program-step'; -export { authenticate } from './authenticate'; -export { FRAMEWORK_REGISTRY } from './registry'; -export { getDetectedWarehouseSources } from './warehouse-source/detect'; -export { AUDIT_CHECKS_KEY } from './audit/types'; diff --git a/src/programs/mcp-analytics/index.ts b/src/programs/mcp-analytics/index.ts index c8b4bba50..faa9afd1c 100644 --- a/src/programs/mcp-analytics/index.ts +++ b/src/programs/mcp-analytics/index.ts @@ -1,7 +1,44 @@ +import type { AbortCase } from '@agent/types'; +import { ErrorCodes } from '@shared/errors'; import { createSkillProgram } from '@programs/agent-skill/index'; -import { MCP_ANALYTICS_OPTIONS } from './run.js'; -export { MCP_ANALYTICS_ABORT_CASES } from './run.js'; +const MCP_ANALYTICS_REPORT_FILE = 'posthog-mcp-analytics-report.md'; + +/** + * `[ABORT]` reasons the mcp-analytics skill emits when the project can't be + * instrumented. Kept in sync with the stop conditions in the skill's + * `description.md` (context-mill `context/skills/mcp-analytics`). + */ +export const MCP_ANALYTICS_ABORT_CASES: AbortCase[] = [ + { + match: /^unsupported language for mcp analytics$/i, + errorCode: ErrorCodes.DetectUnsupportedPlatform, + message: 'Unsupported language for MCP analytics', + body: + 'MCP analytics supports TypeScript/JavaScript (`@posthog/mcp`) and Python ' + + '(`posthog.mcp`, shipped inside the `posthog` package). This project ' + + "doesn't look like either, so there's nothing to instrument. " + + 'See https://posthog.com/docs/mcp-analytics for the supported setups.', + }, + { + match: /^no mcp server found$/i, + message: 'No MCP server found', + body: + 'This command instruments an existing MCP server with PostHog analytics, ' + + 'but no MCP server was found in this project. If you just want PostHog ' + + 'product analytics, run `npx @posthog/wizard` instead.', + }, + { + match: /^could not locate the server entry point$/i, + message: 'Could not locate the MCP server entry point', + body: + "This project has MCP signals, but the agent couldn't find where the " + + "server is constructed or requests are dispatched, so there's nowhere " + + 'safe to add instrumentation. See https://posthog.com/docs/mcp-analytics ' + + 'for the supported server styles, or point the wizard at the package ' + + "that defines the server if it's in a monorepo subdirectory.", + }, +]; /** * `wizard mcp-analytics` — flat skill command. @@ -17,4 +54,23 @@ export { MCP_ANALYTICS_ABORT_CASES } from './run.js'; * 'mcp-analytics'` from context-mill — a deliberate breaking change, done then, * not pre-emptively. */ -export const mcpAnalyticsConfig = createSkillProgram(MCP_ANALYTICS_OPTIONS); +export const mcpAnalyticsConfig = createSkillProgram({ + skillId: 'mcp-analytics', + command: 'mcp-analytics', + id: 'mcp-analytics', + description: 'Add PostHog MCP Analytics to your MCP server', + integrationLabel: 'mcp-analytics', + customPrompt: + "Instrument this project's MCP server with PostHog MCP analytics. Run the " + + '`mcp-analytics` skill end-to-end: detect the server style, install ' + + '`@posthog/mcp` and `posthog-node`, wrap the server (or use `PostHogMCP` ' + + 'for a custom dispatcher), wire the project API key and host, and verify. ' + + 'Make only additive changes — do not alter tool behavior. The final report ' + + `is written to ./${MCP_ANALYTICS_REPORT_FILE}.`, + successMessage: `MCP analytics configured! View the report at ./${MCP_ANALYTICS_REPORT_FILE}`, + reportFile: MCP_ANALYTICS_REPORT_FILE, + docsUrl: 'https://posthog.com/docs/mcp-analytics', + spinnerMessage: 'Setting up MCP analytics...', + estimatedDurationMinutes: 5, + abortCases: MCP_ANALYTICS_ABORT_CASES, +}); diff --git a/src/programs/mcp-analytics/run.ts b/src/programs/mcp-analytics/run.ts deleted file mode 100644 index 690a31291..000000000 --- a/src/programs/mcp-analytics/run.ts +++ /dev/null @@ -1,62 +0,0 @@ -import type { AbortCase } from '@agent/types'; -import { ErrorCodes } from '@shared/errors'; -import type { SkillProgramOptions } from '@programs/agent-skill/run-definition'; - -const MCP_ANALYTICS_REPORT_FILE = 'posthog-mcp-analytics-report.md'; - -/** - * `[ABORT]` reasons the mcp-analytics skill emits when the project can't be - * instrumented. Kept in sync with the stop conditions in the skill's - * `description.md` (context-mill `context/skills/mcp-analytics`). - */ -export const MCP_ANALYTICS_ABORT_CASES: AbortCase[] = [ - { - match: /^unsupported language for mcp analytics$/i, - errorCode: ErrorCodes.DetectUnsupportedPlatform, - message: 'Unsupported language for MCP analytics', - body: - 'MCP analytics supports TypeScript/JavaScript (`@posthog/mcp`) and Python ' + - '(`posthog.mcp`, shipped inside the `posthog` package). This project ' + - "doesn't look like either, so there's nothing to instrument. " + - 'See https://posthog.com/docs/mcp-analytics for the supported setups.', - }, - { - match: /^no mcp server found$/i, - message: 'No MCP server found', - body: - 'This command instruments an existing MCP server with PostHog analytics, ' + - 'but no MCP server was found in this project. If you just want PostHog ' + - 'product analytics, run `npx @posthog/wizard` instead.', - }, - { - match: /^could not locate the server entry point$/i, - message: 'Could not locate the MCP server entry point', - body: - "This project has MCP signals, but the agent couldn't find where the " + - "server is constructed or requests are dispatched, so there's nowhere " + - 'safe to add instrumentation. See https://posthog.com/docs/mcp-analytics ' + - 'for the supported server styles, or point the wizard at the package ' + - "that defines the server if it's in a monorepo subdirectory.", - }, -]; - -export const MCP_ANALYTICS_OPTIONS: SkillProgramOptions = { - skillId: 'mcp-analytics', - command: 'mcp-analytics', - id: 'mcp-analytics', - description: 'Add PostHog MCP Analytics to your MCP server', - integrationLabel: 'mcp-analytics', - customPrompt: - "Instrument this project's MCP server with PostHog MCP analytics. Run the " + - '`mcp-analytics` skill end-to-end: detect the server style, install ' + - '`@posthog/mcp` and `posthog-node`, wrap the server (or use `PostHogMCP` ' + - 'for a custom dispatcher), wire the project API key and host, and verify. ' + - 'Make only additive changes — do not alter tool behavior. The final report ' + - `is written to ./${MCP_ANALYTICS_REPORT_FILE}.`, - successMessage: `MCP analytics configured! View the report at ./${MCP_ANALYTICS_REPORT_FILE}`, - reportFile: MCP_ANALYTICS_REPORT_FILE, - docsUrl: 'https://posthog.com/docs/mcp-analytics', - spinnerMessage: 'Setting up MCP analytics...', - estimatedDurationMinutes: 5, - abortCases: MCP_ANALYTICS_ABORT_CASES, -}; diff --git a/src/programs/mcp/index.ts b/src/programs/mcp/index.ts index e632b013f..5c0e35ce5 100644 --- a/src/programs/mcp/index.ts +++ b/src/programs/mcp/index.ts @@ -9,7 +9,7 @@ */ import type { ProgramConfig } from '@programs/program-step'; -import { McpOutcome } from '@shared/run-state'; +import { McpOutcome } from '@lib/wizard-session'; export const mcpAddConfig: ProgramConfig = { id: 'mcp-add', diff --git a/src/programs/metrics/index.ts b/src/programs/metrics/index.ts index 7d979fb95..c96612ae3 100644 --- a/src/programs/metrics/index.ts +++ b/src/programs/metrics/index.ts @@ -1,11 +1,13 @@ import type { ProgramConfig, ProgramStep } from '@programs/program-step'; import { AGENT_SKILL_STEPS } from '@programs/agent-skill/index'; -import { METRICS_REPORT_FILE, METRICS_RUN } from './run.js'; +import { getContentBlocks } from '@ui/tui/decks/agent-skill/index'; const METRICS_STEPS: ProgramStep[] = AGENT_SKILL_STEPS.map((step) => step.id === 'intro' ? { ...step, screenId: 'metrics-intro' } : step, ); +const METRICS_REPORT_FILE = 'posthog-metrics-report.md'; + /** * `wizard metrics` — instrument the project with PostHog application metrics * (`posthog.metrics` counters, gauges, and histograms). @@ -27,5 +29,46 @@ export const metricsConfig: ProgramConfig = { agentFlow: 'metrics', steps: METRICS_STEPS, reportFile: METRICS_REPORT_FILE, - run: METRICS_RUN, + getContentBlocks, + run: { + integrationLabel: 'metrics', + // No `skillId`: the agent must load the menu and install the right + // variant itself. The prompt below tells it how. + customPrompt: + () => `Instrument this project with PostHog application metrics. + +This flow has no pre-installed skill — you install the right one yourself: + +1. Call \`load_skill_menu\` with \`category: "metrics"\`. The menu is the + source of truth: one variant per platform. + +2. Pick the variant that matches the project: + - Python tooling (\`pyproject.toml\`, \`requirements.txt\`, \`Pipfile\`) → + \`metrics-python\` (needs \`posthog\` >= 7.23.0) + - \`package.json\` with server-side Node code → \`metrics-nodejs\` + (needs \`posthog-node\` >= 5.43.0) + - \`package.json\` that is browser-only → \`metrics-javascript\` + (needs \`posthog-js\` >= 1.399.0) + - Kubernetes manifests / Helm charts and the user wants cluster-level + scraping → \`metrics-kubernetes\` + - Any other language → \`metrics-other\` (plain OTLP exporter) + A full-stack app (e.g. Next.js) usually wants the server variant — metrics + measure service work, not user actions. Genuinely ambiguous → + \`wizard_ask\` with a multi-choice picker. + +3. Call \`install_skill\` with the picked variant id. Then follow that skill's + \`SKILL.md\` and references end-to-end — it covers where to place metrics + (middleware, background jobs, external calls, business commit sites) and + the low-cardinality attribute rules. + +Make only additive changes — reuse an existing PostHog client by adding the +\`metrics\` config to it rather than constructing a second client, and do not +touch existing identify calls, event capture, or dashboards. The final report +is written to ./${METRICS_REPORT_FILE}.`, + successMessage: `Application metrics configured! View the report at ./${METRICS_REPORT_FILE}`, + reportFile: METRICS_REPORT_FILE, + docsUrl: 'https://posthog.com/docs/metrics', + spinnerMessage: 'Setting up application metrics...', + estimatedDurationMinutes: 5, + }, }; diff --git a/src/programs/metrics/run.ts b/src/programs/metrics/run.ts deleted file mode 100644 index 295cc2455..000000000 --- a/src/programs/metrics/run.ts +++ /dev/null @@ -1,44 +0,0 @@ -import type { ProgramRun } from '@programs/program-run'; - -export const METRICS_REPORT_FILE = 'posthog-metrics-report.md'; - -export const METRICS_RUN: ProgramRun = { - integrationLabel: 'metrics', - // No `skillId`: the agent must load the menu and install the right - // variant itself. The prompt below tells it how. - customPrompt: () => `Instrument this project with PostHog application metrics. - -This flow has no pre-installed skill — you install the right one yourself: - -1. Call \`load_skill_menu\` with \`category: "metrics"\`. The menu is the - source of truth: one variant per platform. - -2. Pick the variant that matches the project: - - Python tooling (\`pyproject.toml\`, \`requirements.txt\`, \`Pipfile\`) → - \`metrics-python\` (needs \`posthog\` >= 7.23.0) - - \`package.json\` with server-side Node code → \`metrics-nodejs\` - (needs \`posthog-node\` >= 5.43.0) - - \`package.json\` that is browser-only → \`metrics-javascript\` - (needs \`posthog-js\` >= 1.399.0) - - Kubernetes manifests / Helm charts and the user wants cluster-level - scraping → \`metrics-kubernetes\` - - Any other language → \`metrics-other\` (plain OTLP exporter) - A full-stack app (e.g. Next.js) usually wants the server variant — metrics - measure service work, not user actions. Genuinely ambiguous → - \`wizard_ask\` with a multi-choice picker. - -3. Call \`install_skill\` with the picked variant id. Then follow that skill's - \`SKILL.md\` and references end-to-end — it covers where to place metrics - (middleware, background jobs, external calls, business commit sites) and - the low-cardinality attribute rules. - -Make only additive changes — reuse an existing PostHog client by adding the -\`metrics\` config to it rather than constructing a second client, and do not -touch existing identify calls, event capture, or dashboards. The final report -is written to ./${METRICS_REPORT_FILE}.`, - successMessage: `Application metrics configured! View the report at ./${METRICS_REPORT_FILE}`, - reportFile: METRICS_REPORT_FILE, - docsUrl: 'https://posthog.com/docs/metrics', - spinnerMessage: 'Setting up application metrics...', - estimatedDurationMinutes: 5, -}; diff --git a/src/programs/migration/index.ts b/src/programs/migration/index.ts index 64751e04d..907160f6d 100644 --- a/src/programs/migration/index.ts +++ b/src/programs/migration/index.ts @@ -1,11 +1,28 @@ import type { ProgramConfig } from '@programs/program-step'; +import type { AbortCase } from '@agent/types'; import { WIZARD_TOOL_NAMES } from '@agent'; import { MIGRATION_PROGRAM } from './steps.js'; -import { - DEFAULT_MIGRATE_SKILL_ID, - MIGRATION_REPORT_FILE, - MIGRATION_RUN, -} from './run.js'; +import { getContentBlocks } from '../../ui/tui/decks/migration/index.js'; + +const MIGRATION_REPORT_FILE = 'migration-report.md'; + +const MIGRATION_ABORT_CASES: AbortCase[] = [ + { + match: /^no source-sdk calls found$/i, + message: 'No source-SDK calls found', + body: + 'The migration needs an existing third-party SDK to migrate from. No ' + + 'calls to the source SDK appear anywhere in this project. If you ' + + "haven't installed PostHog yet, you don't need this command — run " + + '`npx @posthog/wizard@latest` to add PostHog from scratch.', + }, +]; + +// Default skill id when nothing else picks one. The `wizard migrate ` +// subcommands override this via skillCommandFactory using each manifest +// entry's skillId, so this default only kicks in for legacy callers (e.g. +// programmatic uses of migrationConfig directly). +const DEFAULT_MIGRATE_SKILL_ID = 'migrate-statsig'; export const migrationConfig: ProgramConfig = { command: 'migrate', @@ -14,9 +31,26 @@ export const migrationConfig: ProgramConfig = { skillId: DEFAULT_MIGRATE_SKILL_ID, steps: MIGRATION_PROGRAM, reportFile: MIGRATION_REPORT_FILE, + getContentBlocks, allowedTools: ['Agent'], disallowedTools: [WIZARD_TOOL_NAMES.wizardAsk], - run: MIGRATION_RUN, + run: { + skillId: DEFAULT_MIGRATE_SKILL_ID, + integrationLabel: 'migration', + customPrompt: () => + 'Migrate this project from its existing third-party analytics, ' + + 'feature-flag, and observability tools to PostHog. Run the `migrate` ' + + 'skill end-to-end: follow the step chain starting at ' + + 'references/1-presence.md. Only replace existing source-SDK call sites ' + + 'with PostHog equivalents — make zero unrelated changes and no ' + + `net-new instrumentation. The final report is written to ./${MIGRATION_REPORT_FILE}.`, + successMessage: `Migration complete! View the report at ./${MIGRATION_REPORT_FILE}`, + reportFile: MIGRATION_REPORT_FILE, + docsUrl: '', + spinnerMessage: 'Migrating to PostHog...', + estimatedDurationMinutes: 8, + abortCases: MIGRATION_ABORT_CASES, + }, requires: ['posthog-integration'], }; diff --git a/src/programs/migration/run.ts b/src/programs/migration/run.ts deleted file mode 100644 index 25a28f14d..000000000 --- a/src/programs/migration/run.ts +++ /dev/null @@ -1,40 +0,0 @@ -import type { AbortCase } from '@agent/types'; -import type { ProgramRun } from '@programs/program-run'; - -export const MIGRATION_REPORT_FILE = 'migration-report.md'; - -const MIGRATION_ABORT_CASES: AbortCase[] = [ - { - match: /^no source-sdk calls found$/i, - message: 'No source-SDK calls found', - body: - 'The migration needs an existing third-party SDK to migrate from. No ' + - 'calls to the source SDK appear anywhere in this project. If you ' + - "haven't installed PostHog yet, you don't need this command — run " + - '`npx @posthog/wizard@latest` to add PostHog from scratch.', - }, -]; - -// Default skill id when nothing else picks one. The `wizard migrate ` -// subcommands override this via skillCommandFactory using each manifest -// entry's skillId, so this default only kicks in for legacy callers (e.g. -// programmatic uses of migrationConfig directly). -export const DEFAULT_MIGRATE_SKILL_ID = 'migrate-statsig'; - -export const MIGRATION_RUN: ProgramRun = { - skillId: DEFAULT_MIGRATE_SKILL_ID, - integrationLabel: 'migration', - customPrompt: () => - 'Migrate this project from its existing third-party analytics, ' + - 'feature-flag, and observability tools to PostHog. Run the `migrate` ' + - 'skill end-to-end: follow the step chain starting at ' + - 'references/1-presence.md. Only replace existing source-SDK call sites ' + - 'with PostHog equivalents — make zero unrelated changes and no ' + - `net-new instrumentation. The final report is written to ./${MIGRATION_REPORT_FILE}.`, - successMessage: `Migration complete! View the report at ./${MIGRATION_REPORT_FILE}`, - reportFile: MIGRATION_REPORT_FILE, - docsUrl: '', - spinnerMessage: 'Migrating to PostHog...', - estimatedDurationMinutes: 8, - abortCases: MIGRATION_ABORT_CASES, -}; diff --git a/src/programs/migration/steps.ts b/src/programs/migration/steps.ts index e317f2da9..cb41d5bc0 100644 --- a/src/programs/migration/steps.ts +++ b/src/programs/migration/steps.ts @@ -1,5 +1,5 @@ import type { ProgramStep } from '@programs/program-step'; -import { RunPhase } from '@shared/run-state'; +import { RunPhase } from '@lib/wizard-session'; import { HEALTH_CHECK_STEP } from '@programs/shared/health-check-step'; export const MIGRATION_PROGRAM: ProgramStep[] = [ diff --git a/src/programs/posthog-integration/__tests__/detect.test.ts b/src/programs/posthog-integration/__tests__/detect.test.ts index fd77283bd..ec2179504 100644 --- a/src/programs/posthog-integration/__tests__/detect.test.ts +++ b/src/programs/posthog-integration/__tests__/detect.test.ts @@ -3,7 +3,6 @@ import * as fs from 'fs'; import * as os from 'os'; import * as path from 'path'; -import { testProgramRunHost } from '../../../../test/program-host'; vi.mock('@utils/analytics', () => ({ analytics: { @@ -45,6 +44,7 @@ import { type WizardSession, } from '@lib/wizard-session'; import type { DetectedSource } from '@programs/warehouse-sources/types'; +import { testProgramRunHost } from '../../../../test/program-host'; function makeTmpDir(): string { return fs.mkdtempSync(path.join(os.tmpdir(), 'warehouse-reporting-')); @@ -113,29 +113,6 @@ describe('detection always runs, independent of consent', () => { }); }); -describe('framework variant label projection', () => { - it('sends the gathered Next.js router label through the program context', async () => { - const installDir = makeTmpDir(); - try { - fs.writeFileSync( - path.join(installDir, 'package.json'), - JSON.stringify({ dependencies: { next: '^15.0.0' } }), - ); - fs.mkdirSync(path.join(installDir, 'app')); - fs.writeFileSync(path.join(installDir, 'app/layout.tsx'), 'export {};'); - - const ctx = makeCtx(buildSession({ installDir })); - await detectPostHogIntegration(ctx); - - expect(ctx.setDetectedFramework).toHaveBeenCalledWith( - 'Next.js app router 📱', - ); - } finally { - cleanup(installDir); - } - }); -}); - describe('a scan failure is distinguishable from a clean zero-source scan', () => { let tmpDir: string; diff --git a/src/programs/posthog-integration/__tests__/index.test.ts b/src/programs/posthog-integration/__tests__/index.test.ts index 70c238611..76b7a2836 100644 --- a/src/programs/posthog-integration/__tests__/index.test.ts +++ b/src/programs/posthog-integration/__tests__/index.test.ts @@ -7,12 +7,10 @@ */ import { posthogIntegrationConfig } from '@programs/posthog-integration/index'; -import type { ProgramRunHost } from '@programs/host-capabilities'; import { buildSession, type WizardSession } from '@lib/wizard-session'; import { analytics } from '@utils/analytics'; import { isUsingTypeScript } from '@utils/setup-utils'; -import { HostResolution } from '@shared/host-resolution'; -import { Integration } from '@shared/constants'; +import { testProgramRunHost } from '../../../../test/program-host'; vi.mock('@utils/analytics', () => ({ analytics: { @@ -50,19 +48,10 @@ function sessionWithFramework(): WizardSession { return s; } -function runHost(): ProgramRunHost { - return { - getFrameworkContext: vi.fn(), - setFrameworkContext: vi.fn(), - warn: vi.fn(), - uploadEnvironmentVariables: vi.fn().mockResolvedValue(['POSTHOG_KEY']), - }; -} - -async function resolveRun(session: WizardSession, host = runHost()) { +async function resolveRun(session: WizardSession) { const { run } = posthogIntegrationConfig; if (typeof run !== 'function') throw new Error('expected a run function'); - return run(session, host); + return run(session, testProgramRunHost(session)); } describe('posthog-integration run() — typescript tag', () => { @@ -89,59 +78,4 @@ describe('posthog-integration run() — typescript tag', () => { expect(session.typescript).toBe(false); expect(analytics.setTag).toHaveBeenCalledWith('typescript', false); }); - - it('routes missing package warnings through the run host', async () => { - (isUsingTypeScript as Mock).mockReturnValue(false); - const session = sessionWithFramework(); - if (!session.frameworkConfig) throw new Error('missing framework config'); - session.frameworkConfig = { - ...session.frameworkConfig, - detection: { - ...session.frameworkConfig.detection, - usesPackageJson: true, - }, - }; - const host = runHost(); - - await resolveRun(session, host); - - expect(host.warn).toHaveBeenCalledWith( - 'Could not find package.json. Continuing anyway — the agent will handle it.', - ); - }); - - it('routes hosting uploads through the run host with the project directory', async () => { - (isUsingTypeScript as Mock).mockReturnValue(false); - const session = sessionWithFramework(); - if (!session.frameworkConfig) throw new Error('missing framework config'); - session.frameworkConfig = { - ...session.frameworkConfig, - metadata: { - ...session.frameworkConfig.metadata, - integration: Integration.nextjs, - }, - environment: { - ...session.frameworkConfig.environment, - uploadToHosting: true, - }, - }; - const host = runHost(); - const run = await resolveRun(session, host); - - await run.postRun?.( - { signup: false, dashboardUrl: null, notebookUrl: null }, - { - accessToken: 'token', - projectApiKey: 'phc_test', - projectId: 123, - host: HostResolution.fromApiHost('https://us.posthog.com'), - }, - ); - - expect(host.uploadEnvironmentVariables).toHaveBeenCalledWith( - { POSTHOG_KEY: 'phc_test' }, - Integration.nextjs, - '/tmp/app', - ); - }); }); diff --git a/src/programs/posthog-integration/ai-sdk-stamp.ts b/src/programs/posthog-integration/ai-sdk-stamp.ts deleted file mode 100644 index 2c0c6e5f1..000000000 --- a/src/programs/posthog-integration/ai-sdk-stamp.ts +++ /dev/null @@ -1,37 +0,0 @@ -/** The organization's wizard_ai_sdk_detected stamp, computed from explicit evidence. */ - -import type { ApiUser } from '@shared/api'; -import { DiscoveredFeature } from '@shared/scan-consent'; -import { analytics } from '@utils/analytics'; -import { AI_SOURCE_KINDS } from '@programs/warehouse-sources/registry'; -import type { DetectedSource } from '@programs/warehouse-sources/types'; - -export type AiSdkStampEvidence = { - apiUser: Pick | null; - discoveredFeatures: readonly DiscoveredFeature[]; - warehouseSources: readonly DetectedSource[]; - /** Scan consent was granted, so local detection results may be reported. */ - mayReportScanResults: boolean; -}; - -function hasAiSdkEvidence(evidence: AiSdkStampEvidence): boolean { - return ( - evidence.warehouseSources.some((s) => AI_SOURCE_KINDS.has(s.kind)) || - evidence.discoveredFeatures.includes(DiscoveredFeature.LLM) - ); -} - -/** - * Boolean only, on the org, never the list of kinds or any non-AI tool: a - * decline must not leak even the shape of what local detection saw. - */ -export function stampAiSdkDetected(evidence: AiSdkStampEvidence): void { - if (!evidence.mayReportScanResults) return; - const organizationId = evidence.apiUser?.organization?.id; - if (!organizationId) return; - if (!hasAiSdkEvidence(evidence)) return; - - analytics.groupIdentify('organization', organizationId, { - wizard_ai_sdk_detected: true, - }); -} diff --git a/src/programs/posthog-integration/detect.ts b/src/programs/posthog-integration/detect.ts index 835a3707d..a7845ad44 100644 --- a/src/programs/posthog-integration/detect.ts +++ b/src/programs/posthog-integration/detect.ts @@ -11,10 +11,11 @@ import type { ProgramReadyContext } from '@programs/program-step'; import { - type DiscoveredFeature, + DiscoveredFeature, mayReportScanResults, ScanConsent, -} from '@shared/scan-consent'; + type WizardSession, +} from '@lib/wizard-session'; import type { ApiUser } from '@shared/api'; import { FRAMEWORK_REGISTRY } from '@programs/registry'; import { @@ -25,12 +26,13 @@ import { } from '@programs/detection/index'; import { analytics } from '@utils/analytics'; import { detectWarehouseSources } from '@programs/warehouse-sources/detect'; +import { AI_SOURCE_KINDS } from '@programs/warehouse-sources/registry'; +import type { DetectedSource } from '@programs/warehouse-sources/types'; import { DETECTED_WAREHOUSE_SOURCES_KEY, getDetectedWarehouseSources, } from '@programs/warehouse-source/detect'; import { findPackageJsons } from '@programs/shared/package-scanning'; -import { stampAiSdkDetected } from '@programs/posthog-integration/ai-sdk-stamp'; export async function detectPostHogIntegration( ctx: ProgramReadyContext, @@ -67,10 +69,7 @@ export async function detectPostHogIntegration( // pre-copy object and the live session would never see it. ctx.setSkillId(detectedIntegration); - const detectedLabel = config.metadata.getDetectedFrameworkLabel?.(context); - if (detectedLabel) { - ctx.setDetectedFramework(detectedLabel); - } else if (!session.detectedFrameworkLabel) { + if (!session.detectedFrameworkLabel) { ctx.setDetectedFramework(config.metadata.name); } @@ -107,20 +106,6 @@ export async function detectPostHogIntegration( const WAREHOUSE_SCAN_STATE_KEY = 'warehouseScanState'; type WarehouseScanState = 'ok' | 'failed'; -type AiSdkDetectionState = { - apiUser: Pick | null; - discoveredFeatures: DiscoveredFeature[]; - frameworkContext: Record; - scanConsent: ScanConsent; - aiSdkStampReported: boolean; -}; - -type WarehouseReportState = { - frameworkContext: Record; - scanConsent: ScanConsent; - warehouseSourcesReported: boolean; -}; - /** * Scan for data warehouse source signals (Postgres, Stripe, Hubspot, …) and, * when found, stash them for a `wizard warehouse` suggestion on the outro. @@ -160,6 +145,37 @@ function detectWarehouseSourcesForSuggestion( } } +/** What the org stamp reads, with no session. */ +export type AiSdkStampEvidence = { + apiUser: Pick | null; + discoveredFeatures: readonly DiscoveredFeature[]; + warehouseSources: readonly DetectedSource[]; + /** Scan consent was granted, so local detection results may be reported. */ + mayReportScanResults: boolean; +}; + +function hasAiSdkEvidence(evidence: AiSdkStampEvidence): boolean { + return ( + evidence.warehouseSources.some((s) => AI_SOURCE_KINDS.has(s.kind)) || + evidence.discoveredFeatures.includes(DiscoveredFeature.LLM) + ); +} + +/** + * Boolean only, on the org, never the list of kinds or any non-AI tool: a + * decline must not leak even the shape of what local detection saw. + */ +export function stampAiSdkDetected(evidence: AiSdkStampEvidence): void { + if (!evidence.mayReportScanResults) return; + const organizationId = evidence.apiUser?.organization?.id; + if (!organizationId) return; + if (!hasAiSdkEvidence(evidence)) return; + + analytics.groupIdentify('organization', organizationId, { + wizard_ai_sdk_detected: true, + }); +} + /** * Fires the org stamp once per session, right after `authenticate()` succeeds * — never from the consent path, since consent on the intro screen resolves @@ -171,7 +187,7 @@ function detectWarehouseSourcesForSuggestion( * anyway) and this only ever runs from the later, idempotent bootstrap.ts * call — still correctly finding no evidence, since CI skips the detect step. */ -export function maybeStampAiSdkDetected(session: AiSdkDetectionState): void { +export function maybeStampAiSdkDetected(session: WizardSession): void { // Direct mutation, not a store setter: unlike `warehouseSourcesReported` // (latched only from TUI-only consent screens), this runs from // `authenticate()`, which also fires in `--ci` mode, where the session is a @@ -204,7 +220,7 @@ export function maybeStampAiSdkDetected(session: AiSdkDetectionState): void { * without sending. */ export function reportWarehouseSourcesDetected( - session: WarehouseReportState, + session: WizardSession, ): boolean { if (session.warehouseSourcesReported) return false; // 'undecided' means come back later, not no. diff --git a/src/programs/posthog-integration/index.ts b/src/programs/posthog-integration/index.ts index 6a53b00de..ad4c5b935 100644 --- a/src/programs/posthog-integration/index.ts +++ b/src/programs/posthog-integration/index.ts @@ -1,14 +1,9 @@ import type { ProgramConfig, ProgramStep } from '@programs/program-step'; +import { runProgramAgent } from '@programs/run-agent-legacy'; import type { ProgramRun } from '@programs/program-run'; -import type { - ProgramCiHost, - ProgramRunHost, -} from '@programs/host-capabilities'; -import type { FrameworkDetectionState } from '@programs/detection/context'; -import { AgentSignals, OutroKind, WIZARD_TOOL_NAMES } from '@agent'; -import { isAskDisabled } from '@shared/ask-policy'; -import { RunPhase } from '@shared/run-state'; -import { mayReportScanResults } from '@shared/scan-consent'; +import { AgentSignals, shouldDisableAsk, WIZARD_TOOL_NAMES } from '@agent'; +import type { WizardSession } from '@lib/wizard-session'; +import { mayReportScanResults, OutroKind, RunPhase } from '@lib/wizard-session'; import { DEFAULT_PACKAGE_INSTALLATION, SPINNER_MESSAGE, @@ -19,42 +14,29 @@ import { detectFramework, gatherFrameworkContext, } from '@programs/detection/index'; -import { - scopeInstallDirToProject, - type ProjectScopeSession, -} from '@programs/detection/project-scope'; +import { scopeInstallDirToProject } from '@programs/detection/project-scope'; +import type { + ProgramCiHost, + ProgramRunHost, +} from '@programs/host-capabilities'; import { FRAMEWORK_REGISTRY } from '@programs/registry'; import { wizardAbort } from '@utils/wizard-abort'; import { ErrorCodes } from '@shared/errors'; import { WIZARD_DEFAULT_AIO_LOGS_FLAG_KEY, WIZARD_INTERACTION_EVENT_NAME, - type AdditionalFeature, - type Integration, } from '@shared/constants'; import { requestDeepLink } from '@utils/provisioning'; import { openTrackedLink, withUtm } from '@utils/links'; import type { HostResolution } from '@shared/host-resolution'; import { getDetectedWarehouseSources } from '@programs/warehouse-source/detect'; import { POSTHOG_INTEGRATION_PROGRAM } from './steps.js'; +import { getContentBlocks } from '../../ui/tui/decks/posthog-integration/index.js'; import { buildCodingAgentPrompt } from './handoff.js'; import { EVENT_PLAN_FILE } from './constants.js'; const DASHBOARD_DEEP_LINK_KEY = 'dashboardDeepLink'; -type IntegrationCiSession = ProjectScopeSession & - FrameworkDetectionState & { - integration: Integration | null; - }; - -type IntegrationRunSession = Pick< - FrameworkDetectionState, - 'installDir' | 'frameworkConfig' | 'frameworkContext' -> & { - typescript: boolean; - additionalFeatureQueue: AdditionalFeature[]; -}; - const WAREHOUSE_SOURCES_DOCS_URL = 'https://posthog.com/docs/data-warehouse/sources'; @@ -62,11 +44,11 @@ const WAREHOUSE_SOURCES_DOCS_URL = const WAREHOUSE_SEED_TASK_TYPE = 'warehouse'; function resolveContinueUrl( - signup: boolean, + sess: WizardSession, host: HostResolution, deepLink: unknown, ): string | undefined { - if (!signup) return undefined; + if (!sess.signup) return undefined; if (typeof deepLink === 'string' && deepLink) return deepLink; return withUtm(`${host.appHost}/products?source=wizard`, 'outro-continue'); } @@ -138,7 +120,7 @@ function warehouseSourceUrl( * past that is still unconnected and still belongs here. */ function buildWarehouseNextSteps( - sess: Pick, + sess: WizardSession, host: HostResolution, projectId: number | string, completedSeededTypes: readonly string[], @@ -173,9 +155,7 @@ function buildWarehouseNextSteps( * because it is a note in a report: the outro `nextSteps` bullet carries the * same information deterministically, so nothing is lost if the agent drops it. */ -function warehouseReportInstruction( - sess: Pick, -): string { +function warehouseReportInstruction(sess: WizardSession): string { const sources = getDetectedWarehouseSources(sess); if (sources.length === 0) return ''; @@ -202,7 +182,7 @@ function warehouseReportInstruction( * links by {@link buildWarehouseNextSteps}. */ const warehouseSeedTasks: NonNullable = (sess) => { - if (isAskDisabled(sess)) return []; + if (shouldDisableAsk(sess)) return []; const sources = getDetectedWarehouseSources(sess); if (sources.length === 0) return []; @@ -267,6 +247,7 @@ export const posthogIntegrationConfig: ProgramConfig = { agentFlow: 'integration-v2', eventPlanFile: EVENT_PLAN_FILE, steps: POSTHOG_INTEGRATION_PROGRAM, + getContentBlocks, // Basic integration runs without structured user input; drop wizard_ask // so the model can't pop modal prompts mid-run. The runner forwards this // list to the general-purpose subagent as well, so dispatched subagents @@ -285,7 +266,7 @@ export const posthogIntegrationConfig: ProgramConfig = { // CI-mode prerequisite work: the headless equivalent of the detect step's // onReady hook. Auto-detect the framework, then gather context. ciPreRun: async ( - session: IntegrationCiSession, + session: WizardSession, host: ProgramCiHost, ): Promise => { await scopeInstallDirToProject(session, host); @@ -312,9 +293,6 @@ export const posthogIntegrationConfig: ProgramConfig = { benchmark: session.benchmark, yaraReport: session.yaraReport, }); - const detectedLabel = - frameworkConfig.metadata.getDetectedFrameworkLabel?.(context); - if (detectedLabel) session.detectedFrameworkLabel = detectedLabel; for (const [key, value] of Object.entries(context)) { if (!(key in session.frameworkContext)) { session.frameworkContext[key] = value; @@ -323,7 +301,7 @@ export const posthogIntegrationConfig: ProgramConfig = { }, run: async ( - session: IntegrationRunSession, + session: WizardSession, host: ProgramRunHost, ): Promise => { const config = session.frameworkConfig!; @@ -464,10 +442,15 @@ ${warehouseReportInstruction(session)} credentials.host.apiHost, ); if (config.environment.uploadToHosting) { - const uploadedEnvVars = await host.uploadEnvironmentVariables( + const { uploadEnvironmentVariablesStep } = await import( + '@steps/index' + ); + const uploadedEnvVars = await uploadEnvironmentVariablesStep( envVars, - config.metadata.integration, - session.installDir, + { + integration: config.metadata.integration, + session: sess, + }, ); if (uploadedEnvVars.length > 0) { analytics.capture(WIZARD_INTERACTION_EVENT_NAME, { @@ -486,7 +469,7 @@ ${warehouseReportInstruction(session)} ); if (deepLink) { const taggedDeepLink = withUtm(deepLink, 'dashboard-deeplink'); - session.frameworkContext[DASHBOARD_DEEP_LINK_KEY] = taggedDeepLink; + sess.frameworkContext[DASHBOARD_DEEP_LINK_KEY] = taggedDeepLink; openTrackedLink(taggedDeepLink, 'dashboard-deeplink', { auto: true, }); @@ -494,9 +477,9 @@ ${warehouseReportInstruction(session)} } }, - buildOutroNextSteps: (_context, credentials, completedSeededTypes) => + buildOutroNextSteps: (sess, credentials, completedSeededTypes) => buildWarehouseNextSteps( - session, + sess, credentials.host, credentials.projectId, completedSeededTypes, @@ -507,9 +490,9 @@ ${warehouseReportInstruction(session)} credentials.projectApiKey, credentials.host.apiHost, ); - const deepLink = session.frameworkContext[DASHBOARD_DEEP_LINK_KEY]; + const deepLink = sess.frameworkContext[DASHBOARD_DEEP_LINK_KEY]; const continueUrl = resolveContinueUrl( - sess.signup, + sess, credentials.host, deepLink, ); @@ -530,7 +513,7 @@ ${warehouseReportInstruction(session)} // The linear sequence seeds no tasks, so nothing here was connected // during the run. `buildOutroNextSteps` carries the orchestrated case. nextSteps: buildWarehouseNextSteps( - session, + sess, credentials.host, credentials.projectId, [], @@ -560,8 +543,10 @@ export const integrationRunStep: ProgramStep = { id: 'run', label: 'Integration', screenId: 'run', - // The host runs this child without its terminal outro or analytics shutdown. - runProgramId: 'posthog-integration', + // composed: runs inside the host program (self-driving), so skip the + // integration's terminal outro + analytics shutdown of the shared client. + run: (session) => + runProgramAgent(posthogIntegrationConfig, session, { composed: true }), isComplete: (session) => session.runPhase === RunPhase.Completed || session.runPhase === RunPhase.Error, diff --git a/src/programs/posthog-integration/steps.ts b/src/programs/posthog-integration/steps.ts index 7a2c422c6..3c4e9dc5b 100644 --- a/src/programs/posthog-integration/steps.ts +++ b/src/programs/posthog-integration/steps.ts @@ -7,15 +7,12 @@ */ import type { ProgramStep } from '@programs/program-step'; -import type { FrameworkConfig } from '@programs/framework-config'; -import { RunPhase } from '@shared/run-state'; +import type { WizardSession } from '@lib/wizard-session'; +import { RunPhase } from '@lib/wizard-session'; import { HEALTH_CHECK_STEP } from '@programs/shared/health-check-step'; import { detectPostHogIntegration } from './detect.js'; -function needsSetup(session: { - frameworkConfig: FrameworkConfig | null; - frameworkContext: Record; -}): boolean { +function needsSetup(session: WizardSession): boolean { const config = session.frameworkConfig; if (!config?.metadata.setup?.questions) return false; diff --git a/src/programs/program-file-watchers.ts b/src/programs/program-file-watchers.ts deleted file mode 100644 index 9d1b8488b..000000000 --- a/src/programs/program-file-watchers.ts +++ /dev/null @@ -1,62 +0,0 @@ -import path from 'node:path'; -import type { FileWatcherHandle } from '@shared/file-watcher'; -import { AUDIT_CHECKS_KEY } from './audit/types.js'; -import { seedAuditLedger } from './audit/seed.js'; -import { watchAuditLedger } from './audit/watch-ledger.js'; -import { ProgramEventPlanWatcher } from './posthog-integration/watch-event-plan.js'; -import type { ProgramStore } from './program-store.js'; -import type { ProgramSettings } from './run-program.js'; - -export type ProgramFileWatchers = { - seedAuditLedger(): void; - refresh(): void; - stop(): void; -}; - -/** Own the files emitted by this invocation until its agent run settles. */ -export function startProgramFileWatchers( - program: Pick< - ProgramSettings, - 'auditLedgerFile' | 'auditSeedChecks' | 'eventPlanFile' - >, - installDir: string, - store: ProgramStore, -): ProgramFileWatchers { - const ledger: FileWatcherHandle | null = program.auditLedgerFile - ? watchAuditLedger(installDir, program.auditLedgerFile, (checks) => - store.setFrameworkContext(AUDIT_CHECKS_KEY, checks), - ) - : null; - const eventPlan: ProgramEventPlanWatcher | null = program.eventPlanFile - ? new ProgramEventPlanWatcher( - path.join(installDir, program.eventPlanFile), - (events) => store.setEventPlan(events), - ) - : null; - - try { - eventPlan?.start(); - } catch (error) { - ledger?.stop(); - throw error; - } - - return { - seedAuditLedger() { - if (!program.auditSeedChecks) return; - // The watcher already took its ignore-initial snapshot. This write is - // part of this invocation and must be visible before the agent starts. - seedAuditLedger(installDir, [...program.auditSeedChecks]); - store.setFrameworkContext(AUDIT_CHECKS_KEY, program.auditSeedChecks); - ledger?.refresh(); - }, - refresh() { - ledger?.refresh(); - eventPlan?.refresh(); - }, - stop() { - ledger?.stop(); - eventPlan?.stop(); - }, - }; -} diff --git a/src/programs/program-registry.ts b/src/programs/program-registry.ts index 105e89e37..be63b4b8a 100644 --- a/src/programs/program-registry.ts +++ b/src/programs/program-registry.ts @@ -24,6 +24,7 @@ import { errorTrackingUploadSourceMapsConfig } from './error-tracking-upload-sou import { errorTrackingConfig } from './error-tracking/index.js'; import { selfDrivingConfig } from './self-driving/index.js'; import { AGENT_SKILL_STEPS } from './agent-skill/index.js'; +import { getContentBlocks as agentSkillContentBlocks } from '../ui/tui/decks/agent-skill/index.js'; import { mcpAddConfig, mcpRemoveConfig, @@ -49,6 +50,7 @@ export const agentSkillConfig: ProgramConfig = { id: 'agent-skill', description: 'Run an arbitrary context-mill skill', steps: AGENT_SKILL_STEPS, + getContentBlocks: agentSkillContentBlocks, allowedTools: ['Agent'], run: (session) => { const skillId = session.skillId ?? 'agent-skill'; diff --git a/src/programs/program-run.ts b/src/programs/program-run.ts index 442f49a18..a3d9a15be 100644 --- a/src/programs/program-run.ts +++ b/src/programs/program-run.ts @@ -1,30 +1,21 @@ /** * A program's run definition: the agent's `AgentRunDefinition` plus the - * completion hooks that read the program's completion data. The agent never calls these — - * `src/lib/runners/run-program-agent.ts` binds them to the run's credentials and hands the + * completion hooks that read the session. The agent never calls these — + * `run-agent-legacy.ts` binds them to the run's credentials and hands the * agent `RunConfig.hooks`. */ -import type { AgentRunDefinition, OutroData } from '@agent/types'; -import type { Credentials } from '@shared/api'; - -export type ProgramCompletionContext = Readonly<{ - signup: boolean; - dashboardUrl: string | null; - notebookUrl: string | null; -}>; +import type { AgentRunDefinition } from '@agent/types'; +import type { Credentials, WizardSession } from '@lib/wizard-session'; export interface ProgramRun extends AgentRunDefinition { /** Runs after agent completes, before outro (e.g. env var upload). */ - postRun?: ( - context: ProgramCompletionContext, - credentials: Credentials, - ) => Promise; + postRun?: (session: WizardSession, credentials: Credentials) => Promise; /** Custom outro data. Omit for default built from successMessage/reportFile/docsUrl. */ buildOutroData?: ( - context: ProgramCompletionContext, + session: WizardSession, credentials: Credentials, - ) => OutroData | null; + ) => WizardSession['outroData']; /** * Outro bullets for a sequence that composes its own outro data. * @@ -41,7 +32,7 @@ export interface ProgramRun extends AgentRunDefinition { * already did — the sequence stays ignorant of what any type means. */ buildOutroNextSteps?: ( - context: ProgramCompletionContext, + session: WizardSession, credentials: Credentials, completedSeededTypes: readonly string[], ) => { heading: string; items: string[] } | undefined; diff --git a/src/programs/program-step.ts b/src/programs/program-step.ts index 10bcec2bd..016524c8f 100644 --- a/src/programs/program-step.ts +++ b/src/programs/program-step.ts @@ -6,8 +6,10 @@ import type { import type { WizardReadinessResult } from '@shared/health-checks/readiness'; import type { ProgramRun } from '@programs/program-run'; import type { Integration } from '@shared/constants'; -import type { AuditCheck } from '@shared/audit-ledger'; import type { FrameworkConfig } from '@programs/framework-config'; +import type { ContentBlock } from '@ui/tui/primitives/index'; +import type { WizardStore } from '@ui/tui/store'; +import type { Tip } from '@ui/tui/components/TipsCard'; // Type-only — erased at compile time, so no runtime cycle with the // registry that imports `ProgramConfig` back from this module. import type { ProgramId } from './program-registry.js'; @@ -75,10 +77,13 @@ export interface ProgramStep { screenId?: string; /** - * For a composed run step (`screenId: 'run'`): identifies the child program - * whose agent the host runs. Omit to run this program's own agent. + * For a run step (`screenId: 'run'`): runs this step's own agent. A program + * exports a self-contained run step and another imports it into its step list + * — e.g. posthog-integration exports a run step that runs its agent, and + * self-driving imports it before its own run step. Omit to run the host + * program's own agent (`config.run`). */ - runProgramId?: ProgramId; + run?: (session: WizardSession) => Promise; /** * For a run step: prepare a derived session before its agent runs — e.g. @@ -290,14 +295,28 @@ export interface ProgramConfig { eventPlanFile?: string; /** Audit ledger to mirror into the session, relative to `installDir`. */ auditLedgerFile?: string; - /** Ledger rows written before the agent starts, so the run screen renders before its first update. */ - auditSeedChecks?: readonly AuditCheck[]; /** * Channel the task stream publishes this run under, when it differs from the * program id. A family leaf runs on the generic skill program, so without * this every `wizard audit ` would report as `agent-skill`. */ streamWorkflowId?: string; + /** + * LearnCard deck rendered in the shared `RunScreen` while the agent + * runs. Lives at `/content/index.tsx` by convention. + * Programs that ship a custom RunScreen variant (audit) or skip the + * run step (posthog-doctor) leave this unset. + */ + getContentBlocks?: (store?: WizardStore) => ContentBlock[]; + /** + * Tips shown in the run screen's right pane (the `Tips` sidebar) once + * the LearnCard finishes. Lets a program supply its own explainer copy + * (e.g. self-driving explaining what signal sources and scouts are) + * instead of the generic onboarding deck. Unset → `RunScreen` falls back + * to `DEFAULT_TIPS`, so every other program is unaffected. Lives at + * `/content/tips.ts` by convention. + */ + getTips?: (store?: WizardStore) => Tip[]; /** * Subcommand-specific CLI options. Spread into yargs `.options(...)` when the * program's subcommand is registered. Program-specific knowledge stays in diff --git a/src/programs/program-store.ts b/src/programs/program-store.ts index 6b4dfb542..eea4544bf 100644 --- a/src/programs/program-store.ts +++ b/src/programs/program-store.ts @@ -4,7 +4,6 @@ import type { RunResult, } from '../agent/types.js'; import type { ApiProject, ApiUser, Credentials } from '../shared/api.js'; -import type { PlannedEvent } from './posthog-integration/watch-event-plan.js'; /** One agent run's progress event, attributed to its run. */ export type ProgramRunProgress = { @@ -35,7 +34,6 @@ export type ProgramInvocationData = { apiProject: ApiProject | null; apiUser: ApiUser | null; detection: { frameworkContext: Record }; - eventPlan: PlannedEvent[]; /** The route of the agent run; null until it resolves. */ binding: ResolvedBinding | null; /** Latched once the organization's AI SDK stamp was considered for this login. */ @@ -78,7 +76,6 @@ export class ProgramStore { apiProject: null, apiUser: null, detection: { frameworkContext: {} }, - eventPlan: [], binding: null, aiSdkStampReported: options.aiSdkStampReported ?? false, }; @@ -100,11 +97,6 @@ export class ProgramStore { this.emitData(); } - setEventPlan(events: PlannedEvent[]): void { - this.data.eventPlan = structuredClone(events); - this.emitData(); - } - setBinding(binding: ResolvedBinding): void { this.data.binding = structuredClone(binding); this.emitData(); diff --git a/src/programs/replay-vision/index.ts b/src/programs/replay-vision/index.ts index 87ce0dbdb..9003c0363 100644 --- a/src/programs/replay-vision/index.ts +++ b/src/programs/replay-vision/index.ts @@ -1,13 +1,10 @@ +import type { AbortCase } from '@agent/types'; import { Integration } from '@shared/constants'; import { detectFramework, gatherFrameworkContext, } from '@programs/detection/index'; -import { - scopeInstallDirToProject, - type ProjectScopeSession, -} from '@programs/detection/project-scope'; -import type { FrameworkDetectionState } from '@programs/detection/context'; +import { scopeInstallDirToProject } from '@programs/detection/project-scope'; import type { ProgramCiHost } from '@programs/host-capabilities'; import { FRAMEWORK_REGISTRY } from '@programs/registry'; import { createSkillProgram } from '@programs/agent-skill/index'; @@ -18,12 +15,12 @@ import type { ProgramReadyContext, ProgramStep, } from '@programs/program-step'; +import type { WizardSession } from '@lib/wizard-session'; import { analytics } from '@utils/analytics'; import { wizardAbort } from '@utils/wizard-abort'; import { ErrorCodes } from '@shared/errors'; -import { REPLAY_VISION_OPTIONS } from './run.js'; -export { REPLAY_VISION_ABORT_CASES } from './run.js'; +const REPLAY_VISION_REPORT_FILE = 'posthog-replay-vision-report.md'; /** * The platforms session replay can actually record on. Replay vision watches @@ -58,12 +55,6 @@ export const REPLAY_VISION_SUPPORTED: ReadonlySet = new Set([ Integration.flutter, ]); -type ReplayVisionCiSession = ProjectScopeSession & - FrameworkDetectionState & { - integration: Integration | null; - skillId: string | null; - }; - async function abortUnsupportedPlatform( integration: Integration, ): Promise { @@ -87,6 +78,22 @@ async function abortUnsupportedPlatform( }); } +/** + * `[ABORT]` reasons the replay-vision skill emits when the run can't proceed. + * Kept in sync with the stop conditions in the skill's `description.md` + * (context-mill `context/skills/replay-vision`). + */ +export const REPLAY_VISION_ABORT_CASES: AbortCase[] = [ + { + match: /^replay vision not available for this project$/i, + message: 'Replay vision is not available for this project', + body: + 'Every Replay vision scanner endpoint reported that the feature is not ' + + 'available here yet. Session replay setup done so far is kept. See ' + + 'https://posthog.com/docs/replay-vision for availability.', + }, +]; + /** * Framework detection ahead of the run, exactly like the default integration * program. The orchestrator requires it: `session.skillId` must hold the @@ -113,7 +120,32 @@ const DETECT_STEP: ProgramStep = { }, }; -const base = createSkillProgram(REPLAY_VISION_OPTIONS); +const base = createSkillProgram({ + // The menu ids this skill `-`, and context-mill's + // `replay-vision/config.yaml` declares a single variant, `setup`. The bare + // `replay-vision` id does not exist — the orchestrator never installs this + // (it resolves per-task mini-skills instead), but the linear path does, and + // aborts `skill-not-found` on a miss. + skillId: 'replay-vision-setup', + command: 'replay-vision', + id: 'replay-vision', + description: 'Set up PostHog Replay Vision scanners for your product', + integrationLabel: 'replay-vision', + customPrompt: + 'Set up PostHog Replay vision. Run the `replay-vision` skill end-to-end: ' + + 'make sure session replay is recording (server-side enable plus a ' + + 'posthog-js init check), then create the vision scanners the skill ' + + "defines, scoped to this product's key flows read out of the repo. If " + + 'PostHog is not integrated yet, install and initialize the SDK first as ' + + 'the skill instructs — do not abort. The final report is written to ' + + `./${REPLAY_VISION_REPORT_FILE}.`, + successMessage: `Replay vision configured! View the report at ./${REPLAY_VISION_REPORT_FILE}`, + reportFile: REPLAY_VISION_REPORT_FILE, + docsUrl: 'https://posthog.com/docs/replay-vision', + spinnerMessage: 'Setting up Replay vision...', + estimatedDurationMinutes: 6, + abortCases: REPLAY_VISION_ABORT_CASES, +}); /** * `wizard replay-vision` — flat skill command on the orchestrator sequence. @@ -141,7 +173,7 @@ export const replayVisionConfig: ProgramConfig = { steps: [DETECT_STEP, ...AGENT_SKILL_STEPS], ciPreRun: async ( - session: ReplayVisionCiSession, + session: WizardSession, host: ProgramCiHost, ): Promise => { await scopeInstallDirToProject(session, host); @@ -173,9 +205,6 @@ export const replayVisionConfig: ProgramConfig = { benchmark: session.benchmark, yaraReport: session.yaraReport, }); - const detectedLabel = - frameworkConfig.metadata.getDetectedFrameworkLabel?.(context); - if (detectedLabel) session.detectedFrameworkLabel = detectedLabel; for (const [key, value] of Object.entries(context)) { if (!(key in session.frameworkContext)) { session.frameworkContext[key] = value; diff --git a/src/programs/replay-vision/run.ts b/src/programs/replay-vision/run.ts deleted file mode 100644 index b8eaca446..000000000 --- a/src/programs/replay-vision/run.ts +++ /dev/null @@ -1,47 +0,0 @@ -import type { AbortCase } from '@agent/types'; -import type { SkillProgramOptions } from '@programs/agent-skill/run-definition'; - -const REPLAY_VISION_REPORT_FILE = 'posthog-replay-vision-report.md'; - -/** - * `[ABORT]` reasons the replay-vision skill emits when the run can't proceed. - * Kept in sync with the stop conditions in the skill's `description.md` - * (context-mill `context/skills/replay-vision`). - */ -export const REPLAY_VISION_ABORT_CASES: AbortCase[] = [ - { - match: /^replay vision not available for this project$/i, - message: 'Replay vision is not available for this project', - body: - 'Every Replay vision scanner endpoint reported that the feature is not ' + - 'available here yet. Session replay setup done so far is kept. See ' + - 'https://posthog.com/docs/replay-vision for availability.', - }, -]; - -export const REPLAY_VISION_OPTIONS: SkillProgramOptions = { - // The menu ids this skill `-`, and context-mill's - // `replay-vision/config.yaml` declares a single variant, `setup`. The bare - // `replay-vision` id does not exist — the orchestrator never installs this - // (it resolves per-task mini-skills instead), but the linear path does, and - // aborts `skill-not-found` on a miss. - skillId: 'replay-vision-setup', - command: 'replay-vision', - id: 'replay-vision', - description: 'Set up PostHog Replay Vision scanners for your product', - integrationLabel: 'replay-vision', - customPrompt: - 'Set up PostHog Replay vision. Run the `replay-vision` skill end-to-end: ' + - 'make sure session replay is recording (server-side enable plus a ' + - 'posthog-js init check), then create the vision scanners the skill ' + - "defines, scoped to this product's key flows read out of the repo. If " + - 'PostHog is not integrated yet, install and initialize the SDK first as ' + - 'the skill instructs — do not abort. The final report is written to ' + - `./${REPLAY_VISION_REPORT_FILE}.`, - successMessage: `Replay vision configured! View the report at ./${REPLAY_VISION_REPORT_FILE}`, - reportFile: REPLAY_VISION_REPORT_FILE, - docsUrl: 'https://posthog.com/docs/replay-vision', - spinnerMessage: 'Setting up Replay vision...', - estimatedDurationMinutes: 6, - abortCases: REPLAY_VISION_ABORT_CASES, -}; diff --git a/src/programs/revenue-analytics/abort-cases.ts b/src/programs/revenue-analytics/abort-cases.ts deleted file mode 100644 index abb917a32..000000000 --- a/src/programs/revenue-analytics/abort-cases.ts +++ /dev/null @@ -1,25 +0,0 @@ -import type { AbortCase } from '@agent/types'; - -/** `[ABORT] ` cases the revenue analytics skill can emit. */ -export const REVENUE_ABORT_CASES: AbortCase[] = [ - { - // Skill emits: [ABORT] Could not find a PostHog distinct_id - match: /^could not find a posthog distinct_id$/i, - message: 'Could not find a PostHog distinct_id', - body: - 'The agent could not find PostHog distinct_id usage in your codebase. ' + - 'Your users must be identified in PostHog before they can be tagged in Stripe. ' + - 'Please identify your users and try again.', - docsUrl: 'https://posthog.com/docs/product-analytics/identify', - }, - { - // Skill emits: [ABORT] Could not find a Stripe integration - match: /^could not find a stripe integration$/i, - message: 'Could not find a Stripe integration', - body: - 'The Wizard could not find an existing Stripe customer, charge, ' + - 'subscription, or other Stripe operations. Please run the Revenue ' + - 'Analytics Wizard on a project with an existing Stripe integration.', - docsUrl: 'https://posthog.com/docs/revenue-analytics', - }, -]; diff --git a/src/programs/revenue-analytics/detect.ts b/src/programs/revenue-analytics/detect.ts index 544ba7a9d..b81090ffd 100644 --- a/src/programs/revenue-analytics/detect.ts +++ b/src/programs/revenue-analytics/detect.ts @@ -6,6 +6,8 @@ */ import { existsSync, statSync } from 'fs'; +import type { WizardSession } from '@lib/wizard-session'; +import type { AbortCase } from '@agent/types'; import { findPackageJsons } from '@programs/shared/package-scanning'; export { @@ -30,7 +32,29 @@ export type RevenueDetectError = | { kind: 'missing-posthog'; foundStripe: string[] } | { kind: 'missing-stripe'; foundPosthog: string[] }; -export { REVENUE_ABORT_CASES } from './abort-cases.js'; +/** `[ABORT] ` cases the revenue analytics skill can emit. */ +export const REVENUE_ABORT_CASES: AbortCase[] = [ + { + // Skill emits: [ABORT] Could not find a PostHog distinct_id + match: /^could not find a posthog distinct_id$/i, + message: 'Could not find a PostHog distinct_id', + body: + 'The agent could not find PostHog distinct_id usage in your codebase. ' + + 'Your users must be identified in PostHog before they can be tagged in Stripe. ' + + 'Please identify your users and try again.', + docsUrl: 'https://posthog.com/docs/product-analytics/identify', + }, + { + // Skill emits: [ABORT] Could not find a Stripe integration + match: /^could not find a stripe integration$/i, + message: 'Could not find a Stripe integration', + body: + 'The Wizard could not find an existing Stripe customer, charge, ' + + 'subscription, or other Stripe operations. Please run the Revenue ' + + 'Analytics Wizard on a project with an existing Stripe integration.', + docsUrl: 'https://posthog.com/docs/revenue-analytics', + }, +]; /** * Scan `session.installDir` for PostHog + Stripe SDKs. Writes detection @@ -40,7 +64,7 @@ export { REVENUE_ABORT_CASES } from './abort-cases.js'; * The skill install happens later in the bootstrap runner, not here. */ export function detectRevenuePrerequisites( - session: { installDir: string }, + session: WizardSession, setFrameworkContext: (key: string, value: unknown) => void, ): void { const fail = (error: RevenueDetectError) => diff --git a/src/programs/revenue-analytics/index.ts b/src/programs/revenue-analytics/index.ts index 31acec8d4..4fae357ac 100644 --- a/src/programs/revenue-analytics/index.ts +++ b/src/programs/revenue-analytics/index.ts @@ -1,7 +1,8 @@ import type { ProgramConfig } from '@programs/program-step'; import { WIZARD_TOOL_NAMES } from '@agent'; import { REVENUE_ANALYTICS_PROGRAM } from './steps.js'; -import { REVENUE_ANALYTICS_RUN } from './run.js'; +import { REVENUE_ABORT_CASES } from './detect.js'; +import { getContentBlocks } from '../../ui/tui/decks/revenue-analytics/index.js'; export const revenueAnalyticsConfig: ProgramConfig = { command: 'revenue-analytics', @@ -9,9 +10,20 @@ export const revenueAnalyticsConfig: ProgramConfig = { id: 'revenue-analytics-setup', skillId: 'revenue-analytics-setup', steps: REVENUE_ANALYTICS_PROGRAM, + getContentBlocks, allowedTools: ['Agent'], disallowedTools: [WIZARD_TOOL_NAMES.wizardAsk], - run: REVENUE_ANALYTICS_RUN, + run: { + skillId: 'revenue-analytics-setup', + integrationLabel: 'revenue-analytics-setup', + customPrompt: () => 'Set up revenue analytics for this project.', + successMessage: 'Revenue analytics configured!', + reportFile: 'posthog-revenue-report.md', + docsUrl: 'https://posthog.com/docs/revenue-analytics', + spinnerMessage: 'Setting up revenue analytics...', + estimatedDurationMinutes: 5, + abortCases: REVENUE_ABORT_CASES, + }, requires: ['posthog-integration'], }; diff --git a/src/programs/revenue-analytics/run.ts b/src/programs/revenue-analytics/run.ts deleted file mode 100644 index 861c75992..000000000 --- a/src/programs/revenue-analytics/run.ts +++ /dev/null @@ -1,14 +0,0 @@ -import type { ProgramRun } from '@programs/program-run'; -import { REVENUE_ABORT_CASES } from './abort-cases.js'; - -export const REVENUE_ANALYTICS_RUN: ProgramRun = { - skillId: 'revenue-analytics-setup', - integrationLabel: 'revenue-analytics-setup', - customPrompt: () => 'Set up revenue analytics for this project.', - successMessage: 'Revenue analytics configured!', - reportFile: 'posthog-revenue-report.md', - docsUrl: 'https://posthog.com/docs/revenue-analytics', - spinnerMessage: 'Setting up revenue analytics...', - estimatedDurationMinutes: 5, - abortCases: REVENUE_ABORT_CASES, -}; diff --git a/src/programs/revenue-analytics/steps.ts b/src/programs/revenue-analytics/steps.ts index 3162fa63e..b53926536 100644 --- a/src/programs/revenue-analytics/steps.ts +++ b/src/programs/revenue-analytics/steps.ts @@ -6,7 +6,7 @@ */ import type { ProgramStep } from '@programs/program-step'; -import { RunPhase } from '@shared/run-state'; +import { RunPhase } from '@lib/wizard-session'; import { HEALTH_CHECK_STEP } from '@programs/shared/health-check-step'; import { detectRevenuePrerequisites } from './detect.js'; diff --git a/src/lib/runners/run-program-agent.ts b/src/programs/run-agent-legacy.ts similarity index 72% rename from src/lib/runners/run-program-agent.ts rename to src/programs/run-agent-legacy.ts index 7057849d1..41bb0b05b 100644 --- a/src/lib/runners/run-program-agent.ts +++ b/src/programs/run-agent-legacy.ts @@ -1,25 +1,25 @@ -/** Runs a ProgramConfig through runProgram with the session, `getUI()` and `wizardAbort` as its host, until Release C replaces it. */ +/** + * The session-driven agent runner every existing caller uses. + * + * `runProgramAgent(programConfig, session)` runs the gates the TUI owns + * (health, settings), then hands the program to `runProgram` with the session + * and `getUI()` as its host: credentials come from `authenticate`, the AI + * opt-in and post-auth gates park on the UI, every progress event maps back + * onto `getUI()`, and the invocation's data projects back onto the session. + * It applies the result — `wizardAbort` with the outcome's terminal status for + * a decided failure, the terminal analytics event for a finished top-level run. + * + * This is the only file that knows about `getUI()`, the session and + * `wizardAbort` on the agent's behalf. Programs replace it in Release B. + */ -import { isDeepStrictEqual } from 'node:util'; -import type { WizardSession } from '@lib/wizard-session'; +import { mayReportScanResults, type WizardSession } from '@lib/wizard-session'; import { analytics } from '@utils/analytics'; -import { createUiReducer, getUI, uiInteraction, type WizardUI } from '@ui'; -import { RunOutcome, TASK_OUTCOMES_KEY } from '@agent'; -import { - AUDIT_CHECKS_KEY, - authenticate, - FRAMEWORK_REGISTRY, - getDetectedWarehouseSources, - postAuthGateSteps, - runProgram, -} from '@programs'; -import type { - ProgramCompletionContext, - ProgramConfig, - ProgramInvocationData, - ProgramRunHost, - WizardFlagSnapshot, -} from '@programs/types'; +import { getUI, type WizardUI } from '@ui'; +import { createUiReducer, uiInteraction } from '@ui/agent-progress'; +import { flushScanReport, RunOutcome, TASK_OUTCOMES_KEY } from '@agent'; +import type { ProgramRun } from './program-run'; +import type { ProgramRunHost } from './host-capabilities'; import { backupAndFixClaudeSettings, checkAllSettingsConflicts, @@ -34,13 +34,22 @@ import { SERVICE_LABELS, } from '@shared/health-checks/readiness'; import { enableDebugLogs, logToFile, initLogFile } from '@utils/debug'; -import { wizardAbort } from '@utils/wizard-abort'; +import { registerCleanup, wizardAbort } from '@utils/wizard-abort'; import { ErrorCodes } from '@shared/errors'; import { isNonInteractiveEnvironment } from '@utils/environment'; import { Sequence, type Integration } from '@shared/constants'; -import { mayReportScanResults } from '@shared/scan-consent'; +import { FRAMEWORK_REGISTRY } from '@programs/registry'; +import { postAuthGateSteps, type ProgramConfig } from './program-step'; +import { authenticate } from './authenticate'; +import { startAuditLedgerWatcher } from './audit/ledger-watcher'; +import { getDetectedWarehouseSources } from './warehouse-source/detect'; +import { runProgram, type WizardFlagSnapshot } from './run-program'; +import type { ProgramInvocationData } from './program-store'; -/** Resolve the program's run from the session, run the gates, run it through runProgram and apply the result. */ +/** + * Resolve a ProgramConfig's agent run definition and execute the pipeline. + * Entry point for the runners and for composed run steps. + */ export async function runProgramAgent( programConfig: ProgramConfig, session: WizardSession, @@ -50,27 +59,47 @@ export async function runProgramAgent( throw new Error(`Program "${programConfig.id}" has no run configuration.`); } - const ui = getUI(); - const runHost: ProgramRunHost = { - getFrameworkContext: (key) => ui.getFrameworkContext(key), - setFrameworkContext: (key, value) => ui.setFrameworkContext(key, value), - warn: (message) => ui.log.warn(message), - uploadEnvironmentVariables: async (envVars, integration, installDir) => { - const { uploadEnvironmentVariablesStep } = await import( - '@steps/upload-environment-variables' - ); - return uploadEnvironmentVariablesStep(envVars, { - integration, - session: { installDir }, - }); - }, + // Before `run()` resolves: an audit seeds the ledger from inside its recipe, + // and a watcher started later would ignore that write as pre-existing. + const ledger = programConfig.auditLedgerFile + ? startAuditLedgerWatcher(session.installDir, programConfig.auditLedgerFile) + : null; + if (ledger) registerCleanup(() => ledger.stop()); + + try { + const runDef = + typeof programConfig.run === 'function' + ? await programConfig.run(session, uiRunHost()) + : programConfig.run; + + await runSessionProgram( + session, + runDef, + programConfig, + options.composed ?? false, + ); + } finally { + ledger?.stop(); + } +} + +/** The run host each program effect reaches `getUI()` through, read at call time. */ +function uiRunHost(): ProgramRunHost { + return { + getFrameworkContext: (key) => getUI().getFrameworkContext(key), + setFrameworkContext: (key, value) => + getUI().setFrameworkContext(key, value), + warn: (message) => getUI().log.warn(message), }; - const run = - typeof programConfig.run === 'function' - ? await programConfig.run(session, runHost) - : programConfig.run; - const composed = options.composed ?? false; +} +/** Gates → runProgram with the session as its host → apply result. */ +async function runSessionProgram( + session: WizardSession, + run: ProgramRun, + programConfig: ProgramConfig, + composed: boolean, +): Promise { // 1. Init logging + debug initLogFile(); session.skillId = run.skillId ?? run.integrationLabel; @@ -90,10 +119,9 @@ export async function runProgramAgent( // 3. Settings conflicts await runSettingsGate(session); + const ui = getUI(); const reduceUi = createUiReducer(ui); - const projectData = projectProgramData(ui, session, () => - restoreClaudeSettings(session.installDir), - ); + const projectData = projectProgramData(ui, session); // runProgram turns a throwing host capability into a failed run; the CLI roots expect the throw. let hostFailure: { error: unknown } | undefined; @@ -104,12 +132,6 @@ export async function runProgramAgent( }); const framework = session.integration ?? session.skillId ?? undefined; - // Each hook reads the session when it runs, so URLs the run emitted reach it. - const completionContext = (): ProgramCompletionContext => ({ - signup: session.signup, - dashboardUrl: session.dashboardUrl, - notebookUrl: session.notebookUrl, - }); const result = await runProgram( programConfig.id, { @@ -148,15 +170,14 @@ export async function runProgramAgent( : undefined, hooks: { postRun: run.postRun - ? (creds) => run.postRun!(completionContext(), creds) + ? (creds) => run.postRun!(session, creds) : undefined, buildOutroData: run.buildOutroData - ? (creds) => - run.buildOutroData!(completionContext(), creds) ?? undefined + ? (creds) => run.buildOutroData!(session, creds) ?? undefined : undefined, buildOutroNextSteps: run.buildOutroNextSteps ? (creds, completed) => - run.buildOutroNextSteps!(completionContext(), creds, completed) + run.buildOutroNextSteps!(session, creds, completed) : undefined, recordTaskOutcomes: (outcomes) => { session.frameworkContext[TASK_OUTCOMES_KEY] = outcomes; @@ -168,9 +189,6 @@ export async function runProgramAgent( allowedTools: programConfig.allowedTools, disallowedTools: programConfig.disallowedTools, excludedTaskTypes: programConfig.excludedTaskTypes, - auditLedgerFile: programConfig.auditLedgerFile, - auditSeedChecks: programConfig.auditSeedChecks, - eventPlanFile: programConfig.eventPlanFile, postAuthGates: postAuthGateSteps(programConfig.steps).map( (step) => step.id, ), @@ -182,19 +200,27 @@ export async function runProgramAgent( }, { credentials: { - // authenticate() is idempotent, so a later run in the same invocation reuses the login. - resolve: (programId) => + // Idempotent within a run: a second agent run in the same invocation + // (self-driving's integration phase) reuses the first login. + resolve: () => keepFailure( - authenticate(session, programId, ui).then(() => ({ + authenticate(session, programConfig.id).then(() => ({ posthog: session.credentials!, - inferenceAuth: session.inferenceAuth, project: session.apiProject, apiUser: session.apiUser, })), ), }, - featureFlags: () => keepFailure(loadWizardFlags()), - // Each step the user settles between auth and run, such as the source-maps project picker. + // The actual AI opt-in gate: it parks while AiOptInRequiredScreen is up, + // before the skill install and agent start, so no source leaves the machine. + awaitAiApproval: async () => { + logToFile('[agent-runner] checking AI opt-in gate'); + await ui.waitForAiOptIn(); + logToFile('[agent-runner] AI opt-in gate cleared'); + return true; + }, + // Each step the user settles between auth and run, such as the source-maps + // project picker, which writes its choice to frameworkContext for the prompt. awaitPostAuthGates: async ({ gates }) => { for (const gate of gates) { logToFile(`[agent-runner] awaiting post-auth gate: ${gate}`); @@ -202,20 +228,12 @@ export async function runProgramAgent( logToFile(`[agent-runner] post-auth gate cleared: ${gate}`); } }, + featureFlags: () => keepFailure(loadWizardFlags()), onProgress: (progress) => { if (progress.kind === 'run') reduceUi(progress.event); else projectData(progress.data); }, interaction: uiInteraction(ui), - // The CLI roots commit new skills at exit, so a later drain still removes them. - deferSkillCommit: true, - // The actual AI opt-in gate: it parks before the skill install and agent start. - awaitAiApproval: async () => { - logToFile('[agent-runner] checking AI opt-in gate'); - await ui.waitForAiOptIn(); - logToFile('[agent-runner] AI opt-in gate cleared'); - return true; - }, }, ); if (hostFailure) throw hostFailure.error; @@ -244,18 +262,12 @@ export async function runProgramAgent( } } -// ── Host capabilities ───────────────────────────────────────────────── - /** Mirror the invocation's data onto the session and the UI the TUI reads. */ function projectProgramData( ui: WizardUI, session: WizardSession, - restoreSettings: () => void, ): (data: ProgramInvocationData) => void { - let outroRestoreRegistered = false; - // Snapshots are copies, so forward by value; the store starts with no plan. - let eventPlan: ProgramInvocationData['eventPlan'] = []; - let auditChecks: unknown; + let bindingSeen = false; return (data) => { const current = session.credentials; if ( @@ -273,19 +285,25 @@ function projectProgramData( ui.setAccessToken(session.credentials); } if (data.aiSdkStampReported) session.aiSdkStampReported = true; - // Registered before the run can reach the outro; the abort path restores through its own cleanup. - if (data.binding?.sequence === Sequence.linear && !outroRestoreRegistered) { - outroRestoreRegistered = true; - ui.onEnterScreen('outro', restoreSettings); - } - if (!isDeepStrictEqual(data.eventPlan, eventPlan)) { - eventPlan = data.eventPlan; - ui.setEventPlan(eventPlan); - } - const checks = data.detection.frameworkContext[AUDIT_CHECKS_KEY]; - if (checks !== undefined && !isDeepStrictEqual(checks, auditChecks)) { - auditChecks = checks; - ui.setFrameworkContext(AUDIT_CHECKS_KEY, checks); + if (!data.binding || bindingSeen) return; + bindingSeen = true; + + // Cleanup coverage for the abort/cancel path: `wizardAbort` runs the + // registered cleanups, and the agent's own `finally` covers completion. + // flushScanReport is idempotent, so the overlap is a harmless no-op. + registerCleanup(() => { + const report = flushScanReport({ yaraReport: session.yaraReport }); + if (report) ui.log.info(report); + }); + + // Linear settings restoration fires on entry to the outro screen, so it + // is registered before the run can reach that screen. The abort path + // still restores through the cleanup `backupAndFixClaudeSettings` + // registered. + if (data.binding.sequence === Sequence.linear) { + ui.onEnterScreen('outro', () => + restoreClaudeSettings(session.installDir), + ); } }; } diff --git a/src/programs/run-program.ts b/src/programs/run-program.ts index 505979c11..8b786efcc 100644 --- a/src/programs/run-program.ts +++ b/src/programs/run-program.ts @@ -1,41 +1,36 @@ /** A caller-owned program invocation. No TUI store or session is required. */ import path from 'path'; import { randomUUID } from 'crypto'; -import { runAgent, RunOutcome } from '@agent'; +import { buildRunTags, resolveBinding, runAgent, RunOutcome } from '@agent'; import type { AgentInteraction, AgentRunDefinition, + ProgramBinding, RunConfig, RunHooks, RunInput, RunResult, + SwitchboardCtx, } from '@agent/types'; -import { getSkillsBaseUrl } from '@shared/constants'; -import type { AuditCheck } from '@shared/audit-ledger'; -import type { Harness, Integration, Sequence } from '@shared/constants'; -import { ErrorCodes } from '@shared/errors'; -import { buildRunTags } from '@shared/run-tags'; import { - registerRunSkillCleanup, - type RunSkillCleanup, -} from '@shared/skill-run-cleanup'; -import type { DiscoveredFeature } from '@shared/scan-consent'; + getSkillsBaseUrl, + Sequence, + WIZARD_ORCHESTRATOR_FLAG_KEY, + WIZARD_SELF_DRIVING_USE_PI_HARNESS_FLAG_KEY, + type Harness, + type Integration, +} from '@shared/constants'; +import { ErrorCodes } from '@shared/errors'; +import type { DiscoveredFeature } from '@lib/wizard-session'; import { analytics, groupsFromUser } from '@utils/analytics'; import { logToFile } from '@utils/debug'; import type { DetectedSource } from './warehouse-sources/types'; -import { - createPosthogInferenceAuthProvider, - type CredentialsProvider, - type ResolvedProgramCredentials, +import type { + CredentialsProvider, + ResolvedProgramCredentials, } from './credentials'; -import { refreshCredentialsIfNeeded } from './token-refresh'; -import { stampAiSdkDetected } from './posthog-integration/ai-sdk-stamp'; -import { startProgramFileWatchers } from './program-file-watchers'; -import { resolveProgramBinding } from './binding'; -import { getProgramCommandments } from './commandments'; -import { areSeededTasksEnabled, resolveStageOverrides } from './experiments'; -import { captureSwitchboardDecision } from './binding-telemetry'; -import { snapshotProgramInput } from './snapshot-program-input'; +import { refreshCredentialsIfNeeded } from './authenticate'; +import { stampAiSdkDetected } from './posthog-integration/detect'; import { ProgramStore, type ProgramDiagnostic, @@ -65,10 +60,6 @@ export type ProgramSettings = { allowedTools?: RunConfig['allowedTools']; disallowedTools?: RunConfig['disallowedTools']; excludedTaskTypes?: RunConfig['excludedTaskTypes']; - auditLedgerFile?: string; - /** Written to the audit ledger before the agent starts. */ - auditSeedChecks?: readonly AuditCheck[]; - eventPlanFile?: string; /** Steps the host settles after auth and before the agent starts. */ postAuthGates?: readonly string[]; }; @@ -119,8 +110,6 @@ export interface ProgramOptions { }) => Promise; /** Evaluate feature flags for a run whose input carries none. */ featureFlags?: () => Promise; - /** Leave new skills armed after success; the host commits them at exit. */ - deferSkillCommit?: boolean; signal?: AbortSignal; } @@ -151,6 +140,22 @@ const DEFAULT_FLAGS: RunInput['flags'] = { yaraReport: false, }; +/** Fields that carry functions or class instances; everything else is data. */ +const KEPT_BY_REFERENCE = [ + 'credentials', + 'run', + 'program', + 'hooks', + 'seedTasks', +] as const satisfies readonly (keyof ProgramInput)[]; + +/** Copy the host's input, so a later host write cannot reach the run or its hooks. */ +function snapshotProgramInput(input: ProgramInput): ProgramInput { + const data: Partial = { ...input }; + for (const key of KEPT_BY_REFERENCE) delete data[key]; + return { ...input, ...structuredClone(data) }; +} + /** Run an existing program from explicit inputs, with invocation-owned state. */ export async function runProgram( programId: string, @@ -162,33 +167,6 @@ export async function runProgram( aiSdkStampReported: input.aiSdkStampReported, onData: options.onProgress, }); - // Registered, so a process drain mid-run (wizardAbort, a signal) removes new skills too. - const skills = registerRunSkillCleanup(input.installDir); - try { - const result = await runWithStore(programId, input, options, store); - if (result.outcome !== RunOutcome.Success) cleanFailedRun(skills); - else if (!options.deferSkillCommit) skills.commit(); - return result; - } catch (error) { - cleanFailedRun(skills); - throw error; - } -} - -function cleanFailedRun(skills: RunSkillCleanup): void { - try { - skills(); - } catch (error) { - logToFile('[programs] failed-run skill cleanup error:', error); - } -} - -async function runWithStore( - programId: string, - input: ProgramInput, - options: ProgramOptions, - store: ProgramStore, -): Promise { const { installDir, run } = input; const program = input.program ?? {}; const artifacts: ProgramRunOutcome['artifacts'] = {}; @@ -311,91 +289,120 @@ async function runWithStore( const wizardFlags = { ...flagSnapshot.flags }; const wizardFlagPayloads = { ...flagSnapshot.payloads }; - const fileWatchers = startProgramFileWatchers(program, installDir, store); - try { - fileWatchers.seedAuditLedger(); + // Resolve which sequence and harness run the program (CLI → PostHog flag → + // per-program binding → default) and tag both axes onto analytics. + const switchboard: SwitchboardCtx = { + program: programId, + composed: input.composed ?? false, + flags: wizardFlags, + flagPayloads: wizardFlagPayloads, + cliHarness: input.overrides?.harness, + cliSequence: input.overrides?.sequence, + cliModel: input.overrides?.model, + }; + const binding = resolveBinding(switchboard); + analytics.setTag('sequence', binding.sequence); + analytics.setTag('harness', binding.harness); + captureSwitchboardDecision(switchboard, binding); + store.setBinding(binding); - const switchboard = { - program: programId, - composed: input.composed ?? false, - flags: wizardFlags, - flagPayloads: wizardFlagPayloads, - cliHarness: input.overrides?.harness, - cliSequence: input.overrides?.sequence, - cliModel: input.overrides?.model, - }; - const binding = resolveProgramBinding(switchboard); - analytics.setTag('sequence', binding.sequence); - analytics.setTag('harness', binding.harness); - captureSwitchboardDecision(switchboard, binding); - store.setBinding(binding); + const wizardMetadata = { + ...buildRunTags({ + programId, + integration: run.integrationLabel, + runId: analytics.runId, + build: analytics.build, + skillId: run.skillId, + }), + SEQUENCE: binding.sequence, + HARNESS: binding.harness, + }; + artifacts.reportFile = path.resolve(installDir, run.reportFile); + const adapter = store.beginRun(runId, options.onProgress); - const inferenceAuth = - credentials.inferenceAuth ?? - createPosthogInferenceAuthProvider(credentials.posthog, programId); - const wizardMetadata = { - ...buildRunTags({ - programId, - integration: run.integrationLabel, - runId: analytics.runId, - build: analytics.build, - skillId: run.skillId, - }), - SEQUENCE: binding.sequence, - HARNESS: binding.harness, - }; - artifacts.reportFile = path.resolve(installDir, run.reportFile); - const adapter = store.beginRun(runId, options.onProgress); + const result = await runAgent( + { + programId, + run, + composed: input.composed ?? false, + binding, + switchboard, + skillsBaseUrl: getSkillsBaseUrl(), + wizardFlags, + wizardFlagPayloads, + wizardMetadata, + allowedTools: program.allowedTools, + disallowedTools: program.disallowedTools, + agentFlow: program.agentFlow, + excludedTaskTypes: program.excludedTaskTypes, + seedTasks: input.seedTasks, + hooks: input.hooks, + }, + { + installDir, + credentials: credentials.posthog, + project: credentials.project, + apiUser: credentials.apiUser, + skillId: input.skillId ?? run.skillId ?? run.integrationLabel, + integration: input.integration, + frameworkDocsUrl: input.frameworkDocsUrl, + flags, + host: { ...input.host }, + }, + { + interaction: options.interaction, + onProgress: (event) => adapter.onProgress(event), + signal: options.signal, + }, + ); + adapter.finish(result); + return settle( + result.outcome, + result.outcome === RunOutcome.Success ? undefined : result.failure, + ); +} - const result = await runAgent( - { - programId, - run, - composed: input.composed ?? false, - binding, - programCommandments: getProgramCommandments(programId), - stageOverrides: resolveStageOverrides( - programId, - wizardFlags, - wizardFlagPayloads, - ), - seededTasksEnabled: areSeededTasksEnabled(wizardFlags), - skillsBaseUrl: getSkillsBaseUrl(), - wizardFlags, - wizardFlagPayloads, - wizardMetadata, - allowedTools: program.allowedTools, - disallowedTools: program.disallowedTools, - agentFlow: program.agentFlow, - excludedTaskTypes: program.excludedTaskTypes, - seedTasks: input.seedTasks, - hooks: input.hooks, - }, - { - installDir, - credentials: credentials.posthog, - inferenceAuth, - project: credentials.project, - apiUser: credentials.apiUser, - skillId: input.skillId ?? run.skillId ?? run.integrationLabel, - integration: input.integration, - frameworkDocsUrl: input.frameworkDocsUrl, - flags, - host: { ...input.host }, - } as RunInput, - { - interaction: options.interaction, - onProgress: (event) => adapter.onProgress(event), - signal: options.signal, - }, - ); - adapter.finish(result); - fileWatchers.refresh(); - return settle( - result.outcome, - result.outcome === RunOutcome.Success ? undefined : result.failure, - ); - } finally { - fileWatchers.stop(); - } +/** + * One event + one log line per run: what entered the switchboard, which + * precedence rung decided each axis, and the final pick. + */ +function captureSwitchboardDecision( + ctx: SwitchboardCtx, + binding: ProgramBinding, +): void { + const trace = ctx.trace ?? {}; + // Unpinned orchestrator runs choose a model per task from the context-mill agent prompts; the orchestrator logs that map once the prompts load. + const perTaskModel = + binding.sequence === Sequence.orchestrator && trace.model === 'binding'; + const model = perTaskModel ? 'chosen-per-task' : binding.model; + const modelSource = perTaskModel ? 'agent-prompts' : trace.model; + analytics.wizardCapture('switchboard resolved', { + program: ctx.program, + flag_self_driving_use_pi_harness: + ctx.flags[WIZARD_SELF_DRIVING_USE_PI_HARNESS_FLAG_KEY], + flag_self_driving_pi_payload: JSON.stringify( + ctx.flagPayloads?.[WIZARD_SELF_DRIVING_USE_PI_HARNESS_FLAG_KEY] ?? null, + ), + flag_orchestrator: ctx.flags[WIZARD_ORCHESTRATOR_FLAG_KEY], + cli_harness: ctx.cliHarness, + cli_sequence: ctx.cliSequence, + cli_model: ctx.cliModel, + harness_source: trace.harness, + model_source: modelSource, + sequence_source: trace.sequence, + harness: binding.harness, + model, + thinking_level: binding.thinkingLevel, + sequence: binding.sequence, + }); + logToFile( + `[switchboard] decision: program=${ctx.program}` + + ` in(orchestrator=${ctx.flags[WIZARD_ORCHESTRATOR_FLAG_KEY] ?? '-'},` + + ` cli=${ctx.cliHarness ?? '-'}/${ctx.cliSequence ?? '-'}/${ + ctx.cliModel ?? '-' + })` + + ` → harness=${binding.harness} (${trace.harness ?? '?'})` + + ` model=${model} (${modelSource ?? '?'})` + + ` sequence=${binding.sequence} (${trace.sequence ?? '?'})`, + ); } diff --git a/src/programs/self-driving/detect-agentic.ts b/src/programs/self-driving/detect-agentic.ts index dd10288a7..2df306bf7 100644 --- a/src/programs/self-driving/detect-agentic.ts +++ b/src/programs/self-driving/detect-agentic.ts @@ -12,8 +12,6 @@ */ import type { - AgenticDetectionContext, - AgenticDetectOptions, AgenticDetectionReport, DetectEvent, } from '@programs/detection/agentic'; @@ -22,8 +20,8 @@ import { toIntegrationCandidates, } from '@programs/detection/project-scope'; import { gatherFrameworkContext } from '@programs/detection/index'; -import type { FrameworkDetectionState } from '@programs/detection/context'; import type { Integration } from '@shared/constants'; +import type { WizardSession } from '@lib/wizard-session'; export type { DetectEvent }; @@ -85,14 +83,12 @@ export function toIntegrationReport( /** Run the Haiku detector over the repo and classify projects for integration. */ export async function detectSelfDrivingIntegrationProjects( - session: AgenticDetectionContext, + session: WizardSession, onEvent?: DetectEvent, - onProgress?: AgenticDetectOptions['onProgress'], ): Promise { const report = await detectIntegrationProjects(session, { programId: 'self-driving', onEvent, - onProgress, }); return toIntegrationReport(report); } @@ -106,7 +102,7 @@ export async function detectSelfDrivingIntegrationProjects( * integrate-run step's `onRunPrep`. */ export async function prepSelfDrivingIntegration( - session: FrameworkDetectionState, + session: WizardSession, ): Promise { // `session` is the phase's derived session — its installDir is already the // picked project (the integrate-run step's `targetDir`), so just gather that @@ -122,9 +118,6 @@ export async function prepSelfDrivingIntegration( benchmark: session.benchmark, yaraReport: session.yaraReport, }); - const detectedLabel = - frameworkConfig.metadata.getDetectedFrameworkLabel?.(context); - if (detectedLabel) session.detectedFrameworkLabel = detectedLabel; for (const [key, value] of Object.entries(context)) { if (!(key in session.frameworkContext)) { session.frameworkContext[key] = value; diff --git a/src/programs/self-driving/detect.ts b/src/programs/self-driving/detect.ts index b3fca72b7..e53598ed7 100644 --- a/src/programs/self-driving/detect.ts +++ b/src/programs/self-driving/detect.ts @@ -28,6 +28,7 @@ import { } from 'fs'; import { join } from 'path'; import { analytics } from '@utils/analytics'; +import type { WizardSession } from '@lib/wizard-session'; import type { AbortCase } from '@agent/types'; import { ErrorCodes } from '@shared/errors'; import { detectWarehouseSources } from '@programs/warehouse-sources/detect'; @@ -48,9 +49,9 @@ export const SELF_DRIVING_INTEGRATE_PATH_KEY = 'selfDrivingIntegratePath'; export const SELF_DRIVING_DETECTED_TOOLS_KEY = 'selfDrivingDetectedTools'; /** Read the detected tools out of frameworkContext. */ -export function getSelfDrivingDetectedTools(session: { - frameworkContext: Record; -}): DetectedSource[] { +export function getSelfDrivingDetectedTools( + session: WizardSession, +): DetectedSource[] { return ( (session.frameworkContext[SELF_DRIVING_DETECTED_TOOLS_KEY] as | DetectedSource[] @@ -289,7 +290,7 @@ export const SELF_DRIVING_ABORT_CASES: AbortCase[] = [ * screen renders it and blocks. */ export function detectSelfDrivingPrerequisites( - session: { installDir: string }, + session: WizardSession, setFrameworkContext: (key: string, value: unknown) => void, ): void { const fail = (error: SelfDrivingDetectError) => diff --git a/src/programs/self-driving/index.ts b/src/programs/self-driving/index.ts index 3e22b6e66..b43e00fce 100644 --- a/src/programs/self-driving/index.ts +++ b/src/programs/self-driving/index.ts @@ -2,7 +2,7 @@ import { join } from 'path'; import { access, rm } from 'node:fs/promises'; import type { ProgramConfig } from '@programs/program-step'; import type { ProgramRun } from '@programs/program-run'; -import { OutroKind } from '@agent'; +import { OutroKind, type WizardSession } from '@lib/wizard-session'; import { createSkillProgram } from '../agent-skill/index.js'; import { SELF_DRIVING_PROGRAM } from './steps.js'; import { @@ -15,7 +15,9 @@ import { NO_DEFAULT_LIMIT, PRICE_PER_PR_USD, PRICING_LONG, -} from '@shared/self-driving-pricing'; +} from '../../ui/tui/decks/self-driving/pricing.js'; +import { getTips } from '../../ui/tui/decks/self-driving/tips.js'; +import { getContentBlocks } from '../../ui/tui/decks/self-driving/index.js'; export const SELF_DRIVING_SKILL_ID = 'self-driving-setup'; const REPORT_FILE = 'posthog-self-driving-report.md'; @@ -45,10 +47,7 @@ async function removeInstalledSkill(installDir: string): Promise { // A session closure (not a static object) so `customPrompt` can read the // tools detected in the codebase — written to frameworkContext by the detect // step — and hand them to the prompt for STEP 4/STEP 5 prioritisation. -const buildRun = (session: { - installDir: string; - frameworkContext: Record; -}): Promise => +const buildRun = (session: WizardSession): Promise => Promise.resolve({ skillId: SELF_DRIVING_SKILL_ID, integrationLabel: SELF_DRIVING_SKILL_ID, @@ -85,7 +84,7 @@ const buildRun = (session: { // conversion, scout enable rate) keeps counting when a run words its tasks differently. resolveStepKey: resolveSelfDrivingStepKey, - postRun: async () => { + postRun: async (session) => { await removeInstalledSkill(session.installDir); }, @@ -129,6 +128,8 @@ export const selfDrivingConfig: ProgramConfig = { }), steps: SELF_DRIVING_PROGRAM, run: buildRun, + getTips, + getContentBlocks, }; export { SELF_DRIVING_PROGRAM } from './steps.js'; diff --git a/src/programs/self-driving/steps.ts b/src/programs/self-driving/steps.ts index 5daf8a382..2bad6952c 100644 --- a/src/programs/self-driving/steps.ts +++ b/src/programs/self-driving/steps.ts @@ -16,7 +16,7 @@ import type { ProgramStep } from '@programs/program-step'; import { resolveProjectDir } from '@programs/detection/agentic'; -import { RunPhase } from '@shared/run-state'; +import { RunPhase, type WizardSession } from '@lib/wizard-session'; import { HEALTH_CHECK_STEP } from '@programs/shared/health-check-step'; import { integrationRunStep } from '@programs/posthog-integration/index'; import { @@ -27,16 +27,11 @@ import { import { prepSelfDrivingIntegration } from './detect-agentic.js'; /** True once detection found PostHog already present in the project. */ -type SelfDrivingStepContext = { - installDir: string; - frameworkContext: Record; -}; - -const postHogPresent = (session: SelfDrivingStepContext): boolean => +const postHogPresent = (session: WizardSession): boolean => session.frameworkContext[POSTHOG_PRESENT_KEY] === true; /** Absolute dir to integrate into: the picked sub-app (LLM output — the shared resolver clamps escapes), else the repo root. */ -const integrationDir = (session: SelfDrivingStepContext): string => +const integrationDir = (session: WizardSession): string => resolveProjectDir( session.installDir, session.frameworkContext[SELF_DRIVING_INTEGRATE_PATH_KEY], diff --git a/src/programs/shared/health-check-step.ts b/src/programs/shared/health-check-step.ts index 79e12d926..d7a147159 100644 --- a/src/programs/shared/health-check-step.ts +++ b/src/programs/shared/health-check-step.ts @@ -12,22 +12,16 @@ */ import type { ProgramStep } from '@programs/program-step'; +import type { WizardSession } from '@lib/wizard-session'; import { evaluateWizardReadiness, WizardReadiness, SIGNUP_WIZARD_READINESS_CONFIG, getBlockingServiceKeys, - type WizardReadinessResult, } from '@shared/health-checks/readiness'; import { logToFile } from '@utils/debug'; -type HealthCheckState = { - readinessResult: WizardReadinessResult | null; - signup: boolean; - outageDismissed: boolean; -}; - -export function healthCheckReady(session: HealthCheckState): boolean { +export function healthCheckReady(session: WizardSession): boolean { if (!session.readinessResult) return false; if (session.signup) { diff --git a/src/programs/shared/posthog-cli-preinstall.ts b/src/programs/shared/posthog-cli-preinstall.ts index e137250f8..18077bcb4 100644 --- a/src/programs/shared/posthog-cli-preinstall.ts +++ b/src/programs/shared/posthog-cli-preinstall.ts @@ -8,7 +8,7 @@ * needs the CLI. */ -import { installOrUpdatePostHogCli } from '@shared/posthog-cli-install'; +import { installOrUpdatePostHogCli } from '@steps/install-cli-steering'; import { analytics } from '@utils/analytics'; let attempted = false; diff --git a/src/programs/snapshot-program-input.ts b/src/programs/snapshot-program-input.ts deleted file mode 100644 index 8572e973e..000000000 --- a/src/programs/snapshot-program-input.ts +++ /dev/null @@ -1,21 +0,0 @@ -import type { ProgramInput } from './run-program'; - -/** Fields that carry functions or class instances; everything else is data. */ -const KEPT_BY_REFERENCE = [ - 'credentials', - 'run', - 'program', - 'hooks', - 'seedTasks', -] as const satisfies readonly (keyof ProgramInput)[]; - -/** - * Copy a host's program input when runProgram receives it, so a later host - * write cannot reach the run or the completion hooks built from it. Data is - * cloned; functions, and the objects that carry them, stay by reference. - */ -export function snapshotProgramInput(input: ProgramInput): ProgramInput { - const data: Partial = { ...input }; - for (const key of KEPT_BY_REFERENCE) delete data[key]; - return { ...input, ...structuredClone(data) }; -} diff --git a/src/programs/task-stream/__tests__/event-plan-watcher.test.ts b/src/programs/task-stream/__tests__/event-plan-watcher.test.ts new file mode 100644 index 000000000..af68b9838 --- /dev/null +++ b/src/programs/task-stream/__tests__/event-plan-watcher.test.ts @@ -0,0 +1,163 @@ +import { + mkdtempSync, + rmSync, + symlinkSync, + unlinkSync, + writeFileSync, +} from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { + EventPlanWatcher, + normalizeEventPlan, +} from '@programs/task-stream/event-plan-watcher'; +import { EVENT_PLAN_FILE } from '@programs/posthog-integration/constants'; +import type { PlannedEvent, WizardStore } from '@ui/tui/store'; + +const wait = (ms: number) => new Promise((resolve) => setTimeout(resolve, ms)); + +function createStore(installDir: string) { + let eventPlan: PlannedEvent[] = []; + return { + session: { installDir }, + get eventPlan() { + return eventPlan; + }, + setEventPlan(events: PlannedEvent[]) { + eventPlan = events; + }, + } as WizardStore; +} + +describe('EventPlanWatcher', () => { + let installDir: string; + let watcher: EventPlanWatcher | undefined; + + beforeEach(() => { + installDir = mkdtempSync(join(tmpdir(), 'wizard-event-plan-')); + }); + + afterEach(() => { + watcher?.stop(); + watcher = undefined; + rmSync(installDir, { recursive: true, force: true }); + }); + + it('normalizes canonical fields and legacy fallbacks', () => { + expect( + normalizeEventPlan([ + { event_name: 'signed_up', event_description: 'User signs up' }, + { name: 'invited_user', description: 'User sends an invite' }, + { event: 'created_team' }, + { event_name: 42, name: 'valid_fallback' }, + { event_name: 'x'.repeat(401) }, + { event_name: ' ' }, + { description: 'missing name' }, + ]), + ).toEqual([ + { name: 'signed_up', description: 'User signs up' }, + { name: 'invited_user', description: 'User sends an invite' }, + { name: 'created_team', description: '' }, + { name: 'valid_fallback', description: '' }, + ]); + }); + + it('caps event count and description length', () => { + const events = Array.from({ length: 60 }, (_, index) => ({ + event_name: `event_${index}`, + event_description: 'x'.repeat(5000), + })); + + const normalized = normalizeEventPlan(events); + + expect(normalized).toHaveLength(50); + expect(normalized?.[0].description).toHaveLength(4000); + }); + + it('captures the plan when the file is written after startup', async () => { + const store = createStore(installDir); + const path = join(installDir, EVENT_PLAN_FILE); + watcher = new EventPlanWatcher(store, path, { + pollIntervalMs: 30, + }); + watcher.start(); + + writeFileSync( + path, + JSON.stringify([{ event_name: 'completed_onboarding' }]), + ); + await wait(120); + + expect(store.eventPlan).toEqual([ + { name: 'completed_onboarding', description: '' }, + ]); + }); + + it('captures the first non-empty plan once', async () => { + const store = createStore(installDir); + const path = join(installDir, EVENT_PLAN_FILE); + watcher = new EventPlanWatcher(store, path, { + pollIntervalMs: 30, + }); + watcher.start(); + + writeFileSync(path, JSON.stringify([])); + await wait(40); + expect(store.eventPlan).toEqual([]); + + writeFileSync(path, JSON.stringify([{ event_name: 'first_event' }])); + await wait(40); + writeFileSync(path, JSON.stringify([{ event_name: 'later_event' }])); + watcher.refresh(); + + expect(store.eventPlan).toEqual([{ name: 'first_event', description: '' }]); + }); + + it('keeps the last captured plan after the file is deleted', async () => { + const path = join(installDir, EVENT_PLAN_FILE); + const store = createStore(installDir); + watcher = new EventPlanWatcher(store, path, { + pollIntervalMs: 30, + }); + watcher.start(); + writeFileSync(path, JSON.stringify([{ event_name: 'created_report' }])); + await wait(40); + + unlinkSync(path); + await wait(80); + + expect(store.eventPlan).toEqual([ + { name: 'created_report', description: '' }, + ]); + }); + + it('ignores a plan file that predates the current run', () => { + const path = join(installDir, EVENT_PLAN_FILE); + writeFileSync(path, JSON.stringify([{ event_name: 'stale_event' }])); + const store = createStore(installDir); + watcher = new EventPlanWatcher(store, path); + + watcher.start(); + watcher.refresh(); + + expect(store.eventPlan).toEqual([]); + }); + + it('rejects oversized and symbolic-link plan files', () => { + const path = join(installDir, EVENT_PLAN_FILE); + const store = createStore(installDir); + watcher = new EventPlanWatcher(store, path); + watcher.start(); + + writeFileSync(path, JSON.stringify([{ event_name: 'x'.repeat(300_000) }])); + watcher.refresh(); + expect(store.eventPlan).toEqual([]); + + unlinkSync(path); + const target = join(installDir, 'external-plan.json'); + writeFileSync(target, JSON.stringify([{ event_name: 'linked_event' }])); + symlinkSync(target, path); + watcher.refresh(); + expect(store.eventPlan).toEqual([]); + }); +}); diff --git a/src/programs/task-stream/__tests__/task-stream-push.test.ts b/src/programs/task-stream/__tests__/task-stream-push.test.ts index d39d94da9..cfae3e18b 100644 --- a/src/programs/task-stream/__tests__/task-stream-push.test.ts +++ b/src/programs/task-stream/__tests__/task-stream-push.test.ts @@ -7,6 +7,10 @@ import type { import type { WizardStore, TaskItem } from '@ui/tui/store'; import { TaskStatus } from '@ui/wizard-ui'; import { RunPhase, type PendingQuestion } from '@lib/wizard-session'; +import { mkdtempSync, rmSync, writeFileSync } from 'node:fs'; +import { tmpdir } from 'node:os'; +import { join } from 'node:path'; +import { EVENT_PLAN_FILE } from '@programs/posthog-integration/constants'; type Listener = () => void; @@ -109,6 +113,7 @@ function createPush( opts: { dest?: ReturnType; enabled?: boolean; + eventPlanPath?: string; auditChecks?: () => unknown; } = {}, ) { @@ -117,6 +122,7 @@ function createPush( store, programId: 'test-program', destinations: [dest], + eventPlanPath: opts.eventPlanPath, auditChecks: opts.auditChecks, enabled: opts.enabled, }); @@ -130,6 +136,47 @@ describe('TaskStreamPush', () => { // ── Existing event-sequencing behaviour ──────────────────────── + it('populates the event plan when destination delivery is disabled', async () => { + const installDir = mkdtempSync(join(tmpdir(), 'wizard-headless-plan-')); + const eventPlanPath = join(installDir, EVENT_PLAN_FILE); + const store = createMockStore({ installDir }); + const { push, dest } = createPush(store, { + enabled: false, + eventPlanPath, + }); + + push.attach(); + writeFileSync( + eventPlanPath, + JSON.stringify([{ event_name: 'created_workspace' }]), + ); + await push.shutdown(2000); + + expect(store.eventPlan).toEqual([ + { name: 'created_workspace', description: '' }, + ]); + expect(dest.calls).toHaveLength(0); + + rmSync(installDir, { recursive: true, force: true }); + }); + + it('does not inspect event-plan artifacts unless explicitly configured', () => { + const installDir = mkdtempSync(join(tmpdir(), 'wizard-unrelated-plan-')); + writeFileSync( + join(installDir, EVENT_PLAN_FILE), + JSON.stringify([{ event_name: 'stale_event' }]), + ); + const store = createMockStore({ installDir }); + const { push } = createPush(store); + + push.attach(); + + expect(store.eventPlan).toEqual([]); + + push.detach(); + rmSync(installDir, { recursive: true, force: true }); + }); + describe('event ordering (imperative push)', () => { it('first push sends CREATE', async () => { const store = createMockStore(); @@ -616,23 +663,40 @@ describe('TaskStreamPush', () => { describe('spec: shutdown flushes terminal phase', () => { it('includes the captured event plan in the final Completed push', async () => { + const installDir = mkdtempSync(join(tmpdir(), 'wizard-final-plan-')); + const eventPlanPath = join(installDir, EVENT_PLAN_FILE); const plan = [ { name: 'created_dashboard', description: 'User creates a dashboard' }, ]; - const store = createMockStore({ runPhase: RunPhase.Running }); - const { push, dest } = createPush(store); - - push.attach(); - store._emit(); - await flushMicrotasks(); + const store = createMockStore({ + installDir, + runPhase: RunPhase.Running, + }); + const { push, dest } = createPush(store, { eventPlanPath }); - // The program's plan reaches the store through the host's UI projection. - store.setEventPlan(plan); - store._setAndEmit({ runPhase: RunPhase.Completed }); - await push.shutdown(2000); + try { + push.attach(); + store._emit(); + await flushMicrotasks(); - expect(dest.calls.at(-1)?.[0]).toBe(StreamEvent.Complete); - expect(dest.calls.at(-1)?.[1].event_plan).toEqual({ events: plan }); + writeFileSync( + eventPlanPath, + JSON.stringify([ + { + event_name: plan[0].name, + event_description: plan[0].description, + }, + ]), + ); + store._setAndEmit({ runPhase: RunPhase.Completed }); + await push.shutdown(2000); + + expect(dest.calls.at(-1)?.[0]).toBe(StreamEvent.Complete); + expect(dest.calls.at(-1)?.[1].event_plan).toEqual({ events: plan }); + } finally { + push.detach(); + rmSync(installDir, { recursive: true, force: true }); + } }); it('shutdown awaits one final push when phase is terminal', async () => { diff --git a/src/programs/task-stream/destinations/posthog.ts b/src/programs/task-stream/destinations/posthog.ts index 71e469997..061a0d8a8 100644 --- a/src/programs/task-stream/destinations/posthog.ts +++ b/src/programs/task-stream/destinations/posthog.ts @@ -22,7 +22,7 @@ import type { TaskStreamUpdate, StreamEvent, } from '@programs/task-stream/types'; -import type { Credentials } from '@shared/api'; +import type { Credentials } from '@lib/wizard-session'; import { logToFile } from '@utils/debug'; export interface PostHogDestinationOptions { diff --git a/src/programs/posthog-integration/watch-event-plan.ts b/src/programs/task-stream/event-plan-watcher.ts similarity index 83% rename from src/programs/posthog-integration/watch-event-plan.ts rename to src/programs/task-stream/event-plan-watcher.ts index abfd7c851..d04cee77c 100644 --- a/src/programs/posthog-integration/watch-event-plan.ts +++ b/src/programs/task-stream/event-plan-watcher.ts @@ -1,10 +1,9 @@ +import type { PlannedEvent, WizardStore } from '@ui/tui/store'; import { startFileWatcher, type FileWatcherHandle, type FileWatcherOptions, -} from '@shared/file-watcher'; - -export type PlannedEvent = { name: string; description: string }; +} from '@lib/file-watcher'; const MAX_EVENT_PLAN_FILE_BYTES = 256 * 1024; const MAX_EVENT_COUNT = 50; @@ -41,14 +40,13 @@ export function normalizeEventPlan(parsed: unknown): PlannedEvent[] | null { return events; } -/** Capture the first non-empty event plan emitted by this run. */ -export class ProgramEventPlanWatcher { +export class EventPlanWatcher { private handle: FileWatcherHandle | null = null; private captured = false; constructor( + private readonly store: WizardStore, private readonly path: string, - private readonly onEvents: (events: PlannedEvent[]) => void, private readonly options: FileWatcherOptions = {}, ) {} @@ -62,11 +60,8 @@ export class ProgramEventPlanWatcher { if (!events || events.length === 0) return; this.captured = true; - try { - this.onEvents(events); - } finally { - this.stop(); - } + this.store.setEventPlan(events); + this.stop(); }, { ignoreInitialFile: true, diff --git a/src/programs/task-stream/task-stream-push.ts b/src/programs/task-stream/task-stream-push.ts index d46da705d..d4bce7af9 100644 --- a/src/programs/task-stream/task-stream-push.ts +++ b/src/programs/task-stream/task-stream-push.ts @@ -16,9 +16,14 @@ * latest state once the current one settles. */ -import { RunPhase, TaskStatus } from '@shared/run-state'; -import { OutroKind } from '@agent'; -import type { OutroData, PendingQuestion } from '@agent/types'; +import type { WizardStore, TaskItem } from '@ui/tui/store'; +import { TaskStatus } from '@ui/wizard-ui'; +import { + RunPhase, + OutroKind, + type OutroData, + type PendingQuestion, +} from '@lib/wizard-session'; import { type TaskStreamDestination, type TaskStreamUpdate, @@ -28,7 +33,7 @@ import { StreamTaskStatus, StreamEvent, } from './types'; -import type { PlannedEvent } from '../posthog-integration/watch-event-plan.js'; +import { EventPlanWatcher } from './event-plan-watcher'; import { rollUpAuditAreas } from './audit-areas'; import { logToFile } from '@utils/debug'; import { sanitizeErrorDetail } from '@shared/errors'; @@ -46,9 +51,7 @@ const STATUS_MAP: Record = { [TaskStatus.Skipped]: StreamTaskStatus.Completed, }; -function buildTasks( - items: ReadonlyArray<{ label: string; status: TaskStatus }>, -): StreamTask[] { +function buildTasks(items: TaskItem[]): StreamTask[] { return items.map((item, i) => ({ id: String(i), title: item.label, @@ -107,23 +110,12 @@ function buildPendingInput( }; } -export interface TaskStreamSource { - readonly session: { - skillId: string | null; - runPhase: RunPhase; - outroData: OutroData | null; - pendingQuestion: PendingQuestion | null; - }; - readonly tasks: ReadonlyArray<{ label: string; status: TaskStatus }>; - readonly eventPlan: PlannedEvent[]; - readonly handoffText: string | null; - subscribe(callback: () => void): () => void; -} - export interface TaskStreamPushOptions { - store: TaskStreamSource; + store: WizardStore; programId: string; destinations: TaskStreamDestination[]; + /** Optional absolute event-plan path to load into the store once. */ + eventPlanPath?: string; /** The run's audit ledger, when it has one. The runner owns the watcher. */ auditChecks?: () => unknown; /** When false, destination subscription/delivery remains disabled. */ @@ -131,11 +123,12 @@ export interface TaskStreamPushOptions { } export class TaskStreamPush { - private readonly store: TaskStreamSource; + private readonly store: WizardStore; private readonly destinations: TaskStreamDestination[]; private readonly startedAt: string; private readonly programId: string; private readonly sessionId: string; + private readonly eventPlanWatcher: EventPlanWatcher | null; private readonly auditChecks: (() => unknown) | null; private enabled: boolean; @@ -155,6 +148,9 @@ export class TaskStreamPush { this.destinations = opts.destinations; this.enabled = opts.enabled ?? true; const startedAt = new Date(); + this.eventPlanWatcher = opts.eventPlanPath + ? new EventPlanWatcher(this.store, opts.eventPlanPath) + : null; this.auditChecks = opts.auditChecks ?? null; this.startedAt = secondPrecisionIso(startedAt); // skillId may not be set yet — fall back to programId so the @@ -166,8 +162,13 @@ export class TaskStreamPush { this.sessionId = `${this.programId}-${skillId}-${this.startedAt}`; } - /** Subscribe to store changes, unless destination delivery is disabled. */ - attach(store?: TaskStreamSource): void { + /** + * Load the event plan and subscribe to store changes. Destination delivery + * remains disabled when `enabled === false`, but the plan still populates the + * store for local and headless consumers. + */ + attach(store?: WizardStore): void { + this.eventPlanWatcher?.start(); if (!this.enabled) return; if (this.unsubscribe) return; const target = store ?? this.store; @@ -176,6 +177,7 @@ export class TaskStreamPush { /** Stop subscribing. Does not flush. */ detach(): void { + this.eventPlanWatcher?.stop(); if (this.unsubscribe) { this.unsubscribe(); this.unsubscribe = null; @@ -195,6 +197,7 @@ export class TaskStreamPush { timeoutMs: number = DEFAULT_SHUTDOWN_TIMEOUT_MS, ): Promise { this.shuttingDown = true; + this.eventPlanWatcher?.refresh(); if (this.debounceTimer) { clearTimeout(this.debounceTimer); this.debounceTimer = null; diff --git a/src/programs/task-stream/types.ts b/src/programs/task-stream/types.ts index d8d38b42f..373e8a587 100644 --- a/src/programs/task-stream/types.ts +++ b/src/programs/task-stream/types.ts @@ -12,7 +12,7 @@ * on TaskStreamPush but serialised to `workflow_id` here. */ -import type { RunPhase } from '@shared/run-state'; +import type { RunPhase } from '@lib/wizard-session'; export enum StreamTaskStatus { Pending = 'pending', diff --git a/src/programs/token-refresh.ts b/src/programs/token-refresh.ts deleted file mode 100644 index 0909cb3e5..000000000 --- a/src/programs/token-refresh.ts +++ /dev/null @@ -1,63 +0,0 @@ -/** The pre-run OAuth token refresh, owned by programs and free of any session. */ - -import type { Credentials } from '@shared/api'; -import { markGrantRevoked } from '@shared/auth-session-state'; -import { analytics } from '@utils/analytics'; -import { logToFile } from '@utils/debug'; -import { OAuthError } from '@utils/oauth-errors'; -import { refreshAccessToken } from '@utils/oauth-token'; - -// Below this remaining lifetime a run risks outliving its token; just-minted and 7-day tokens skip. -// Only a second agent run in one invocation can be this old — see self-driving's chained phases. -const REFRESH_WHEN_REMAINING_MS = 50 * 60 * 1000; - -/** - * Grants the token endpoint refuses permanently. A dead grant means the login - * is gone, not that the network blipped, so only these mark the session. - */ -const DEAD_GRANT_CODES = new Set(['invalid_grant', 'invalid_client']); - -/** Best-effort pre-run refresh; the same object comes back unless the token was refreshed. */ -export async function refreshCredentialsIfNeeded( - credentials: Credentials, - options: { baseUrl?: string }, -): Promise { - if (!credentials.refreshToken) return credentials; - - // No expiry means we cannot tell how much life is left, so leave it alone — - // refreshing every run would spend a rotation for nothing. - if (credentials.expiresAt === undefined) return credentials; - if (credentials.expiresAt - Date.now() >= REFRESH_WHEN_REMAINING_MS) { - return credentials; - } - - try { - const token = await refreshAccessToken( - credentials.refreshToken, - options.baseUrl, - credentials.oauthClientId, - ); - // Replaced, not mutated: readers hold this object, and a new one keeps the - // store and the (possibly shallow-copied) session explicitly in step. - return { - ...credentials, - accessToken: token.access_token, - // Rotation: keep the returned refresh token or the old one stops working. - refreshToken: token.refresh_token ?? credentials.refreshToken, - expiresAt: Date.now() + token.expires_in * 1000, - }; - } catch (error) { - // A dead grant is recorded but not thrown: the current token may still have - // minutes of life, and failing here would break runs that would have worked. - // If a 401 does follow, the auth-error screen can finally name the cause. - if (error instanceof OAuthError && DEAD_GRANT_CODES.has(error.code)) { - markGrantRevoked(); - analytics.wizardCapture('auth session expired', { reason: error.code }); - } - logToFile( - '[oauth] pre-run token refresh failed, continuing with the existing token:', - error instanceof Error ? error.message : error, - ); - return credentials; - } -} diff --git a/src/programs/types.ts b/src/programs/types.ts index 074ff8e8e..cfa82b2d0 100644 --- a/src/programs/types.ts +++ b/src/programs/types.ts @@ -6,7 +6,6 @@ export type { ProgramReadyContext, StoreInitContext, } from './program-step'; -export type { ProgramCompletionContext } from './program-run'; export type { FrameworkConfig, SetupQuestion } from './framework-config'; export type { ProgramCiHost, ProgramRunHost } from './host-capabilities'; export type { @@ -25,4 +24,3 @@ export type { ProgramRunProgress, SettledProgramRun, } from './program-store'; -export type { ProgramSwitchboardCtx } from './binding'; diff --git a/src/programs/warehouse-source/detect.ts b/src/programs/warehouse-source/detect.ts index 5d2ce8c31..0fa74c37e 100644 --- a/src/programs/warehouse-source/detect.ts +++ b/src/programs/warehouse-source/detect.ts @@ -8,6 +8,7 @@ import { existsSync, statSync } from 'fs'; import { analytics } from '@utils/analytics'; +import type { WizardSession } from '@lib/wizard-session'; import type { AbortCase } from '@agent/types'; import { detectWarehouseSources } from '@programs/warehouse-sources/detect'; import type { DetectedSource } from '@programs/warehouse-sources/types'; @@ -28,9 +29,9 @@ export const DETECTED_WAREHOUSE_SOURCES_KEY = 'detectedWarehouseSources'; * Read the detected sources out of frameworkContext. Single accessor shared by * the intro screen and the prompt builder so the key + cast live in one place. */ -export function getDetectedWarehouseSources(session: { - frameworkContext: Record; -}): DetectedSource[] { +export function getDetectedWarehouseSources( + session: WizardSession, +): DetectedSource[] { return ( (session.frameworkContext[DETECTED_WAREHOUSE_SOURCES_KEY] as | DetectedSource[] @@ -68,7 +69,7 @@ export const WAREHOUSE_ABORT_CASES: AbortCase[] = [ * sources (or a `detectError`) into frameworkContext for the intro screen. */ export function detectWarehousePrerequisites( - session: { installDir: string }, + session: WizardSession, setFrameworkContext: (key: string, value: unknown) => void, ): void { const fail = (error: WarehouseDetectError) => diff --git a/src/programs/warehouse-source/index.ts b/src/programs/warehouse-source/index.ts index 42f19de90..fb23dd28c 100644 --- a/src/programs/warehouse-source/index.ts +++ b/src/programs/warehouse-source/index.ts @@ -1,20 +1,20 @@ import type { ProgramConfig } from '@programs/program-step'; import type { ProgramRun } from '@programs/program-run'; -import { LONGER_ASK_TIMEOUT_MS } from '@shared/ask-policy'; +import type { WizardSession } from '@lib/wizard-session'; +import { LONGER_ASK_TIMEOUT_MS } from '@agent'; import { WAREHOUSE_SOURCE_PROGRAM } from './steps.js'; import { WAREHOUSE_ABORT_CASES, getDetectedWarehouseSources, } from './detect.js'; - -type WarehouseRunState = Parameters[0]; +import { getContentBlocks } from '../../ui/tui/decks/warehouse-source/index.js'; /** * Inject the detected sources (and their creation mode) into the prompt so the * skill knows what to set up. The *how* — in-CLI creation vs deep-link, field * collection, validation — lives in the skill, not here. */ -function buildPrompt(session: WarehouseRunState): string { +function buildPrompt(session: WizardSession): string { const sources = getDetectedWarehouseSources(session); if (sources.length === 0) { return 'Set up a data warehouse source for this project.'; @@ -47,9 +47,10 @@ export const warehouseSourceConfig: ProgramConfig = { id: 'warehouse-source', skillId: 'data-warehouse-source-setup', steps: WAREHOUSE_SOURCE_PROGRAM, + getContentBlocks, reportFile: 'posthog-warehouse-report.md', allowedTools: ['Agent'], - run: (session: WarehouseRunState): Promise => + run: (session: WizardSession): Promise => Promise.resolve({ skillId: 'data-warehouse-source-setup', integrationLabel: 'data-warehouse-source-setup', diff --git a/src/programs/warehouse-source/steps.ts b/src/programs/warehouse-source/steps.ts index ae5431091..0a8ed69db 100644 --- a/src/programs/warehouse-source/steps.ts +++ b/src/programs/warehouse-source/steps.ts @@ -7,7 +7,7 @@ */ import type { ProgramStep } from '@programs/program-step'; -import { RunPhase } from '@shared/run-state'; +import { RunPhase } from '@lib/wizard-session'; import { detectWarehousePrerequisites } from './detect.js'; export const WAREHOUSE_SOURCE_PROGRAM: ProgramStep[] = [ diff --git a/src/programs/web-analytics-doctor/abort-cases.ts b/src/programs/web-analytics-doctor/abort-cases.ts deleted file mode 100644 index 7b5fbd129..000000000 --- a/src/programs/web-analytics-doctor/abort-cases.ts +++ /dev/null @@ -1,34 +0,0 @@ -import type { AbortCase } from '@agent/types'; -import { ErrorCodes } from '@shared/errors'; - -export const WEB_ANALYTICS_ABORT_CASES: AbortCase[] = [ - { - match: /^no web analytics events$/i, - message: 'No web analytics events', - body: - 'The doctor found no $pageview events in the last 30 days, so there is ' + - 'nothing to audit yet. Make sure PostHog is initialized and capturing ' + - 'pageviews, then run the doctor again.', - docsUrl: 'https://posthog.com/docs/web-analytics/getting-started', - }, - { - match: /^insufficient permissions$/i, - errorCode: ErrorCodes.AuthMissingScope, - message: 'Insufficient permissions', - body: - 'The doctor could not query your project — the authenticated token is ' + - 'missing query access. Re-run the wizard to sign in again, or use a key ' + - 'with read access to your events.', - docsUrl: 'https://posthog.com/docs/web-analytics', - }, - { - match: /^posthog sdk not installed$/i, - errorCode: ErrorCodes.DetectNoPosthogSdk, - message: 'PostHog SDK not installed', - body: - 'The doctor could not find a PostHog SDK in this project. Install and ' + - 'configure PostHog first (run `npx @posthog/wizard`), then run the ' + - 'doctor to check your web analytics setup.', - docsUrl: 'https://posthog.com/docs/libraries/js', - }, -]; diff --git a/src/programs/web-analytics-doctor/detect.ts b/src/programs/web-analytics-doctor/detect.ts index d89d9ca82..e7df63673 100644 --- a/src/programs/web-analytics-doctor/detect.ts +++ b/src/programs/web-analytics-doctor/detect.ts @@ -1,4 +1,7 @@ import { existsSync, statSync } from 'fs'; +import type { WizardSession } from '@lib/wizard-session'; +import type { AbortCase } from '@agent/types'; +import { ErrorCodes } from '@shared/errors'; import { findPackageJsons } from '@programs/shared/package-scanning'; export type WebAnalyticsDetectError = @@ -10,10 +13,40 @@ export type WebAnalyticsDetectError = | { kind: 'no-package-json' } | { kind: 'no-posthog'; scannedCount: number }; -export { WEB_ANALYTICS_ABORT_CASES } from './abort-cases.js'; +export const WEB_ANALYTICS_ABORT_CASES: AbortCase[] = [ + { + match: /^no web analytics events$/i, + message: 'No web analytics events', + body: + 'The doctor found no $pageview events in the last 30 days, so there is ' + + 'nothing to audit yet. Make sure PostHog is initialized and capturing ' + + 'pageviews, then run the doctor again.', + docsUrl: 'https://posthog.com/docs/web-analytics/getting-started', + }, + { + match: /^insufficient permissions$/i, + errorCode: ErrorCodes.AuthMissingScope, + message: 'Insufficient permissions', + body: + 'The doctor could not query your project — the authenticated token is ' + + 'missing query access. Re-run the wizard to sign in again, or use a key ' + + 'with read access to your events.', + docsUrl: 'https://posthog.com/docs/web-analytics', + }, + { + match: /^posthog sdk not installed$/i, + errorCode: ErrorCodes.DetectNoPosthogSdk, + message: 'PostHog SDK not installed', + body: + 'The doctor could not find a PostHog SDK in this project. Install and ' + + 'configure PostHog first (run `npx @posthog/wizard`), then run the ' + + 'doctor to check your web analytics setup.', + docsUrl: 'https://posthog.com/docs/libraries/js', + }, +]; export function detectWebAnalyticsPrerequisites( - session: { installDir: string }, + session: WizardSession, setFrameworkContext: (key: string, value: unknown) => void, ): void { const fail = (error: WebAnalyticsDetectError) => diff --git a/src/programs/web-analytics-doctor/index.ts b/src/programs/web-analytics-doctor/index.ts index a7c2863ce..4d11413c2 100644 --- a/src/programs/web-analytics-doctor/index.ts +++ b/src/programs/web-analytics-doctor/index.ts @@ -1,10 +1,33 @@ import type { ProgramConfig } from '@programs/program-step'; import { createSkillProgram } from '../agent-skill/index.js'; import { WEB_ANALYTICS_DOCTOR_PROGRAM } from './steps.js'; -import { WEB_ANALYTICS_DOCTOR_OPTIONS } from './run.js'; +import { WEB_ANALYTICS_ABORT_CASES } from './detect.js'; + +const REPORT_FILE = 'posthog-web-analytics-report.md'; +const DOCS_URL = 'https://posthog.com/docs/web-analytics'; export const webAnalyticsDoctorConfig: ProgramConfig = { - ...createSkillProgram(WEB_ANALYTICS_DOCTOR_OPTIONS), + ...createSkillProgram({ + skillId: 'web-analytics-doctor', + command: 'web-analytics', + id: 'web-analytics-doctor', + description: 'Audit and fix your PostHog web analytics setup', + integrationLabel: 'web-analytics-doctor', + customPrompt: + "Run the web-analytics-doctor skill to check this project's PostHog web " + + 'analytics setup. Audit read-only first, then present the findings to the ' + + 'user with a single wizard_ask multi-select and apply only the fixes they ' + + 'choose — editing project code and/or PostHog project settings via the ' + + 'MCP — before writing the report.', + successMessage: + 'Web analytics check complete! You can view the report at ./posthog-web-analytics-report.md', + reportFile: REPORT_FILE, + docsUrl: DOCS_URL, + spinnerMessage: 'Checking your web analytics setup...', + estimatedDurationMinutes: 5, + requires: ['posthog-integration'], + abortCases: WEB_ANALYTICS_ABORT_CASES, + }), steps: WEB_ANALYTICS_DOCTOR_PROGRAM, parentCommand: 'audit', }; diff --git a/src/programs/web-analytics-doctor/run.ts b/src/programs/web-analytics-doctor/run.ts deleted file mode 100644 index cd6bf6b44..000000000 --- a/src/programs/web-analytics-doctor/run.ts +++ /dev/null @@ -1,27 +0,0 @@ -import type { SkillProgramOptions } from '@programs/agent-skill/run-definition'; -import { WEB_ANALYTICS_ABORT_CASES } from './abort-cases.js'; - -const REPORT_FILE = 'posthog-web-analytics-report.md'; -const DOCS_URL = 'https://posthog.com/docs/web-analytics'; - -export const WEB_ANALYTICS_DOCTOR_OPTIONS: SkillProgramOptions = { - skillId: 'web-analytics-doctor', - command: 'web-analytics', - id: 'web-analytics-doctor', - description: 'Audit and fix your PostHog web analytics setup', - integrationLabel: 'web-analytics-doctor', - customPrompt: - "Run the web-analytics-doctor skill to check this project's PostHog web " + - 'analytics setup. Audit read-only first, then present the findings to the ' + - 'user with a single wizard_ask multi-select and apply only the fixes they ' + - 'choose — editing project code and/or PostHog project settings via the ' + - 'MCP — before writing the report.', - successMessage: - 'Web analytics check complete! You can view the report at ./posthog-web-analytics-report.md', - reportFile: REPORT_FILE, - docsUrl: DOCS_URL, - spinnerMessage: 'Checking your web analytics setup...', - estimatedDurationMinutes: 5, - requires: ['posthog-integration'], - abortCases: WEB_ANALYTICS_ABORT_CASES, -}; diff --git a/src/shared/__tests__/claude-settings-backup.test.ts b/src/shared/__tests__/claude-settings-backup.test.ts index 884963cbf..5f5085f55 100644 --- a/src/shared/__tests__/claude-settings-backup.test.ts +++ b/src/shared/__tests__/claude-settings-backup.test.ts @@ -1,7 +1,6 @@ import * as fs from 'fs'; import * as os from 'os'; import * as path from 'path'; -import { clearCleanup, runCleanups } from '@utils/cleanup-registry'; const { captureException, wizardCapture } = vi.hoisted(() => ({ captureException: vi.fn(), @@ -36,14 +35,12 @@ describe('claude settings backup/restore', () => { let claudeDir: string; beforeEach(() => { - clearCleanup(); captureException.mockClear(); wizardCapture.mockClear(); ({ dir, claudeDir } = makeProject()); }); afterEach(() => { - clearCleanup(); fs.rmSync(dir, { recursive: true, force: true }); }); @@ -63,16 +60,6 @@ describe('claude settings backup/restore', () => { expect(captureException).not.toHaveBeenCalled(); }); - it('registers restoration with the process cleanup registry', () => { - fs.writeFileSync(path.join(claudeDir, SETTINGS), '{"apiKeyHelper":"x"}'); - - expect(backupAndFixClaudeSettings(dir)).toBe(true); - runCleanups(); - - expect(read(SETTINGS)).toBe('{"apiKeyHelper":"x"}'); - expect(exists(BACKUP)).toBe(false); - }); - it('returns false and stays silent when there is nothing to back up', () => { const ok = backupAndFixClaudeSettings(dir); diff --git a/src/shared/__tests__/skill-run-cleanup.test.ts b/src/shared/__tests__/skill-run-cleanup.test.ts deleted file mode 100644 index 8cde918c4..000000000 --- a/src/shared/__tests__/skill-run-cleanup.test.ts +++ /dev/null @@ -1,38 +0,0 @@ -import fs from 'node:fs'; -import os from 'node:os'; -import path from 'node:path'; -import { - commitRegisteredRunSkillCleanups, - registerRunSkillCleanup, -} from '../skill-run-cleanup'; -import { - clearCleanup, - registerCleanup, - runCleanups, -} from '@utils/cleanup-registry'; - -it('commits every registered skill directory without disarming unrelated cleanup', () => { - const root = fs.mkdtempSync(path.join(os.tmpdir(), 'wizard-cleanup-')); - const extraDir = path.join(root, 'nested-project'); - const unrelated = vi.fn(); - try { - const skillDirs = [root, extraDir].map((installDir) => { - registerRunSkillCleanup(installDir); - const skillDir = path.join(installDir, '.claude', 'skills', 'installed'); - fs.mkdirSync(skillDir, { recursive: true }); - fs.writeFileSync(path.join(skillDir, '.posthog-wizard'), ''); - return skillDir; - }); - registerCleanup(unrelated); - - commitRegisteredRunSkillCleanups(); - runCleanups(); - - expect(skillDirs.every((skillDir) => fs.existsSync(skillDir))).toBe(true); - expect(unrelated).toHaveBeenCalledOnce(); - } finally { - commitRegisteredRunSkillCleanups(); - clearCleanup(); - fs.rmSync(root, { recursive: true, force: true }); - } -}); diff --git a/src/shared/ask-policy.ts b/src/shared/ask-policy.ts deleted file mode 100644 index 5266e499c..000000000 --- a/src/shared/ask-policy.ts +++ /dev/null @@ -1,31 +0,0 @@ -/** When a run may put a `wizard_ask` question to a person, and how long it waits. */ - -/** - * Whether the `wizard_ask` overlay stays unwired for this run. Non-interactive - * modes (CI, signup) have no human to answer. Per-program disabling adds - * WIZARD_ASK_TOOL_NAME to the program's `disallowedTools` instead, so the SDK - * rejects calls outright. - * - * `e2eAsk` is the one escape hatch. The e2e harness runs a `ci` session, but it - * does have an answerer: the driver loop answers each `wizard_ask` batch from - * the program's e2e profile. Without the flag the agent-in-the-loop layer (the - * ask bridge in both sequence arms, and the orchestrator's seeded warehouse - * task) stays unreachable from a test. - * - * Only the e2e TUI host sets the flag, from the `E2E_ASK` env var. No CLI flag - * populates it, so plain `--ci` and `--signup` runs keep the overlay unwired. - */ -export function isAskDisabled(flags: { - ci: boolean; - signup: boolean; - e2eAsk: boolean; -}): boolean { - return (flags.ci || flags.signup) && !flags.e2eAsk; -} - -/** - * The longer per-question timeout, for asks that send the user on an errand — - * open a database console, mint a restricted API key. The default is sized for - * a question answerable from memory and expires long before an errand is done. - */ -export const LONGER_ASK_TIMEOUT_MS = 20 * 60 * 1000; diff --git a/src/shared/ci-gateway-auth.ts b/src/shared/ci-gateway-auth.ts deleted file mode 100644 index 96b1d4b79..000000000 --- a/src/shared/ci-gateway-auth.ts +++ /dev/null @@ -1,27 +0,0 @@ -/** Build a fixed CI bearer without mutating a gateway mint session. */ -import { IS_PRODUCTION_BUILD } from '@env'; -import { isTrustedGatewayUrl, type GatewayAuth } from './gateway-auth'; - -export function createCiGatewayAuth( - token: string, - projectId: number, - gatewayUrl: string, -): GatewayAuth { - if (IS_PRODUCTION_BUILD) - throw new Error('CI gateway auth requires a non-production build'); - if (!token.trim() || !Number.isSafeInteger(projectId) || projectId <= 0) { - throw new Error('CI gateway auth requires a token and valid project ID'); - } - if ( - !/^https?:\/\//.test(gatewayUrl) || - !isTrustedGatewayUrl(gatewayUrl, '') - ) { - throw new Error('CI gateway auth requires a trusted gateway origin'); - } - return { - token: token.trim(), - teamId: projectId, - gatewayUrl: gatewayUrl.replace(/\/+$/, ''), - refreshAtMs: Infinity, - }; -} diff --git a/src/shared/claude-settings.ts b/src/shared/claude-settings.ts index bff862005..43da0b418 100644 --- a/src/shared/claude-settings.ts +++ b/src/shared/claude-settings.ts @@ -11,7 +11,7 @@ import path from 'path'; import * as fs from 'fs'; import * as os from 'os'; import { analytics } from '@utils/analytics'; -import { registerCleanup } from '@utils/cleanup-registry'; +import { registerCleanup } from '@utils/wizard-abort'; import { BLOCKED_AGENT_ENV_KEYS, BLOCKED_AGENT_ENV_PATTERNS, diff --git a/src/shared/errors/__tests__/run-failure.test.ts b/src/shared/errors/__tests__/run-failure.test.ts index 8c82dbbf4..9682327b6 100644 --- a/src/shared/errors/__tests__/run-failure.test.ts +++ b/src/shared/errors/__tests__/run-failure.test.ts @@ -2,6 +2,7 @@ import { describe, expect, it } from 'vitest'; import { classifyRunFailure } from '../run-failure'; import { ErrorCodes } from '../codes'; import { WizardError } from '@utils/wizard-abort'; +import { GatewayMintRefused } from '@agent/gateway-session'; vi.mock('@utils/analytics', () => ({ analytics: { wizardCapture: vi.fn(), captureException: vi.fn() }, @@ -11,11 +12,7 @@ describe('classifyRunFailure', () => { it('keeps a mint refusal as its own code and message', () => { // The runners print this message alone, without the unhandled framing. const failure = classifyRunFailure( - new WizardError( - 'This account is blocked.', - { status: 403, outcome: 'blocked' }, - ErrorCodes.GatewayMintRefused, - ), + new GatewayMintRefused(403, 'This account is blocked.', 'blocked'), ); expect(failure).toEqual({ code: ErrorCodes.GatewayMintRefused, diff --git a/src/shared/gateway-auth.ts b/src/shared/gateway-auth.ts deleted file mode 100644 index f673c5eb2..000000000 --- a/src/shared/gateway-auth.ts +++ /dev/null @@ -1,99 +0,0 @@ -/** Gateway bearer types and checks shared by the agent and its hosts. */ - -export interface GatewayAuth { - /** Base URL for model calls (no `/v1`; transports append their route). */ - gatewayUrl: string; - /** Gateway bearer, minted normally or supplied directly by CI. */ - token: string; - /** Team verified by the mint, or explicitly supplied for CI attribution. */ - teamId?: number; - /** - * Instant past which a 401 on this bearer is age rather than a bad - * credential: the cache re-mints past it, and a session still holding the - * old bearer may re-mint once. Before it the mint has to be trusted. - */ - refreshAtMs: number; -} - -/** Whether a 401 on this bearer may be age (past its refresh instant) rather than a bad credential. */ -export function isPastRefresh(auth: GatewayAuth, now = Date.now()): boolean { - return now >= auth.refreshAtMs; -} - -/** - * Whether a server-supplied origin may receive a bearer and prompt content: - * https (loopback excepted), and either a current cloud gateway or the host the run - * authenticated against. - */ -export function isTrustedGatewayUrl(value: string, apiHost: string): boolean { - let url: URL; - try { - url = new URL(value); - } catch { - return false; - } - // Consumers append routes to this value, so anything beyond an origin - // (path, query, fragment, userinfo) would build a malformed endpoint. - if ( - url.pathname !== '/' || - url.search || - url.hash || - url.username || - url.password - ) { - return false; - } - const localhost = - url.hostname === 'localhost' || - url.hostname === '127.0.0.1' || - url.hostname === 'host.docker.internal'; - // Loopback is the dev gateway, and is the one case allowed over http. - if (localhost) return true; - if (url.protocol !== 'https:') return false; - if (url.hostname.endsWith('.posthog.com')) { - return ( - url.origin === 'https://ai-gateway.us.posthog.com' || - url.origin === 'https://ai-gateway.eu.posthog.com' - ); - } - try { - return url.hostname === new URL(apiHost).hostname; - } catch { - return false; - } -} - -/** - * The v2 run-metadata carrier: one JSON blob for the `X-PostHog-Properties` - * header. Plain keys only, since the gateway strips `$`-prefixed keys as reserved, - * so feature-flag variants land as `wizard_flag_` instead of the legacy - * `$feature/` (dashboards keying on `$feature/wizard-*` read the new key - * post-cutover). - */ -export function buildWizardPropertiesBlob( - wizardMetadata: Record, - wizardFlags: Record, - teamId?: number, -): string { - // The gateway pins `$ai_product` to `wizard:`, and rejects a legacy - // product override on a scoped token, so the unprefixed key every cost and - // error consumer reads is only present if this blob declares it. - const props: Record = { ai_product: 'wizard' }; - if (teamId !== undefined) props.team_id = teamId; - for (const [key, value] of Object.entries(wizardMetadata)) { - props[stripPropertyPrefix(key)] = value; - } - for (const [flagKey, variant] of Object.entries(wizardFlags)) { - if (!flagKey.toLowerCase().startsWith('wizard')) continue; - props[`wizard_flag_${flagKey.toLowerCase()}`] = variant; - } - return JSON.stringify(props); -} - -const LEGACY_PROPERTY_PREFIX = 'X-POSTHOG-PROPERTY-'; - -function stripPropertyPrefix(key: string): string { - return key.toUpperCase().startsWith(LEGACY_PROPERTY_PREFIX) - ? key.slice(LEGACY_PROPERTY_PREFIX.length).toLowerCase() - : key; -} diff --git a/src/shared/posthog-cli-install.ts b/src/shared/posthog-cli-install.ts deleted file mode 100644 index 10854d1dd..000000000 --- a/src/shared/posthog-cli-install.ts +++ /dev/null @@ -1,46 +0,0 @@ -import { spawnSync } from 'node:child_process'; - -import { debug } from '@utils/debug'; - -export interface CliInstallResult { - success: boolean; - error?: string; - /** The spawn failure or a synthesized non-zero-exit error. */ - errorObject?: Error; -} - -export const cliSpawnOptions = { - encoding: 'utf-8' as const, - // npm/posthog-cli are npm.cmd/posthog-cli.cmd on Windows. - shell: process.platform === 'win32', -}; - -/** Install or update the PostHog CLI with npm in the user's environment. */ -export function installOrUpdatePostHogCli(): CliInstallResult { - const args = ['install', '--global', '@posthog/cli@latest']; - debug(`Running npm ${args.join(' ')}`); - - const result = spawnSync('npm', args, cliSpawnOptions); - - if (result.error) { - return { - success: false, - error: `Failed to run npm: ${result.error.message}. Is Node.js installed?`, - errorObject: result.error, - }; - } - if (result.status !== 0) { - const detail = (result.stderr || result.stdout || '').trim(); - const message = - detail || - `npm install --global @posthog/cli@latest exited with status ${ - result.status ?? 'unknown' - }`; - return { - success: false, - error: message, - errorObject: new Error(message), - }; - } - return { success: true }; -} diff --git a/src/shared/run-state.ts b/src/shared/run-state.ts deleted file mode 100644 index c148bff79..000000000 --- a/src/shared/run-state.ts +++ /dev/null @@ -1,27 +0,0 @@ -/** Lifecycle phase of the main program work (agent run, MCP install, etc.). */ -export enum RunPhase { - Idle = 'idle', - Running = 'running', - Completed = 'completed', - Error = 'error', -} - -/** Outcome of the MCP server installation program step. */ -export enum McpOutcome { - NoClients = 'no_clients', - Skipped = 'skipped', - Installed = 'installed', - Failed = 'failed', -} - -/** Task state shared by program progress and its TUI projection. */ -export enum TaskStatus { - Pending = 'pending', - InProgress = 'in_progress', - Completed = 'completed', - Skipped = 'skipped', -} - -export function isTaskStatus(value: string): value is TaskStatus { - return (Object.values(TaskStatus) as string[]).includes(value); -} diff --git a/src/shared/run-tags.ts b/src/shared/run-tags.ts deleted file mode 100644 index 144d4b638..000000000 --- a/src/shared/run-tags.ts +++ /dev/null @@ -1,26 +0,0 @@ -import { CallType } from '@shared/constants'; - -/** - * Global identifiers attached to every LLM gateway trace for a run. They ride on - * each `$ai_generation` the gateway emits (in the `X-PostHog-Properties` blob - * `buildAgentEnv` builds), so traces are filterable by program, framework, run, - * and build type for cost attribution and dashboards. `skill_id` is omitted when - * the run has none. - */ -export function buildRunTags(args: { - programId: string; - integration: string; - runId: string; - build: string; - skillId?: string; -}): Record { - return { - program_id: args.programId, - integration: args.integration, - run_id: args.runId, - build: args.build, - // Triage and detection spread these tags and override this one. - call_type: CallType.agent, - ...(args.skillId ? { skill_id: args.skillId } : {}), - }; -} diff --git a/src/shared/scan-consent.ts b/src/shared/scan-consent.ts deleted file mode 100644 index 62c7227af..000000000 --- a/src/shared/scan-consent.ts +++ /dev/null @@ -1,33 +0,0 @@ -/** Features discovered by the feature-discovery subagent */ -export enum DiscoveredFeature { - Stripe = 'stripe', - LLM = 'llm', -} - -/** Consent to report what local detection found. */ -export enum ScanConsent { - Undecided = 'undecided', - Granted = 'granted', - Declined = 'declined', -} - -type ScanConsentState = { scanConsent: string }; - -/** One place to ask, so a new consent state does not need three edits. */ -export function mayReportScanResults(session: ScanConsentState): boolean { - return session.scanConsent === 'granted'; -} - -/** Lives here so analytics infrastructure never learns what consent means. */ -export function reportableDiscoveredFeatures( - session: ScanConsentState & { discoveredFeatures: DiscoveredFeature[] }, -): DiscoveredFeature[] | undefined { - return mayReportScanResults(session) ? session.discoveredFeatures : undefined; -} - -/** Also a scan result, so it travels under the same consent as the rest. */ -export function reportablePosthogSdkDetected( - session: ScanConsentState & { posthogSdkDetected: boolean }, -): boolean | undefined { - return mayReportScanResults(session) ? session.posthogSdkDetected : undefined; -} diff --git a/src/shared/skill-run-cleanup.ts b/src/shared/skill-run-cleanup.ts deleted file mode 100644 index 1d986db40..000000000 --- a/src/shared/skill-run-cleanup.ts +++ /dev/null @@ -1,62 +0,0 @@ -import { lstatSync, readdirSync, rmSync } from 'node:fs'; -import { join } from 'node:path'; -import { logToFile } from '@utils/debug'; -import { registerCleanup } from '@utils/cleanup-registry'; - -export type RunSkillCleanup = (() => void) & { commit: () => void }; -const registeredSkillCleanups = new Set(); - -/** An absent directory is an empty snapshot; a symlink is never a skill root. */ -function skillEntries(root: string) { - try { - if (!lstatSync(root).isDirectory()) return null; - return readdirSync(root, { withFileTypes: true }); - } catch (error) { - if ((error as NodeJS.ErrnoException).code === 'ENOENT') return []; - throw error; - } -} - -/** Preserve every entry that existed before the run, including older Wizard installs. */ -export function captureRunSkillCleanup(installDir: string): RunSkillCleanup { - const root = join(installDir, '.claude', 'skills'); - const before = skillEntries(root); - const preexisting = new Set(before?.map((entry) => entry.name)); - let committed = false; - - const cleanup = (() => { - registeredSkillCleanups.delete(cleanup); - if (committed || !before) return; - const current = skillEntries(root); - if (!current) return; - for (const entry of current) { - if (!entry.isDirectory() || preexisting.has(entry.name)) continue; - const skillDir = join(root, entry.name); - try { - if (!lstatSync(join(skillDir, '.posthog-wizard')).isFile()) continue; - } catch (error) { - if ((error as NodeJS.ErrnoException).code === 'ENOENT') continue; - throw error; - } - rmSync(skillDir, { recursive: true, force: true }); - logToFile(`[agent-runner] removed failed-run skill ${entry.name}`); - } - }) as RunSkillCleanup; - cleanup.commit = () => { - committed = true; - registeredSkillCleanups.delete(cleanup); - }; - return cleanup; -} - -export function registerRunSkillCleanup(installDir: string): RunSkillCleanup { - const cleanup = captureRunSkillCleanup(installDir); - registeredSkillCleanups.add(cleanup); - registerCleanup(cleanup); - return cleanup; -} - -/** Disarm only skill callbacks; other abort cleanup remains registered. */ -export function commitRegisteredRunSkillCleanups(): void { - for (const cleanup of registeredSkillCleanups) cleanup.commit(); -} diff --git a/src/shared/utils/__tests__/environment.test.ts b/src/shared/utils/__tests__/environment.test.ts index 933862b01..0830d12c5 100644 --- a/src/shared/utils/__tests__/environment.test.ts +++ b/src/shared/utils/__tests__/environment.test.ts @@ -10,7 +10,7 @@ import { readEnvironment } from '@utils/environment'; import { buildSession } from '@lib/wizard-session'; -import { isAskDisabled } from '@shared/ask-policy'; +import { shouldDisableAsk } from '@agent/agent-runner'; /** Every var this file sets, cleared between cases. */ const TOUCHED = [ @@ -61,6 +61,6 @@ describe('readEnvironment', () => { ...readEnvironment(), }); expect(session.e2eAsk).toBe(false); - expect(isAskDisabled(session)).toBe(true); + expect(shouldDisableAsk(session)).toBe(true); }); }); diff --git a/src/shared/utils/__tests__/oauth-refresh.test.ts b/src/shared/utils/__tests__/oauth-refresh.test.ts index 4e0ad8891..525ea9400 100644 --- a/src/shared/utils/__tests__/oauth-refresh.test.ts +++ b/src/shared/utils/__tests__/oauth-refresh.test.ts @@ -1,5 +1,5 @@ import axios from 'axios'; -import { refreshAccessToken } from '@utils/oauth-token'; +import { refreshAccessToken } from '@utils/oauth'; import { POSTHOG_PROXY_CLIENT_ID } from '@shared/constants'; vi.mock('axios'); diff --git a/src/shared/utils/__tests__/oauth.test.ts b/src/shared/utils/__tests__/oauth.test.ts index b2d1435e2..a77b35c49 100644 --- a/src/shared/utils/__tests__/oauth.test.ts +++ b/src/shared/utils/__tests__/oauth.test.ts @@ -3,9 +3,9 @@ import { extractOAuthCode, isAuthorizationTimeout, missingOAuthScopes, + OAuthTokenResponseSchema, parseOAuthScopes, } from '@utils/oauth'; -import { OAuthTokenResponseSchema } from '@utils/oauth-token'; import { WIZARD_OAUTH_SCOPES, WIZARD_PROVISIONING_SCOPES, diff --git a/src/shared/utils/analytics.ts b/src/shared/utils/analytics.ts index 350bec9c6..c5a67fe3f 100644 --- a/src/shared/utils/analytics.ts +++ b/src/shared/utils/analytics.ts @@ -5,11 +5,11 @@ import { ANALYTICS_TEAM_TAG, WIZARD_FLAG_KEYS, } from '@shared/constants'; -import type { WizardSession } from '@lib/wizard-session'; import { reportableDiscoveredFeatures, reportablePosthogSdkDetected, -} from '@shared/scan-consent'; + type WizardSession, +} from '@lib/wizard-session'; import type { ApiUser } from '@shared/api'; import { v4 as uuidv4 } from 'uuid'; import { IS_PRODUCTION_BUILD, RUN_SURFACE, TASK_ID, TASK_RUN_ID } from '@env'; diff --git a/src/shared/utils/cleanup-registry.ts b/src/shared/utils/cleanup-registry.ts deleted file mode 100644 index 071955574..000000000 --- a/src/shared/utils/cleanup-registry.ts +++ /dev/null @@ -1,22 +0,0 @@ -/** Process-local cleanup callbacks shared by the CLI and agent. */ -const cleanupFns: Array<() => void> = []; - -export function registerCleanup(fn: () => void): void { - cleanupFns.push(fn); -} - -export function clearCleanup(): void { - cleanupFns.length = 0; -} - -/** Runs all registered cleanup functions and drains the array. */ -export function runCleanups(): void { - const fns = cleanupFns.splice(0); - for (const fn of fns) { - try { - fn(); - } catch { - /* cleanup should not prevent exit */ - } - } -} diff --git a/src/shared/utils/environment.ts b/src/shared/utils/environment.ts index f3941fb8b..e639afa26 100644 --- a/src/shared/utils/environment.ts +++ b/src/shared/utils/environment.ts @@ -28,7 +28,7 @@ export function isNonInteractiveEnvironment(): boolean { * `e2eAsk` re-wires the `wizard_ask` bridge in an otherwise non-interactive * run. Only the e2e TUI host may set it: a real `--ci` run has nobody to answer, * so every question would stall for the bridge timeout instead of failing fast - * with an actionable error. See `isAskDisabled`. + * with an actionable error. See `shouldDisableAsk`. */ const NEVER_FROM_ENV = ['e2eAsk']; diff --git a/src/shared/utils/oauth-token.ts b/src/shared/utils/oauth-token.ts deleted file mode 100644 index 92443d5c5..000000000 --- a/src/shared/utils/oauth-token.ts +++ /dev/null @@ -1,84 +0,0 @@ -/** OAuth token-endpoint helpers that need no UI: the token response and the refresh grant. */ -import axios from 'axios'; -import { z } from 'zod'; -import { - POSTHOG_DEV_CLIENT_ID, - POSTHOG_PROXY_CLIENT_ID, - WIZARD_USER_AGENT, -} from '@shared/constants'; -import { logToFile } from './debug'; -import { getOAuthUrl, resolveBaseUrl } from './urls'; -import { oauthErrorFromTokenBody } from './oauth-errors'; - -export const OAuthTokenResponseSchema = z.object({ - access_token: z.string(), - expires_in: z.number(), - token_type: z.string(), - scope: z.string(), - refresh_token: z.string().optional(), - scoped_teams: z.array(z.number()).optional(), - scoped_organizations: z.array(z.string()).optional(), - // Sent by PostHog Cloud (and passed through the oauth.posthog.com proxy); absent on - // self-hosted. `.catch(undefined)` so an unrecognized value degrades to the probe - // fallback instead of failing the whole login. - posthog_region: z.enum(['us', 'eu']).optional().catch(undefined), - posthog_base_url: z.string().optional().catch(undefined), -}); - -export type OAuthTokenResponse = z.infer; - -/** - * OAuth client ID for the current target. A pinned base URL (`--base-url`, or - * IS_DEV's implicit localhost) means we're talking to a dev-seeded stack, which - * registers the dev client; prod uses the proxy client. - * - * TODO: this assumes any pinned base URL is a dev-seeded instance that - * registers POSTHOG_DEV_CLIENT_ID. If we ever point `--base-url` at a non-dev - * instance with its own OAuth app, make the client ID configurable (e.g. a - * `--oauth-client-id` flag) instead of always falling back to the dev client. - */ -export function getOAuthClientId(baseUrl?: string): string { - return resolveBaseUrl(baseUrl) - ? POSTHOG_DEV_CLIENT_ID - : POSTHOG_PROXY_CLIENT_ID; -} - -// Refresh-token grant (RFC 6749 §6); the server rotates, so callers must store the returned refresh_token. -export async function refreshAccessToken( - refreshToken: string, - baseUrl?: string, - clientId?: string, -): Promise { - const oauthUrl = getOAuthUrl(baseUrl); - logToFile(`[oauth] refreshing access token at ${oauthUrl}/oauth/token`); - try { - const response = await axios.post( - `${oauthUrl}/oauth/token`, - { - grant_type: 'refresh_token', - refresh_token: refreshToken, - // The grant only refreshes under its minting app — provisioning signups pass their regional client. - client_id: clientId ?? getOAuthClientId(baseUrl), - }, - { - headers: { - 'Content-Type': 'application/json', - 'User-Agent': WIZARD_USER_AGENT, - }, - timeout: 30_000, - }, - ); - const token = OAuthTokenResponseSchema.parse(response.data); - logToFile('[oauth] access token refreshed'); - return token; - } catch (e) { - logToFile( - '[oauth] token refresh failed:', - e instanceof Error ? e.message : e, - ); - const refreshError = axios.isAxiosError(e) - ? oauthErrorFromTokenBody(e.response?.data) - : null; - throw refreshError ?? e; - } -} diff --git a/src/shared/utils/oauth.ts b/src/shared/utils/oauth.ts index e3f0cd226..dda07db81 100644 --- a/src/shared/utils/oauth.ts +++ b/src/shared/utils/oauth.ts @@ -3,10 +3,13 @@ import * as http from 'node:http'; import { execSync } from 'node:child_process'; import axios from 'axios'; import { logToFile } from './debug'; +import { z } from 'zod'; import { getUI } from '@ui'; import { OAUTH_PORTS, OAUTH_TIMEOUT_MS, + POSTHOG_DEV_CLIENT_ID, + POSTHOG_PROXY_CLIENT_ID, WIZARD_USER_AGENT, } from '@shared/constants'; import { getOAuthUrl, resolveBaseUrl } from './urls'; @@ -20,11 +23,6 @@ import { oauthErrorFromCallbackParams, oauthErrorFromTokenBody, } from './oauth-errors'; -import { - getOAuthClientId, - OAuthTokenResponseSchema, - type OAuthTokenResponse, -} from './oauth-token'; const OAUTH_CALLBACK_STYLES = ` `; +export const OAuthTokenResponseSchema = z.object({ + access_token: z.string(), + expires_in: z.number(), + token_type: z.string(), + scope: z.string(), + refresh_token: z.string().optional(), + scoped_teams: z.array(z.number()).optional(), + scoped_organizations: z.array(z.string()).optional(), + // Sent by PostHog Cloud (and passed through the oauth.posthog.com proxy); absent on + // self-hosted. `.catch(undefined)` so an unrecognized value degrades to the probe + // fallback instead of failing the whole login. + posthog_region: z.enum(['us', 'eu']).optional().catch(undefined), + posthog_base_url: z.string().optional().catch(undefined), +}); + +export type OAuthTokenResponse = z.infer; + export const WIZARD_COMPLETION_SCOPE = 'event_definition:write'; export function parseOAuthScopes(scope: string): string[] { @@ -102,6 +117,22 @@ interface OAuthConfig { baseUrl?: string; } +/** + * OAuth client ID for the current target. A pinned base URL (`--base-url`, or + * IS_DEV's implicit localhost) means we're talking to a dev-seeded stack, which + * registers the dev client; prod uses the proxy client. + * + * TODO: this assumes any pinned base URL is a dev-seeded instance that + * registers POSTHOG_DEV_CLIENT_ID. If we ever point `--base-url` at a non-dev + * instance with its own OAuth app, make the client ID configurable (e.g. a + * `--oauth-client-id` flag) instead of always falling back to the dev client. + */ +function getOAuthClientId(baseUrl?: string): string { + return resolveBaseUrl(baseUrl) + ? POSTHOG_DEV_CLIENT_ID + : POSTHOG_PROXY_CLIENT_ID; +} + function getLocalOAuthOrigin(port: number): string { return `http://localhost:${port}`; } @@ -436,6 +467,46 @@ async function exchangeCodeForToken( return token; } +// Refresh-token grant (RFC 6749 §6); the server rotates, so callers must store the returned refresh_token. +export async function refreshAccessToken( + refreshToken: string, + baseUrl?: string, + clientId?: string, +): Promise { + const oauthUrl = getOAuthUrl(baseUrl); + logToFile(`[oauth] refreshing access token at ${oauthUrl}/oauth/token`); + try { + const response = await axios.post( + `${oauthUrl}/oauth/token`, + { + grant_type: 'refresh_token', + refresh_token: refreshToken, + // The grant only refreshes under its minting app — provisioning signups pass their regional client. + client_id: clientId ?? getOAuthClientId(baseUrl), + }, + { + headers: { + 'Content-Type': 'application/json', + 'User-Agent': WIZARD_USER_AGENT, + }, + timeout: 30_000, + }, + ); + const token = OAuthTokenResponseSchema.parse(response.data); + logToFile('[oauth] access token refreshed'); + return token; + } catch (e) { + logToFile( + '[oauth] token refresh failed:', + e instanceof Error ? e.message : e, + ); + const refreshError = axios.isAxiosError(e) + ? oauthErrorFromTokenBody(e.response?.data) + : null; + throw refreshError ?? e; + } +} + /** * Warn — at login, while the user is still watching — when the grant came back * narrower than the request, and record the gap so narrowed runs are countable. diff --git a/src/shared/utils/package-manager.ts b/src/shared/utils/package-manager.ts index 28daf2890..cd95772e1 100644 --- a/src/shared/utils/package-manager.ts +++ b/src/shared/utils/package-manager.ts @@ -2,6 +2,8 @@ import * as fs from 'fs'; import * as path from 'path'; import { readFileHead } from './bounded-fs'; import { withProgress } from './telemetry'; +import { getPackageDotJson, updatePackageDotJson } from './setup-utils'; +import type { PackageJson } from './package-json'; import { analytics } from './analytics'; import type { WizardRunOptions } from './types'; @@ -16,6 +18,11 @@ export interface PackageManager { runScriptCommand: string; flags: string; detect: (opts: InstallDirOpt) => boolean; + addOverride: ( + pkgName: string, + pkgVersion: string, + opts: InstallDirOpt, + ) => Promise; } function hasLockfile(installDir: string, file: string): boolean { @@ -33,6 +40,38 @@ function lockfileHeaderContains( ); } +type OverrideSlot = 'npm' | 'yarn' | 'pnpm'; + +async function writeOverride( + slot: OverrideSlot, + pkgName: string, + pkgVersion: string, + { installDir }: InstallDirOpt, +): Promise { + const pkg = await getPackageDotJson({ installDir }); + let next: PackageJson; + if (slot === 'yarn') { + next = { + ...pkg, + resolutions: { ...(pkg.resolutions ?? {}), [pkgName]: pkgVersion }, + }; + } else if (slot === 'pnpm') { + next = { + ...pkg, + pnpm: { + ...(pkg.pnpm ?? {}), + overrides: { ...(pkg.pnpm?.overrides ?? {}), [pkgName]: pkgVersion }, + }, + }; + } else { + next = { + ...pkg, + overrides: { ...(pkg.overrides ?? {}), [pkgName]: pkgVersion }, + }; + } + await updatePackageDotJson(next, { installDir }); +} + export const BUN: PackageManager = { name: 'bun', label: 'Bun', @@ -42,6 +81,8 @@ export const BUN: PackageManager = { flags: '', detect: ({ installDir }) => hasLockfile(installDir, 'bun.lockb') || hasLockfile(installDir, 'bun.lock'), + addOverride: (pkgName, pkgVersion, opts) => + writeOverride('npm', pkgName, pkgVersion, opts), }; export const YARN_V1: PackageManager = { @@ -53,6 +94,8 @@ export const YARN_V1: PackageManager = { flags: '--ignore-workspace-root-check', detect: ({ installDir }) => lockfileHeaderContains(installDir, 'yarn.lock', 'yarn lockfile v1'), + addOverride: (pkgName, pkgVersion, opts) => + writeOverride('yarn', pkgName, pkgVersion, opts), }; /** YARN V2/3/4 */ @@ -65,6 +108,8 @@ export const YARN_V2: PackageManager = { flags: '', detect: ({ installDir }) => lockfileHeaderContains(installDir, 'yarn.lock', '__metadata'), + addOverride: (pkgName, pkgVersion, opts) => + writeOverride('yarn', pkgName, pkgVersion, opts), }; export const PNPM: PackageManager = { @@ -75,6 +120,8 @@ export const PNPM: PackageManager = { runScriptCommand: 'pnpm', flags: '--ignore-workspace-root-check', detect: ({ installDir }) => hasLockfile(installDir, 'pnpm-lock.yaml'), + addOverride: (pkgName, pkgVersion, opts) => + writeOverride('pnpm', pkgName, pkgVersion, opts), }; export const NPM: PackageManager = { @@ -85,6 +132,8 @@ export const NPM: PackageManager = { runScriptCommand: 'npm run', flags: '', detect: ({ installDir }) => hasLockfile(installDir, 'package-lock.json'), + addOverride: (pkgName, pkgVersion, opts) => + writeOverride('npm', pkgName, pkgVersion, opts), }; // Expo is selected by upstream config (app.json / app.config.*) rather than @@ -97,6 +146,8 @@ export const EXPO: PackageManager = { runScriptCommand: 'npx expo run', flags: '', detect: () => false, + addOverride: (pkgName, pkgVersion, opts) => + writeOverride('npm', pkgName, pkgVersion, opts), }; export const packageManagers: PackageManager[] = [ diff --git a/src/shared/utils/provisioning.ts b/src/shared/utils/provisioning.ts index 959ab89d7..ad24cd7ec 100644 --- a/src/shared/utils/provisioning.ts +++ b/src/shared/utils/provisioning.ts @@ -46,7 +46,7 @@ const getProvisioningBaseUrl = ( * that registers the dev client; prod uses the client registered for the target * region (the wizard OAuth app is registered separately per region). * - * TODO: same assumption as `getOAuthClientId` in oauth-token.ts — a pinned base URL is + * TODO: same assumption as `getOAuthClientId` in oauth.ts — a pinned base URL is * treated as a dev-seeded instance. Make configurable if we ever point * `--base-url` at a non-dev instance with its own OAuth app. */ diff --git a/src/shared/utils/setup-utils.ts b/src/shared/utils/setup-utils.ts index f44fcca32..0b327017d 100644 --- a/src/shared/utils/setup-utils.ts +++ b/src/shared/utils/setup-utils.ts @@ -306,6 +306,39 @@ export async function installPackage({ }); } +/** + * Get package.json or abort the wizard if not found. + * Only use where package.json is required (e.g., package install, overrides). + * For detection/version-checks, use tryGetPackageJson() instead. + */ +export async function getPackageDotJson({ + installDir, +}: Pick): Promise { + const pkgPath = join(installDir, 'package.json'); + + let raw: string; + try { + raw = await fs.promises.readFile(pkgPath, 'utf8'); + } catch { + getUI().log.error( + 'Could not find package.json. Make sure to run the wizard in the root of your app!', + ); + await abort(); + return {}; + } + + try { + const parsed = JSON.parse(raw) as PackageJson | null; + return parsed ?? {}; + } catch { + getUI().log.error( + `Unable to parse your package.json. Make sure it has a valid format!`, + ); + await abort(); + return {}; + } +} + /** * Try to get package.json, returning null if it doesn't exist. * Use this for detection purposes where missing package.json is expected (e.g., Python projects). @@ -324,6 +357,25 @@ export async function tryGetPackageJson({ } } +export async function updatePackageDotJson( + packageDotJson: PackageJson, + { installDir }: Pick, +): Promise { + const pkgPath = join(installDir, 'package.json'); + const serialized = JSON.stringify(packageDotJson, null, 2); + + try { + await fs.promises.writeFile(pkgPath, serialized, { + encoding: 'utf8', + flag: 'w', + }); + return; + } catch { + getUI().log.error(`Unable to update your package.json.`); + await abort(); + } +} + /** * Detect and return the package manager. Pure — no prompts. * Falls back to first detected or npm if ambiguous. diff --git a/src/shared/utils/wizard-abort.ts b/src/shared/utils/wizard-abort.ts index ec5275ac8..0a2b8ae30 100644 --- a/src/shared/utils/wizard-abort.ts +++ b/src/shared/utils/wizard-abort.ts @@ -17,9 +17,6 @@ import { emitWizardError, sanitizeErrorDetail, } from '@shared/errors'; -import { runCleanups } from './cleanup-registry'; - -export { registerCleanup, clearCleanup, runCleanups } from './cleanup-registry'; // Still importable from here; the class lives with the error codes. export { WizardError }; @@ -36,6 +33,28 @@ interface WizardAbortOptions { status?: 'error' | 'cancelled'; } +const cleanupFns: Array<() => void> = []; + +export function registerCleanup(fn: () => void): void { + cleanupFns.push(fn); +} + +export function clearCleanup(): void { + cleanupFns.length = 0; +} + +/** Runs all registered cleanup functions and drains the array. */ +export function runCleanups(): void { + const fns = cleanupFns.splice(0); + for (const fn of fns) { + try { + fn(); + } catch { + /* cleanup should not prevent exit */ + } + } +} + function resolveErrorCode( options: WizardAbortOptions, error: Error | WizardError | undefined, diff --git a/src/steps/install-cli-steering/index.ts b/src/steps/install-cli-steering/index.ts index ad844293d..45e133219 100644 --- a/src/steps/install-cli-steering/index.ts +++ b/src/steps/install-cli-steering/index.ts @@ -4,11 +4,6 @@ import * as os from 'node:os'; import * as path from 'node:path'; import { debug } from '@utils/debug'; -import { cliSpawnOptions } from '@shared/posthog-cli-install'; -export { - installOrUpdatePostHogCli, - type CliInstallResult, -} from '@shared/posthog-cli-install'; /** * A coding agent whose global instructions file the PostHog CLI steering @@ -68,6 +63,59 @@ export interface SteeringInstallResult { error?: string; } +export interface CliInstallResult { + success: boolean; + error?: string; + /** + * The underlying failure as an Error, for callers that report to error + * tracking. Carries the real spawn error (with its stack) when npm couldn't + * launch; a synthesized Error for a non-zero exit. Set whenever success is + * false. + */ + errorObject?: Error; +} + +const spawnOptions = { + encoding: 'utf-8' as const, + // npm/posthog-cli are npm.cmd/posthog-cli.cmd on Windows; spawnSync only + // resolves them through a shell. + shell: process.platform === 'win32', +}; + +/** + * Install or update the PostHog CLI in the user's environment. `npm install + * --global @posthog/cli@latest` covers both first-time installs and upgrades + * for existing npm-installed CLIs. + */ +export function installOrUpdatePostHogCli(): CliInstallResult { + const args = ['install', '--global', '@posthog/cli@latest']; + debug(`Running npm ${args.join(' ')}`); + + const result = spawnSync('npm', args, spawnOptions); + + if (result.error) { + return { + success: false, + error: `Failed to run npm: ${result.error.message}. Is Node.js installed?`, + errorObject: result.error, + }; + } + if (result.status !== 0) { + const detail = (result.stderr || result.stdout || '').trim(); + const message = + detail || + `npm install --global @posthog/cli@latest exited with status ${ + result.status ?? 'unknown' + }`; + return { + success: false, + error: message, + errorObject: new Error(message), + }; + } + return { success: true }; +} + /** * Delegate the actual write to the installed `posthog-cli api agents-md * install`. The steering snippet lives in the CLI (its single source of truth), @@ -81,7 +129,7 @@ export function installSteeringSnippet( debug(`Running posthog-cli ${args.join(' ')}`); const result = spawnSync('posthog-cli', args, { - ...cliSpawnOptions, + ...spawnOptions, env: { ...process.env, POSTHOG_CLI_EXPERIMENTAL_API: '1' }, }); diff --git a/src/steps/upload-environment-variables/index.ts b/src/steps/upload-environment-variables/index.ts index 6ea4e0d61..51b61bc22 100644 --- a/src/steps/upload-environment-variables/index.ts +++ b/src/steps/upload-environment-variables/index.ts @@ -2,6 +2,7 @@ import type { Integration } from '@shared/constants'; import { withProgress } from '@utils/telemetry'; import { analytics } from '@utils/analytics'; import { getUI } from '@ui'; +import type { WizardSession } from '@lib/wizard-session'; import { EnvironmentProvider } from './EnvironmentProvider'; import { VercelEnvironmentProvider } from './providers/vercel'; @@ -12,7 +13,7 @@ export const uploadEnvironmentVariablesStep = async ( session, }: { integration: Integration; - session: { installDir: string }; + session: WizardSession; }, ): Promise => { const providers: EnvironmentProvider[] = [ diff --git a/src/ui/__tests__/headless-ui.test.ts b/src/ui/__tests__/headless-ui.test.ts index 16281c313..916c72405 100644 --- a/src/ui/__tests__/headless-ui.test.ts +++ b/src/ui/__tests__/headless-ui.test.ts @@ -30,14 +30,4 @@ describe('HeadlessUI', () => { logSpy.mockRestore(); }); - - it('keeps the event plan in its store for the task stream', () => { - const setEventPlan = vi.fn(); - const ui = new HeadlessUI({ setEventPlan } as unknown as WizardStore); - const plan = [{ name: 'signed_up', description: 'User signs up' }]; - - ui.setEventPlan(plan); - - expect(setEventPlan).toHaveBeenCalledExactlyOnceWith(plan); - }); }); diff --git a/src/ui/headless-ui.ts b/src/ui/headless-ui.ts index efe2630b7..7029865f1 100644 --- a/src/ui/headless-ui.ts +++ b/src/ui/headless-ui.ts @@ -1,14 +1,5 @@ import { LoggingUI } from './logging-ui'; - -interface HeadlessRunStore { - syncTodos( - todos: Array<{ content: string; status: string; activeForm?: string }>, - ): void; - setHandoffText(text: string): void; - setFrameworkContext(key: string, value: unknown): void; - setEventPlan(events: Array<{ name: string; description: string }>): void; - session: { frameworkContext: Record }; -} +import type { WizardStore } from './tui/store'; /** * `LoggingUI` plus it feeds run state into a `WizardStore` so the background @@ -19,7 +10,7 @@ interface HeadlessRunStore { * ledger arrives through `setFrameworkContext`, the seam every UI implements. */ export class HeadlessUI extends LoggingUI { - constructor(private readonly store: HeadlessRunStore) { + constructor(private readonly store: WizardStore) { super(); } @@ -38,10 +29,6 @@ export class HeadlessUI extends LoggingUI { this.store.setFrameworkContext(key, value); } - setEventPlan(events: Array<{ name: string; description: string }>): void { - this.store.setEventPlan(events); - } - getFrameworkContext(key: string): unknown { return this.store.session.frameworkContext[key]; } diff --git a/src/ui/index.ts b/src/ui/index.ts index 1e5652ed2..dc374144a 100644 --- a/src/ui/index.ts +++ b/src/ui/index.ts @@ -22,4 +22,3 @@ export function setUI(ui: WizardUI): void { } export type { WizardUI, SpinnerHandle } from './wizard-ui'; -export { createUiReducer, uiInteraction } from './agent-progress'; diff --git a/src/ui/tui/__tests__/keyboard-equivalence.test.tsx b/src/ui/tui/__tests__/keyboard-equivalence.test.tsx index 824ec3412..43e9d24cb 100644 --- a/src/ui/tui/__tests__/keyboard-equivalence.test.tsx +++ b/src/ui/tui/__tests__/keyboard-equivalence.test.tsx @@ -64,10 +64,6 @@ vi.mock('opn', () => ({ default: vi.fn() })); vi.mock('@shared/api', async (importOriginal) => ({ ...(await importOriginal()), fetchSlackConnected: vi.fn().mockResolvedValue(false), - // This test compares the handoff commit, not the following GitHub screen's - // polling effect. A real request can settle between the keyboard and action - // snapshots and add githubConnected only to the mounted keyboard path. - fetchGithubConnected: vi.fn(() => new Promise(() => undefined)), fetchUserData: vi.fn(() => new Promise(() => undefined)), })); vi.mock('@shared/skill-menu', async (importOriginal) => ({ diff --git a/src/ui/tui/__tests__/store-invariants.test.ts b/src/ui/tui/__tests__/store-invariants.test.ts index c217466e8..2c42514bf 100644 --- a/src/ui/tui/__tests__/store-invariants.test.ts +++ b/src/ui/tui/__tests__/store-invariants.test.ts @@ -168,11 +168,6 @@ const MUTATIONS: MutationCase[] = [ invoke: (s) => s.setCredentials(CREDENTIALS), emits: 1, }, - { - name: 'setInferenceAuth', - invoke: (s) => s.setInferenceAuth({ resolve: vi.fn() }), - emits: 1, - }, { name: 'setAccessToken', invoke: (s) => s.setAccessToken(CREDENTIALS), diff --git a/src/ui/tui/decks/__tests__/registry.test.ts b/src/ui/tui/decks/__tests__/registry.test.ts deleted file mode 100644 index e3024d26f..000000000 --- a/src/ui/tui/decks/__tests__/registry.test.ts +++ /dev/null @@ -1,55 +0,0 @@ -import { PROGRAM_REGISTRY } from '@programs'; -import type { ProgramId } from '@programs/types'; -import { - getProgramContentBlocks, - getProgramTips, -} from '@ui/tui/decks/registry'; -import { getLearnDeckPrograms } from '@ui/tui/playground/demos/LearnDeckDemo'; - -const runDecks = [ - ['posthog-integration', 'It handles the entire PostHog setup process'], - ['revenue-analytics-setup', 'Welcome.'], - ['warehouse-source', 'Welcome.'], - ['error-tracking-upload-source-maps', 'When you ship to production'], - ['error-tracking', "I'm wiring PostHog Error Tracking"], - ['migration', 'making a plan to migrate from Statsig'], - ['self-driving', "It's setting up PostHog Self-driving"], - ['agent-skill', 'Welcome.'], - ['mcp-analytics', 'Welcome.'], - ['replay-vision', 'Welcome.'], - ['web-analytics-doctor', 'Welcome.'], - ['ai-observability', 'Welcome.'], - ['metrics', 'Welcome.'], -] as const satisfies ReadonlyArray; - -it('keeps an outcome for every program with the standard run screen', () => { - const actualRunPrograms = PROGRAM_REGISTRY.filter((program) => - program.steps.some((step) => step.screenId === 'run'), - ).map((program) => program.id); - expect(actualRunPrograms.sort()).toEqual(runDecks.map(([id]) => id).sort()); -}); - -it.each(runDecks)('selects the %s learn deck', (id, expectedCopy) => { - expect(JSON.stringify(getProgramContentBlocks(id))).toContain(expectedCopy); -}); - -it('selects program tips only for the two custom tips decks', () => { - const expectedTips = new Map([ - ['error-tracking', 'session-replay'], - ['self-driving', 'signal-source'], - ]); - for (const [id] of runDecks) { - expect(getProgramTips(id)?.[0]?.id).toBe(expectedTips.get(id)); - } -}); - -it('makes every standard run deck reviewable in the playground', () => { - const ids = getLearnDeckPrograms().map((program) => program.id); - expect([...ids].sort()).toEqual(runDecks.map(([id]) => id).sort()); - for (const id of ['mcp-analytics', 'replay-vision', 'web-analytics-doctor']) { - expect(ids).toContain(id); - } - for (const id of ['mcp-add', 'slack-connect', 'audit']) { - expect(ids).not.toContain(id); - } -}); diff --git a/src/ui/tui/decks/registry.ts b/src/ui/tui/decks/registry.ts deleted file mode 100644 index 78231905c..000000000 --- a/src/ui/tui/decks/registry.ts +++ /dev/null @@ -1,48 +0,0 @@ -import type { ProgramId } from '@programs/types'; -import type { WizardStore } from '@ui/tui/store'; -import type { ContentBlock } from '@ui/tui/primitives/content-types'; -import type { Tip } from '@ui/tui/components/TipsCard'; -import { getContentBlocks as agentSkillBlocks } from './agent-skill/index.js'; -import { getContentBlocks as errorTrackingBlocks } from './error-tracking/index.js'; -import { getTips as errorTrackingTips } from './error-tracking/tips.js'; -import { getContentBlocks as sourceMapsBlocks } from './error-tracking-upload-source-maps/index.js'; -import { getContentBlocks as migrationBlocks } from './migration/index.js'; -import { getContentBlocks as integrationBlocks } from './posthog-integration/index.js'; -import { getContentBlocks as selfDrivingBlocks } from './self-driving/index.js'; -import { getTips as selfDrivingTips } from './self-driving/tips.js'; - -type LearnDeck = (store?: WizardStore) => ContentBlock[]; -type TipsDeck = (store?: WizardStore) => Tip[]; - -// Listed programs get this deck. Any other program gets the agent-skill deck. -const LEARN_DECKS: Partial> = { - 'posthog-integration': integrationBlocks, - 'revenue-analytics-setup': agentSkillBlocks, - 'warehouse-source': agentSkillBlocks, - 'error-tracking-upload-source-maps': sourceMapsBlocks, - 'error-tracking': errorTrackingBlocks, - migration: migrationBlocks, - 'self-driving': selfDrivingBlocks, - 'agent-skill': agentSkillBlocks, - 'ai-observability': agentSkillBlocks, - metrics: agentSkillBlocks, -}; - -const TIPS_DECKS: Partial> = { - 'error-tracking': errorTrackingTips, - 'self-driving': selfDrivingTips, -}; - -export function getProgramContentBlocks( - programId: ProgramId, - store?: WizardStore, -): ContentBlock[] { - return (LEARN_DECKS[programId] ?? agentSkillBlocks)(store); -} - -export function getProgramTips( - programId: ProgramId, - store?: WizardStore, -): Tip[] | undefined { - return TIPS_DECKS[programId]?.(store); -} diff --git a/src/ui/tui/decks/revenue-analytics/index.tsx b/src/ui/tui/decks/revenue-analytics/index.tsx new file mode 100644 index 000000000..974ebc431 --- /dev/null +++ b/src/ui/tui/decks/revenue-analytics/index.tsx @@ -0,0 +1,7 @@ +/** + * Revenue-analytics learn-deck. Currently delegates to the generic skill + * deck; replace this re-export with a program-specific script when the + * revenue narrative grows its own diagrams or talking points. + */ + +export { getContentBlocks } from '@ui/tui/decks/agent-skill/index'; diff --git a/src/ui/tui/decks/self-driving/index.tsx b/src/ui/tui/decks/self-driving/index.tsx index 6f0b7cf0d..53d7c7a01 100644 --- a/src/ui/tui/decks/self-driving/index.tsx +++ b/src/ui/tui/decks/self-driving/index.tsx @@ -1,4 +1,4 @@ -import { NO_DEFAULT_LIMIT, PRICING_LONG } from '@shared/self-driving-pricing'; +import { NO_DEFAULT_LIMIT, PRICING_LONG } from './pricing.js'; /** * Self-driving learn-deck — the narrative script played while the agent * sets up Self-driving. Teaches the vocabulary ladder (signal source → diff --git a/src/shared/self-driving-pricing.ts b/src/ui/tui/decks/self-driving/pricing.ts similarity index 100% rename from src/shared/self-driving-pricing.ts rename to src/ui/tui/decks/self-driving/pricing.ts diff --git a/src/ui/tui/decks/self-driving/tips.ts b/src/ui/tui/decks/self-driving/tips.ts index 21a226f81..cdb303025 100644 --- a/src/ui/tui/decks/self-driving/tips.ts +++ b/src/ui/tui/decks/self-driving/tips.ts @@ -1,4 +1,4 @@ -import { NO_DEFAULT_LIMIT, PRICING_LONG } from '@shared/self-driving-pricing'; +import { NO_DEFAULT_LIMIT, PRICING_LONG } from './pricing.js'; /** * Sidebar tips for the self-driving run — short footnotes on the diff --git a/src/ui/tui/decks/warehouse-source/index.tsx b/src/ui/tui/decks/warehouse-source/index.tsx new file mode 100644 index 000000000..0f364bf00 --- /dev/null +++ b/src/ui/tui/decks/warehouse-source/index.tsx @@ -0,0 +1,7 @@ +/** + * Warehouse-source learn-deck. Delegates to the generic skill deck; replace + * this re-export with a program-specific script when the data warehouse + * narrative grows its own diagrams or talking points. + */ + +export { getContentBlocks } from '@ui/tui/decks/agent-skill/index'; diff --git a/src/ui/tui/hooks/file-watcher.ts b/src/ui/tui/hooks/file-watcher.ts index 974466508..5dd272aa0 100644 --- a/src/ui/tui/hooks/file-watcher.ts +++ b/src/ui/tui/hooks/file-watcher.ts @@ -1,14 +1,8 @@ import { useEffect } from 'react'; -import { - startFileWatcher, - type FileWatcherOptions, -} from '@shared/file-watcher'; +import { startFileWatcher, type FileWatcherOptions } from '@lib/file-watcher'; -export { startFileWatcher } from '@shared/file-watcher'; -export type { - FileWatcherHandle, - FileWatcherOptions, -} from '@shared/file-watcher'; +export { startFileWatcher } from '@lib/file-watcher'; +export type { FileWatcherHandle, FileWatcherOptions } from '@lib/file-watcher'; /** React hook wrapping `startFileWatcher`. Starts on mount, stops on unmount * or when `path` changes. `onUpdate` and `options` are captured at mount diff --git a/src/ui/tui/playground/demos/LearnDeckDemo.tsx b/src/ui/tui/playground/demos/LearnDeckDemo.tsx index 5885f3410..b94b38796 100644 --- a/src/ui/tui/playground/demos/LearnDeckDemo.tsx +++ b/src/ui/tui/playground/demos/LearnDeckDemo.tsx @@ -10,8 +10,11 @@ * Arrow keys are reserved for the playground's tab switcher, so this demo * uses letter keys. * - * Decks are pulled from `PROGRAM_REGISTRY` so every program with the - * standard run screen is reviewable here, including generic fallback decks. + * Decks are pulled from `PROGRAM_REGISTRY` so every program that ships a + * deck is reviewable here. Migration also gets per-variant entries (one + * per `--product=` choice) so the variant composer in + * `migration/content/index.tsx` can be exercised side-by-side with the + * generic deck. */ import { Box, Text, useInput } from 'ink'; @@ -27,7 +30,6 @@ import { Colors } from '@ui/tui/styles'; import type { WizardStore } from '@ui/tui/store'; import { PROGRAM_REGISTRY } from '@programs'; import { AUDIT_AREA_SLIDES } from '@ui/tui/screens/audit/slides/index'; -import { getProgramContentBlocks } from '@ui/tui/decks/registry'; import type { AreaSlide } from '@ui/tui/screens/audit/slides/shared'; interface Deck { @@ -83,20 +85,16 @@ interface LearnDeckDemoProps { store: WizardStore; } -export const getLearnDeckPrograms = () => - PROGRAM_REGISTRY.filter((program) => - program.steps.some((step) => step.screenId === 'run'), - ); - export const LearnDeckDemo = ({ store }: LearnDeckDemoProps) => { const decks: Deck[] = useMemo(() => { const all: Deck[] = []; - // Every program with the standard run screen. Seed the store's + // Every program in the registry that ships a deck. Seed the store's // skillId from the program config so decks that template the skill // name (e.g. agent-skill's "Running the skill...") render the // real value instead of "unknown". - for (const program of getLearnDeckPrograms()) { + for (const program of PROGRAM_REGISTRY) { + if (!program.getContentBlocks) continue; const stub = program.skillId ? withSessionOverride(store, { skillId: program.skillId }) : store; @@ -105,7 +103,7 @@ export const LearnDeckDemo = ({ store }: LearnDeckDemoProps) => { label: `${program.id} (${program.command ?? 'default'})${ program.skillId ? ` · skill: ${program.skillId}` : '' }`, - blocks: getProgramContentBlocks(program.id, stub), + blocks: program.getContentBlocks(stub), }); } diff --git a/src/ui/tui/playground/demos/RunScreenDemo.tsx b/src/ui/tui/playground/demos/RunScreenDemo.tsx index 20f6dc30c..60a92516f 100644 --- a/src/ui/tui/playground/demos/RunScreenDemo.tsx +++ b/src/ui/tui/playground/demos/RunScreenDemo.tsx @@ -35,7 +35,8 @@ import type { ProgressItem, TabDefinition } from '@ui/tui/primitives/index'; import { LearnCard } from '@ui/tui/components/LearnCard'; import { TipsCard } from '@ui/tui/components/TipsCard'; import { VisualizerTab } from '@ui/tui/components/PhaseVisuals'; -import { getProgramContentBlocks } from '@ui/tui/decks/registry'; +import { getProgramConfig } from '@programs'; +import { getContentBlocks as getSkillContentBlocks } from '@ui/tui/decks/agent-skill/index'; import { Colors } from '@ui/tui/styles'; import { WIZARD_LOG_FILE } from '@utils/paths'; @@ -221,7 +222,10 @@ export const RunScreenDemo = ({ store }: RunScreenDemoProps) => { })); const learnBlocks = useMemo(() => { - return getProgramContentBlocks(store.router.activeProgram, store); + const getBlocks = + getProgramConfig(store.router.activeProgram).getContentBlocks ?? + getSkillContentBlocks; + return getBlocks(store); }, [store]); const leftPane = store.learnCardComplete ? ( diff --git a/src/ui/tui/screens/ErrorTrackingDetectScreen.tsx b/src/ui/tui/screens/ErrorTrackingDetectScreen.tsx index cab674797..4cf416e1d 100644 --- a/src/ui/tui/screens/ErrorTrackingDetectScreen.tsx +++ b/src/ui/tui/screens/ErrorTrackingDetectScreen.tsx @@ -9,7 +9,6 @@ import { useEffect, useRef, useState, useSyncExternalStore } from 'react'; import type { WizardStore } from '@ui/tui/store'; import { LoadingBox, PickerMenu } from '@ui/tui/primitives/index'; import { Colors, Icons } from '@ui/tui/styles'; -import { createUiReducer, getUI } from '@ui'; import { FRAMEWORK_REGISTRY } from '@programs/registry'; import { detectErrorTrackingProjects, @@ -63,7 +62,6 @@ export const ErrorTrackingDetectScreen = ({ setActivity((prev) => [...prev, line].slice(-MAX_ACTIVITY_LINES)); } }, - createUiReducer(getUI()), ); if (!cancelled) setState({ kind: 'ready', report }); } catch (err) { diff --git a/src/ui/tui/screens/RunScreen.tsx b/src/ui/tui/screens/RunScreen.tsx index 7786af304..db3cc8cd8 100644 --- a/src/ui/tui/screens/RunScreen.tsx +++ b/src/ui/tui/screens/RunScreen.tsx @@ -24,10 +24,8 @@ import { VisualizerTab } from '@ui/tui/components/PhaseVisuals'; import { TipsCard } from '@ui/tui/components/TipsCard'; import { useStdoutDimensions } from '@ui/tui/hooks/useStdoutDimensions'; -import { - getProgramContentBlocks, - getProgramTips, -} from '@ui/tui/decks/registry'; +import { getProgramConfig } from '@programs'; +import { getContentBlocks as getSkillContentBlocks } from '@ui/tui/decks/agent-skill/index'; import { WIZARD_LOG_FILE } from '@utils/paths'; @@ -66,15 +64,20 @@ export const RunScreen = ({ store }: RunScreenProps) => { const statuses = store.statusMessages.length > 0 ? store.statusMessages : undefined; + // Each program owns its content deck (program/content/index.tsx) + // and wires it onto its ProgramConfig.getContentBlocks. Fall back to the + // agent-skill deck for runtime-created configs (e.g. `wizard skill `) + // that aren't in the static registry. const activeProgram = store.router.activeProgram; - const learnBlocks = useMemo( - () => getProgramContentBlocks(activeProgram, store), - [store, activeProgram], - ); + const learnBlocks = useMemo(() => { + const getBlocks = + getProgramConfig(activeProgram).getContentBlocks ?? getSkillContentBlocks; + return getBlocks(store); + }, [store, activeProgram]); // Program-supplied tips for the right pane; undefined falls back to // DEFAULT_TIPS inside TipsCard, so non-self-driving programs are unaffected. - const programTips = getProgramTips(activeProgram, store); + const programTips = getProgramConfig(activeProgram).getTips?.(store); const leftPane = store.learnCardComplete ? ( diff --git a/src/ui/tui/screens/SelfDrivingIntegrationDetectScreen.tsx b/src/ui/tui/screens/SelfDrivingIntegrationDetectScreen.tsx index 16b669a31..6db94d501 100644 --- a/src/ui/tui/screens/SelfDrivingIntegrationDetectScreen.tsx +++ b/src/ui/tui/screens/SelfDrivingIntegrationDetectScreen.tsx @@ -14,7 +14,6 @@ import { useEffect, useRef, useState, useSyncExternalStore } from 'react'; import type { WizardStore } from '@ui/tui/store'; import { LoadingBox, PickerMenu } from '@ui/tui/primitives/index'; import { Colors, Icons } from '@ui/tui/styles'; -import { createUiReducer, getUI } from '@ui'; import { Integration } from '@shared/constants'; import { FRAMEWORK_REGISTRY } from '@programs/registry'; import { SELF_DRIVING_INTEGRATE_PATH_KEY } from '@programs/self-driving/detect'; @@ -88,7 +87,6 @@ export const SelfDrivingIntegrationDetectScreen = ({ setActivity((prev) => [...prev, line].slice(-MAX_ACTIVITY_LINES)); } }, - createUiReducer(getUI()), ); if (!cancelled) setState({ kind: 'ready', report }); } catch (err) { diff --git a/src/ui/tui/screens/SelfDrivingIntroScreen.tsx b/src/ui/tui/screens/SelfDrivingIntroScreen.tsx index d42b97755..99ed43f9d 100644 --- a/src/ui/tui/screens/SelfDrivingIntroScreen.tsx +++ b/src/ui/tui/screens/SelfDrivingIntroScreen.tsx @@ -18,7 +18,7 @@ import { NO_DEFAULT_LIMIT, PRICING_LONG, PRICING_SHORT, -} from '@shared/self-driving-pricing'; +} from '@ui/tui/decks/self-driving/pricing.js'; import type { SelfDrivingDetectError } from '@programs/self-driving/index'; interface SelfDrivingIntroScreenProps { diff --git a/src/ui/tui/screens/SourceMapsDetectScreen.tsx b/src/ui/tui/screens/SourceMapsDetectScreen.tsx index 4083399b4..645cd61cd 100644 --- a/src/ui/tui/screens/SourceMapsDetectScreen.tsx +++ b/src/ui/tui/screens/SourceMapsDetectScreen.tsx @@ -12,7 +12,6 @@ import { useEffect, useRef, useState, useSyncExternalStore } from 'react'; import type { WizardStore } from '@ui/tui/store'; import { LoadingBox, PickerMenu } from '@ui/tui/primitives/index'; import { Colors, Icons } from '@ui/tui/styles'; -import { createUiReducer, getUI } from '@ui'; import { SOURCE_MAPS_CONTEXT_KEYS, VARIANT_DISPLAY_NAME, @@ -63,15 +62,11 @@ export const SourceMapsDetectScreen = ({ let cancelled = false; void (async () => { try { - const report = await detectSourceMapsProjects( - store.session, - (line) => { - if (!cancelled) { - setActivity((prev) => [...prev, line].slice(-MAX_ACTIVITY_LINES)); - } - }, - createUiReducer(getUI()), - ); + const report = await detectSourceMapsProjects(store.session, (line) => { + if (!cancelled) { + setActivity((prev) => [...prev, line].slice(-MAX_ACTIVITY_LINES)); + } + }); if (!cancelled) setState({ kind: 'ready', report }); } catch (err) { if (!cancelled) { diff --git a/src/ui/tui/screens/audit/AuditRunScreen.tsx b/src/ui/tui/screens/audit/AuditRunScreen.tsx index b6f7a522d..1ef1c0a46 100644 --- a/src/ui/tui/screens/audit/AuditRunScreen.tsx +++ b/src/ui/tui/screens/audit/AuditRunScreen.tsx @@ -27,7 +27,8 @@ export const AuditRunScreen = ({ store }: AuditRunScreenProps) => { () => store.getSnapshot(), ); - // The runner watches the ledger in headless runs too; render the store value. + // The ledger reaches the store through `AuditLedgerWatcher`, which runs for + // headless runs too. This screen only renders what the store holds. const statuses = store.statusMessages.length > 0 ? store.statusMessages : undefined; diff --git a/src/ui/tui/services/__tests__/mcp-suggested-prompts-services.test.ts b/src/ui/tui/services/__tests__/mcp-suggested-prompts-services.test.ts deleted file mode 100644 index f84e9b226..000000000 --- a/src/ui/tui/services/__tests__/mcp-suggested-prompts-services.test.ts +++ /dev/null @@ -1,70 +0,0 @@ -import { runMcpPromptViaSdk } from '@agent'; -import { createPosthogInferenceAuthProvider } from '@programs'; -import { HostResolution } from '@shared/host-resolution'; -import type { Credentials } from '@lib/wizard-session'; -import { WizardStore } from '@ui/tui/store'; -import { createMcpSuggestedPromptsServices } from '../mcp-suggested-prompts-services'; - -vi.mock('@agent', async (original) => ({ - ...(await original()), - runMcpPromptViaSdk: vi.fn(), -})); -vi.mock('@programs', async (original) => ({ - ...(await original()), - createPosthogInferenceAuthProvider: vi.fn(), -})); - -const credentials: Credentials = { - accessToken: 'phx_test', - projectApiKey: 'phc_test', - projectId: 42, - host: HostResolution.fromRegion('us'), -}; - -async function consumePrompt(store: WizardStore): Promise { - const services = createMcpSuggestedPromptsServices(store); - for await (const chunk of services.runPromptStreaming({ - prompt: 'Show recent events', - credentials, - signal: new AbortController().signal, - })) { - void chunk; - } -} - -beforeEach(() => { - vi.clearAllMocks(); - vi.mocked(runMcpPromptViaSdk).mockImplementation(async function* () { - await Promise.resolve(); - yield* []; - }); -}); - -it('uses the fixed provider from the TUI store for MCP prompt inference', async () => { - const store = new WizardStore('mcp-tutorial'); - const fixed = { resolve: vi.fn() }; - store.setInferenceAuth(fixed); - - await consumePrompt(store); - - expect(vi.mocked(runMcpPromptViaSdk).mock.calls[0][0].inferenceAuth).toBe( - fixed, - ); - expect(createPosthogInferenceAuthProvider).not.toHaveBeenCalled(); -}); - -it('creates the ordinary PostHog provider when the store has none', async () => { - const store = new WizardStore('mcp-tutorial'); - const fallback = { resolve: vi.fn() }; - vi.mocked(createPosthogInferenceAuthProvider).mockReturnValue(fallback); - - await consumePrompt(store); - - expect(createPosthogInferenceAuthProvider).toHaveBeenCalledWith( - credentials, - store.analyticsProgramId, - ); - expect(vi.mocked(runMcpPromptViaSdk).mock.calls[0][0].inferenceAuth).toBe( - fallback, - ); -}); diff --git a/src/ui/tui/services/mcp-suggested-prompts-services.ts b/src/ui/tui/services/mcp-suggested-prompts-services.ts index 2c5f69c5c..813596516 100644 --- a/src/ui/tui/services/mcp-suggested-prompts-services.ts +++ b/src/ui/tui/services/mcp-suggested-prompts-services.ts @@ -12,7 +12,7 @@ import type { Credentials } from '@lib/wizard-session'; import { getOrAskForProjectData } from '@utils/setup-utils'; -import { Program, createPosthogInferenceAuthProvider } from '@programs'; +import { Program } from '@programs'; import type { WizardStore } from '@ui/tui/store'; import type { ApiUser } from '@shared/api'; import { @@ -130,12 +130,6 @@ export function createMcpSuggestedPromptsServices( // trace tags are built where the headers are, keeping the agent module // out of the TUI's startup graph. programId: store.analyticsProgramId, - inferenceAuth: - store.session.inferenceAuth ?? - createPosthogInferenceAuthProvider( - args.credentials, - store.analyticsProgramId, - ), integration: store.session.integration ?? undefined, }), @@ -159,7 +153,6 @@ export function createMcpSuggestedPromptsServices( async function* runProductionPromptStreaming(args: { prompt: string; credentials: Credentials; - inferenceAuth: import('@agent/types').InferenceAuthProvider; signal: AbortSignal; resumeSessionId?: string; programId?: string; diff --git a/src/ui/tui/store.ts b/src/ui/tui/store.ts index 03d14c820..cbe3a8688 100644 --- a/src/ui/tui/store.ts +++ b/src/ui/tui/store.ts @@ -453,11 +453,6 @@ export class WizardStore { this.emitChange(); } - setInferenceAuth(provider: WizardSession['inferenceAuth']): void { - this.$session.setKey('inferenceAuth', provider); - this.emitChange(); - } - /** Post-refresh credential swap. No `auth complete` — see WizardUI. */ setAccessToken(credentials: WizardSession['credentials']): void { this.$session.setKey('credentials', credentials); diff --git a/src/ui/wizard-ui.ts b/src/ui/wizard-ui.ts index 56ef24e8a..e972a5063 100644 --- a/src/ui/wizard-ui.ts +++ b/src/ui/wizard-ui.ts @@ -17,7 +17,17 @@ import type { OutroData, PendingQuestion, } from '@lib/wizard-session'; -export { TaskStatus, isTaskStatus } from '@shared/run-state'; + +export enum TaskStatus { + Pending = 'pending', + InProgress = 'in_progress', + Completed = 'completed', + Skipped = 'skipped', +} + +export function isTaskStatus(value: string): value is TaskStatus { + return (Object.values(TaskStatus) as string[]).includes(value); +} // Progress payloads are the agent's contract; re-exported so UI code keeps its import path. import type { diff --git a/test/module-graph.ts b/test/module-graph.ts deleted file mode 100644 index 2ea666c1e..000000000 --- a/test/module-graph.ts +++ /dev/null @@ -1,123 +0,0 @@ -/** The repo's module resolver, shared by the architecture test and the entry closure checks. */ -import * as fs from 'fs'; -import * as path from 'path'; -import { fileURLToPath } from 'url'; -import * as ts from 'typescript'; - -export const REPO_ROOT = path.resolve( - path.dirname(fileURLToPath(import.meta.url)), - '..', -); - -export type Aliases = ReadonlyArray; - -export function toRepoRelative(abs: string): string { - return path.relative(REPO_ROOT, abs).split(path.sep).join('/'); -} - -function isFile(abs: string): boolean { - return fs.statSync(abs, { throwIfNoEntry: false })?.isFile() ?? false; -} - -export function loadAliases(): Aliases { - const tsconfig = JSON.parse( - fs.readFileSync(path.join(REPO_ROOT, 'tsconfig.build.json'), 'utf8'), - ) as { compilerOptions?: { paths?: Record } }; - return Object.entries(tsconfig.compilerOptions?.paths ?? {}).map( - ([pattern, targets]) => [pattern, targets[0]] as const, - ); -} - -export function aliasTarget(spec: string, aliases: Aliases): string | null { - for (const [pattern, target] of aliases) { - if (pattern.endsWith('*')) { - const prefix = pattern.slice(0, -1); - if (spec.startsWith(prefix)) { - return path.resolve( - REPO_ROOT, - target.slice(0, -1) + spec.slice(prefix.length), - ); - } - } else if (spec === pattern) { - return path.resolve(REPO_ROOT, target); - } - } - return null; -} - -export function probe(base: string): string | null { - const candidates: string[] = []; - if (base.endsWith('.js')) { - const stem = base.slice(0, -3); - candidates.push(`${stem}.ts`, `${stem}.tsx`); - } - candidates.push( - `${base}.ts`, - `${base}.tsx`, - path.join(base, 'index.ts'), - path.join(base, 'index.tsx'), - base, - ); - return candidates.find(isFile) ?? null; -} - -/** A file's transpiled output: what runs, with type-only imports erased. */ -export function transpiled(file: string): string { - const source = fs.readFileSync(path.join(REPO_ROOT, file), 'utf8'); - return ts.transpileModule(source, { - fileName: file, - compilerOptions: { - module: ts.ModuleKind.ESNext, - target: ts.ScriptTarget.ES2022, - jsx: ts.JsxEmit.ReactJSX, - }, - }).outputText; -} - -/** Repo files that load with `entry`: static imports and re-exports, plus `import()` when `includeDynamic`. */ -export function staticImportClosure( - entry: string, - includeDynamic = false, -): string[] { - const aliases = loadAliases(); - const pending = [entry]; - const visited = new Set(); - - while (pending.length > 0) { - const file = pending.pop(); - if (!file || visited.has(file)) continue; - visited.add(file); - const specs: string[] = []; - const visit = (node: ts.Node): void => { - if ( - (ts.isImportDeclaration(node) || ts.isExportDeclaration(node)) && - node.moduleSpecifier && - ts.isStringLiteral(node.moduleSpecifier) - ) { - specs.push(node.moduleSpecifier.text); - } else if ( - ts.isCallExpression(node) && - node.expression.kind === ts.SyntaxKind.ImportKeyword && - node.arguments[0] && - ts.isStringLiteral(node.arguments[0]) - ) { - specs.push(node.arguments[0].text); - } - if (includeDynamic) ts.forEachChild(node, visit); - }; - ts.createSourceFile( - `${file}.js`, - transpiled(file), - ts.ScriptTarget.ES2022, - ).statements.forEach(visit); - for (const spec of specs) { - const base = spec.startsWith('.') - ? path.resolve(REPO_ROOT, path.dirname(file), spec) - : aliasTarget(spec, aliases); - const target = base && probe(base); - if (target) pending.push(toRepoRelative(target)); - } - } - - return [...visited].sort(); -} diff --git a/test/program-host.ts b/test/program-host.ts index 328c1d0d5..135cacda0 100644 --- a/test/program-host.ts +++ b/test/program-host.ts @@ -10,21 +10,14 @@ export function testProgramRunHost(session?: { context[key] = value; }, warn: () => undefined, - uploadEnvironmentVariables: () => Promise.resolve([]), }; } export function testProgramCiHost(): ProgramCiHost { return { - auth: { - setCredentials: () => undefined, - setRoleAtOrganization: () => undefined, - setApiUser: () => undefined, - }, log: { info: () => undefined, warn: () => undefined, }, - onProgress: () => undefined, }; } diff --git a/tsconfig.json b/tsconfig.json index ed78df9c4..b3b6d448f 100644 --- a/tsconfig.json +++ b/tsconfig.json @@ -17,8 +17,6 @@ "vitest.config.ts", "spec/**/*", "src/**/*", - "scripts/**/*.ts", - "scripts/**/*.tsx", "test/**/*", "e2e-harness/**/*", "types/**/*" From a083676c22a0fee0b279bcc805c0a7ad2435daca Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Thu, 24 Sep 2026 19:25:18 -0400 Subject: [PATCH 14/29] docs(programs): document runProgram, its host and the agent contracts Describe the callable program and agent interfaces as the tree at cbbe9fa2 implements them. - docs/developer-interfaces.md: what runProgram is for and who hosts it, the signatures, every ProgramInput, ProgramSettings and ProgramOptions field, the outcome, the pipeline order, cancellation, the failure outcomes, how the legacy adapter builds run, the host capabilities, and runAgent for detection and standalone callers. - src/programs/README.md: the host, the store, the adapter, the host capabilities and the current limits. - src/agent/README.md and src/agent/runner/README.md: prompt, collectTranscript, requestRemark, the transcript tail, the activity event, scanReport: 'defer', and who calls runAgent. - AGENTS.md: DEFAULT_BINDING is Pi + linear, program callbacks use their host, and the programs boundary links its docs. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- .claude/skills/adding-skill-program/SKILL.md | 4 +- .claude/skills/wizard-development/SKILL.md | 6 +- .../references/ARCHITECTURE.md | 18 +- AGENTS.md | 46 +- README.md | 28 +- docs/benchmarking.md | 2 +- docs/developer-interfaces.md | 760 +++++++++++++----- docs/local-dev.md | 3 - e2e-harness/ARCHITECTURE.md | 2 +- src/agent/README.md | 274 ++++--- src/agent/runner/README.md | 200 ++--- src/agent/runner/harness/anthropic/README.md | 4 +- src/agent/runner/sequence/README.md | 9 +- src/programs/README.md | 475 ++++------- src/shared/README.md | 2 +- src/shared/health-checks/testme.md | 2 +- 16 files changed, 1027 insertions(+), 808 deletions(-) diff --git a/.claude/skills/adding-skill-program/SKILL.md b/.claude/skills/adding-skill-program/SKILL.md index 28e4ed022..3aedd8778 100644 --- a/.claude/skills/adding-skill-program/SKILL.md +++ b/.claude/skills/adding-skill-program/SKILL.md @@ -62,8 +62,8 @@ changing execution behavior. 3. Register the config in [PROGRAM_REGISTRY](../../../src/programs/program-registry.ts) and add its Pi/orchestrator entry to - [PROGRAM_BINDINGS](../../../src/programs/binding.ts). - [Existing binding checks](../../../src/programs/__tests__/switchboard.test.ts) + [PROGRAM_BINDINGS](../../../src/agent/runner/switchboard/index.ts). + [Existing binding checks](../../../src/agent/runner/__tests__/switchboard.test.ts) enforce coverage; `ProgramId` currently widens to `string`. 4. For a standalone native command, create a command module with [nativeCommandFactory](../../../src/commands/factories/native-command-factory.ts) diff --git a/.claude/skills/wizard-development/SKILL.md b/.claude/skills/wizard-development/SKILL.md index c9819acd4..5fee5c7a5 100644 --- a/.claude/skills/wizard-development/SKILL.md +++ b/.claude/skills/wizard-development/SKILL.md @@ -22,7 +22,7 @@ infrastructure should consume those boundaries. | Framework detection, context, env conventions | [FrameworkConfig](../../../src/programs/framework-config.ts) and [framework configs](../../../src/programs/frameworks/) | | Integration instructions and orchestrator flows/tasks | [context-mill](https://github.com/PostHog/context-mill) | | Programs, steps, prerequisites and outcomes | [programs](../../../src/programs/) | -| Sequence, harness, model and effort selection | [program bindings](../../../src/programs/binding.ts) and [agent clamps](../../../src/agent/runner/switchboard/) | +| Sequence, harness, model and effort selection | [switchboard](../../../src/agent/runner/switchboard/) | | Local tool permissions and scanner adapters | [agent-interface](../../../src/agent/agent-interface.ts), [YARA hooks](../../../src/agent/yara-hooks.ts), [Pi security](../../../src/agent/runner/harness/pi/security.ts) | | Scanner rules | [warlock](https://github.com/PostHog/warlock) | | Token admission and budgets | [PostHog mint endpoint](https://github.com/PostHog/posthog/blob/master/posthog/llm/wizard_gateway_token.py) and [ai-gateway](https://github.com/PostHog/ai-gateway) | @@ -43,8 +43,8 @@ infrastructure should consume those boundaries. new Anthropic models. Existing routing has not all migrated: -[DEFAULT_AGENT_BINDING](../../../src/agent/default-binding.ts) selects Pi + linear -for standalone runs; programs apply their own binding and flag overrides. Set new +[DEFAULT_BINDING](../../../src/agent/runner/switchboard/index.ts) still +selects Anthropic + linear, with per-program and flag overrides. Set new bindings explicitly. Migrating an existing program requires checking its flow, tasks, and lifecycle hooks; changing the default constant alone is insufficient. Both harnesses implement `run` and `runTask`. diff --git a/.claude/skills/wizard-development/references/ARCHITECTURE.md b/.claude/skills/wizard-development/references/ARCHITECTURE.md index 72a0228b4..570d1e5e1 100644 --- a/.claude/skills/wizard-development/references/ARCHITECTURE.md +++ b/.claude/skills/wizard-development/references/ARCHITECTURE.md @@ -60,11 +60,10 @@ example. Native command modules still need registration in ## Switchboard contract -`resolveProgramBinding(ctx)` in [programs](../../../../src/programs/binding.ts) -is the routing seam. It receives the program, flag snapshot/payloads, -composition state and development overrides, returning a resolved binding and -stamping a trace of the selected precedence rungs. The agent receives this -binding and any pre-resolved task-role routes as data. +`resolveBinding(ctx, role?)` is the routing seam. Read its exported types rather +than copying their fields into another document. It receives the program, flag +snapshot/payloads, composition state, and development overrides, returning the +binding and stamping a trace of the selected precedence rungs. - Harness/model: development CLI override, declared flag route, per-program binding, default. @@ -73,8 +72,7 @@ binding and any pre-resolved task-role routes as data. [Harness](../../../../src/agent/runner/switchboard/harness.ts) and [sequence](../../../../src/agent/runner/switchboard/sequence.ts) contain the -generic precedence and clamp chains over caller-supplied policy. Published -builds omit CLI overrides. `RUN_SURFACE` can disable +exact chains. Published builds omit CLI overrides. `RUN_SURFACE` can disable harness experiments; static bindings and harness capabilities also affect resolution. Composed sub-runs remain linear even when a CLI override requests orchestration. @@ -86,11 +84,11 @@ admitted by the minted token and gateway. Local routing cannot bypass that external policy; see the [model admission checklist](../SKILL.md#execution-policy-and-model-admission). -Program flags belong in -[experiments](../../../../src/programs/experiments/). +Flags belong in +[switchboard/flags](../../../../src/agent/runner/switchboard/flags/). Experiments declare their program scope; malformed payloads yield no experiment route. Reuse the -[switchboard tests](../../../../src/programs/__tests__/switchboard.test.ts) +[switchboard tests](../../../../src/agent/runner/__tests__/switchboard.test.ts) and experiment tests to check full bindings and isolation of unrelated programs. Do not add a second flag-reading path inside a harness or sequence. diff --git a/AGENTS.md b/AGENTS.md index bff00c6db..e21a1667b 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -32,11 +32,12 @@ Each domain has a dedicated boundary: `docs/runbooks/warlock-kill-switch.md`. ONLY USE THIS IF ABSOLUTELY NECESSARY. - **Agent** → `src/agent/`, imported only through `@agent` (values) and `@agent/types` (types); see [src/agent/README.md](src/agent/README.md) -- **Shared** → `src/shared/`, stateless library code with no upward imports; - see [src/shared/README.md](src/shared/README.md) -- **Programs** → configs, detection, framework registry and task stream in - `src/programs/`; runtime and type entries are `@programs` and - `@programs/types` +- **Shared** → `src/shared/`, stateless library code with no upward imports; see + [src/shared/README.md](src/shared/README.md) +- **Programs** → configs, detection, framework registry, task stream and + `runProgram` in `src/programs/`. The runtime and type entries are `@programs` + and `@programs/types`. See [src/programs/README.md](src/programs/README.md) + and the [developer interfaces](docs/developer-interfaces.md) - **TUI** → screens, primitives and content decks in `src/ui/tui/` Adding a new concern means finding the narrowest existing surface, not adding @@ -76,9 +77,8 @@ Agent SDK is a supported legacy fallback, deprecated as the default; retain it for major Pi vulnerabilities or gaps in support for new Anthropic models. This is the contribution policy, not a claim that every existing binding has -migrated. The default, `DEFAULT_AGENT_BINDING`, is Pi + linear. Set new bindings -explicitly and check sequence-specific hooks before migrating existing flows. -See +migrated: `DEFAULT_BINDING` is Pi + linear. Set new bindings explicitly and +check sequence-specific hooks before migrating existing flows. See [execution policy and model admission](.claude/skills/wizard-development/SKILL.md#execution-policy-and-model-admission) for the gateway allowlists, required system prompt, and composition constraints. @@ -103,8 +103,8 @@ aliases. | Subcommand | What it audits | | ----------------------------- | ---------------------------------------------------- | -| `wizard audit events` | event capture quality + cost | -| `wizard audit all` | comprehensive audit across every area (**default**) | +| `wizard audit events` | event capture quality + cost | +| `wizard audit all` | comprehensive audit across every area (**default**) | | `wizard audit autocapture` | autocapture setup + cost | | `wizard audit feature-flags` | feature flag usage + cost | | `wizard audit identify` | `$identify` implementation | @@ -133,9 +133,8 @@ confuse it with the top-level `wizard skill` command. ([`src/commands/factories/native-command-factory.ts`](src/commands/factories/native-command-factory.ts)). - **Family commands** (e.g. `audit`) resolve subcommands at runtime against the `cliEntries` in `skill-menu.json`. Logic lives in - [`src/commands/dispatch-family.ts`](src/commands/dispatch-family.ts). - Adding a skill-backed subcommand is a **context-mill** release, not a wizard - change. + [`src/programs/dispatch-family.ts`](src/programs/dispatch-family.ts). Adding a + skill-backed subcommand is a **context-mill** release, not a wizard change. ### Commands vs. programs (don't confuse these) @@ -163,15 +162,12 @@ pnpm try --install-dir= # Run the wizard locally against a test proje pnpm build # Compile TypeScript pnpm test # Unit tests (builds first) pnpm test:watch # Unit tests in watch mode -pnpm test:e2e # Jest E2E suite on recorded fixtures (builds first) -pnpm test:e2e:tui # Live, credentialed: full TUI on a workbench app copy +pnpm test:e2e # End-to-end tests pnpm lint # Prettier + ESLint checks pnpm fix # Auto-fix lint issues pnpm dev # Build, link globally, watch for changes ``` -`test:e2e:tui` needs `APP_DIR`, `PROJECT_ID`, a personal key (`POSTHOG_PERSONAL_API_KEY` or `POSTHOG_KEY_FILE`) and `WIZARD_CI_GATEWAY_TOKEN_FILE`; see the Testing section of the [README](README.md). Headless `runProgram` and `runAgent` runs live in the [wizard-workbench](https://github.com/PostHog/wizard-workbench) harness, pointed at this checkout by `WIZARD_REPO`. - Choose verification for the change: check links and formatting for docs; run `pnpm typecheck` and focused existing tests for code. Build when bundling or runtime behavior changes. `pnpm test` already builds; avoid building twice. Use @@ -179,8 +175,8 @@ nonmutating lint checks, and scope formatting fixes to edited files. Do not add tests for prose, compiler-enforced shapes, or duplicated implementation. Keep new code comments to one line; put longer explanations in linked docs. -Local `--ci`, smoke-test, and full headless runs require two separate secrets: -a PostHog personal API key and an already-issued gateway token supplied through +Local `--ci`, smoke-test, and full headless runs require two separate secrets: a +PostHog personal API key and an already-issued gateway token supplied through `WIZARD_CI_GATEWAY_TOKEN_FILE`, plus the target project ID. Follow the [credential setup](docs/local-dev.md#credentials-for-local-ci-and-headless-runs). @@ -205,12 +201,14 @@ wizard run points. Full catalog: [`docs/local-dev.md`](docs/local-dev.md). - TypeScript everywhere. Use `type` (not `interface`) for framework context types so they satisfy `Record`. - All UI calls go through `getUI()` (returns `WizardUI` interface). Never import - the store directly from business logic. -- Shared helpers never call `getUI()`; they take a sink or return data. `debug()` - reaches the UI through the sink `src/ui/index.ts` installs. + the store directly from business logic. A program's `run` and `ciPreRun` + callbacks use the host they receive (`ProgramRunHost`, `ProgramCiHost`), not + `getUI()`. +- Shared helpers never call `getUI()`; they take a sink or return data. + `debug()` reaches the UI through the sink `src/ui/index.ts` installs. - Outside `src/agent`, import the agent through `@agent` or `@agent/types`. Add - to those entry modules rather than deep-importing; lint and - `pnpm test:arch` reject `@agent/*` paths elsewhere. + to those entry modules rather than deep-importing; lint and `pnpm test:arch` + reject `@agent/*` paths elsewhere. - Session mutations go through explicit store setters that call `emitChange()`. Never mutate `session` directly — nanostore holds a shallow copy. - The router resolves the active screen from session state. No imperative diff --git a/README.md b/README.md index 75239f93c..d1470fd36 100644 --- a/README.md +++ b/README.md @@ -167,7 +167,9 @@ route review to their owning team instead. | `src/programs/warehouse-source/` | `@PostHog/team-warehouse-sources` | | `src/programs/web-analytics-doctor/` | `@PostHog/team-web-analytics` | | `src/ui/tui/decks/error-tracking-upload-source-maps/` | `@PostHog/team-error-tracking` | +| `src/ui/tui/decks/revenue-analytics/` | `@PostHog/team-web-analytics` | | `src/ui/tui/decks/self-driving/` | `@PostHog/team-self-driving` | +| `src/ui/tui/decks/warehouse-source/` | `@PostHog/team-warehouse-sources` | Ownership is by directory. Programs not listed above (`agent-skill`, `audit`, `events-audit`, `mcp`, `migration`, `posthog-doctor`, @@ -393,11 +395,6 @@ that conventional code implies. If you want to use this code as a starting place for your own project, here's a quick explainer on its structure. -For code that runs without the terminal UI, see the -[non-interactive developer interfaces](docs/developer-interfaces.md), including -the [standalone agent](src/agent/README.md) and -[callable programs](src/programs/README.md). - ## Entrypoint: `run.ts` The entrypoint for this tool is `run.ts`. Use this file to interpret arguments @@ -565,29 +562,14 @@ To run unit tests, run: bin/test ``` -To run the jest E2E suite, which replays recorded LLM calls, run: +To run E2E tests run: ```bash bin/test-e2e ``` -See [`e2e-tests/README.md`](e2e-tests/README.md) to add or re-record tests. - -Live end-to-end runs are credentialed. Point `APP_DIR` at an app copy from -[wizard-workbench](https://github.com/PostHog/wizard-workbench), which owns the -fixture apps and the assertions: - -```bash -pnpm test:e2e:tui # the full TUI in a PTY, frames to SNAP_OUT -``` - -It reads `PROJECT_ID`, `POSTHOG_PERSONAL_API_KEY` or `POSTHOG_KEY_FILE`, and -`WIZARD_CI_GATEWAY_TOKEN_FILE`, and writes its result to `E2E_RESULT_JSON` when -set. - -The workbench also runs one program through `runProgram`, or one agent through -`runAgent`, with no TUI: `pnpm wizard-program` and `pnpm wizard-agent` there, -with `WIZARD_REPO` set to this checkout. +E2E tests are a bit more complicated to create and adjust due to to their mocked +LLM calls. See the `e2e-tests/README.md` for more information. #### Explore with an agent diff --git a/docs/benchmarking.md b/docs/benchmarking.md index 8d4a15bd9..e5fadf4f7 100644 --- a/docs/benchmarking.md +++ b/docs/benchmarking.md @@ -102,7 +102,7 @@ Everything below ships in this repo (`wizard/`) and its workbench orchestrator on pi, per-task models from context-mill frontmatter; off → the linear anthropic default). Per-stage variations ride `wizard-orchestrator-override` payloads (`{stage: {model?, effort?}}`, - variant keys in `wizard/src/programs/experiments/schemes.ts`). + variant keys in `wizard/src/agent/runner/switchboard/flags/schemes.ts`). The baseline is `{"wizard-orchestrator":"false"}` — never an empty override, or live remote flags leak into the baseline. diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index 105ce3490..baab17f8f 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -1,228 +1,604 @@ -# Non-interactive developer interfaces - -Wizard has two repository-local TypeScript call surfaces and one development CLI -mode for running without a terminal UI. The TypeScript aliases below are -internal to this repository. `@posthog/wizard` publishes a CLI, not these -functions as a stable package API. - -| Surface | Use it for | Detailed contract | -| --------------------------------------- | -------------------------------------------------------- | ------------------------------------------------------------------------------------------ | -| `runAgent(config, input, options)` | One already-configured AI run | [Agent reference](../src/agent/README.md) | -| `runProgram(programId, input, options)` | One program run the caller builds from a `ProgramConfig` | [Programs reference](../src/programs/README.md) | -| Development `--ci` | A process-owned, non-interactive CLI run | [Local CI credentials and recipe](local-dev.md#credentials-for-local-ci-and-headless-runs) | - -## Standalone agent - -`runAgent` takes a resolved `RunConfig` (run definition, binding, tools and -policy), `RunInput` (project, credentials, required inference-auth provider, -flags and host), and optional `onProgress`, `interaction`, and `signal` options. -Import the function from `@agent` and types from `@agent/types`. It returns a -`RunResult` with a `success`, `aborted`, `failed`, or `crashed` outcome and a -final task/status/usage snapshot. Progress is delivered in emission order. -Observer throws and rejections from thenables returned by `onProgress` are -logged without failing the run. Callbacks must handle errors from detached -asynchronous work they start. Without `interaction`, questions have no answer -bridge and optional task notices are declined. Non-success results carry a -`failure` with a code and message, and may have an attached `Error`. The caller -chooses how to log or present a failure. +# Developer interfaces + +Wizard has two TypeScript call surfaces for code that runs without the terminal +UI. `runProgram` runs one program. `runAgent` runs one agent. Both live in this +repository and import through its path aliases. The `@posthog/wizard` npm +package publishes the CLI, not these functions. + +| Surface | Import from | Use it for | +| --------------------------------------- | ----------------------------------------- | ----------------------------------------------------------------------- | +| `runProgram(programId, input, options)` | `@programs`, types from `@programs/types` | One program's agent run, with the program's policy, route and telemetry | +| `runAgent(config, input, options)` | `@agent`, types from `@agent/types` | One agent run from a resolved config, with no program policy | + +The [programs reference](../src/programs/README.md) covers the store, the +adapter and the current limits. The [agent reference](../src/agent/README.md) +covers the run contract. + +## What `runProgram` is for + +`runProgram` runs one program's agent from explicit inputs. It needs no +`WizardSession`, no TUI store and no `getUI()` call. The caller is the host. The +host supplies the run definition, the credentials, consent, the gate waits and +the answers. `runProgram` applies the program policy around the agent run. It +resolves credentials, checks AI-processing approval, waits for post-auth gates, +loads flags, refreshes the token and resolves the route. Then it calls +`runAgent` and returns one outcome. + +Two hosts call it: + +- **The legacy adapter.** `runProgramAgent` in + [`run-agent-legacy.ts`](../src/programs/run-agent-legacy.ts) serves the TUI + and the `--ci` runner. See + [how the legacy adapter builds `run`](#how-the-legacy-adapter-builds-run). +- **The workbench.** + [PostHog/wizard-workbench#4190](https://github.com/PostHog/wizard-workbench/pull/4190) + adds `services/wizard-program/`. Its `pnpm wizard-program` imports + `runProgram` from `@programs` in the wizard checkout that `WIZARD_REPO` names. + It runs one program against an app copy with no TUI. + +## Signatures ```ts -import { runAgent } from '@agent'; -import type { RunConfig, RunInput } from '@agent/types'; - -export async function runStandalone( - config: RunConfig, - input: RunInput, - signal?: AbortSignal, -) { - return runAgent(config, input, { - signal, - onProgress: (event) => { - if (event.kind === 'status') console.log(event.message); - }, - }); -} +import { runProgram } from '@programs'; +import type { + ProgramInput, + ProgramOptions, + ProgramRunOutcome, +} from '@programs/types'; + +// The shape `@programs` exports. +export const signature: ( + programId: string, + input: ProgramInput, + options?: ProgramOptions, +) => Promise = runProgram; ``` -The caller prepares `config` and `input`. The agent doesn't authenticate the -PostHog user or detect the project. The caller must supply -`input.inferenceAuth`, whose `resolve()` returns gateway authentication and can -refresh it during a long run. There is no session control protocol on this API. -An aborted signal returns an `aborted` result. It doesn't pause the run. The -agent returns a caught coded error as `failed` and an uncoded throw as -`crashed`. Both retain the caught `Error` (or an `Error` wrapper for a -non-`Error` throw), which the host can rethrow when it needs exception -semantics. `runAgent` never sends the terminal `setup wizard finished` event. -The host sends it when its process is done. - -A runnable reference host is the -[wizard-workbench](https://github.com/PostHog/wizard-workbench) harness, -`pnpm wizard-agent` with `WIZARD_REPO` set to a wizard checkout. It runs one -agent through `runAgent`, with no TUI and no programs. - -### Inference authentication - -For first-party inference authentication, import -`createPosthogInferenceAuthProvider` from `@programs` and pass authenticated -PostHog credentials and the run's program ID (for example, `config.programId`, -`'audit'`, or `'metrics'`). Its provider mints a gateway token and refreshes it -near expiry. The host still handles user login and project selection. -`runProgram` builds this provider itself when resolved credentials leave -`inferenceAuth` out, so you need it only when you call `runAgent` directly. -Development [CI](#development-ci-and-experimental-headless-runner) instead uses -an already-issued fixed token. - -## Callable program - -`runProgram` takes a program ID, a `ProgramInput` and optional `ProgramOptions`. -Import it from `@programs` and types from `@programs/types`. It doesn't look the -ID up in a registry. The ID picks the binding policy, the commandments, the -stage overrides and the analytics attribution. The caller builds the rest from -the program's `ProgramConfig`: - -- **`input.run`.** The required `AgentRunDefinition`. A static - `ProgramConfig.run` passes as is. A dynamic one reads a session and a - `ProgramRunHost`, so the caller resolves it first. -- **`input.program`.** The `ProgramSettings`: `requiresAi`, `agentFlow`, the - tool allow and deny lists, `excludedTaskTypes`, the audit ledger and its seed - checks, the event plan file and the post-auth gates. - -The host supplies credentials in one of two ways: - -- **Resolved.** `input.credentials` carries - `{ posthog, inferenceAuth?, project, apiUser }`. -- **A provider.** `options.credentials.resolve(programId, { signal })`. - `runProgram` calls it once, only when `input.credentials` is absent. - -Either way, `runProgram` identifies the user for analytics, stamps the -organization's AI SDK evidence, and refreshes an OAuth token that is close to -expiry before the agent starts. Launch choices go in `input.overrides` as -`{ harness?, sequence?, model? }`. `runProgram` resolves the binding from them -and the flag snapshot, and captures the switchboard decision. It copies the -input when it receives it, so a later host write can't reach the run. - -The other options are host capabilities: - -- **`awaitAiApproval({ programId, signal })`.** Answers the AI-processing - approval. Without it, a run that needs approval fails. -- **`awaitPostAuthGates({ programId, gates, signal })`.** Settles the post-auth - gates, such as a project picker. -- **`featureFlags()`.** Evaluates flags when the input has none. It gets no - signal. -- **`interaction`, `onProgress`, `deferSkillCommit` and `signal`.** Answer the - agent's questions, observe the run, leave new skills for the host to commit, - and cancel the invocation. - -A rejection from `credentials`, `awaitAiApproval`, `awaitPostAuthGates`, -`featureFlags` or the token refresh resolves as `failed`, or as `aborted` once -the signal has aborted. The promise rejects only when the invocation itself -breaks, such as input that can't be copied. Read the outcome, and still catch a -rejection. - -`runProgram` returns a `ProgramRunOutcome`: the outcome and failure, -`settledRuns` with the agent run's result, `diagnostics` for observer failures -and late events, `artifacts.reportFile`, and the invocation `data`, including a -captured event plan. The data contains credentials, so don't log it. Agent -failures retain an attached `Error` when one exists. - -`onProgress` receives two kinds of `ProgramProgress`. A run event is -`{ kind: 'run', runId, event }`, where `event` is the agent's progress. A -program-data event is `{ kind: 'program', data }`, a copy of the invocation data -after each write. Narrow on `kind` first: +`programId` names the program for analytics, the route, the gateway spend pin +and the commandments. `runProgram` doesn't look it up in the program registry. +An ID with no entry in `PROGRAM_BINDINGS` runs on `DEFAULT_BINDING`. + +`@programs/types` exports the input, settings, option, outcome and progress +types. The credential types are the exception. Name them as +`ProgramInput['credentials']` and `ProgramOptions['credentials']`. They live in +[`credentials.ts`](../src/programs/credentials.ts) as +`ResolvedProgramCredentials` and `CredentialsProvider`. + +## `ProgramInput` + +`installDir` and `run` are required. Every other field is optional. + +| Field | What `runProgram` does with it | +| ---------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `installDir` | The agent's working directory. `artifacts.reportFile` resolves against it. | +| `run` | The program's `AgentRunDefinition`, built by the host from its `ProgramConfig`. It becomes `RunConfig.run`. Its `reportFile` names the report. Its `integrationLabel` and `skillId` label analytics and traces. | +| `program` | The `ProgramSettings` read from the same `ProgramConfig`. See [`ProgramSettings`](#programsettings). Defaults to `{}`. | +| `credentials` | A resolved login: `{ posthog, project, apiUser }`. It wins over `options.credentials`. | +| `runId` | Labels the agent run in `settledRuns` and in run progress events. A random UUID when absent. It isn't the analytics run ID. | +| `overrides` | `{ harness?, sequence?, model? }`, the launch overrides such as `--harness`. The switchboard applies them in development and test builds and ignores them in published builds. | +| `composed` | `true` for a sub-run inside a host program. The switchboard clamps it to linear, and the agent leaves the terminal outro to the host. Defaults to `false`. | +| `skillId` | The run's skill, for question attribution and the result's `skillId`. The orchestrator uses it as the framework key when `integration` is absent. Defaults to `run.skillId`, then `run.integrationLabel`. | +| `integration` | The detected framework. | +| `frameworkDocsUrl` | The framework's docs page, for the orchestrator's preflight message. | +| `flags` | Run flags: `ci`, `signup`, `debug`, `e2eAsk`, `localMcp`, `captureAio`, `benchmark` and `yaraReport`. A missing flag is `false`. | +| `host` | Where PostHog is: `baseUrl`, `region`, `email`, `projectId` and `apiKey`. `baseUrl` also points the token refresh. | +| `wizardFlags` | An evaluated feature flag snapshot. When it's absent, `runProgram` calls `options.featureFlags`. | +| `wizardFlagPayloads` | The flag payloads from the same snapshot. | +| `seedTasks` | Returns the tasks the orchestrator queues before its planner runs. | +| `hooks` | `RunHooks`, each called with the run's credentials: `postRun`, `buildOutroData`, `buildOutroNextSteps` and `recordTaskOutcomes`. | +| `discoveredFeatures`, `warehouseSources`, `mayReportScanResults` | Evidence for the organization's AI SDK stamp. When `mayReportScanResults` is absent or `false`, no stamp is sent. | +| `aiSdkStampReported` | `true` when the host already considered the stamp for this login. `runProgram` then skips it. | + +The agent calls completion hooks only through `hooks`. It never calls the +`postRun`, `buildOutroData` or `buildOutroNextSteps` of a `ProgramRun`, because +those take a session. Bind them into `hooks` yourself. The linear sequence +installs `run.skillId`, not `input.skillId`. + +`runProgram` copies the input when it receives it. `credentials`, `run`, +`program`, `hooks` and `seedTasks` stay by reference, because they can carry +functions or class instances. `structuredClone` copies every other field. A +later host write to the input doesn't reach the run. A copied field that +`structuredClone` can't copy, such as a function in `host`, rejects the call. + +## `ProgramSettings` + +`input.program` carries the program-level settings from `ProgramConfig`. + +| Field | What `runProgram` does with it | +| ------------------- | -------------------------------------------------------------------------------------------------------------- | +| `requiresAi` | `false` skips the AI-processing approval. Any other value, including absent, keeps the check. | +| `agentFlow` | The context-mill flow the orchestrator loads. Defaults to the program ID. | +| `allowedTools` | Tools added on top of the base allowed set. | +| `disallowedTools` | Tools removed from the base allowed set. | +| `excludedTaskTypes` | `(flags) => types`. The task types this run excludes for its flag snapshot. | +| `postAuthGates` | Step IDs the host settles after auth and before the agent starts. `runProgram` passes them to the gate option. | + +## `ProgramOptions` + +Every option is optional. An awaited capability receives the invocation's +`signal`. Without `options.signal`, it receives a signal that never aborts. + +| Option | What `runProgram` does with it | When it's absent | +| -------------------- | --------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------- | +| `credentials` | Calls `resolve(programId, { signal })` once, only when `input.credentials` is absent. | Without `input.credentials`, the run fails. | +| `awaitAiApproval` | Calls it with `{ programId, signal }` when the run needs approval. `true` proceeds. `false` aborts. | A run that needs approval fails. | +| `awaitPostAuthGates` | Calls it once with `{ programId, gates, signal }` when `program.postAuthGates` isn't empty. | The run doesn't wait. | +| `featureFlags` | Calls it once when `input.wizardFlags` is absent. It receives no signal. | The run uses the input's flags, or none. | +| `interaction` | Passes it to `runAgent`, which asks it the agent's questions and task notices. | The agent installs no ask bridge and declines optional task notices. | +| `onProgress` | Receives run events and data snapshots. See [progress](#progress). It is never awaited. | The outcome still carries the settled run and the final data. | +| `signal` | Cancels the invocation. See [cancellation](#cancellation). | Nothing cancels it. | + +A run needs approval when all three hold: + +- `program.requiresAi` isn't `false`. +- Neither `flags.ci` nor `flags.signup` is set. +- The organization's `is_ai_data_processing_approved` isn't `true`. + +`flags.ci` and `flags.signup` skip the check. Set them only when the host has +already handled consent. + +## The outcome + +`runProgram` resolves with a `ProgramRunOutcome`. + +| Field | What it holds | +| ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | +| `programId` | The ID the call ran. | +| `outcome` | A `RunOutcome`: `success`, `aborted`, `failed` or `crashed`. | +| `failure` | Present on every outcome except `success`. An `AgentFailure` with a `code` and a `message`. An agent failure can also carry `outroData`, `error`, `exitCode`, `detail` and `authErrorDetail`. A crash always carries the original `error`. | +| `settledRuns` | `{ runId, result }` for the agent run, once `runAgent` returned. Empty when the invocation ended before the agent. | +| `data` | A copy of the invocation's `ProgramInvocationData` when it settled. | +| `diagnostics` | Up to 10 observer failures and late events, newest last. | +| `artifacts.reportFile` | The absolute report path: `run.reportFile` resolved against `installDir`. Set just before the agent starts. It doesn't prove the agent wrote the file. | + +`ProgramInvocationData` has these fields: + +| Field | What it holds | +| -------------------------------------- | -------------------------------------------------------------------------------------------------------------- | +| `credentials`, `apiProject`, `apiUser` | The resolved login, updated when the token refresh returns new credentials. `null` until credentials resolve. | +| `detection.frameworkContext` | Always `{}` from `runProgram`. Framework context reaches a program through its [run host](#host-capabilities). | +| `binding` | The resolved route: `sequence`, `harness`, `model` and `thinkingLevel`. `null` until the route resolves. | +| `aiSdkStampReported` | `true` once the organization's AI SDK stamp was considered. | + +Data and progress are `structuredClone` copies. A class instance comes back as a +plain object: `data.credentials.host` has the `HostResolution` fields but isn't +an instance. `data` holds PostHog tokens, so don't log it. + +`runProgram` writes to the process-wide analytics client. It sends +`agent started` and `switchboard resolved`, sets the `sequence` and `harness` +tags, and identifies the user. It never sends `setup wizard finished`. The host +sends it from the outcome with `analytics.shutdown(status)`. + +### Progress + +`onProgress` receives a `ProgramProgress`, a union of two kinds. Narrow on +`kind` first. + +- **`{ kind: 'run', runId, event }`.** One agent event. `event` is a copied + [`AgentProgress`](../src/agent/README.md#signatures). +- **`{ kind: 'program', data }`.** A copy of `ProgramInvocationData`, sent after + each write. A write happens when credentials resolve, when the AI SDK stamp + latches, when the refresh returns new credentials, and when the route + resolves. + +`runProgram` never awaits the observer. A throw or a rejection becomes a +diagnostic. So does a run event that arrives after the agent run finished. The +outcome copies `diagnostics` when it settles, so a rejection that lands later +isn't in it. + +## Pipeline order + +1. Copy the input. If `signal` is already aborted, return `aborted`. +2. Send the `agent started` analytics event. +3. Resolve credentials from `input.credentials`, else from + `options.credentials`. Without either, fail. +4. Store the login. Identify the user for analytics and set the organization + group. +5. Consider the AI SDK stamp once, unless `aiSdkStampReported` is `true`. +6. When the run needs approval, await `awaitAiApproval`. +7. When `program.postAuthGates` isn't empty, await `awaitPostAuthGates`. +8. When `input.wizardFlags` is absent, await `featureFlags`. +9. Refresh the OAuth token when it has a refresh token, an expiry, and less than + 50 minutes left. A failed refresh keeps the current token. +10. Resolve the route from the program ID, `composed`, the flags and + `overrides`. Tag analytics with the sequence and harness, and send + `switchboard resolved`. +11. Build the run tags, set `artifacts.reportFile`, and call `runAgent` with + `interaction`, the run's progress and `signal`. +12. Record the agent's result and settle. + +The [runner reference](../src/agent/runner/README.md) describes what `runAgent` +does from step 11. + +## Cancellation + +`runProgram` checks `signal` before it starts and after each await in steps 3 +to 9. An abort at any of those points returns `aborted` with the message +`Run cancelled by host.`, even when the capability already resolved. A +capability that rejects after the abort also returns `aborted`. + +`runProgram` doesn't race a capability against the signal. It waits for the +capability to settle, so a capability must settle when its signal aborts. +`featureFlags` receives no signal. + +After step 9, `runProgram` doesn't check the signal itself. `runAgent` receives +it. An abort before the agent's setup returns `aborted` right away. An abort +during the run returns the agent's `aborted` result. Both carry the message +`Agent run cancelled`. + +## Failure outcomes + +Most endings resolve. A rejected promise means the call itself broke. + +| Cause | `outcome` | `failure.code` | `failure.message` | +| -------------------------------------------------------- | ------------------- | ------------------------ | ----------------------------------------------------------------- | +| No `input.credentials` and no `options.credentials` | `failed` | `PHW_INTERNAL_UNHANDLED` | `Credentials are required to run .` | +| The run needs approval and there is no `awaitAiApproval` | `failed` | `PHW_INTERNAL_UNHANDLED` | `AI processing approval is required before this program can run.` | +| `awaitAiApproval` resolves `false` | `aborted` | `PHW_AGENT_ABORT` | `AI processing approval declined.` | +| An await in steps 3 to 9 rejects | `failed` | `PHW_INTERNAL_UNHANDLED` | The error's message | +| The host's signal aborts before the agent | `aborted` | `PHW_AGENT_ABORT` | `Run cancelled by host.` | +| The agent run doesn't succeed | The agent's outcome | The agent's code | The agent's message | +| A copied input field can't be cloned | The promise rejects | | | + +The agent returns `failed` for a coded error, a decided failure or an agent that +stops itself with `[ABORT]`. It returns `aborted` only for the host's signal. It +returns `crashed` for an uncoded throw, with the original `Error` attached. See +the [agent reference](../src/agent/README.md#signatures). + +## Example + +This host runs the `metrics` program with a login it already holds. It asks its +own consent flow for approval: ```ts +import { RunOutcome } from '@agent'; import { getProgramConfig, runProgram } from '@programs'; -import type { ProgramOptions } from '@programs/types'; +import type { ProgramInput, ProgramOptions } from '@programs/types'; + +type Login = NonNullable; +type AskForApproval = NonNullable; export async function runMetrics( installDir: string, - credentials: NonNullable, - awaitAiApproval: NonNullable, + login: Login, + askForApproval: AskForApproval, signal?: AbortSignal, -) { +): Promise { const config = getProgramConfig('metrics'); if (!config.run || typeof config.run === 'function') { throw new Error('metrics has a static run definition'); } + const result = await runProgram( config.id, { installDir, run: config.run, - program: { agentFlow: config.agentFlow, requiresAi: config.requiresAi }, + credentials: login, + program: { + requiresAi: config.requiresAi, + agentFlow: config.agentFlow, + allowedTools: config.allowedTools, + disallowedTools: config.disallowedTools, + excludedTaskTypes: config.excludedTaskTypes, + }, }, { - credentials, - awaitAiApproval, + awaitAiApproval: askForApproval, signal, onProgress: (progress) => { if (progress.kind !== 'run') return; - if (progress.event.kind === 'tasks') { - console.log(progress.runId, progress.event.tasks); + if (progress.event.kind === 'status') { + console.log(progress.event.message); } }, }, ); - if (result.outcome !== 'success') { - if (result.failure?.error) throw result.failure.error; - throw new Error(result.failure?.message ?? `Metrics ${result.outcome}`); + + if (result.outcome !== RunOutcome.Success) { + throw result.failure?.error ?? new Error(result.failure?.message); } return result.artifacts.reportFile; } ``` -The [program reference](../src/programs/README.md#inputs) lists every input, -setting and option. There is no live store or step-control handle. `runProgram` -never sends the terminal `setup wizard finished` event. A long-lived host -decides when to send it, from the outcome. - -A runnable reference host is the workbench harness's `pnpm wizard-program`. It -builds the run and settings from the `ProgramConfig`, the way the session -adapter does, and runs one program against the app in `APP_DIR`, with resolved -credentials and no TUI. - -### What stays with the host - -`runProgram` runs one agent. It doesn't check service readiness or Claude -settings, walk composed steps, or run a program with no agent. The session -adapter, `src/lib/runners/run-program-agent.ts`, runs the readiness and settings -gates before it calls `runProgram`. The TUI walks each composed step as its own -call with `composed: true`, and runs the steps of programs with no agent, such -as `posthog-doctor`, `mcp-add` and `slack`. See the -[program reference](../src/programs/README.md#current-limits). - -## Development CI and experimental headless runner - -Development/test builds accept `--ci`. This is a whole-process CLI path, not an -awaitable function returning `ProgramRunOutcome`. It requires an install -directory, a PostHog personal API key, a project ID, and an already-issued -gateway token in the file named by `WIZARD_CI_GATEWAY_TOKEN_FILE`: - -```bash -WIZARD_CI_GATEWAY_TOKEN_FILE="$HOME/.config/posthog/wizard-gateway-token" \ -pnpm try --ci --api-key "$POSTHOG_PERSONAL_API_KEY" \ - --project-id "$POSTHOG_WIZARD_PROJECT_ID" \ - --region us --install-dir /absolute/path/to/test-app +## How the legacy adapter builds `run` + +`runProgramAgent(programConfig, session, { composed? })` is the host for the TUI +and for `--ci`. It builds the input from the `ProgramConfig` and the session, +then calls `runProgram` once: + +1. It throws when the config has no `run`. +2. It starts the audit ledger watcher when the config names `auditLedgerFile`, + before `run` resolves. It stops the watcher when the run returns. +3. It resolves `run`. A static `ProgramRun` passes as is. A function is called + as `run(session, host)`, with a [`ProgramRunHost`](#host-capabilities) over + `getUI()`. +4. It sets `session.skillId`, starts the log file, and runs the health gate and + the Claude settings gate. +5. It calls `runProgram` with the input and options below. +6. It applies the outcome. + +The input comes from the config and the session: + +| `ProgramInput` field | Source | +| -------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- | +| `run` | `ProgramConfig.run`, resolved in step 3. | +| `program` | `requiresAi`, `agentFlow`, `allowedTools`, `disallowedTools` and `excludedTaskTypes` from the config. | +| `program.postAuthGates` | The IDs of `postAuthGateSteps(config.steps)`, the gated steps between the `auth` screen and the `run` screen. | +| `hooks` | `run.postRun`, `run.buildOutroData` and `run.buildOutroNextSteps`, bound to the session. | +| `hooks.recordTaskOutcomes` | Writes the drained queue's outcomes to `session.frameworkContext[TASK_OUTCOMES_KEY]`. | +| `seedTasks` | `config.seedTasks`, bound to the session. | +| `installDir`, `composed`, `skillId`, `integration` | The session and the `composed` option. | +| `overrides` | `session.harness`, `session.sequence` and `session.model`. | +| `frameworkDocsUrl` | The `FRAMEWORK_REGISTRY` docs URL for `session.integration`, else for `session.skillId`. | +| `flags`, `host` | The session's run flags, and its `baseUrl`, `region`, `email`, `projectId` and `apiKey`. | +| `aiSdkStampReported`, `discoveredFeatures`, `warehouseSources`, `mayReportScanResults` | The session's stamp latch, detection results and scan consent. | + +The options come from the UI: + +| `ProgramOptions` field | Source | +| ---------------------- | ------------------------------------------------------------------------------------- | +| `credentials` | `authenticate(session, programId)`, then the session's credentials, project and user. | +| `awaitAiApproval` | Waits for the AI opt-in screen to clear, then resolves `true`. | +| `awaitPostAuthGates` | Waits for each gate in order with `ui.waitForGate`. | +| `featureFlags` | The analytics client's flags and payloads. | +| `onProgress` | Run events go to the UI reducer. Data snapshots project onto the session and the UI. | +| `interaction` | The UI's ask overlay and task notice modal. | +| `signal` | Not set. | + +The projection copies a refreshed token onto the session and the UI, and latches +the AI SDK stamp on the session. When the route first resolves, it registers a +cleanup that flushes the scan report. On a linear route, it also restores the +Claude settings when the outro screen opens. + +The adapter applies the outcome the way the CLI roots expect: + +- A login failure or a flag load failure rethrows after `runProgram` resolves. +- `crashed` rethrows `failure.error`. +- Any other non-success shows the auth error screen when `authErrorDetail` is + set. Then it calls `wizardAbort` with the failure, with status `cancelled` for + `aborted` and `error` otherwise. +- A `success` that isn't composed sends `setup wizard finished` through + `analytics.shutdown('success')`. A failed flush is logged, and the run stays a + success. + +A host with no TUI can build the input the same way. A dynamic `run` and the +program hooks take a `WizardSession`, so the host builds one and fills what the +program reads: + +```ts +import { buildSession } from '@lib/wizard-session'; +import { getProgramConfig } from '@programs'; +import { postAuthGateSteps } from '@programs/program-step'; +import type { ProgramId, ProgramInput, ProgramRunHost } from '@programs/types'; + +export async function buildProgramInput( + programId: ProgramId, + installDir: string, + host: ProgramRunHost, +): Promise { + const config = getProgramConfig(programId); + if (!config.run) throw new Error(`${programId} has no run`); + + const session = buildSession({ installDir, ci: true }); + const run = + typeof config.run === 'function' + ? await config.run(session, host) + : config.run; + const { postRun, buildOutroData, buildOutroNextSteps } = run; + const { seedTasks } = config; + + return { + installDir, + run, + program: { + requiresAi: config.requiresAi, + agentFlow: config.agentFlow, + allowedTools: config.allowedTools, + disallowedTools: config.disallowedTools, + excludedTaskTypes: config.excludedTaskTypes, + postAuthGates: postAuthGateSteps(config.steps).map((step) => step.id), + }, + seedTasks: seedTasks && (() => seedTasks(session)), + hooks: { + postRun: postRun && ((credentials) => postRun(session, credentials)), + buildOutroData: + buildOutroData && + ((credentials) => buildOutroData(session, credentials) ?? undefined), + buildOutroNextSteps: + buildOutroNextSteps && + ((credentials, completed) => + buildOutroNextSteps(session, credentials, completed)), + }, + }; +} +``` + +Some programs read session fields that a TUI step or `ciPreRun` fills. For +example, the `posthog-integration` run reads `session.frameworkConfig`, which +its detect step sets. + +## Host capabilities + +A program's `run` and `ciPreRun` callbacks take their effects from a host. They +don't call `getUI()`. Both types come from `@programs/types`: + +```ts +import type { ProgramCiHost, ProgramRunHost } from '@programs/types'; + +const frameworkContext = new Map(); + +export const runHost: ProgramRunHost = { + getFrameworkContext: (key) => frameworkContext.get(key), + setFrameworkContext: (key, value) => { + frameworkContext.set(key, value); + }, + warn: (message) => console.warn(message), +}; + +export const ciHost: ProgramCiHost = { + log: { + info: (message) => console.log(message), + warn: (message) => console.warn(message), + }, +}; +``` + +| Capability | Received by | What programs use it for | Supplied by | +| ---------------- | --------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------- | +| `ProgramRunHost` | `ProgramConfig.run(session, host)` | `posthog-integration` warns when the SDK package or `package.json` is missing. `error-tracking` warns from the `posthog-cli` preinstall. `error-tracking-upload-source-maps` reads the picked project and writes the completed variant. `audit` passes it to its base run. | The legacy adapter. Each call reads `getUI()` when it runs. | +| `ProgramCiHost` | `ProgramConfig.ciPreRun(session, host)` | Project scoping logs its scan progress and its fallbacks. `posthog-integration`, `error-tracking` and `replay-vision` scope the project this way. | The non-interactive runner. Each call reads `getUI().log` when it runs. | + +A run definition keeps its host. The source maps program reads the picked +project when the agent's prompt is built, after the picker screen, and writes +`sourceMapsCompletedVariant` in `postRun`. So a `ProgramRunHost` must stay +usable until the run returns. + +Neither capability is a `runProgram` option. Both belong to building the run +from a `ProgramConfig`, before `runProgram` starts. + +## `runAgent` for detection and standalone callers + +Call `runAgent` directly when you have a resolved route and no program policy to +apply. It doesn't authenticate, check consent, load flags, resolve a route or +send `setup wizard finished`. The caller does those. + +```ts +import { runAgent } from '@agent'; +import type { + AgentInteraction, + AgentProgress, + RunConfig, + RunInput, + RunResult, +} from '@agent/types'; + +// The shape `@agent` exports. +export const signature: ( + config: RunConfig, + input: RunInput, + options?: { + onProgress?: (event: AgentProgress) => unknown; + interaction?: AgentInteraction; + signal?: AbortSignal; + }, +) => Promise = runAgent; +``` + +These callers use it: + +| Caller | What it runs | +| --------------------------------------------------------------------------------- | ------------------------------------------------------------------ | +| `runProgram` | One program's agent run. | +| `detectProjectsWithAgent` in [`agentic.ts`](../src/programs/detection/agentic.ts) | The agentic project scan, one `runAgent` call per attempt. | +| [`a3-fault-probe.no-jest.ts`](../scripts/a3-fault-probe.no-jest.ts) | A fault probe against a local gateway, with synthetic credentials. | + +`runAgent` mints a gateway token from `input.credentials` for `config.programId` +before any agent starts, and re-mints near expiry. A refused mint returns +`failed`. In development and test builds, +`configureGatewayFromCIEnvironment(projectId, region)` from `@agent` loads a +fixed gateway token from the file that `WIZARD_CI_GATEWAY_TOKEN_FILE` names. +`runAgent` then uses that token and doesn't mint. The gateway auth is +process-wide. + +### Detection + +Agentic detection scans a repository for its projects through `runAgent`. +`scopeInstallDirToProject`, the self-driving and error tracking project scans, +and the source maps scan call it. It needs `session.credentials` and throws +without them. Each attempt passes this run: + +| Setting | Value | +| ----------------------- | --------------------------------------------------------------------------- | +| `binding` | Linear, Anthropic, Haiku. | +| `run.prompt` | The scan prompt. It replaces the assembled project prompt. | +| `run.collectTranscript` | `true`. The report is read from `snapshot.transcriptTail`. | +| `run.requestRemark` | `false`. | +| `composed` | `true`. | +| `allowedTools` | `Read`, `Grep` and `Glob`. | +| `scanReport` | `'defer'`. The scan's security scans count toward the program run's report. | +| `programId` | The caller's program, so the scan's spend is attributed to it. | +| `wizardMetadata` | The run tags plus `call_type: detection`. | + +Each attempt has its own deadline signal. A first attempt that times out retries +once. A second timeout throws `AgenticDetectionTimeoutError`. Any other +non-success throws the failure's error. The scan parses verdict lines from the +transcript tail. A tail with no verdicts that contains `[ABORT]` is an empty +report. A tail with no verdicts retries once, then throws. + +Detection forwards each `activity` line to its `onEvent` callback. The UI +receives every other event except `lifecycle`, `completion`, `spinner`, and log +lines below `warn`. + +### Standalone example + +```ts +import { runAgent, RunOutcome } from '@agent'; +import type { RunConfig, RunInput } from '@agent/types'; +import { + getSkillsBaseUrl, + Harness, + HAIKU_MODEL, + Sequence, +} from '@shared/constants'; + +export async function listProjectFiles( + programId: string, + input: RunInput, + signal?: AbortSignal, +): Promise { + const config: RunConfig = { + programId, + run: { + integrationLabel: 'list-files', + prompt: () => 'List the files in the working directory. Change nothing.', + collectTranscript: true, + requestRemark: false, + spinnerMessage: 'Listing files...', + successMessage: 'Listed files', + estimatedDurationMinutes: 1, + reportFile: '', + docsUrl: 'https://posthog.com/docs', + }, + composed: true, + binding: { + sequence: Sequence.linear, + harness: Harness.anthropic, + model: HAIKU_MODEL, + }, + switchboard: { program: programId, composed: true, flags: {} }, + skillsBaseUrl: getSkillsBaseUrl(), + wizardFlags: {}, + wizardFlagPayloads: {}, + wizardMetadata: {}, + allowedTools: ['Read', 'Glob'], + }; + + const result = await runAgent(config, input, { + signal, + onProgress: (event) => { + if (event.kind === 'activity') console.log(event.line); + }, + }); + if (result.outcome !== RunOutcome.Success) { + throw result.failure.error ?? new Error(result.failure.message); + } + return result.snapshot.transcriptTail ?? ''; +} ``` -The runner logs progress and writes a local task-stream JSONL dump. Callers -observe the process exit and its logs, rather than a returned result. The -gateway token file is read into a fixed provider for CI. Pre-run detection and -the program's agent run use that same provider. This path doesn't mint or -refresh the token. Published builds reject `--ci`. The internal -`runWizardCI(config, options): void` entry point uses the session adapter, -`src/lib/runners/run-program-agent.ts`. The adapter builds the run from the -`ProgramConfig`, runs the readiness and settings gates, then calls `runProgram` -once. Agentic detection runs before that call, through its own `runAgent` call. -MCP suggested prompts use a separate SDK path with their own progress and -cancellation. - -An experimental published-build headless path exists internally as -`runWizardHeadless(config, options): void`. It shares the process-owned runner, -logs progress, and can push task-stream updates to PostHog when telemetry is -enabled. Its selector is deliberately hidden and is not a supported invocation -recipe. Neither internal function returns a structured, awaitable outcome. - -There is no controlled headless mode and no socket control API. No supported -route, command, or event protocol pauses a run, supplies an answer later, or -reads its live state from another process. +`collectTranscript` and `requestRemark: false` take effect on the linear +sequence with the Anthropic harness. See the +[agent reference](../src/agent/README.md#run-definition). + +## Process-owned CLI runs + +Development and test builds accept `--ci`. It is a whole-process run, not a +function that returns an outcome. The non-interactive runner installs the +`LoggingUI`, builds a session, and loads the fixed CI gateway token. It runs the +program's `ciPreRun` with a `ProgramCiHost`, or the steps' `onReady` hooks. Then +it calls `runProgramAgent(config, session)`. The process exit code and the logs +report the result. Published builds don't accept `--ci`. For the credentials it +needs, see +[local credentials](local-dev.md#credentials-for-local-ci-and-headless-runs). diff --git a/docs/local-dev.md b/docs/local-dev.md index 5fc65f934..93f949cdd 100644 --- a/docs/local-dev.md +++ b/docs/local-dev.md @@ -3,9 +3,6 @@ Running the wizard against local servers. Four things can independently be local, and this doc is the catalog of how to control each. -For the callable agent and program contracts, see the -[non-interactive developer interfaces](developer-interfaces.md). - ## Credentials for local CI and headless runs Local `--ci` runs, smoke tests, and full headless/snapshot agent runs need diff --git a/e2e-harness/ARCHITECTURE.md b/e2e-harness/ARCHITECTURE.md index 2c62f9be4..5b17c8d56 100644 --- a/e2e-harness/ARCHITECTURE.md +++ b/e2e-harness/ARCHITECTURE.md @@ -137,7 +137,7 @@ Two decision points ask a person to act, and the harness stands in for them. `wizard_ask` overlay. A `ci` session normally has no ask bridge at all. The host sets `session.e2eAsk` from `E2E_ASK=true`, which keeps the bridge wired (see -`isAskDisabled`). In the fixed route, the profile answers every question: +`shouldDisableAsk`). In the fixed route, the profile answers every question: `askAnswers` routes a question to a value, else the first option, else the `'e2e'` sentinel. Route credentials with `${ENV_VAR}` values, never literals. diff --git a/src/agent/README.md b/src/agent/README.md index a4f387d90..8fba26f1b 100644 --- a/src/agent/README.md +++ b/src/agent/README.md @@ -1,10 +1,8 @@ # Agent -The agent runs one program's AI pipeline against a project directory. It takes -resolved data in, reports through progress events, asks through an injected -answerer, and returns a result. It never reads a session, a store or a UI. For -the callable program host and development CI runner, see the -[non-interactive developer interfaces](../../docs/developer-interfaces.md). +The agent runs one AI pipeline against a project directory. It takes resolved +data in, reports through progress events, asks through an injected answerer, and +returns a result. It never reads a session, a store or a UI. ## Signatures @@ -12,6 +10,7 @@ Import runtime values from `@agent` and types from `@agent/types`. Nothing outside `src/agent` imports deeper. Lint and the architecture test reject it. ```ts +import { runAgent } from '@agent'; import type { AgentInteraction, AgentProgress, @@ -20,8 +19,8 @@ import type { RunResult, } from '@agent/types'; -// The shape of `runAgent`, exported from `@agent`. -declare function runAgent( +// The shape `@agent` exports. +export const signature: ( config: RunConfig, input: RunInput, options?: { @@ -29,158 +28,178 @@ declare function runAgent( interaction?: AgentInteraction; signal?: AbortSignal; }, -): Promise; +) => Promise = runAgent; ``` -- `RunConfig`: the opaque program id, its `AgentRunDefinition` (prompt, skill, - tools, copy), the resolved `binding` (sequence, harness, model and task-role - routes), supplied program commandments and stage policy, the skills origin, - flag snapshot, trace tags, tool allow and deny lists, seed tasks, bound - completion `hooks` and `scanReport`. `scanReport: 'defer'` leaves this run's - scans to the host run's report. The default, `'flush'`, writes the report when - this run ends. -- Two `AgentRunDefinition` options shape the run's output. - `collectTranscript: true` keeps the last 256 KiB of assistant text as - `snapshot.transcriptTail` and reports each step as `activity` progress. It - works on the linear sequence with the Anthropic harness. - `requestRemark: false` skips the end-of-run reflection remark, which is on by - default. -- `RunInput`: install directory, resolved PostHog credentials, required - `inferenceAuth`, project and user payloads, skill id, detected integration, - `flags` (`ci`, `signup`, `debug`, `e2eAsk`, `localMcp`, `captureAio`, - `benchmark`, `yaraReport`) and the host the CLI was told. The caller supplies - an `InferenceAuthProvider` whose `resolve()` returns gateway authentication. - The agent resolves it before execution and again when the harness needs - refreshed auth. See the - [first-party provider](../../docs/developer-interfaces.md#inference-authentication) - for gateway token minting and refresh. +- `RunConfig`: the program ID, its `AgentRunDefinition` as `run`, `composed`, + the resolved `binding` (sequence, harness, model, effort), the switchboard + inputs, the skills origin, the flag snapshot and its payloads, the trace tags, + the tool allow and deny lists, `agentFlow`, `excludedTaskTypes`, `seedTasks`, + the bound completion `hooks`, and `scanReport`. +- `RunInput`: the install directory, resolved credentials, the project and user + payloads, the skill ID, the detected integration and its docs URL, `flags` + (`ci`, `signup`, `debug`, `e2eAsk`, `localMcp`, `captureAio`, `benchmark`, + `yaraReport`), and the host the CLI was told. - `RunResult`: `outcome` is `RunOutcome.Success | Aborted | Failed | Crashed`. - Success may carry an `outro`. The other three carry a `failure` - (`AgentFailure`: required code and message, optional outro data, `Error`, exit - code, detail, and authentication detail). `failure.error` may be attached, and - `Crashed` requires one. A failed result need not have an attached `Error`. - Every result carries a `snapshot` of what the run reported: tasks, status - lines, stage, token usage totals, final cost, dashboard and notebook URLs, - handoff text, and the transcript tail when the run definition sets - `collectTranscript`. It may also carry `skillId`. + `Success` can carry an `outro`. The other three carry a `failure` + (`AgentFailure`: message, outro data, error, exit code, error code, detail, + auth error detail). Every result carries `skillId` and a `snapshot` of what + the run reported: tasks, status lines, stage, token usage totals, final cost, + dashboard and notebook URLs, handoff text, and the transcript tail when the + run collected one. - `AgentProgress`: one event per thing the run reports, in emission order. Kinds: `lifecycle`, `spinner`, `log`, `status`, `tasks`, `stage`, `url`, - `usage`, `finalCost`, `authError`, `handoff`, `completion`, and `activity` - (one line per step, only from a run that collects its transcript). Payloads - are copies, never live objects. -- `AgentInteraction`: every member optional. `ask(question, { signal })` - resolves with answers, and `taskNotice(notice, { signal })` resolves with - whether to keep an optional task. Each request has its own signal, which - aborts when that request times out, the host aborts the run, or another task - fails the run. On abort the host dismisses that request alone, without - throwing. -- `signal`: an optional `AbortSignal` from the host. A pre-aborted signal - returns `Aborted` before execution. An abort during execution reaches the - active harness and returns `Aborted` with the current snapshot. It does not - pause or resume a run. `Aborted` means only that the host's signal cancelled - the run. -- Errors: the agent does not exit the process or throw for a decided failure. A - caught coded error returns `Failed`, and an uncoded throw returns `Crashed`. - Both retain the caught `Error` (or an `Error` wrapper for a non-`Error` - throw). An agent that stops itself with `[ABORT]` returns `Failed` with its - abort code. A gateway 401 returns an authentication failure with detail for - the host to present. The host decides how to present a returned failure, set - an exit code, or rethrow an attached error. Final scan-report flushing is best - effort and does not replace the run result. -- Skills: a run that does not end in `Success` removes the skill directories it - added under `/.claude/skills` that carry the `.posthog-wizard` - marker. Directories that were there before the run stay. -- Analytics shutdown is host-owned: the agent never sends the terminal - `setup wizard finished` event, and `runProgram` doesn't either. The host sends - it from the outcome: `Success` is `success`, `Aborted` is `cancelled`, - `Failed` and `Crashed` are `error`. - -`@agent` exports eleven runtime names, and -`src/agent/__tests__/public-entry.test.ts` holds that list: - -- **`runAgent` and `RunOutcome`.** The run and its outcome enum. -- **`OutroKind`.** The kind of an outro in `completion` progress and in failure - outro data. -- **`DEFAULT_AGENT_BINDING`.** The Pi and linear binding for standalone callers. -- **`resolveHarness` and `harnessRunsTasks`.** What programs resolve a binding - with. `harnessRunsTasks` says which harnesses the orchestrator can drive. -- **`AgentSignals` and `WIZARD_TOOL_NAMES`.** The marker strings that program - prompts embed, and the tool ids that go in tool allow and deny lists. -- **`downloadSkill` and `runMcpPromptViaSdk`.** The skill installer and the - suggested-prompts stream. Each loads its module on first call. -- **`TASK_OUTCOMES_KEY`.** The `frameworkContext` key under which the session - adapter stores an orchestrator run's task outcomes. + `usage`, `finalCost`, `authError`, `handoff`, `completion` and `activity`. + Payloads are copies, never live objects. +- `AgentInteraction`: every member is optional. `ask(question, { signal })` + resolves with answers. `taskNotice(notice, { signal })` resolves with whether + to keep an optional task. Each request has its own signal. It aborts when that + request times out, the host aborts the run, or another task fails the run. On + abort the host dismisses that request alone, without throwing. +- Errors: the agent doesn't exit the process and doesn't reject. A caught coded + error, such as a refused gateway mint, becomes `Failed`. An uncoded throw + becomes `Crashed` with the error attached. A gateway 401 returns an auth + failure, and the host decides whether to show auth UI. `Aborted` means the + host's signal cancelled the run. An agent that stops itself with `[ABORT]` + returns `Failed` with its abort code. +- Analytics shutdown is host-owned. The agent never sends the terminal + `setup wizard finished` event. The host sends it from the outcome: `Success` + is `success`, `Aborted` is `cancelled`, `Failed` and `Crashed` are `error`. + +Other runtime exports: `RunOutcome`, `AgentSignals`, `WIZARD_TOOL_NAMES`, +`resolveBinding`, `shouldDisableAsk`, `LONGER_ASK_TIMEOUT_MS`, `buildRunTags`, +`configureGatewayFromCIEnvironment`, `flushScanReport`, `downloadSkill`, +`TASK_OUTCOMES_KEY`, and `runMcpPromptViaSdk`, which loads the streaming module +on first call. + +## Run definition + +`RunConfig.run` is an `AgentRunDefinition`. These fields shape the prompt and +what the run collects: + +| Field | What it does | +| ------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `skillId` | The skill the linear sequence installs before the agent starts. Omit it to let the agent discover skills. | +| `customPrompt` | Instructions appended after the default project prompt. | +| `prompt` | Replaces the assembled project prompt. It receives the same `PromptContext`: project ID and key, host, skill path, and the organization and team opt-ins. | +| `collectTranscript` | Keeps a transcript tail on `snapshot.transcriptTail` and reports each agent step as an `activity` event. | +| `requestRemark` | Asks for the end-of-run reflection remark. Defaults to `true`. `false` skips it. | +| `abortCases` | Known `[ABORT] ` cases and the outro each one renders. | + +`prompt`, `customPrompt` and `abortCases` apply on the linear sequence. The +orchestrator builds each task's prompt from its context-mill flow. +`collectTranscript` and `requestRemark` take effect on the linear sequence with +the Anthropic harness. The Pi harness always asks for the remark on a linear +run. The orchestrator never asks for it. + +The other fields carry the run's copy, report file, docs URL, extra MCP servers, +question limits and step analytics. See +[`shared/types.ts`](runner/shared/types.ts). + +### Transcript tail + +With `collectTranscript`, the run observes every SDK message: + +- Each assistant text block joins the tail. The tail keeps the newest blocks, up + to 256 × 1024 characters, and drops the oldest first. A single block over the + cap stays whole. +- The run's final result text follows the kept blocks. +- `snapshot.transcriptTail` is the kept blocks, one per line, then the final + result. +- Each non-empty text block also emits `{ kind: 'activity', line }`, trimmed and + cut at 100 characters. +- Each tool call emits `{ kind: 'activity', line }` with the tool name and its + file path, pattern or path. + +Only the caller that set `collectTranscript` wants `activity` lines. The TUI's +progress reducer ignores them. + +### Scan report + +The agent counts its security scans in process-wide state. At the end of a run, +`runAgent` flushes them. It sends the scan telemetry and resets the counts. With +`flags.yaraReport`, it also writes the local report file and emits its path as +an `info` log event. `RunConfig.scanReport: 'defer'` skips the flush, so the +run's scans count toward the next flush. Agentic detection defers, so its scans +land in the program run's report. + +## Who calls `runAgent` + +- **`runProgram`**, for every program's agent run. It resolves credentials, + consent, flags and the route first. See the + [developer interfaces](../../docs/developer-interfaces.md). +- **Agentic detection**, `detectProjectsWithAgent` in + [`src/programs/detection/agentic.ts`](../programs/detection/agentic.ts). It + sets `prompt`, `collectTranscript: true`, `requestRemark: false` and + `scanReport: 'defer'`, and reads its report from the transcript tail. +- **The fault probe**, + [`scripts/a3-fault-probe.no-jest.ts`](../../scripts/a3-fault-probe.no-jest.ts), + against a local gateway with synthetic credentials. + +A standalone caller supplies everything `runProgram` would resolve. `runAgent` +doesn't authenticate the user, check consent, load flags or pick a route. It +does mint gateway auth: before any agent starts, it mints a scoped gateway token +from `input.credentials` for `config.programId`, and re-mints near expiry. In +development and test builds, `configureGatewayFromCIEnvironment` loads a fixed +token from `WIZARD_CI_GATEWAY_TOKEN_FILE` instead, and `runAgent` uses it +without minting. The gateway auth cache is process-wide. Minimal invocation: ```ts import { runAgent, RunOutcome } from '@agent'; -import type { AgentInteraction, RunConfig, RunInput } from '@agent/types'; +import type { + AskAnswers, + PendingQuestion, + RunConfig, + RunInput, +} from '@agent/types'; -export async function runOnce( +export async function runWithAnswers( config: RunConfig, input: RunInput, - ask: NonNullable, -) { + answersFor: (question: PendingQuestion) => Promise, +): Promise { const result = await runAgent(config, input, { onProgress: (event) => { if (event.kind === 'log') console.log(event.message); }, - interaction: { ask }, + interaction: { + ask: (question) => answersFor(question), + }, }); if (result.outcome !== RunOutcome.Success) { - console.error(result.failure.error ?? result.failure.message); process.exitCode = result.failure.exitCode ?? 1; } - return result; } ``` -The result is the agent's termination report. Check `outcome` first and then -read `failure`. A failed result can have no attached `Error`. If a higher layer -uses exceptions, it can rethrow `failure.error` when present and construct an -error from `failure.message` otherwise. Preserve the original `Error` object -when rethrowing so its stack and cause remain available. - -`src/agent/__tests__/run-agent-standalone.test.ts` runs this with no UI, no -store and no registry. - -Pass an `AbortController` signal in the options and call `controller.abort()` to -cancel an active run. The result then has `RunOutcome.Aborted`. +`src/agent/__tests__/run-agent-standalone.test.ts` runs the agent this way with +no UI, no store and no registry. ## Intent -Programs call the agent to do the work a skill describes. `runProgram` builds -the `RunConfig` and `RunInput` for every program run. The TUI and the `--ci` -runner reach it through `src/lib/runners/run-program-agent.ts`, which supplies -session capabilities and maps progress back onto `getUI()`. +Programs call the agent to do the work a skill describes. The TUI and the +non-interactive runner observe the run through `onProgress` and answer it +through `interaction`. `src/programs/run-agent-legacy.ts` does both on top of +the session, through `runProgram`. -Agentic detection calls `runAgent` itself, before the program runs. It uses a -linear Haiku run on the Anthropic harness, with `collectTranscript`, -`requestRemark: false` and `scanReport: 'defer'`. It reads its report from the -transcript tail and makes up to two attempts, with deadlines of 60 and 90 -seconds. A standalone host builds the config and input itself, as the -[wizard-workbench](https://github.com/PostHog/wizard-workbench) harness does -with `pnpm wizard-agent`. - -Without `onProgress` the run completes and its snapshot still comes back in the +Without `onProgress` the run completes, and its snapshot still comes back in the result. Without `interaction` the agent installs no ask bridge: `wizard_ask` -returns its "not available" error and optional task notices are declined. Plain -`--ci` runs also disable the ask bridge and decline notices. A throwing observer -is logged and the run continues. Progress callbacks are not awaited. Throws and -rejections from returned thenables are logged. Observers must handle errors from -detached work they start. +returns its "not available" error and optional task notices are declined, which +is what a `--ci` run does. A throwing or rejecting observer is logged and the +run continues. ## Architecture The agent owns run state for one invocation: the task queue, phase, status, resolved skill, handoff text, usage and the final result. It depends on -`src/shared` and `src/env.ts`. Callers supply program routing and policy through -`RunConfig`. +`src/shared` and on `src/env.ts`. It depends on program types only until the +bindings table moves to programs. ```text caller ── RunConfig + RunInput ──▶ runAgent - │ prepareRun: supplied gateway auth, triage provider + │ prepareRun: gateway mint, triage provider ▼ sequence (linear | orchestrator) │ @@ -188,10 +207,13 @@ caller ── RunConfig + RunInput ──▶ runAgent │ onProgress ◀── events ───┤──── questions ──▶ interaction ▼ + scan report flush (unless deferred) + ▼ RunResult ``` -`runner/` holds the dispatcher, sequences, harnesses and the switchboard. -`tools/` holds the wizard tools shared by both harnesses. `middleware/` holds -the benchmark pipeline. `progress.ts` defines the event and interaction -contracts. `yara-hooks.ts` scans what the run installs. +`runner/` holds the dispatcher, sequences, harnesses and the switchboard. See +the [runner reference](runner/README.md). `tools/` holds the wizard tools shared +by both harnesses. `middleware/` holds the benchmark pipeline. `progress.ts` +defines the event and interaction contracts. `yara-hooks.ts` scans what the run +installs. diff --git a/src/agent/runner/README.md b/src/agent/runner/README.md index 6dba64608..763419e36 100644 --- a/src/agent/runner/README.md +++ b/src/agent/runner/README.md @@ -1,18 +1,18 @@ # agent runner -How an agent run is assembled. Everything under this directory is plumbing — the +How an agent run is assembled. Everything under this directory is plumbing: the pieces that decide _how_ a program runs (which query shape, which agent SDK, which model) and the pieces that then actually run it. ``` ┌──────────────┐ ┌─────────────┐ ┌────────────────────────────┐ │ │ │ │────▶│ sequence (query shape) │ - │ programs │────▶│ binding │ │ linear | orchestrator │ + │ programs │────▶│ switchboard │ │ linear | orchestrator │ │ │ │ │ └────────────────────────────┘ - │ integration │ │ selects the │ - │ audit │ │ sequence, │ ┌────────────────────────────┐ - │ migration │ │ harness and │────▶│ harness (SDK adapter) │ - │ ... │ │ model │ │ anthropic | pi | ... │ + │ integration │ │ binds each │ + │ audit │ │ program to │ ┌────────────────────────────┐ + │ migration │ │ a pair │────▶│ harness (SDK adapter) │ + │ ... │ │ │ │ anthropic | pi | ... │ └──────────────┘ └─────────────┘ └────────────────────────────┘ ``` @@ -23,10 +23,10 @@ retained for very simple tasks and legacy support. The Anthropic Agent SDK is a supported legacy fallback, deprecated as the default, retained for major Pi vulnerabilities or gaps in support for new Anthropic models. -`DEFAULT_AGENT_BINDING`, the standalone default, is Pi + linear. Explicit -program bindings and flags determine actual behavior. Both harnesses implement -`run` and `runTask`. Composed sub-runs are clamped to linear, and linear-only -post-run/outro hooks do not automatically transfer to an orchestrated flow. +Existing `DEFAULT_BINDING` is Pi + linear. Explicit program bindings and flags +determine actual behavior. Both harnesses implement `run` and `runTask`. +Composed sub-runs are clamped to linear, and linear-only post-run/outro hooks do +not automatically transfer to an orchestrated flow. New models require Wizard capabilities **and** mint model/effort allowlists, gateway provider/transport support, and compatibility with required Wizard and @@ -43,54 +43,60 @@ Five layers, each with its own job. Nothing crosses layers unless it has to. `runAgent(config, input, {onProgress?, interaction?, signal?}) → RunResult`. It takes resolved execution data and an invocation snapshot (`shared/types.ts`), reports through `onProgress` and asks through `interaction` (`../progress.ts`), -and returns decided outcomes and caught run-body crashes as results. It never -renders, reads a session or exits. Its caller does the host work. For a program -run, `runProgram` in `src/programs/run-program.ts` calls the host's credentials -provider once, identifies the user, stamps the AI SDK evidence, awaits the -host's approval and post-auth gates, loads flags, refreshes an OAuth token near -expiry, and resolves the binding. Its caller builds the run definition and the -program settings from the `ProgramConfig`. The legacy -`src/lib/runners/run-program-agent.ts` does that from the session, supplies the -capabilities, and maps progress back onto `getUI()`. Agentic detection builds -its own config and input, and calls `runAgent` directly. +and returns every ending as a result. It never renders, reads a session or +exits. The caller resolves credentials, consent, flags and the route first. For +programs, `runProgram` in `src/programs/run-program.ts` does that, and the +legacy adapter in `src/programs/run-agent-legacy.ts` runs the readiness and +settings gates and maps progress back onto `getUI()`. + +The entry point also owns two per-run switches. When `config.run` sets +`collectTranscript`, it creates the transcript tail and passes it to the +sequence. After the sequence returns, it adds the tail to +`snapshot.transcriptTail`. When `config.scanReport` is `'defer'`, it skips the +end-of-run scan report flush. See +[transcript tail](../README.md#transcript-tail) and +[scan report](../README.md#scan-report). **Prepare** (`shared/bootstrap.ts`) is the on-ramp inside the agent: logging -targets, caller-supplied inference auth and the scan-triage classifier. Whether -the run turns out to be linear or orchestrator, anthropic or pi, the setup is -the same. - -**The switchboard** (`switchboard/`) holds the sequence and harness registries -and the harness and model resolution helper, `resolveHarness` (CLI > flag > -program config > default). The program layer owns the sequence precedence -(`resolveProgramBinding`) and turns its program ID, validated flag route and CLI -overrides into a resolved binding before it calls `runAgent`. The CLI overrides -arrive as `ProgramInput.overrides`. `harnessRunsTasks` tells it which harnesses -the orchestrator can drive. Agent code uses that binding to select a sequence -and harness. It doesn't read the program registry or parse feature flags. - -**Sequences** (`sequence/`) are LLM query shapes. Once the binding has picked -one, that sequence takes over the run and owns _how the LLM's work is shaped_. -See `sequence/README.md`. - -- **linear** — one long conversation with the model, start to finish. -- **orchestrator** — many focused conversations coordinated by a task queue, - each with its own prompt, tools, and model. +targets, the gateway mint and the scan-triage classifier. Whether the run turns +out to be linear or orchestrator, anthropic or pi, the setup is the same. + +**The switchboard** (`switchboard/`) is the router. Given a program id + the +fetched flags + any CLI overrides, it returns a `ProgramBinding`: which query +shape (sequence), which agent SDK (harness), which model. Two independent +middleware chains, one per axis, apply precedence rules (CLI > flag > program +config > default). This is the only layer that makes routing decisions. + +**Sequences** (`sequence/`) are LLM query shapes. Once the switchboard has +picked one, that sequence takes over the run and owns _how the LLM's work is +shaped_. See `sequence/README.md`. + +- **linear**: one long conversation with the model, start to finish. It builds + the prompt from `run.prompt` when set, else from the project prompt plus + `run.customPrompt`. It puts the transcript tail in front of the benchmark + middleware. +- **orchestrator**: many focused conversations coordinated by a task queue, each + with its own prompt, tools, and model. **Harnesses** (`harness/`) are SDK adapters. Sequences don't call Anthropic's or -pi.dev's SDKs directly — they go through a harness, which knows how to translate +pi.dev's SDKs directly. They go through a harness, which knows how to translate a run request into that SDK's shape. All harnesses drive the PostHog LLM gateway. -- **anthropic** — wraps Anthropic's official Claude Agent SDK. See - `harness/anthropic/README.md`. -- **pi** — wraps pi.dev's coding-agent library. See `harness/pi/README.md`. +- **anthropic**: wraps Anthropic's official Claude Agent SDK. See + `harness/anthropic/README.md`. Its linear run feeds every SDK message to the + middleware, so the transcript tail and its `activity` events come from here. + It passes `run.requestRemark` to the stop hook. +- **pi**: wraps pi.dev's coding-agent library. See `harness/pi/README.md`. Its + linear run always asks for the remark and takes no middleware. + +The orchestrator asks neither harness for a remark on its tasks. ## How they connect -- Programs supply inference auth. Prepare resolves it and builds triage for the - resolved harness. -- The program layer resolves the binding with the switchboard helpers. Agent - code dispatches the selected sequence and harness. +- Prepare mints the gateway token and builds triage for the resolved harness. +- The switchboard knows which sequences and harnesses exist (via its two + registries), but not what they do. - A sequence knows how to shape a conversation, but delegates the actual model call to a harness. - A harness adapts its SDK, gateway transport, security hooks, and tool surface. @@ -103,73 +109,79 @@ Each layer is replaceable. %%{init: {"block": {"padding": 20}}}%% block-beta columns 11 - hostBand["Host: program or caller"]:11 - runProgram["runProgram"]:3 space:1 programOutcome["ProgramRunOutcome"]:3 space:1 hostSignal["Host AbortSignal"]:3 + hostBand["Host: legacy adapter and UI"]:11 + runProgramAgent["runProgramAgent"]:3 space:1 wizardAbort["wizardAbort"]:3 space:4 + space:11 + programsBand["Programs"]:11 + runProgram["runProgram"]:3 space:1 programOutcome["ProgramRunOutcome"]:3 space:4 space:11 runnerBand["Agent runner"]:11 - runAgent["runAgent"]:3 space:1 runResult["RunResult"]:3 space:1 runnerSignal["RunAgentOptions.signal"]:3 + runAgent["runAgent"]:3 space:1 runResult["RunResult"]:3 space:4 + space:11 + sequenceBand["Orchestrator sequence"]:11 + runOrchestrator["runOrchestrator"]:3 space:1 sequenceResult["SequenceResult"]:3 space:4 space:11 - sequenceBand["Selected sequence"]:11 - sequence["linear | orchestrator"]:3 space:1 sequenceResult["SequenceResult"]:3 space:1 sequenceSignal["signal"]:3 + drainQueue["drainQueue"]:3 space:5 runAbort["AbortController"]:3 space:11 harnessBand["Selected harness"]:11 - agentHarness["AgentHarness"]:3 space:1 agentResult["AgentResult"]:3 space:1 harnessSignal["harness input signal"]:3 + agentHarness["AgentHarness"]:3 space:1 agentResult["AgentResult"]:3 space:1 signal["TaskRunInputs.signal"]:3 space:11 sdkBand["External model SDK"]:11 sdk["Selected SDK"]:3 space:8 + runProgramAgent --> runProgram runProgram --> runAgent - runProgram --> programOutcome - runAgent --> sequence - runAgent --> runResult - sequence --> agentHarness + runAgent --> runOrchestrator + runOrchestrator --> drainQueue + drainQueue --> agentHarness agentHarness --> sdk agentHarness --> agentResult agentResult --> sequenceResult sequenceResult --> runResult runResult --> programOutcome - hostSignal --> runnerSignal - runnerSignal --> sequenceSignal - sequenceSignal --> harnessSignal - harnessSignal --> agentHarness + programOutcome --> wizardAbort + drainQueue --> runAbort + runAbort --> signal classDef owner fill:#9ca3af1f,stroke:#9ca3af,stroke-width:1.5px - classDef contract fill:#3b82f626,stroke:#3b82f6,stroke-width:2px - class hostBand,runnerBand,sequenceBand,harnessBand,sdkBand owner - class programOutcome,runResult,sequenceResult,agentResult contract + classDef changed fill:#3b82f626,stroke:#3b82f6,stroke-width:2px + class hostBand,programsBand,runnerBand,sequenceBand,harnessBand,sdkBand owner + class programOutcome,runResult,agentResult,runAbort changed ``` -Calls descend on the left, results return through the middle, and a host-owned -abort signal descends on the right. A standalone caller invokes `runAgent` -without `runProgram`, and so does agentic detection. A program may return a -pre-run failure without starting the agent, and a caught preparation error -produces `RunResult` without a `SequenceResult`. The orchestrator stops -scheduling on the first fatal task result, cancels active siblings and pending -asks, and waits for them to settle before returning that failure. A host signal -can also cancel active harness work. +Calls descend on the left, results return through the middle, and cancellation +moves down the right. Blue marks the result contracts and run-scoped abort. On +the first fatal task result, `drainQueue` stops scheduling, cancels active work +and pending asks, joins siblings, then preserves that failure for the host to +present. ## Flow -1. The host builds the run and settings from the `ProgramConfig`, and runs any - readiness or settings gates it needs. `runProgram` then resolves credentials, - awaits the host's gates, loads PostHog flags, refreshes a token near expiry, - starts the file watchers and resolves a `ResolvedBinding` (sequence, harness - and model) from `input.overrides` and the flags. It tags the run and captures - the switchboard decision. -2. `runAgent(config, input, options)` resolves the supplied inference auth and - prepares triage. -3. The sequence takes over. It shapes the LLM's work into one conversation - (linear) or many (orchestrator), and reports through `onProgress`. +1. The caller resolves credentials, consent, flags and a + `ProgramBinding { sequence, harness, model }`, and analytics tags the run. + For programs, `runProgram` does this. +2. `runAgent(config, input, options)` creates the transcript tail when the run + collects one, then prepares (mint, triage). +3. Sequence takes over. It shapes the LLM's work into one conversation (linear) + or many (orchestrator), reporting through `onProgress`. 4. Harness drives each conversation through its SDK, using the bound model, on the PostHog LLM gateway. -5. The scan report flushes once, on a best-effort basis: as the run ends, or - earlier when a process drain runs the cleanups, unless `RunConfig.scanReport` - defers it to the host run. Its line arrives as `log` progress. `runAgent` - resolves a `RunResult` with an outcome and progress snapshot. A non-success - result carries a code and message, and a caught error stays attached. -6. `runProgram` records the result in its outcome, and the host applies it. The - legacy runner sends a decided failure to `wizardAbort` with the terminal - status its outcome names. For a crash, it rethrows the attached `Error` when - present. After a non-composed success, it sends the terminal success - analytics. Other hosts can log, present, or rethrow the failure as they need. - Neither `runAgent` nor `runProgram` sends terminal analytics. +5. The scan report flushes, unless `config.scanReport` is `'defer'`. `runAgent` + returns a `RunResult` whose snapshot carries the transcript tail when the run + collected one. +6. The caller applies it. `runProgram` settles it into a `ProgramRunOutcome`. + The legacy adapter sends a decided failure to `wizardAbort` with the terminal + status its outcome names, rethrows a crash for the runner's own handling, and + sends the terminal success analytics for a non-composed success. The agent + sends no terminal analytics. + +## Who calls `runAgent` + +- `runProgram`, for every program's agent run. +- Agentic detection, in `src/programs/detection/agentic.ts`. Each attempt is a + linear Haiku run on the Anthropic harness with its own deadline signal. It + sets `run.prompt`, `collectTranscript: true`, `requestRemark: false` and + `scanReport: 'defer'`, and reads its report from `snapshot.transcriptTail`. +- The fault probe in `scripts/a3-fault-probe.no-jest.ts`, and any standalone + host. See the + [developer interfaces](../../../docs/developer-interfaces.md#runagent-for-detection-and-standalone-callers). diff --git a/src/agent/runner/harness/anthropic/README.md b/src/agent/runner/harness/anthropic/README.md index 5d7a6aaa2..d52a8d3d2 100644 --- a/src/agent/runner/harness/anthropic/README.md +++ b/src/agent/runner/harness/anthropic/README.md @@ -11,8 +11,8 @@ choose this harness. supported: `run()` for linear conversations and `runTask()` for orchestrator seed/task calls. Pi also implements both entry points. -The SDK subprocess uses the scoped token supplied by programs through -[gateway-session.ts](../../../../programs/gateway-session.ts). Wizard explicitly sets the +The SDK subprocess uses the scoped token minted by +[gateway-session.ts](../../../gateway-session.ts). Wizard explicitly sets the gateway URL and authentication environment and isolates stored Claude logins. Model selection must satisfy local routing, the SDK's supported transport, mint model/effort allowlists, and the gateway's required prompt policy. The SDK is diff --git a/src/agent/runner/sequence/README.md b/src/agent/runner/sequence/README.md index 9eadfe595..ac4436f9a 100644 --- a/src/agent/runner/sequence/README.md +++ b/src/agent/runner/sequence/README.md @@ -12,8 +12,8 @@ prompt assembly, error routing, post-run work, and outro construction. Retain it for very simple tasks and legacy support. Its context is subject to the harness's compaction behavior. -`AgentRunDefinition.customPrompt`, `prompt`, `collectTranscript`, `abortCases`, -and the program's `postRun` and `buildOutroData` hooks are linear hooks. The orchestrator does not invoke them. Composed program sub-runs +`AgentRunDefinition.customPrompt`, `abortCases`, and the program's `postRun` +and `buildOutroData` hooks are linear hooks. The orchestrator does not invoke them. Composed program sub-runs are also clamped to linear because an orchestrator owns its full lifecycle and cannot nest through the composition seam. @@ -49,7 +49,6 @@ frontmatter contract rather than copying an old manifest example: task types after planning. It is not a universal automatic graph built from frontmatter. -Per-role routes in the resolved binding can adjust a task's harness, model or -effort. Programs resolve `PROGRAM_BINDINGS[id].contextMillOverride` before the -agent starts. New task types usually belong in context-mill; new native +Per-role `PROGRAM_BINDINGS[id].contextMillOverride` can adjust a task's harness, +model or effort. New task types usually belong in context-mill; new native behavior still requires the appropriate Wizard configuration or implementation. diff --git a/src/programs/README.md b/src/programs/README.md index 2b1db5fb1..b93d80b8a 100644 --- a/src/programs/README.md +++ b/src/programs/README.md @@ -1,325 +1,160 @@ # Programs -`runProgram` runs one Wizard program invocation from data the caller supplies. -The caller builds the run definition and the program settings from the program's -`ProgramConfig`. `runProgram` owns the invocation's state, resolves the binding, -and passes the agent run to [`runAgent`](../agent/README.md). It needs no -`WizardSession` or UI store. This is a repository-local TypeScript interface. -The npm package doesn't export it as a public library API. - -## Signatures - -Import runtime values from `@programs` and types from `@programs/types`: - -```ts -import type { - ProgramInput, - ProgramOptions, - ProgramRunOutcome, -} from '@programs/types'; - -// The shape of `runProgram`, exported from `@programs`. -declare function runProgram( - programId: string, - input: ProgramInput, - options?: ProgramOptions, -): Promise; -``` - -`programId` picks the binding policy, the commandments, the stage overrides and -the analytics attribution. `runProgram` doesn't look it up in a registry. An ID -with no binding runs on the default agent binding. The exact shapes are in -[`run-program.ts`](run-program.ts) and [`program-store.ts`](program-store.ts). - -### Exports - -| Export | What it's for | -| --------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------- | -| `runProgram` | Run one program invocation. | -| [`createPosthogInferenceAuthProvider`](../../docs/developer-interfaces.md#inference-authentication) | Mint first-party gateway auth from a PostHog login, for a host that calls `runAgent` itself. | -| `postAuthGateSteps`, `authenticate`, `FRAMEWORK_REGISTRY`, `getDetectedWarehouseSources`, `AUDIT_CHECKS_KEY` | Step-based host helpers. The session adapter uses them to build the input and project the data. | -| `PROGRAM_REGISTRY`, `Program`, `getProgramConfig`, `getSubcommandPrograms`, `getCommandPath`, `getLaunchablePrograms` | The step-based `ProgramConfig` registry that the TUI and the CLI commands use. | - -The type entry adds the input, option, settings, outcome and progress types -named below. It also carries the `ProgramConfig` step types, the -`ProgramCompletionContext` the completion hooks read, the switchboard context -type, and the `ProgramCiHost` and `ProgramRunHost` capability types. - -The entry doesn't re-export the policy that `runProgram` applies. Tests and -tools import it from its module: - -- `PROGRAM_BINDINGS` and `resolveProgramBinding` from [`binding.ts`](binding.ts) -- `captureSwitchboardDecision` from - [`binding-telemetry.ts`](binding-telemetry.ts) -- `getProgramCommandments` from [`commandments.ts`](commandments.ts) -- `resolveStageOverrides` and `areSeededTasksEnabled` from - [`experiments/`](experiments/index.ts) - -### Inputs - -`ProgramInput` needs `installDir` and `run`. Everything else is optional data -for this invocation. - -| Field | What it carries | -| ---------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `installDir` | The project directory. Report paths resolve against it. | -| `run` | The program's `AgentRunDefinition`, built by the caller from its `ProgramConfig`. | -| `program` | `ProgramSettings` read from the same `ProgramConfig`. See [program settings](#program-settings). | -| `credentials` | Resolved `{ posthog, inferenceAuth?, project, apiUser }`. It wins over `options.credentials`. Without `inferenceAuth`, `runProgram` mints first-party gateway auth from the refreshed login. The development `--ci` runner passes a fixed gateway token instead. | -| `runId` | Stable attribution. Generated when absent. | -| `overrides` | `{ harness?, sequence?, model? }` from launch flags such as `--harness`, `--sequence` and `--model`. | -| `composed` | Marks a sub-run of a host program. The binding clamps it to linear, and the agent leaves the outro to its host. | -| `skillId`, `integration`, `frameworkDocsUrl` | The skill to install and the detected framework. | -| `flags`, `host` | Run flags (`ci`, `signup`, `debug` and the rest default to `false`), and where PostHog is. | -| `wizardFlags`, `wizardFlagPayloads` | An evaluated flag snapshot. When `wizardFlags` is absent, `runProgram` asks `options.featureFlags`. | -| `seedTasks`, `hooks` | Tasks the orchestrator queues before the planner, and the completion hooks bound to the run's credentials. | -| `warehouseSources`, `discoveredFeatures`, `mayReportScanResults` | Evidence for the organization's AI SDK stamp. | -| `aiSdkStampReported` | The host already considered the stamp for this login. | - -`runProgram` copies the input when it receives it, so a later host write can't -reach the run. Data fields are structured-cloned. `credentials`, `run`, -`program`, `hooks` and `seedTasks` stay by reference, because they can carry -functions. Any other field that can't be cloned rejects the call. - -### Program settings - -`input.program` carries the program-level data from `ProgramConfig`: - -| Field | What `runProgram` does with it | -| ------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------- | -| `requiresAi` | `false` skips the AI-processing approval. | -| `agentFlow`, `allowedTools`, `disallowedTools`, `excludedTaskTypes` | Passed to the agent run. | -| `auditLedgerFile`, `auditSeedChecks` | The ledger to watch, and the rows written to it before the agent starts. | -| `eventPlanFile` | The event plan to watch. | -| `postAuthGates` | Step IDs the host settles after auth and before the agent starts, through `awaitPostAuthGates`. | - -### Options - -Every option is optional. Each awaited capability receives the invocation's -signal. Without `options.signal`, it gets a signal that never aborts. - -| Option | What `runProgram` does with it | When it's absent | -| -------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------- | -| `credentials` | Calls `resolve(programId, { signal })` once, only when `input.credentials` is absent. | Without `input.credentials`, the run fails. | -| `awaitAiApproval` | Calls it with `{ programId, signal }` when `program.requiresAi` isn't `false`, the organization hasn't approved AI data processing, and neither `flags.ci` nor `flags.signup` is set. `true` proceeds, `false` aborts. | A run that needs approval fails. | -| `awaitPostAuthGates` | Calls it with `{ programId, gates, signal }` when `program.postAuthGates` is not empty. | The run doesn't pause. | -| `featureFlags` | Calls it once when the input has no `wizardFlags`. It gets no signal. | The run uses empty flags. | -| `interaction` | Passes it to the agent run. | The agent installs no ask bridge and declines optional task notices. | -| `onProgress` | Sends run events and program-data snapshots. See [progress](#progress). | The outcome still carries the settled run and the final data. | -| `deferSkillCommit` | Leaves the skills a successful run installed armed for removal. The host commits them later with `commitRegisteredRunSkillCleanups()` from `@shared/skill-run-cleanup`. | `runProgram` commits them on success. | -| `signal` | Cancels the invocation. | Nothing cancels it. | - -`ProgramOptions['credentials']` and `ProgramInput['credentials']` name the -provider and resolved types. Their named types live in -[`credentials.ts`](credentials.ts). - -### Outcome - -`ProgramRunOutcome.outcome` is `success`, `aborted`, `failed` or `crashed`. - -- **`settledRuns`.** The finished agent run, with its `runId` and `RunResult`. -- **`diagnostics`.** Up to 10 observer failures and late events, newest last. -- **`data`.** Invocation data: credentials, project and user, the framework - context, the event plan, the binding and the AI SDK stamp latch. -- **`artifacts.reportFile`.** The absolute report path, set just before the - agent starts. It doesn't prove the file was written. -- **`failure`.** A code and message on a failed, crashed or aborted outcome. An - agent failure may carry the original `Error`. - -Don't log `data`. It holds PostHog credentials. - -### Progress - -`ProgramProgress` is a union of two kinds: - -- **`{ kind: 'run', runId, event }`.** One agent event. `event` is a copied - [`AgentProgress`](../agent/README.md#signatures) value. -- **`{ kind: 'program', data }`.** A copied `ProgramInvocationData` snapshot, - sent after each store write. It holds credentials too. - -Narrow on `kind` before reading `event`. `runProgram` never awaits the observer. -A throw or a rejection becomes a bounded diagnostic in `outcome.diagnostics`. So -does an event that arrives after the run finished. - -### Errors and cancellation - -Most endings resolve. A rejected promise means the invocation itself broke. - -- **Failed.** Missing credentials, or a run that needs approval without - `awaitAiApproval`. A rejection from `credentials`, `awaitAiApproval`, - `awaitPostAuthGates`, `featureFlags` or the token refresh also resolves as - `failed`, with the error's message and code `PHW_INTERNAL_UNHANDLED`. -- **Aborted.** A declined approval returns `aborted` with code - `PHW_AGENT_ABORT`. So does the host's signal, with the message "Run cancelled - by host." -- **Crashed.** An agent crash. The failure carries the original `Error`. -- **Rejected.** Input that can't be cloned. - -`runProgram` checks the signal before it starts and after each awaited -capability. An abort at any of those points returns `aborted`, even when the -capability has already completed. The agent run gets the signal too, and an -abort during the run returns the agent's `aborted` result. - -### Example - -The caller builds the run from the program's config, and implements -`credentials` and `awaitAiApproval` with its own login and consent flow: - -```ts -import { getProgramConfig, runProgram } from '@programs'; -import type { ProgramOptions } from '@programs/types'; - -export async function runMetrics( - installDir: string, - credentials: NonNullable, - awaitAiApproval: NonNullable, - signal?: AbortSignal, -) { - // A dynamic run resolves from session-shaped data and a ProgramRunHost. - const config = getProgramConfig('metrics'); - if (!config.run || typeof config.run === 'function') { - throw new Error('metrics has a static run definition'); - } - - const result = await runProgram( - config.id, - { - installDir, - run: config.run, - program: { agentFlow: config.agentFlow, requiresAi: config.requiresAi }, - }, - { - credentials, - awaitAiApproval, - signal, - onProgress: (progress) => { - if (progress.kind !== 'run') return; - const { runId, event } = progress; - if (event.kind === 'status') console.log(runId, event.message); - }, - }, - ); - - if (result.outcome !== 'success') { - if (result.failure?.error) throw result.failure.error; - throw new Error(result.failure?.message ?? `Metrics ${result.outcome}`); - } - return result; -} -``` - -## Intent - -A host calls `runProgram` to run a program without a TUI session. The host -decides where credentials, answers and consent come from, and which run it -passes. `runProgram` decides how that run goes: its policy, binding, files and -progress. - -Today's callers: - -- **The session adapter.** `src/lib/runners/run-program-agent.ts` serves the TUI - and the `--ci` runner. It resolves `ProgramConfig.run` against the session and - a `ProgramRunHost`, runs the health and settings gates, and reads the program - settings from the config. It supplies the session's login as the credentials - provider, answers approval and post-auth gates from the TUI, and maps progress - back onto `getUI()`. -- **The workbench harness.** `pnpm wizard-program` in - [wizard-workbench](https://github.com/PostHog/wizard-workbench) is a reference - host with no TUI. - -A missing capability never hangs the run and never invents consent. A program -that needs approval fails without `awaitAiApproval`. A run with no `interaction` -asks no questions. A run with no `onProgress` still returns its settled run and -final data in the outcome. - -`flags.ci` and `flags.signup` skip the approval check. Set them only when your -CI authorization or signup flow already handled consent. The flags don't prove -consent. - -## Architecture - -One invocation owns one `ProgramStore`. The host observes a copy of the store -through `onProgress` and the outcome, never the store itself. - -```mermaid -%%{init: {"block": {"padding": 20}}}%% -block-beta - columns 11 - hostBand["Host: session adapter, reference harness or embedder"]:11 - hostCall["build run and settings, then runProgram"]:3 space:1 hostView["onProgress and the outcome"]:3 space:1 hostAnswer["provider, approval, post-auth gates, answerer"]:3 - programsBand["runProgram: one invocation"]:11 - pipeline["credentials, approval, post-auth gates, flags, refresh, watchers, binding"]:3 space:1 store["ProgramStore: data and settled run"]:3 space:1 awaited["awaited capabilities, interaction passed through"]:3 - agentBand["Agent"]:11 - agent["runAgent"]:3 space:1 agentProgress["AgentProgress and RunResult"]:3 space:1 agentAsk["interaction, per-request signal"]:3 - - hostCall -- "ProgramInput, ProgramOptions" --> pipeline - pipeline -- "RunConfig, RunInput" --> agent - agentProgress --> store - store -- "run events, data snapshots" --> hostView - agentAsk --> awaited - awaited --> hostAnswer - - classDef band fill:#9ca3af1f,stroke:#9ca3af,stroke-width:1.5px - classDef contract fill:#3b82f626,stroke:#3b82f6,stroke-width:2px - class hostBand,programsBand,agentBand band - class store,agentProgress contract -``` - -Calls go down the left. Progress comes up the middle. Questions go up the right. - -`runProgram` works in this order: - -1. Copy the input, register the skill cleanup, and check the signal. -2. Resolve credentials, identify the user for analytics, and stamp the AI SDK - evidence once per invocation. -3. Await AI approval and the post-auth gates when they apply. -4. Load flags, and refresh the OAuth token when it has less than 50 minutes - left. -5. Start the file watchers and seed the audit ledger. -6. Resolve the binding from `overrides` and flags, and capture the switchboard - decision. -7. Call `runAgent`, record the result, and stop the watchers. - -`runProgram` is the only owner of the program file watchers. It watches the -audit ledger (`.posthog-audit-checks.json`) and the event plan -(`.posthog-events.json`) while the agent runs. A ledger that an earlier run left -behind is ignored until this run writes it. Updates land in the store as the -`auditChecks` framework context and the event plan. - -Each invocation registers a skill cleanup for its install directory. A -non-success outcome, a rejection or a process drain removes the Wizard skills -the invocation added. Skills that existed before stay. - -`runProgram` sends the `agent started` event and the switchboard decision. It -doesn't send the terminal `setup wizard finished` event. The host sends it from -the outcome. - -`ProgramCiHost` and `ProgramRunHost` belong to the step-based `ProgramConfig` -path, not to `runProgram`. `ProgramCiHost` supplies logging, auth and progress -while a CI pre-run scopes the project. `ProgramRunHost` supplies the live UI -effects a legacy recipe reads while its run definition resolves. The completion -hooks read a `ProgramCompletionContext` built when each hook runs, so URLs the -run emitted reach them. - -The programs layer imports the agent only through `@agent` and `@agent/types`. -Its other imports come from `src/shared` and `src/env.ts`. The one exception is -a type import in `program-step.ts`, where the step types still name the legacy -`WizardSession`. +A program is a `ProgramConfig`: the steps the TUI walks, the agent run it +performs, and its program-level settings. `src/programs/` holds the program +configs, detection, the framework registry, the task stream, and `runProgram`, +which runs one program's agent from explicit inputs. + +The [developer interfaces](../../docs/developer-interfaces.md) document the +callable contract: every `ProgramInput`, `ProgramSettings` and `ProgramOptions` +field, the outcome, the pipeline order, cancellation and the failure outcomes. +This page covers who owns what around that call. + +## Entry points + +Import runtime values from `@programs` and types from `@programs/types`. + +| Entry | Exports | +| ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `@programs` | `runProgram`, and the registry: `PROGRAM_REGISTRY`, `Program`, `getProgramConfig`, `getSubcommandPrograms`, `getCommandPath` and `getLaunchablePrograms`. | +| `@programs/types` | The call types: `ProgramInput`, `ProgramSettings`, `ProgramOptions`, `ProgramOverrides`, `WizardFlagSnapshot` and `ProgramRunOutcome`. The store types: `ProgramProgress`, `ProgramRunProgress`, `ProgramDataProgress`, `ProgramInvocationData`, `ProgramDiagnostic` and `SettledProgramRun`. The host capabilities: `ProgramRunHost` and `ProgramCiHost`. The config types: `ProgramId`, `SubcommandProgram`, `ProgramConfig`, `ProgramStep`, `ProgramReadyContext`, `StoreInitContext`, `FrameworkConfig` and `SetupQuestion`. | + +Code inside the repository reaches deeper modules through `@programs/*`: + +- `runProgramAgent` in [`run-agent-legacy.ts`](run-agent-legacy.ts), the legacy + adapter. +- `postAuthGateSteps` in [`program-step.ts`](program-step.ts), which lists the + gated steps between the `auth` and `run` screens. +- `ResolvedProgramCredentials` and `CredentialsProvider` in + [`credentials.ts`](credentials.ts). +- `ProgramStore` in [`program-store.ts`](program-store.ts). +- The detection tools in [`detection/`](detection/index.ts). + +## The host + +The caller of `runProgram` is the host. `runProgram` owns the policy around one +agent run. The host owns everything around the invocation. + +| `runProgram` owns | The host owns | +| -------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- | +| One `ProgramStore` for the invocation. | Building `run` and the settings from a `ProgramConfig`. | +| Calling the credentials provider, then identifying the user for analytics. | Login and project choice, inside its credentials provider. | +| The AI SDK stamp, once per invocation. | The consent screen behind `awaitAiApproval`. | +| Deciding when approval and post-auth gates apply. | The gate screens behind `awaitPostAuthGates`. | +| Loading flags when the input has none, through the host's `featureFlags`. | Answering questions and task notices through `interaction`. | +| The token refresh right before the agent starts. | Rendering progress from `onProgress`. | +| The route, the `switchboard resolved` event and the run tags. | The readiness and Claude settings gates, before the call. | +| Calling `runAgent` and settling one outcome. | File watchers, such as the audit ledger and the task stream's event plan. | +| | Walking composed steps, and running programs that have no agent. | +| | Applying the outcome: exit code, auth error screen, rethrowing a crash, and `setup wizard finished`. | +| | Cancelling through `signal`. | + +`runProgram` never calls `getUI()` and never reads a session. It still touches +process-wide state: the analytics client, the gateway auth cache and the debug +log. A dead OAuth grant during the token refresh marks the login revoked, so a +later 401 can name the cause. + +A missing capability never hangs the run and never invents consent. A run that +needs approval fails without `awaitAiApproval`. A run with no `interaction` asks +no questions. A run with no `onProgress` still returns its settled run and final +data. + +## The store + +[`ProgramStore`](program-store.ts) holds one invocation's data and its run +ledger. `runProgram` creates one per call. The host never touches the store. It +sees copies through `onProgress` and the outcome. + +- **Data.** `ProgramInvocationData`: the login, the framework context, the route + and the AI SDK stamp latch. `setAuthenticated`, `setBinding`, + `setAiSdkStampReported` and `setFrameworkContext` write it. + `setAiSdkStampReported` writes only the first time. `runProgram` never calls + `setFrameworkContext`, so `detection.frameworkContext` stays `{}`. +- **Copies.** Every write stores a `structuredClone` of its value. `readData()` + returns a fresh clone. After each write, the store sends a + `{ kind: 'program', data }` snapshot, itself a fresh clone. +- **Runs.** `beginRun(runId, observer)` returns an adapter with `onProgress` and + `finish`. `onProgress` forwards a clone of each agent event as + `{ kind: 'run', runId, event }`. `finish(result)` records the result, and + `settledRuns()` lists every finished run. `runProgram` begins one run per + call. +- **Diagnostics.** The store never waits for an observer. A throw or a rejected + promise becomes a `ProgramDiagnostic` with its source and message. A run event + after `finish` becomes a `progress after finish` diagnostic. The source is + `{ runId, eventKind }` for a run event and `{ eventKind: 'data' }` for a data + snapshot. The store keeps the 10 newest diagnostics. + +## The adapter + +`runProgramAgent(programConfig, session, { composed? })` in +[`run-agent-legacy.ts`](run-agent-legacy.ts) is the host for every existing +runner. It is the only program code that reads the session and `getUI()` on the +agent's behalf. It resolves `run` against the session and a `ProgramRunHost`, +runs the readiness and Claude settings gates, and builds the input from the +session and the config. It answers the capabilities from the UI, maps run events +onto `getUI()`, and projects data snapshots onto the session. Then it applies +the outcome with `wizardAbort` and the terminal analytics event. +[How the legacy adapter builds `run`](../../docs/developer-interfaces.md#how-the-legacy-adapter-builds-run) +lists every field it maps. + +Three places call it: + +- **The TUI runner.** [`run-wizard.ts`](../lib/runners/run-wizard.ts) calls it + on the program's `run` screen. A composed step list calls it once per run + step. Each call gets a session scoped to the step's `targetDir` when the step + has one. +- **A composed run step.** `integrationRunStep` in + [`posthog-integration/index.ts`](posthog-integration/index.ts) runs the + integration with `composed: true`. Self-driving splices it into its own steps. +- **The non-interactive runner.** + [`run-non-interactive.ts`](../lib/runners/run-non-interactive.ts) calls it + after `ciPreRun`, for `--ci` and for headless runs. + +## Host capabilities + +A program's `run` and `ciPreRun` callbacks take their effects from a host +instead of calling `getUI()`. Both types live in +[`host-capabilities.ts`](host-capabilities.ts). + +- **`ProgramRunHost`** reaches `ProgramConfig.run(session, host)`. It has + `getFrameworkContext(key)`, `setFrameworkContext(key, value)` and + `warn(message)`. The legacy adapter builds it over `getUI()`, read at each + call. A run definition can keep the host and call it until the run returns. + The source maps program reads the picked project in its prompt builder and + writes `sourceMapsCompletedVariant` in `postRun`. +- **`ProgramCiHost`** reaches `ProgramConfig.ciPreRun(session, host)`. It has + `log.info(message)` and `log.warn(message)`. The non-interactive runner builds + it over `getUI().log`, read at each call. Project scoping logs its scan and + its fallbacks through it. + +Neither capability is a `runProgram` option. They belong to building the run +from a `ProgramConfig`, which happens before `runProgram`. The +[developer interfaces](../../docs/developer-interfaces.md#host-capabilities) +list which programs use each capability. ## Current limits -`runProgram` doesn't discover credentials, choose a project or render questions. -The host supplies those through data and capabilities. - -One call runs one agent. Composition belongs to the host: the TUI walks a -composed step, marked with `runProgramId`, as its own call with -`composed: true`. Programs with no agent, such as `posthog-doctor`, `mcp-add` -and `slack`, don't go through `runProgram`. The TUI runs their steps. - -Agentic detection runs before `runProgram`, as its own `runAgent` call with a -deadline per attempt. The MCP suggested-prompts screen streams through -`runMcpPromptViaSdk`, a separate SDK path. There is no socket controller and no -step-by-step control API. - -For the process-owned `--ci` runner, see the -[non-interactive developer interfaces](../../docs/developer-interfaces.md). +- **One agent per call.** Composition belongs to the host. The TUI walks a + composed step list and calls the adapter once per run step. An imported run + step, such as the integration inside self-driving, runs with `composed: true`. +- **Programs with no agent.** A config with no `run`, such as `posthog-doctor`, + `mcp-add` or `slack`, never reaches `runProgram`. The TUI runs its steps. +- **Building `run` needs a session.** A dynamic `run`, the `ProgramRun` + completion hooks and `seedTasks` all take a `WizardSession`. Some programs + read session fields that a TUI step or `ciPreRun` fills, such as + `session.frameworkConfig`. +- **The module graph.** `@programs` loads the program registry, and with it + every program config, the TUI decks and `@ui`. `run-program.ts` also loads + `@ui` through `authenticate.ts`, although `runProgram` never calls `getUI()`. +- **Other `getUI()` calls.** `authenticate`, the audit ledger watcher, agentic + detection and some framework helpers still call `getUI()` directly. +- **No framework context in the data.** `data.detection.frameworkContext` stays + `{}`. Framework context reaches the program through its run host. +- **Gates and cleanup stay with the host.** `runProgram` runs no readiness or + Claude settings gate, starts no file watcher, and doesn't remove the skills a + run installed. +- **Other agent paths.** Agentic detection makes its own `runAgent` calls, each + with its own deadline. The MCP suggested-prompts screen streams through + `runMcpPromptViaSdk`, a separate SDK path. +- **No live control.** There is no live store handle and no step-by-step control + API. A host observes through `onProgress` and cancels through `signal`. diff --git a/src/shared/README.md b/src/shared/README.md index 0ccd0dcb0..4bbd24aab 100644 --- a/src/shared/README.md +++ b/src/shared/README.md @@ -40,4 +40,4 @@ Shared exists so the agent, programs, TUI, headless and CLI code can use one imp ## Architecture -Shared imports `src/env.ts` and itself. The architecture test classifies `src/shared` as its own surface and lists the remaining upward edges in `src/__tests__/architecture/known-violations.json`; each has an owner in the stack plan. `utils/setup-utils.ts`, `utils/oauth.ts` and `utils/wizard-abort.ts` are TUI and CLI flow code that leave in Release C. `utils/analytics.ts` accepts the legacy session shape as a type only; scan consent and cleanup registration live in shared modules so callable programs load no session or UI code. The only agent import left is type-only: `errors/skill-map.ts` takes `InstallSkillResult` from `@agent/types`. `claude-settings.ts` now uses `agent-env-isolation.ts` in shared, and the agent error map lives in `src/agent/error-map.ts`. +Shared imports `src/env.ts` and itself. The architecture test classifies `src/shared` as its own surface and lists the remaining upward edges in `src/__tests__/architecture/known-violations.json`; each has an owner in the stack plan. `utils/setup-utils.ts`, `utils/oauth.ts` and `utils/wizard-abort.ts` are TUI and CLI flow code that leave in Release C; `utils/analytics.ts` reads the session until C2a. The only agent import left is type-only: `errors/skill-map.ts` takes `InstallSkillResult` from `@agent/types`. `claude-settings.ts` now uses `agent-env-isolation.ts` in shared, and the agent error map lives in `src/agent/error-map.ts`. diff --git a/src/shared/health-checks/testme.md b/src/shared/health-checks/testme.md index ff6681b7f..31df3aca7 100644 --- a/src/shared/health-checks/testme.md +++ b/src/shared/health-checks/testme.md @@ -3,7 +3,7 @@ Run the existing health and gateway tests with mocked HTTP requests: ```bash -pnpm exec vitest run src/shared/health-checks/__tests__/health-checks.test.ts src/programs/__tests__/gateway-session.test.ts +pnpm exec vitest run src/shared/health-checks/__tests__/health-checks.test.ts src/agent/__tests__/gateway-session.test.ts ``` The checks cover gateway readiness, endpoint retries, and skills downloads from From e0db572847b492940e8b347dd8fc8c5d7b2d5bb5 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Thu, 24 Sep 2026 20:17:11 -0400 Subject: [PATCH 15/29] docs(programs): reshape developer interfaces around runProgram and runAgent Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- docs/developer-interfaces.md | 636 +++++++++---------------------- docs/images/wizard-run-paths.svg | 53 +++ src/agent/runner/README.md | 2 +- src/programs/README.md | 4 +- 4 files changed, 230 insertions(+), 465 deletions(-) create mode 100644 docs/images/wizard-run-paths.svg diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index baab17f8f..6888f1185 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -1,42 +1,43 @@ # Developer interfaces -Wizard has two TypeScript call surfaces for code that runs without the terminal -UI. `runProgram` runs one program. `runAgent` runs one agent. Both live in this -repository and import through its path aliases. The `@posthog/wizard` npm -package publishes the CLI, not these functions. - -| Surface | Import from | Use it for | -| --------------------------------------- | ----------------------------------------- | ----------------------------------------------------------------------- | -| `runProgram(programId, input, options)` | `@programs`, types from `@programs/types` | One program's agent run, with the program's policy, route and telemetry | -| `runAgent(config, input, options)` | `@agent`, types from `@agent/types` | One agent run from a resolved config, with no program policy | - -The [programs reference](../src/programs/README.md) covers the store, the -adapter and the current limits. The [agent reference](../src/agent/README.md) -covers the run contract. - -## What `runProgram` is for - -`runProgram` runs one program's agent from explicit inputs. It needs no -`WizardSession`, no TUI store and no `getUI()` call. The caller is the host. The -host supplies the run definition, the credentials, consent, the gate waits and -the answers. `runProgram` applies the program policy around the agent run. It -resolves credentials, checks AI-processing approval, waits for post-auth gates, -loads flags, refreshes the token and resolves the route. Then it calls -`runAgent` and returns one outcome. - -Two hosts call it: - -- **The legacy adapter.** `runProgramAgent` in - [`run-agent-legacy.ts`](../src/programs/run-agent-legacy.ts) serves the TUI - and the `--ci` runner. See - [how the legacy adapter builds `run`](#how-the-legacy-adapter-builds-run). -- **The workbench.** - [PostHog/wizard-workbench#4190](https://github.com/PostHog/wizard-workbench/pull/4190) - adds `services/wizard-program/`. Its `pnpm wizard-program` imports - `runProgram` from `@programs` in the wizard checkout that `WIZARD_REPO` names. - It runs one program against an app copy with no TUI. - -## Signatures +There are four ways to run the wizard. Most end users run it from the TUI. The +headless runner runs the same flow without a terminal UI. `runProgram` runs just +one program, and `runAgent` runs only the agent. + +| Way | Entry | Use it for | Reference | +| ------------ | --------------------------------------------------------------------------------------------------------- | ----------------------------------------------------- | ---------------------------------------------------------------------------- | +| TUI | `npx @posthog/wizard`, through [`run-wizard.ts`](../src/lib/runners/run-wizard.ts) | End users setting up PostHog in a terminal | [README](../README.md) | +| Headless | `npx @posthog/wizard --ci`, through [`run-non-interactive.ts`](../src/lib/runners/run-non-interactive.ts) | CI and scripts, with no prompts | [Local credentials](local-dev.md#credentials-for-local-ci-and-headless-runs) | +| `runProgram` | `runProgram(programId, input, options)` from `@programs` | One program, from code, with its policy and telemetry | [`runProgram`](#runprogram), [programs reference](../src/programs/README.md) | +| `runAgent` | `runAgent(config, input, options)` from `@agent` | Only the agent, from a resolved config | [`runAgent`](#runagent), [agent reference](../src/agent/README.md) | + +![The TUI and the headless runner go through their own state to runProgram. The workbench calls runProgram directly. runProgram calls runAgent, and detection calls runAgent directly](images/wizard-run-paths.svg) + +The dashed boxes are the state the TUI and the headless runner keep above +`runProgram`. The workbench and other integrations call `runProgram` directly. + +`runProgram` and `runAgent` live in this repository and import through its path +aliases. The `@posthog/wizard` npm package publishes the CLI, not these +functions. + +## `runProgram` + +### What `runProgram` is for + +`runProgram` is a clean interface between the TUI, which is the interface, and +the programmatic parts of what a wizard program does. Three things use it: + +1. **The TUI and the headless runner.** Both call `runProgram` to run a full + wizard program. They reach it through `runProgramAgent` in + [`run-agent-legacy.ts`](../src/programs/run-agent-legacy.ts). This is a + temporary adapter, and we will remove it in the full program. +2. **Testing.** The test workbench calls `runProgram` directly, with no TUI. See + `services/wizard-program/` in + [PostHog/wizard-workbench#4190](https://github.com/PostHog/wizard-workbench/pull/4190). +3. **Other integrations.** If you want to integrate with the wizard more + directly, without everything on top, call `runProgram` yourself. + +### Signatures ```ts import { runProgram } from '@programs'; @@ -54,216 +55,83 @@ export const signature: ( ) => Promise = runProgram; ``` -`programId` names the program for analytics, the route, the gateway spend pin -and the commandments. `runProgram` doesn't look it up in the program registry. -An ID with no entry in `PROGRAM_BINDINGS` runs on `DEFAULT_BINDING`. - -`@programs/types` exports the input, settings, option, outcome and progress -types. The credential types are the exception. Name them as -`ProgramInput['credentials']` and `ProgramOptions['credentials']`. They live in -[`credentials.ts`](../src/programs/credentials.ts) as -`ResolvedProgramCredentials` and `CredentialsProvider`. - -## `ProgramInput` - -`installDir` and `run` are required. Every other field is optional. - -| Field | What `runProgram` does with it | -| ---------------------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `installDir` | The agent's working directory. `artifacts.reportFile` resolves against it. | -| `run` | The program's `AgentRunDefinition`, built by the host from its `ProgramConfig`. It becomes `RunConfig.run`. Its `reportFile` names the report. Its `integrationLabel` and `skillId` label analytics and traces. | -| `program` | The `ProgramSettings` read from the same `ProgramConfig`. See [`ProgramSettings`](#programsettings). Defaults to `{}`. | -| `credentials` | A resolved login: `{ posthog, project, apiUser }`. It wins over `options.credentials`. | -| `runId` | Labels the agent run in `settledRuns` and in run progress events. A random UUID when absent. It isn't the analytics run ID. | -| `overrides` | `{ harness?, sequence?, model? }`, the launch overrides such as `--harness`. The switchboard applies them in development and test builds and ignores them in published builds. | -| `composed` | `true` for a sub-run inside a host program. The switchboard clamps it to linear, and the agent leaves the terminal outro to the host. Defaults to `false`. | -| `skillId` | The run's skill, for question attribution and the result's `skillId`. The orchestrator uses it as the framework key when `integration` is absent. Defaults to `run.skillId`, then `run.integrationLabel`. | -| `integration` | The detected framework. | -| `frameworkDocsUrl` | The framework's docs page, for the orchestrator's preflight message. | -| `flags` | Run flags: `ci`, `signup`, `debug`, `e2eAsk`, `localMcp`, `captureAio`, `benchmark` and `yaraReport`. A missing flag is `false`. | -| `host` | Where PostHog is: `baseUrl`, `region`, `email`, `projectId` and `apiKey`. `baseUrl` also points the token refresh. | -| `wizardFlags` | An evaluated feature flag snapshot. When it's absent, `runProgram` calls `options.featureFlags`. | -| `wizardFlagPayloads` | The flag payloads from the same snapshot. | -| `seedTasks` | Returns the tasks the orchestrator queues before its planner runs. | -| `hooks` | `RunHooks`, each called with the run's credentials: `postRun`, `buildOutroData`, `buildOutroNextSteps` and `recordTaskOutcomes`. | -| `discoveredFeatures`, `warehouseSources`, `mayReportScanResults` | Evidence for the organization's AI SDK stamp. When `mayReportScanResults` is absent or `false`, no stamp is sent. | -| `aiSdkStampReported` | `true` when the host already considered the stamp for this login. `runProgram` then skips it. | - -The agent calls completion hooks only through `hooks`. It never calls the -`postRun`, `buildOutroData` or `buildOutroNextSteps` of a `ProgramRun`, because -those take a session. Bind them into `hooks` yourself. The linear sequence -installs `run.skillId`, not `input.skillId`. - -`runProgram` copies the input when it receives it. `credentials`, `run`, -`program`, `hooks` and `seedTasks` stay by reference, because they can carry -functions or class instances. `structuredClone` copies every other field. A -later host write to the input doesn't reach the run. A copied field that -`structuredClone` can't copy, such as a function in `host`, rejects the call. - -## `ProgramSettings` - -`input.program` carries the program-level settings from `ProgramConfig`. - -| Field | What `runProgram` does with it | -| ------------------- | -------------------------------------------------------------------------------------------------------------- | -| `requiresAi` | `false` skips the AI-processing approval. Any other value, including absent, keeps the check. | -| `agentFlow` | The context-mill flow the orchestrator loads. Defaults to the program ID. | -| `allowedTools` | Tools added on top of the base allowed set. | -| `disallowedTools` | Tools removed from the base allowed set. | -| `excludedTaskTypes` | `(flags) => types`. The task types this run excludes for its flag snapshot. | -| `postAuthGates` | Step IDs the host settles after auth and before the agent starts. `runProgram` passes them to the gate option. | - -## `ProgramOptions` - -Every option is optional. An awaited capability receives the invocation's -`signal`. Without `options.signal`, it receives a signal that never aborts. - -| Option | What `runProgram` does with it | When it's absent | -| -------------------- | --------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------- | -| `credentials` | Calls `resolve(programId, { signal })` once, only when `input.credentials` is absent. | Without `input.credentials`, the run fails. | -| `awaitAiApproval` | Calls it with `{ programId, signal }` when the run needs approval. `true` proceeds. `false` aborts. | A run that needs approval fails. | -| `awaitPostAuthGates` | Calls it once with `{ programId, gates, signal }` when `program.postAuthGates` isn't empty. | The run doesn't wait. | -| `featureFlags` | Calls it once when `input.wizardFlags` is absent. It receives no signal. | The run uses the input's flags, or none. | -| `interaction` | Passes it to `runAgent`, which asks it the agent's questions and task notices. | The agent installs no ask bridge and declines optional task notices. | -| `onProgress` | Receives run events and data snapshots. See [progress](#progress). It is never awaited. | The outcome still carries the settled run and the final data. | -| `signal` | Cancels the invocation. See [cancellation](#cancellation). | Nothing cancels it. | - -A run needs approval when all three hold: - -- `program.requiresAi` isn't `false`. -- Neither `flags.ci` nor `flags.signup` is set. -- The organization's `is_ai_data_processing_approved` isn't `true`. - -`flags.ci` and `flags.signup` skip the check. Set them only when the host has -already handled consent. - -## The outcome - -`runProgram` resolves with a `ProgramRunOutcome`. - -| Field | What it holds | -| ---------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | -| `programId` | The ID the call ran. | -| `outcome` | A `RunOutcome`: `success`, `aborted`, `failed` or `crashed`. | -| `failure` | Present on every outcome except `success`. An `AgentFailure` with a `code` and a `message`. An agent failure can also carry `outroData`, `error`, `exitCode`, `detail` and `authErrorDetail`. A crash always carries the original `error`. | -| `settledRuns` | `{ runId, result }` for the agent run, once `runAgent` returned. Empty when the invocation ended before the agent. | -| `data` | A copy of the invocation's `ProgramInvocationData` when it settled. | -| `diagnostics` | Up to 10 observer failures and late events, newest last. | -| `artifacts.reportFile` | The absolute report path: `run.reportFile` resolved against `installDir`. Set just before the agent starts. It doesn't prove the agent wrote the file. | - -`ProgramInvocationData` has these fields: - -| Field | What it holds | -| -------------------------------------- | -------------------------------------------------------------------------------------------------------------- | -| `credentials`, `apiProject`, `apiUser` | The resolved login, updated when the token refresh returns new credentials. `null` until credentials resolve. | -| `detection.frameworkContext` | Always `{}` from `runProgram`. Framework context reaches a program through its [run host](#host-capabilities). | -| `binding` | The resolved route: `sequence`, `harness`, `model` and `thinkingLevel`. `null` until the route resolves. | -| `aiSdkStampReported` | `true` once the organization's AI SDK stamp was considered. | - -Data and progress are `structuredClone` copies. A class instance comes back as a -plain object: `data.credentials.host` has the `HostResolution` fields but isn't -an instance. `data` holds PostHog tokens, so don't log it. - -`runProgram` writes to the process-wide analytics client. It sends -`agent started` and `switchboard resolved`, sets the `sequence` and `harness` -tags, and identifies the user. It never sends `setup wizard finished`. The host -sends it from the outcome with `analytics.shutdown(status)`. - -### Progress - -`onProgress` receives a `ProgramProgress`, a union of two kinds. Narrow on -`kind` first. - -- **`{ kind: 'run', runId, event }`.** One agent event. `event` is a copied - [`AgentProgress`](../src/agent/README.md#signatures). -- **`{ kind: 'program', data }`.** A copy of `ProgramInvocationData`, sent after - each write. A write happens when credentials resolve, when the AI SDK stamp - latches, when the refresh returns new credentials, and when the route - resolves. - -`runProgram` never awaits the observer. A throw or a rejection becomes a -diagnostic. So does a run event that arrives after the agent run finished. The -outcome copies `diagnostics` when it settles, so a rejection that lands later -isn't in it. - -## Pipeline order - -1. Copy the input. If `signal` is already aborted, return `aborted`. -2. Send the `agent started` analytics event. -3. Resolve credentials from `input.credentials`, else from - `options.credentials`. Without either, fail. -4. Store the login. Identify the user for analytics and set the organization - group. -5. Consider the AI SDK stamp once, unless `aiSdkStampReported` is `true`. -6. When the run needs approval, await `awaitAiApproval`. -7. When `program.postAuthGates` isn't empty, await `awaitPostAuthGates`. -8. When `input.wizardFlags` is absent, await `featureFlags`. -9. Refresh the OAuth token when it has a refresh token, an expiry, and less than - 50 minutes left. A failed refresh keeps the current token. -10. Resolve the route from the program ID, `composed`, the flags and - `overrides`. Tag analytics with the sequence and harness, and send - `switchboard resolved`. -11. Build the run tags, set `artifacts.reportFile`, and call `runAgent` with - `interaction`, the run's progress and `signal`. -12. Record the agent's result and settle. - -The [runner reference](../src/agent/runner/README.md) describes what `runAgent` -does from step 11. - -## Cancellation - -`runProgram` checks `signal` before it starts and after each await in steps 3 -to 9. An abort at any of those points returns `aborted` with the message -`Run cancelled by host.`, even when the capability already resolved. A -capability that rejects after the abort also returns `aborted`. - -`runProgram` doesn't race a capability against the signal. It waits for the -capability to settle, so a capability must settle when its signal aborts. -`featureFlags` receives no signal. - -After step 9, `runProgram` doesn't check the signal itself. `runAgent` receives -it. An abort before the agent's setup returns `aborted` right away. An abort -during the run returns the agent's `aborted` result. Both carry the message -`Agent run cancelled`. - -## Failure outcomes - -Most endings resolve. A rejected promise means the call itself broke. - -| Cause | `outcome` | `failure.code` | `failure.message` | -| -------------------------------------------------------- | ------------------- | ------------------------ | ----------------------------------------------------------------- | -| No `input.credentials` and no `options.credentials` | `failed` | `PHW_INTERNAL_UNHANDLED` | `Credentials are required to run .` | -| The run needs approval and there is no `awaitAiApproval` | `failed` | `PHW_INTERNAL_UNHANDLED` | `AI processing approval is required before this program can run.` | -| `awaitAiApproval` resolves `false` | `aborted` | `PHW_AGENT_ABORT` | `AI processing approval declined.` | -| An await in steps 3 to 9 rejects | `failed` | `PHW_INTERNAL_UNHANDLED` | The error's message | -| The host's signal aborts before the agent | `aborted` | `PHW_AGENT_ABORT` | `Run cancelled by host.` | -| The agent run doesn't succeed | The agent's outcome | The agent's code | The agent's message | -| A copied input field can't be cloned | The promise rejects | | | - -The agent returns `failed` for a coded error, a decided failure or an agent that -stops itself with `[ABORT]`. It returns `aborted` only for the host's signal. It -returns `crashed` for an uncoded throw, with the original `Error` attached. See -the [agent reference](../src/agent/README.md#signatures). - -## Example - -This host runs the `metrics` program with a login it already holds. It asks its -own consent flow for approval: +- **`programId`** says which program runs. +- **`input`** says what to run and where: the program's run definition built + from its `ProgramConfig`, the install directory and the login. +- **`options`** is how the caller plugs in: credentials, questions, progress, + the approval and gate waits, flags and cancellation. + +`runProgram` resolves to one `ProgramRunOutcome`. + +### Field definitions + +Each field has a one-line comment next to it in the code. + +- **`programId`**: which program runs. It names analytics, the route and the + gateway spend. +- **[`input`](../src/programs/run-program.ts#L66)** (`ProgramInput`): what to + run and where. `installDir` and `run` are required. +- **[`input.program`](../src/programs/run-program.ts#L56)** (`ProgramSettings`): + the program's settings from its `ProgramConfig`. +- **[`options`](../src/programs/run-program.ts#L89)** (`ProgramOptions`): how + you plug in. The login, questions, the approval and gate waits, flags, + progress and cancellation. +- **[`options.onProgress`](../src/programs/program-store.ts#L9)** + (`ProgramProgress`): what's happening, as agent events and program data + snapshots. It's never awaited. +- **[Outcome](../src/programs/run-program.ts#L106)** (`ProgramRunOutcome`): how + the run ended, with the agent's result, the final data and the report path. + `data` holds tokens, so don't log it. + +### Example + +This runs the `metrics` program with a login you already hold. It asks your own +consent flow for approval, and logs what the program is doing: ```ts import { RunOutcome } from '@agent'; import { getProgramConfig, runProgram } from '@programs'; -import type { ProgramInput, ProgramOptions } from '@programs/types'; +import type { + ProgramInput, + ProgramOptions, + ProgramProgress, +} from '@programs/types'; type Login = NonNullable; type AskForApproval = NonNullable; +// Log what the program is doing. runProgram never waits for this. +function logProgress(progress: ProgramProgress): void { + // Program data changed, for example the route resolved. + if (progress.kind === 'program') { + const route = progress.data.binding; + if (route) console.log(`route: ${route.sequence} on ${route.harness}`); + return; + } + // Otherwise it's one agent event. + const { event } = progress; + switch (event.kind) { + case 'status': + console.log(event.message); + break; + case 'tasks': { + const done = event.tasks.filter((task) => task.status === 'completed'); + console.log(`tasks: ${done.length}/${event.tasks.length} done`); + break; + } + case 'url': + console.log(`${event.which}: ${event.url}`); + break; + } +} + export async function runMetrics( installDir: string, login: Login, askForApproval: AskForApproval, signal?: AbortSignal, ): Promise { + // Build the run from the program's own config. const config = getProgramConfig('metrics'); if (!config.run || typeof config.run === 'function') { throw new Error('metrics has a static run definition'); @@ -272,9 +140,9 @@ export async function runMetrics( const result = await runProgram( config.id, { - installDir, + installDir, // the project to change run: config.run, - credentials: login, + credentials: login, // skip the login step program: { requiresAi: config.requiresAi, agentFlow: config.agentFlow, @@ -284,185 +152,49 @@ export async function runMetrics( }, }, { - awaitAiApproval: askForApproval, - signal, - onProgress: (progress) => { - if (progress.kind !== 'run') return; - if (progress.event.kind === 'status') { - console.log(progress.event.message); - } - }, + awaitAiApproval: askForApproval, // your consent flow + onProgress: logProgress, + signal, // abort to cancel; the run resolves to aborted }, ); + // Endings resolve to an outcome. Check it instead of catching. if (result.outcome !== RunOutcome.Success) { throw result.failure?.error ?? new Error(result.failure?.message); } - return result.artifacts.reportFile; + return result.artifacts.reportFile; // where the agent wrote its report } ``` -## How the legacy adapter builds `run` - -`runProgramAgent(programConfig, session, { composed? })` is the host for the TUI -and for `--ci`. It builds the input from the `ProgramConfig` and the session, -then calls `runProgram` once: - -1. It throws when the config has no `run`. -2. It starts the audit ledger watcher when the config names `auditLedgerFile`, - before `run` resolves. It stops the watcher when the run returns. -3. It resolves `run`. A static `ProgramRun` passes as is. A function is called - as `run(session, host)`, with a [`ProgramRunHost`](#host-capabilities) over - `getUI()`. -4. It sets `session.skillId`, starts the log file, and runs the health gate and - the Claude settings gate. -5. It calls `runProgram` with the input and options below. -6. It applies the outcome. - -The input comes from the config and the session: - -| `ProgramInput` field | Source | -| -------------------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------- | -| `run` | `ProgramConfig.run`, resolved in step 3. | -| `program` | `requiresAi`, `agentFlow`, `allowedTools`, `disallowedTools` and `excludedTaskTypes` from the config. | -| `program.postAuthGates` | The IDs of `postAuthGateSteps(config.steps)`, the gated steps between the `auth` screen and the `run` screen. | -| `hooks` | `run.postRun`, `run.buildOutroData` and `run.buildOutroNextSteps`, bound to the session. | -| `hooks.recordTaskOutcomes` | Writes the drained queue's outcomes to `session.frameworkContext[TASK_OUTCOMES_KEY]`. | -| `seedTasks` | `config.seedTasks`, bound to the session. | -| `installDir`, `composed`, `skillId`, `integration` | The session and the `composed` option. | -| `overrides` | `session.harness`, `session.sequence` and `session.model`. | -| `frameworkDocsUrl` | The `FRAMEWORK_REGISTRY` docs URL for `session.integration`, else for `session.skillId`. | -| `flags`, `host` | The session's run flags, and its `baseUrl`, `region`, `email`, `projectId` and `apiKey`. | -| `aiSdkStampReported`, `discoveredFeatures`, `warehouseSources`, `mayReportScanResults` | The session's stamp latch, detection results and scan consent. | - -The options come from the UI: - -| `ProgramOptions` field | Source | -| ---------------------- | ------------------------------------------------------------------------------------- | -| `credentials` | `authenticate(session, programId)`, then the session's credentials, project and user. | -| `awaitAiApproval` | Waits for the AI opt-in screen to clear, then resolves `true`. | -| `awaitPostAuthGates` | Waits for each gate in order with `ui.waitForGate`. | -| `featureFlags` | The analytics client's flags and payloads. | -| `onProgress` | Run events go to the UI reducer. Data snapshots project onto the session and the UI. | -| `interaction` | The UI's ask overlay and task notice modal. | -| `signal` | Not set. | - -The projection copies a refreshed token onto the session and the UI, and latches -the AI SDK stamp on the session. When the route first resolves, it registers a -cleanup that flushes the scan report. On a linear route, it also restores the -Claude settings when the outro screen opens. - -The adapter applies the outcome the way the CLI roots expect: - -- A login failure or a flag load failure rethrows after `runProgram` resolves. -- `crashed` rethrows `failure.error`. -- Any other non-success shows the auth error screen when `authErrorDetail` is - set. Then it calls `wizardAbort` with the failure, with status `cancelled` for - `aborted` and `error` otherwise. -- A `success` that isn't composed sends `setup wizard finished` through - `analytics.shutdown('success')`. A failed flush is logged, and the run stays a - success. - -A host with no TUI can build the input the same way. A dynamic `run` and the -program hooks take a `WizardSession`, so the host builds one and fills what the -program reads: +### Cancellation -```ts -import { buildSession } from '@lib/wizard-session'; -import { getProgramConfig } from '@programs'; -import { postAuthGateSteps } from '@programs/program-step'; -import type { ProgramId, ProgramInput, ProgramRunHost } from '@programs/types'; +Pass a `signal` to cancel the run. Every wait receives it, and a wait must +settle when it aborts. A cancelled run resolves to `aborted`, not a rejection. -export async function buildProgramInput( - programId: ProgramId, - installDir: string, - host: ProgramRunHost, -): Promise { - const config = getProgramConfig(programId); - if (!config.run) throw new Error(`${programId} has no run`); - - const session = buildSession({ installDir, ci: true }); - const run = - typeof config.run === 'function' - ? await config.run(session, host) - : config.run; - const { postRun, buildOutroData, buildOutroNextSteps } = run; - const { seedTasks } = config; - - return { - installDir, - run, - program: { - requiresAi: config.requiresAi, - agentFlow: config.agentFlow, - allowedTools: config.allowedTools, - disallowedTools: config.disallowedTools, - excludedTaskTypes: config.excludedTaskTypes, - postAuthGates: postAuthGateSteps(config.steps).map((step) => step.id), - }, - seedTasks: seedTasks && (() => seedTasks(session)), - hooks: { - postRun: postRun && ((credentials) => postRun(session, credentials)), - buildOutroData: - buildOutroData && - ((credentials) => buildOutroData(session, credentials) ?? undefined), - buildOutroNextSteps: - buildOutroNextSteps && - ((credentials, completed) => - buildOutroNextSteps(session, credentials, completed)), - }, - }; -} -``` - -Some programs read session fields that a TUI step or `ciPreRun` fills. For -example, the `posthog-integration` run reads `session.frameworkConfig`, which -its detect step sets. +### Failures -## Host capabilities +Most endings resolve to an outcome instead of throwing. Check `outcome` and read +`failure`. The promise rejects only when the call itself can't run, such as an +input field that can't be copied. The cases are in +[`run-program.ts`](../src/programs/run-program.ts#L160). -A program's `run` and `ciPreRun` callbacks take their effects from a host. They -don't call `getUI()`. Both types come from `@programs/types`: +### Program callbacks -```ts -import type { ProgramCiHost, ProgramRunHost } from '@programs/types'; - -const frameworkContext = new Map(); - -export const runHost: ProgramRunHost = { - getFrameworkContext: (key) => frameworkContext.get(key), - setFrameworkContext: (key, value) => { - frameworkContext.set(key, value); - }, - warn: (message) => console.warn(message), -}; - -export const ciHost: ProgramCiHost = { - log: { - info: (message) => console.log(message), - warn: (message) => console.warn(message), - }, -}; -``` +A program's `run` and `ciPreRun` receive a host, `ProgramRunHost` or +`ProgramCiHost`, instead of calling `getUI()`. Both types come from +`@programs/types` and are defined in +[`host-capabilities.ts`](../src/programs/host-capabilities.ts). Build the host +when you build the run from a `ProgramConfig`, before you call `runProgram`. -| Capability | Received by | What programs use it for | Supplied by | -| ---------------- | --------------------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------- | -| `ProgramRunHost` | `ProgramConfig.run(session, host)` | `posthog-integration` warns when the SDK package or `package.json` is missing. `error-tracking` warns from the `posthog-cli` preinstall. `error-tracking-upload-source-maps` reads the picked project and writes the completed variant. `audit` passes it to its base run. | The legacy adapter. Each call reads `getUI()` when it runs. | -| `ProgramCiHost` | `ProgramConfig.ciPreRun(session, host)` | Project scoping logs its scan progress and its fallbacks. `posthog-integration`, `error-tracking` and `replay-vision` scope the project this way. | The non-interactive runner. Each call reads `getUI().log` when it runs. | +## `runAgent` -A run definition keeps its host. The source maps program reads the picked -project when the agent's prompt is built, after the picker screen, and writes -`sourceMapsCompletedVariant` in `postRun`. So a `ProgramRunHost` must stay -usable until the run returns. +### What `runAgent` is for -Neither capability is a `runProgram` option. Both belong to building the run -from a `ProgramConfig`, before `runProgram` starts. +`runAgent` runs only the agent. Call it when you have a resolved route and no +program policy to apply. It doesn't log in, ask for consent, load flags or +resolve a route. The caller does those. -## `runAgent` for detection and standalone callers - -Call `runAgent` directly when you have a resolved route and no program policy to -apply. It doesn't authenticate, check consent, load flags, resolve a route or -send `setup wizard finished`. The caller does those. +### Signatures ```ts import { runAgent } from '@agent'; @@ -486,52 +218,42 @@ export const signature: ( ) => Promise = runAgent; ``` -These callers use it: - -| Caller | What it runs | -| --------------------------------------------------------------------------------- | ------------------------------------------------------------------ | -| `runProgram` | One program's agent run. | -| `detectProjectsWithAgent` in [`agentic.ts`](../src/programs/detection/agentic.ts) | The agentic project scan, one `runAgent` call per attempt. | -| [`a3-fault-probe.no-jest.ts`](../scripts/a3-fault-probe.no-jest.ts) | A fault probe against a local gateway, with synthetic credentials. | - -`runAgent` mints a gateway token from `input.credentials` for `config.programId` -before any agent starts, and re-mints near expiry. A refused mint returns -`failed`. In development and test builds, -`configureGatewayFromCIEnvironment(projectId, region)` from `@agent` loads a -fixed gateway token from the file that `WIZARD_CI_GATEWAY_TOKEN_FILE` names. -`runAgent` then uses that token and doesn't mint. The gateway auth is -process-wide. - -### Detection - -Agentic detection scans a repository for its projects through `runAgent`. -`scopeInstallDirToProject`, the self-driving and error tracking project scans, -and the source maps scan call it. It needs `session.credentials` and throws -without them. Each attempt passes this run: - -| Setting | Value | -| ----------------------- | --------------------------------------------------------------------------- | -| `binding` | Linear, Anthropic, Haiku. | -| `run.prompt` | The scan prompt. It replaces the assembled project prompt. | -| `run.collectTranscript` | `true`. The report is read from `snapshot.transcriptTail`. | -| `run.requestRemark` | `false`. | -| `composed` | `true`. | -| `allowedTools` | `Read`, `Grep` and `Glob`. | -| `scanReport` | `'defer'`. The scan's security scans count toward the program run's report. | -| `programId` | The caller's program, so the scan's spend is attributed to it. | -| `wizardMetadata` | The run tags plus `call_type: detection`. | - -Each attempt has its own deadline signal. A first attempt that times out retries -once. A second timeout throws `AgenticDetectionTimeoutError`. Any other -non-success throws the failure's error. The scan parses verdict lines from the -transcript tail. A tail with no verdicts that contains `[ABORT]` is an empty -report. A tail with no verdicts retries once, then throws. - -Detection forwards each `activity` line to its `onEvent` callback. The UI -receives every other event except `lifecycle`, `completion`, `spinner`, and log -lines below `warn`. - -### Standalone example +- **`config`** says what the agent runs: the run definition, the route and the + tools. +- **`input`** says where and as whom: the project, the login and the flags. +- **`options`** carries progress, questions and cancellation. + +`runAgent` resolves to one `RunResult`. It mints its own gateway token from +`input.credentials`. + +### Field definitions + +Each field has a comment in the code. The +[agent reference](../src/agent/README.md#signatures) covers the rest. + +- **[`config`](../src/agent/runner/shared/types.ts#L153)** (`RunConfig`): what + the agent runs, with its route and tools. +- **[`config.run`](../src/agent/runner/shared/types.ts#L49)** + (`AgentRunDefinition`): the prompt and run options, such as `prompt`, + `collectTranscript` and `requestRemark`. +- **[`input`](../src/agent/runner/shared/types.ts#L207)** (`RunInput`): where + and as whom. The project, the login and the flags. +- **`options`**: `onProgress` for agent events, `interaction` for questions, and + `signal` to cancel. +- **[Result](../src/agent/runner/shared/types.ts#L309)** (`RunResult`): how the + run ended, with a snapshot of its tasks and transcript. + +### Callers + +| Caller | What it runs | +| --------------------------------------------------------------------------------- | ----------------------------------------------- | +| `runProgram` | One program's agent run. | +| `detectProjectsWithAgent` in [`agentic.ts`](../src/programs/detection/agentic.ts) | The agentic project scan, one call per attempt. | +| [`a3-fault-probe.no-jest.ts`](../scripts/a3-fault-probe.no-jest.ts) | A fault probe against a local gateway. | + +### Example + +This lists a project's files with a read-only agent and returns what it said: ```ts import { runAgent, RunOutcome } from '@agent'; @@ -545,23 +267,24 @@ import { export async function listProjectFiles( programId: string, - input: RunInput, + input: RunInput, // the project, the login and the flags signal?: AbortSignal, ): Promise { const config: RunConfig = { - programId, + programId, // attributes the gateway spend run: { integrationLabel: 'list-files', prompt: () => 'List the files in the working directory. Change nothing.', - collectTranscript: true, - requestRemark: false, + collectTranscript: true, // keep the agent's output to return below + requestRemark: false, // no closing remark spinnerMessage: 'Listing files...', successMessage: 'Listed files', estimatedDurationMinutes: 1, reportFile: '', docsUrl: 'https://posthog.com/docs', }, - composed: true, + composed: true, // a sub-run: the caller owns the outro + // You pick the route yourself. runAgent doesn't resolve one. binding: { sequence: Sequence.linear, harness: Harness.anthropic, @@ -572,11 +295,12 @@ export async function listProjectFiles( wizardFlags: {}, wizardFlagPayloads: {}, wizardMetadata: {}, - allowedTools: ['Read', 'Glob'], + allowedTools: ['Read', 'Glob'], // read-only }; const result = await runAgent(config, input, { signal, + // Each step the agent takes, as one line. onProgress: (event) => { if (event.kind === 'activity') console.log(event.line); }, @@ -588,17 +312,5 @@ export async function listProjectFiles( } ``` -`collectTranscript` and `requestRemark: false` take effect on the linear -sequence with the Anthropic harness. See the -[agent reference](../src/agent/README.md#run-definition). - -## Process-owned CLI runs - -Development and test builds accept `--ci`. It is a whole-process run, not a -function that returns an outcome. The non-interactive runner installs the -`LoggingUI`, builds a session, and loads the fixed CI gateway token. It runs the -program's `ciPreRun` with a `ProgramCiHost`, or the steps' `onReady` hooks. Then -it calls `runProgramAgent(config, session)`. The process exit code and the logs -report the result. Published builds don't accept `--ci`. For the credentials it -needs, see -[local credentials](local-dev.md#credentials-for-local-ci-and-headless-runs). +`collectTranscript` and `requestRemark` take effect on the linear sequence with +the Anthropic harness. diff --git a/docs/images/wizard-run-paths.svg b/docs/images/wizard-run-paths.svg new file mode 100644 index 000000000..9b2be6fbf --- /dev/null +++ b/docs/images/wizard-run-paths.svg @@ -0,0 +1,53 @@ + + How each way to run the wizard reaches runAgent + + + + + + + + + + TUI + + + Headless runner + + + + WizardStore + session and Ink screens + + + WizardSession + and LoggingUI + + + Workbench and + other integrations + + + Detection and + standalone callers + + + + runProgram + ProgramStore + + + + runAgent + + + + + + + + + + + + diff --git a/src/agent/runner/README.md b/src/agent/runner/README.md index 763419e36..5bcde7638 100644 --- a/src/agent/runner/README.md +++ b/src/agent/runner/README.md @@ -184,4 +184,4 @@ present. `scanReport: 'defer'`, and reads its report from `snapshot.transcriptTail`. - The fault probe in `scripts/a3-fault-probe.no-jest.ts`, and any standalone host. See the - [developer interfaces](../../../docs/developer-interfaces.md#runagent-for-detection-and-standalone-callers). + [developer interfaces](../../../docs/developer-interfaces.md#runagent). diff --git a/src/programs/README.md b/src/programs/README.md index b93d80b8a..4455cff29 100644 --- a/src/programs/README.md +++ b/src/programs/README.md @@ -94,7 +94,7 @@ runs the readiness and Claude settings gates, and builds the input from the session and the config. It answers the capabilities from the UI, maps run events onto `getUI()`, and projects data snapshots onto the session. Then it applies the outcome with `wizardAbort` and the terminal analytics event. -[How the legacy adapter builds `run`](../../docs/developer-interfaces.md#how-the-legacy-adapter-builds-run) +[How the legacy adapter builds `run`](../../docs/developer-interfaces.md#what-runprogram-is-for) lists every field it maps. Three places call it: @@ -129,7 +129,7 @@ instead of calling `getUI()`. Both types live in Neither capability is a `runProgram` option. They belong to building the run from a `ProgramConfig`, which happens before `runProgram`. The -[developer interfaces](../../docs/developer-interfaces.md#host-capabilities) +[developer interfaces](../../docs/developer-interfaces.md#program-callbacks) list which programs use each capability. ## Current limits From 5651941fc08b4a90ab8ab3d835b1f448dcd84f7f Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Thu, 24 Sep 2026 20:17:37 -0400 Subject: [PATCH 16/29] docs(programs): link the workbench repo, not a pull request Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- docs/developer-interfaces.md | 6 +++--- 1 file changed, 3 insertions(+), 3 deletions(-) diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index 6888f1185..af131d86f 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -31,9 +31,9 @@ the programmatic parts of what a wizard program does. Three things use it: wizard program. They reach it through `runProgramAgent` in [`run-agent-legacy.ts`](../src/programs/run-agent-legacy.ts). This is a temporary adapter, and we will remove it in the full program. -2. **Testing.** The test workbench calls `runProgram` directly, with no TUI. See - `services/wizard-program/` in - [PostHog/wizard-workbench#4190](https://github.com/PostHog/wizard-workbench/pull/4190). +2. **Testing.** The test workbench calls `runProgram` directly, with no TUI, + from `services/wizard-program/` in + [wizard-workbench](https://github.com/PostHog/wizard-workbench). 3. **Other integrations.** If you want to integrate with the wizard more directly, without everything on top, call `runProgram` yourself. From 784a84bb6f24fac805705744954efa742bf83f77 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Thu, 24 Sep 2026 20:18:00 -0400 Subject: [PATCH 17/29] docs(programs): field definitions as tables Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- docs/developer-interfaces.md | 46 ++++++++++++------------------------ 1 file changed, 15 insertions(+), 31 deletions(-) diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index af131d86f..01d944ce8 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -65,23 +65,14 @@ export const signature: ( ### Field definitions -Each field has a one-line comment next to it in the code. - -- **`programId`**: which program runs. It names analytics, the route and the - gateway spend. -- **[`input`](../src/programs/run-program.ts#L66)** (`ProgramInput`): what to - run and where. `installDir` and `run` are required. -- **[`input.program`](../src/programs/run-program.ts#L56)** (`ProgramSettings`): - the program's settings from its `ProgramConfig`. -- **[`options`](../src/programs/run-program.ts#L89)** (`ProgramOptions`): how - you plug in. The login, questions, the approval and gate waits, flags, - progress and cancellation. -- **[`options.onProgress`](../src/programs/program-store.ts#L9)** - (`ProgramProgress`): what's happening, as agent events and program data - snapshots. It's never awaited. -- **[Outcome](../src/programs/run-program.ts#L106)** (`ProgramRunOutcome`): how - the run ended, with the agent's result, the final data and the report path. - `data` holds tokens, so don't log it. +| Field | Type | What it's for | +| ----------------------------------------------------------- | ------------------- | -------------------------------------------------------------------------------- | +| `programId` | `string` | Which program runs. It names analytics, the route and the gateway spend. | +| [`input`](../src/programs/run-program.ts#L66) | `ProgramInput` | What to run and where. `installDir` and `run` are required. | +| [`input.program`](../src/programs/run-program.ts#L56) | `ProgramSettings` | The program's settings from its `ProgramConfig`. | +| [`options`](../src/programs/run-program.ts#L89) | `ProgramOptions` | The login, questions, approval and gate waits, flags, progress and cancellation. | +| [`options.onProgress`](../src/programs/program-store.ts#L9) | `ProgramProgress` | Agent events and program data snapshots. Never awaited. | +| [Outcome](../src/programs/run-program.ts#L106) | `ProgramRunOutcome` | How the run ended, the agent's result, the final data and the report path. | ### Example @@ -228,20 +219,13 @@ export const signature: ( ### Field definitions -Each field has a comment in the code. The -[agent reference](../src/agent/README.md#signatures) covers the rest. - -- **[`config`](../src/agent/runner/shared/types.ts#L153)** (`RunConfig`): what - the agent runs, with its route and tools. -- **[`config.run`](../src/agent/runner/shared/types.ts#L49)** - (`AgentRunDefinition`): the prompt and run options, such as `prompt`, - `collectTranscript` and `requestRemark`. -- **[`input`](../src/agent/runner/shared/types.ts#L207)** (`RunInput`): where - and as whom. The project, the login and the flags. -- **`options`**: `onProgress` for agent events, `interaction` for questions, and - `signal` to cancel. -- **[Result](../src/agent/runner/shared/types.ts#L309)** (`RunResult`): how the - run ended, with a snapshot of its tasks and transcript. +| Field | Type | What it's for | +| ------------------------------------------------------- | -------------------- | ------------------------------------------------------------------------------- | +| [`config`](../src/agent/runner/shared/types.ts#L153) | `RunConfig` | What the agent runs, with its route and tools. | +| [`config.run`](../src/agent/runner/shared/types.ts#L49) | `AgentRunDefinition` | The prompt and run options, such as `collectTranscript`. | +| [`input`](../src/agent/runner/shared/types.ts#L207) | `RunInput` | Where and as whom: the project, the login and the flags. | +| `options` | | `onProgress` for agent events, `interaction` for questions, `signal` to cancel. | +| [Result](../src/agent/runner/shared/types.ts#L309) | `RunResult` | How the run ended, with a snapshot of its tasks and transcript. | ### Callers From e2668feada2e9c6db3496ffe80cf7e1cac13c56a Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Thu, 24 Sep 2026 20:18:24 -0400 Subject: [PATCH 18/29] docs(programs): name files, not paths Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- docs/developer-interfaces.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index 01d944ce8..36ca66e1d 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -32,7 +32,7 @@ the programmatic parts of what a wizard program does. Three things use it: [`run-agent-legacy.ts`](../src/programs/run-agent-legacy.ts). This is a temporary adapter, and we will remove it in the full program. 2. **Testing.** The test workbench calls `runProgram` directly, with no TUI, - from `services/wizard-program/` in + from the `wizard-program` service in [wizard-workbench](https://github.com/PostHog/wizard-workbench). 3. **Other integrations.** If you want to integrate with the wizard more directly, without everything on top, call `runProgram` yourself. From a0f17989f01bcdf13ef8cdb333771106e832e495 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Thu, 24 Sep 2026 20:40:32 -0400 Subject: [PATCH 19/29] docs(programs): short programs and agent READMEs, runner not host Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- AGENTS.md | 4 +- docs/developer-interfaces.md | 21 ++-- src/agent/README.md | 235 ++++------------------------------- src/agent/runner/README.md | 6 +- src/programs/README.md | 193 +++++++--------------------- 5 files changed, 86 insertions(+), 373 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index e21a1667b..44b4b660c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -202,8 +202,8 @@ wizard run points. Full catalog: [`docs/local-dev.md`](docs/local-dev.md). types so they satisfy `Record`. - All UI calls go through `getUI()` (returns `WizardUI` interface). Never import the store directly from business logic. A program's `run` and `ciPreRun` - callbacks use the host they receive (`ProgramRunHost`, `ProgramCiHost`), not - `getUI()`. + callbacks use the runner context they receive (`RunnerContext`, + `CiRunnerContext`), not `getUI()`. - Shared helpers never call `getUI()`; they take a sink or return data. `debug()` reaches the UI through the sink `src/ui/index.ts` installs. - Outside `src/agent`, import the agent through `@agent` or `@agent/types`. Add diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index 36ca66e1d..ee0ff626c 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -25,12 +25,17 @@ functions. ### What `runProgram` is for `runProgram` is a clean interface between the TUI, which is the interface, and -the programmatic parts of what a wizard program does. Three things use it: +the programmatic parts of what a wizard program does. + +> ⚠️ **Temporary adapter.** The TUI and the headless runner reach `runProgram` +> through `runProgramAgent` in +> [`run-agent-legacy.ts`](../src/programs/run-agent-legacy.ts). This is a +> temporary adapter, and we will remove it in the full program. + +Three things use it: 1. **The TUI and the headless runner.** Both call `runProgram` to run a full - wizard program. They reach it through `runProgramAgent` in - [`run-agent-legacy.ts`](../src/programs/run-agent-legacy.ts). This is a - temporary adapter, and we will remove it in the full program. + wizard program. 2. **Testing.** The test workbench calls `runProgram` directly, with no TUI, from the `wizard-program` service in [wizard-workbench](https://github.com/PostHog/wizard-workbench). @@ -171,11 +176,11 @@ input field that can't be copied. The cases are in ### Program callbacks -A program's `run` and `ciPreRun` receive a host, `ProgramRunHost` or -`ProgramCiHost`, instead of calling `getUI()`. Both types come from +A program's `run` and `ciPreRun` receive a runner context, `RunnerContext` or +`CiRunnerContext`, instead of calling `getUI()`. Both types come from `@programs/types` and are defined in -[`host-capabilities.ts`](../src/programs/host-capabilities.ts). Build the host -when you build the run from a `ProgramConfig`, before you call `runProgram`. +[`runner-context.ts`](../src/programs/runner-context.ts). Build it when you +build the run from a `ProgramConfig`, before you call `runProgram`. ## `runAgent` diff --git a/src/agent/README.md b/src/agent/README.md index 8fba26f1b..518609aaa 100644 --- a/src/agent/README.md +++ b/src/agent/README.md @@ -1,219 +1,36 @@ # Agent -The agent runs one AI pipeline against a project directory. It takes resolved -data in, reports through progress events, asks through an injected answerer, and -returns a result. It never reads a session, a store or a UI. +The agent runs one AI pipeline against a project. It takes resolved data in, +reports through progress events, asks through an answerer you pass, and returns +a result. It never reads a session, a store or a UI. -## Signatures +To call it from code, use `runAgent`. The +[developer interfaces](../../docs/developer-interfaces.md#runagent) cover it. -Import runtime values from `@agent` and types from `@agent/types`. Nothing -outside `src/agent` imports deeper. Lint and the architecture test reject it. - -```ts -import { runAgent } from '@agent'; -import type { - AgentInteraction, - AgentProgress, - RunConfig, - RunInput, - RunResult, -} from '@agent/types'; - -// The shape `@agent` exports. -export const signature: ( - config: RunConfig, - input: RunInput, - options?: { - onProgress?: (event: AgentProgress) => unknown; - interaction?: AgentInteraction; - signal?: AbortSignal; - }, -) => Promise = runAgent; -``` - -- `RunConfig`: the program ID, its `AgentRunDefinition` as `run`, `composed`, - the resolved `binding` (sequence, harness, model, effort), the switchboard - inputs, the skills origin, the flag snapshot and its payloads, the trace tags, - the tool allow and deny lists, `agentFlow`, `excludedTaskTypes`, `seedTasks`, - the bound completion `hooks`, and `scanReport`. -- `RunInput`: the install directory, resolved credentials, the project and user - payloads, the skill ID, the detected integration and its docs URL, `flags` - (`ci`, `signup`, `debug`, `e2eAsk`, `localMcp`, `captureAio`, `benchmark`, - `yaraReport`), and the host the CLI was told. -- `RunResult`: `outcome` is `RunOutcome.Success | Aborted | Failed | Crashed`. - `Success` can carry an `outro`. The other three carry a `failure` - (`AgentFailure`: message, outro data, error, exit code, error code, detail, - auth error detail). Every result carries `skillId` and a `snapshot` of what - the run reported: tasks, status lines, stage, token usage totals, final cost, - dashboard and notebook URLs, handoff text, and the transcript tail when the - run collected one. -- `AgentProgress`: one event per thing the run reports, in emission order. - Kinds: `lifecycle`, `spinner`, `log`, `status`, `tasks`, `stage`, `url`, - `usage`, `finalCost`, `authError`, `handoff`, `completion` and `activity`. - Payloads are copies, never live objects. -- `AgentInteraction`: every member is optional. `ask(question, { signal })` - resolves with answers. `taskNotice(notice, { signal })` resolves with whether - to keep an optional task. Each request has its own signal. It aborts when that - request times out, the host aborts the run, or another task fails the run. On - abort the host dismisses that request alone, without throwing. -- Errors: the agent doesn't exit the process and doesn't reject. A caught coded - error, such as a refused gateway mint, becomes `Failed`. An uncoded throw - becomes `Crashed` with the error attached. A gateway 401 returns an auth - failure, and the host decides whether to show auth UI. `Aborted` means the - host's signal cancelled the run. An agent that stops itself with `[ABORT]` - returns `Failed` with its abort code. -- Analytics shutdown is host-owned. The agent never sends the terminal - `setup wizard finished` event. The host sends it from the outcome: `Success` - is `success`, `Aborted` is `cancelled`, `Failed` and `Crashed` are `error`. - -Other runtime exports: `RunOutcome`, `AgentSignals`, `WIZARD_TOOL_NAMES`, -`resolveBinding`, `shouldDisableAsk`, `LONGER_ASK_TIMEOUT_MS`, `buildRunTags`, -`configureGatewayFromCIEnvironment`, `flushScanReport`, `downloadSkill`, -`TASK_OUTCOMES_KEY`, and `runMcpPromptViaSdk`, which loads the streaming module -on first call. - -## Run definition - -`RunConfig.run` is an `AgentRunDefinition`. These fields shape the prompt and -what the run collects: - -| Field | What it does | -| ------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `skillId` | The skill the linear sequence installs before the agent starts. Omit it to let the agent discover skills. | -| `customPrompt` | Instructions appended after the default project prompt. | -| `prompt` | Replaces the assembled project prompt. It receives the same `PromptContext`: project ID and key, host, skill path, and the organization and team opt-ins. | -| `collectTranscript` | Keeps a transcript tail on `snapshot.transcriptTail` and reports each agent step as an `activity` event. | -| `requestRemark` | Asks for the end-of-run reflection remark. Defaults to `true`. `false` skips it. | -| `abortCases` | Known `[ABORT] ` cases and the outro each one renders. | - -`prompt`, `customPrompt` and `abortCases` apply on the linear sequence. The -orchestrator builds each task's prompt from its context-mill flow. -`collectTranscript` and `requestRemark` take effect on the linear sequence with -the Anthropic harness. The Pi harness always asks for the remark on a linear -run. The orchestrator never asks for it. - -The other fields carry the run's copy, report file, docs URL, extra MCP servers, -question limits and step analytics. See -[`shared/types.ts`](runner/shared/types.ts). - -### Transcript tail - -With `collectTranscript`, the run observes every SDK message: - -- Each assistant text block joins the tail. The tail keeps the newest blocks, up - to 256 × 1024 characters, and drops the oldest first. A single block over the - cap stays whole. -- The run's final result text follows the kept blocks. -- `snapshot.transcriptTail` is the kept blocks, one per line, then the final - result. -- Each non-empty text block also emits `{ kind: 'activity', line }`, trimmed and - cut at 100 characters. -- Each tool call emits `{ kind: 'activity', line }` with the tool name and its - file path, pattern or path. +## What goes in and out -Only the caller that set `collectTranscript` wants `activity` lines. The TUI's -progress reducer ignores them. +| Direction | What the agent takes or gives | What never crosses | +| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------- | +| In | `RunConfig`: the program ID, the run definition, a resolved route, tools and a flag snapshot. `RunInput`: the project directory, the login, the project and user, the flags and the PostHog host. `onProgress`, `interaction` and `signal`. | A `WizardSession`, a TUI or headless store, `getUI()` or a `ProgramConfig` | +| Out | Progress events as copies, questions through `interaction`, and one `RunResult`. | Live objects, a thrown error or a process exit | -### Scan report +## Where things live -The agent counts its security scans in process-wide state. At the end of a run, -`runAgent` flushes them. It sends the scan telemetry and resets the counts. With -`flags.yaraReport`, it also writes the local report file and emits its path as -an `info` log event. `RunConfig.scanReport: 'defer'` skips the flush, so the -run's scans count toward the next flush. Agentic detection defers, so its scans -land in the program run's report. +| What | Where | +| ----------------------------------------- | ------------------------------------------------- | +| The entry points | [`index.ts`](index.ts) and [`types.ts`](types.ts) | +| The run types: config, input and result | [`shared/types.ts`](runner/shared/types.ts) | +| Progress events and the answerer | [`progress.ts`](progress.ts) | +| Sequences, harnesses and route resolution | [`runner`](runner/README.md) | +| The wizard tools both harnesses share | [`tools`](tools) | +| The benchmark pipeline | [`middleware`](middleware) | +| Security scans of what the run installs | [`yara-hooks.ts`](yara-hooks.ts) | -## Who calls `runAgent` - -- **`runProgram`**, for every program's agent run. It resolves credentials, - consent, flags and the route first. See the - [developer interfaces](../../docs/developer-interfaces.md). -- **Agentic detection**, `detectProjectsWithAgent` in - [`src/programs/detection/agentic.ts`](../programs/detection/agentic.ts). It - sets `prompt`, `collectTranscript: true`, `requestRemark: false` and - `scanReport: 'defer'`, and reads its report from the transcript tail. -- **The fault probe**, - [`scripts/a3-fault-probe.no-jest.ts`](../../scripts/a3-fault-probe.no-jest.ts), - against a local gateway with synthetic credentials. - -A standalone caller supplies everything `runProgram` would resolve. `runAgent` -doesn't authenticate the user, check consent, load flags or pick a route. It -does mint gateway auth: before any agent starts, it mints a scoped gateway token -from `input.credentials` for `config.programId`, and re-mints near expiry. In -development and test builds, `configureGatewayFromCIEnvironment` loads a fixed -token from `WIZARD_CI_GATEWAY_TOKEN_FILE` instead, and `runAgent` uses it -without minting. The gateway auth cache is process-wide. - -Minimal invocation: - -```ts -import { runAgent, RunOutcome } from '@agent'; -import type { - AskAnswers, - PendingQuestion, - RunConfig, - RunInput, -} from '@agent/types'; - -export async function runWithAnswers( - config: RunConfig, - input: RunInput, - answersFor: (question: PendingQuestion) => Promise, -): Promise { - const result = await runAgent(config, input, { - onProgress: (event) => { - if (event.kind === 'log') console.log(event.message); - }, - interaction: { - ask: (question) => answersFor(question), - }, - }); - if (result.outcome !== RunOutcome.Success) { - process.exitCode = result.failure.exitCode ?? 1; - } -} -``` - -`src/agent/__tests__/run-agent-standalone.test.ts` runs the agent this way with -no UI, no store and no registry. - -## Intent - -Programs call the agent to do the work a skill describes. The TUI and the -non-interactive runner observe the run through `onProgress` and answer it -through `interaction`. `src/programs/run-agent-legacy.ts` does both on top of -the session, through `runProgram`. - -Without `onProgress` the run completes, and its snapshot still comes back in the -result. Without `interaction` the agent installs no ask bridge: `wizard_ask` -returns its "not available" error and optional task notices are declined, which -is what a `--ci` run does. A throwing or rejecting observer is logged and the -run continues. - -## Architecture - -The agent owns run state for one invocation: the task queue, phase, status, -resolved skill, handoff text, usage and the final result. It depends on -`src/shared` and on `src/env.ts`. It depends on program types only until the -bindings table moves to programs. +Import runtime values from `@agent` and types from `@agent/types`. Nothing +outside the agent imports deeper, and lint rejects it. -```text -caller ── RunConfig + RunInput ──▶ runAgent - │ prepareRun: gateway mint, triage provider - ▼ - sequence (linear | orchestrator) - │ - harness (anthropic | pi) ── tools (MCP or pi-native) - │ - onProgress ◀── events ───┤──── questions ──▶ interaction - ▼ - scan report flush (unless deferred) - ▼ - RunResult -``` +## Future work -`runner/` holds the dispatcher, sequences, harnesses and the switchboard. See -the [runner reference](runner/README.md). `tools/` holds the wizard tools shared -by both harnesses. `middleware/` holds the benchmark pipeline. `progress.ts` -defines the event and interaction contracts. `yara-hooks.ts` scans what the run -installs. +- **Bindings move to programs.** The program bindings table still lives in the + agent. It will move to programs, and the agent will take a resolved route + only. diff --git a/src/agent/runner/README.md b/src/agent/runner/README.md index 5bcde7638..d62313350 100644 --- a/src/agent/runner/README.md +++ b/src/agent/runner/README.md @@ -109,7 +109,7 @@ Each layer is replaceable. %%{init: {"block": {"padding": 20}}}%% block-beta columns 11 - hostBand["Host: legacy adapter and UI"]:11 + hostBand["Caller: legacy adapter and UI"]:11 runProgramAgent["runProgramAgent"]:3 space:1 wizardAbort["wizardAbort"]:3 space:4 space:11 programsBand["Programs"]:11 @@ -152,7 +152,7 @@ block-beta Calls descend on the left, results return through the middle, and cancellation moves down the right. Blue marks the result contracts and run-scoped abort. On the first fatal task result, `drainQueue` stops scheduling, cancels active work -and pending asks, joins siblings, then preserves that failure for the host to +and pending asks, joins siblings, then preserves that failure for the caller to present. ## Flow @@ -183,5 +183,5 @@ present. sets `run.prompt`, `collectTranscript: true`, `requestRemark: false` and `scanReport: 'defer'`, and reads its report from `snapshot.transcriptTail`. - The fault probe in `scripts/a3-fault-probe.no-jest.ts`, and any standalone - host. See the + caller. See the [developer interfaces](../../../docs/developer-interfaces.md#runagent). diff --git a/src/programs/README.md b/src/programs/README.md index 4455cff29..d936abd1a 100644 --- a/src/programs/README.md +++ b/src/programs/README.md @@ -1,160 +1,51 @@ # Programs -A program is a `ProgramConfig`: the steps the TUI walks, the agent run it -performs, and its program-level settings. `src/programs/` holds the program -configs, detection, the framework registry, the task stream, and `runProgram`, -which runs one program's agent from explicit inputs. +A program is one thing the wizard does for a user, such as adding PostHog or +setting up error tracking. Each program is a `ProgramConfig`: the screens the +TUI walks, the agent run it performs, and its settings. -The [developer interfaces](../../docs/developer-interfaces.md) document the -callable contract: every `ProgramInput`, `ProgramSettings` and `ProgramOptions` -field, the outcome, the pipeline order, cancellation and the failure outcomes. -This page covers who owns what around that call. +To run a program from code, call `runProgram`. The +[developer interfaces](../../docs/developer-interfaces.md) cover it. -## Entry points +## What goes in and out of `runProgram` -Import runtime values from `@programs` and types from `@programs/types`. - -| Entry | Exports | -| ----------------- | -------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -| `@programs` | `runProgram`, and the registry: `PROGRAM_REGISTRY`, `Program`, `getProgramConfig`, `getSubcommandPrograms`, `getCommandPath` and `getLaunchablePrograms`. | -| `@programs/types` | The call types: `ProgramInput`, `ProgramSettings`, `ProgramOptions`, `ProgramOverrides`, `WizardFlagSnapshot` and `ProgramRunOutcome`. The store types: `ProgramProgress`, `ProgramRunProgress`, `ProgramDataProgress`, `ProgramInvocationData`, `ProgramDiagnostic` and `SettledProgramRun`. The host capabilities: `ProgramRunHost` and `ProgramCiHost`. The config types: `ProgramId`, `SubcommandProgram`, `ProgramConfig`, `ProgramStep`, `ProgramReadyContext`, `StoreInitContext`, `FrameworkConfig` and `SetupQuestion`. | - -Code inside the repository reaches deeper modules through `@programs/*`: - -- `runProgramAgent` in [`run-agent-legacy.ts`](run-agent-legacy.ts), the legacy - adapter. -- `postAuthGateSteps` in [`program-step.ts`](program-step.ts), which lists the - gated steps between the `auth` and `run` screens. -- `ResolvedProgramCredentials` and `CredentialsProvider` in - [`credentials.ts`](credentials.ts). -- `ProgramStore` in [`program-store.ts`](program-store.ts). -- The detection tools in [`detection/`](detection/index.ts). - -## The host - -The caller of `runProgram` is the host. `runProgram` owns the policy around one -agent run. The host owns everything around the invocation. - -| `runProgram` owns | The host owns | -| -------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------------- | -| One `ProgramStore` for the invocation. | Building `run` and the settings from a `ProgramConfig`. | -| Calling the credentials provider, then identifying the user for analytics. | Login and project choice, inside its credentials provider. | -| The AI SDK stamp, once per invocation. | The consent screen behind `awaitAiApproval`. | -| Deciding when approval and post-auth gates apply. | The gate screens behind `awaitPostAuthGates`. | -| Loading flags when the input has none, through the host's `featureFlags`. | Answering questions and task notices through `interaction`. | -| The token refresh right before the agent starts. | Rendering progress from `onProgress`. | -| The route, the `switchboard resolved` event and the run tags. | The readiness and Claude settings gates, before the call. | -| Calling `runAgent` and settling one outcome. | File watchers, such as the audit ledger and the task stream's event plan. | -| | Walking composed steps, and running programs that have no agent. | -| | Applying the outcome: exit code, auth error screen, rethrowing a crash, and `setup wizard finished`. | -| | Cancelling through `signal`. | - -`runProgram` never calls `getUI()` and never reads a session. It still touches -process-wide state: the analytics client, the gateway auth cache and the debug -log. A dead OAuth grant during the token refresh marks the login revoked, so a -later 401 can name the cause. - -A missing capability never hangs the run and never invents consent. A run that -needs approval fails without `awaitAiApproval`. A run with no `interaction` asks -no questions. A run with no `onProgress` still returns its settled run and final -data. - -## The store +| Direction | What `runProgram` takes or gives | What never crosses | +| --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------- | +| In | `ProgramInput`: the run and settings built from a `ProgramConfig`, the project directory, a login and flags. `ProgramOptions`: the login, questions, approval and gate waits, flags, progress and `signal`. | A `WizardSession`, a TUI store or `getUI()` | +| Out | Progress as copies, and one `ProgramRunOutcome`. | The `ProgramStore` itself, or live objects | -[`ProgramStore`](program-store.ts) holds one invocation's data and its run -ledger. `runProgram` creates one per call. The host never touches the store. It -sees copies through `onProgress` and the outcome. +## Where things live -- **Data.** `ProgramInvocationData`: the login, the framework context, the route - and the AI SDK stamp latch. `setAuthenticated`, `setBinding`, - `setAiSdkStampReported` and `setFrameworkContext` write it. - `setAiSdkStampReported` writes only the first time. `runProgram` never calls - `setFrameworkContext`, so `detection.frameworkContext` stays `{}`. -- **Copies.** Every write stores a `structuredClone` of its value. `readData()` - returns a fresh clone. After each write, the store sends a - `{ kind: 'program', data }` snapshot, itself a fresh clone. -- **Runs.** `beginRun(runId, observer)` returns an adapter with `onProgress` and - `finish`. `onProgress` forwards a clone of each agent event as - `{ kind: 'run', runId, event }`. `finish(result)` records the result, and - `settledRuns()` lists every finished run. `runProgram` begins one run per - call. -- **Diagnostics.** The store never waits for an observer. A throw or a rejected - promise becomes a `ProgramDiagnostic` with its source and message. A run event - after `finish` becomes a `progress after finish` diagnostic. The source is - `{ runId, eventKind }` for a run event and `{ eventKind: 'data' }` for a data - snapshot. The store keeps the 10 newest diagnostics. +| What | Where | +| --------------------------------------------- | ----------------------------------------------------------- | +| One program's config, steps and prompt | Its own folder, such as [`metrics`](metrics/index.ts) | +| The list of every program | [`program-registry.ts`](program-registry.ts) | +| The `ProgramConfig` and step types | [`program-step.ts`](program-step.ts) | +| Running one program from explicit inputs | [`run-program.ts`](run-program.ts) | +| What a program's `run` receives from a runner | [`runner-context.ts`](runner-context.ts) | +| Framework detection and project scoping | [`detection`](detection/index.ts) | +| Framework integrations | [`frameworks`](frameworks) and [`registry.ts`](registry.ts) | +| Commands that pick a program by skill | [`dispatch-family.ts`](dispatch-family.ts) | -## The adapter - -`runProgramAgent(programConfig, session, { composed? })` in -[`run-agent-legacy.ts`](run-agent-legacy.ts) is the host for every existing -runner. It is the only program code that reads the session and `getUI()` on the -agent's behalf. It resolves `run` against the session and a `ProgramRunHost`, -runs the readiness and Claude settings gates, and builds the input from the -session and the config. It answers the capabilities from the UI, maps run events -onto `getUI()`, and projects data snapshots onto the session. Then it applies -the outcome with `wizardAbort` and the terminal analytics event. -[How the legacy adapter builds `run`](../../docs/developer-interfaces.md#what-runprogram-is-for) -lists every field it maps. - -Three places call it: - -- **The TUI runner.** [`run-wizard.ts`](../lib/runners/run-wizard.ts) calls it - on the program's `run` screen. A composed step list calls it once per run - step. Each call gets a session scoped to the step's `targetDir` when the step - has one. -- **A composed run step.** `integrationRunStep` in - [`posthog-integration/index.ts`](posthog-integration/index.ts) runs the - integration with `composed: true`. Self-driving splices it into its own steps. -- **The non-interactive runner.** - [`run-non-interactive.ts`](../lib/runners/run-non-interactive.ts) calls it - after `ciPreRun`, for `--ci` and for headless runs. - -## Host capabilities - -A program's `run` and `ciPreRun` callbacks take their effects from a host -instead of calling `getUI()`. Both types live in -[`host-capabilities.ts`](host-capabilities.ts). - -- **`ProgramRunHost`** reaches `ProgramConfig.run(session, host)`. It has - `getFrameworkContext(key)`, `setFrameworkContext(key, value)` and - `warn(message)`. The legacy adapter builds it over `getUI()`, read at each - call. A run definition can keep the host and call it until the run returns. - The source maps program reads the picked project in its prompt builder and - writes `sourceMapsCompletedVariant` in `postRun`. -- **`ProgramCiHost`** reaches `ProgramConfig.ciPreRun(session, host)`. It has - `log.info(message)` and `log.warn(message)`. The non-interactive runner builds - it over `getUI().log`, read at each call. Project scoping logs its scan and - its fallbacks through it. - -Neither capability is a `runProgram` option. They belong to building the run -from a `ProgramConfig`, which happens before `runProgram`. The -[developer interfaces](../../docs/developer-interfaces.md#program-callbacks) -list which programs use each capability. - -## Current limits +Import runtime values from `@programs` and types from `@programs/types`. -- **One agent per call.** Composition belongs to the host. The TUI walks a - composed step list and calls the adapter once per run step. An imported run - step, such as the integration inside self-driving, runs with `composed: true`. -- **Programs with no agent.** A config with no `run`, such as `posthog-doctor`, - `mcp-add` or `slack`, never reaches `runProgram`. The TUI runs its steps. -- **Building `run` needs a session.** A dynamic `run`, the `ProgramRun` - completion hooks and `seedTasks` all take a `WizardSession`. Some programs - read session fields that a TUI step or `ciPreRun` fills, such as - `session.frameworkConfig`. -- **The module graph.** `@programs` loads the program registry, and with it - every program config, the TUI decks and `@ui`. `run-program.ts` also loads - `@ui` through `authenticate.ts`, although `runProgram` never calls `getUI()`. -- **Other `getUI()` calls.** `authenticate`, the audit ledger watcher, agentic - detection and some framework helpers still call `getUI()` directly. -- **No framework context in the data.** `data.detection.frameworkContext` stays - `{}`. Framework context reaches the program through its run host. -- **Gates and cleanup stay with the host.** `runProgram` runs no readiness or - Claude settings gate, starts no file watcher, and doesn't remove the skills a - run installed. -- **Other agent paths.** Agentic detection makes its own `runAgent` calls, each - with its own deadline. The MCP suggested-prompts screen streams through - `runMcpPromptViaSdk`, a separate SDK path. -- **No live control.** There is no live store handle and no step-by-step control - API. A host observes through `onProgress` and cancels through `signal`. +## Add a program + +Follow the +[adding-skill-program](../../.claude/skills/adding-skill-program/SKILL.md) +skill. A framework integration uses +[adding-framework-support](../../.claude/skills/adding-framework-support/SKILL.md) +instead. + +## Future work + +- **One folder per program.** Each program will move into its own folder with + everything it owns, so a team can own a full program through `CODEOWNERS`. +- **No legacy adapter.** The TUI and the headless runner reach `runProgram` + through `runProgramAgent` in [`run-agent-legacy.ts`](run-agent-legacy.ts). + This is a temporary adapter, and we will remove it in the full program. +- **No session in program configs.** A program's `run` and `ciPreRun` still take + a `WizardSession`. This is temporary, and we will remove it in the full + program. +- **No direct `getUI()` calls.** Some program code still calls `getUI()`. This + is temporary, and we will remove it in the full program. From c48f9024d3f528dc93cd1fe44ac591f6243744a0 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Thu, 24 Sep 2026 20:46:10 -0400 Subject: [PATCH 20/29] docs(programs): in and out as a short list Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- src/agent/README.md | 7 +++---- src/programs/README.md | 8 ++++---- 2 files changed, 7 insertions(+), 8 deletions(-) diff --git a/src/agent/README.md b/src/agent/README.md index 518609aaa..0bf33ac32 100644 --- a/src/agent/README.md +++ b/src/agent/README.md @@ -9,10 +9,9 @@ To call it from code, use `runAgent`. The ## What goes in and out -| Direction | What the agent takes or gives | What never crosses | -| --------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | -------------------------------------------------------------------------- | -| In | `RunConfig`: the program ID, the run definition, a resolved route, tools and a flag snapshot. `RunInput`: the project directory, the login, the project and user, the flags and the PostHog host. `onProgress`, `interaction` and `signal`. | A `WizardSession`, a TUI or headless store, `getUI()` or a `ProgramConfig` | -| Out | Progress events as copies, questions through `interaction`, and one `RunResult`. | Live objects, a thrown error or a process exit | +- **In.** A `RunConfig`, a `RunInput`, and callbacks for progress and questions. +- **Out.** One `RunResult` at the end. +- **Never in.** A session, a store or the UI. ## Where things live diff --git a/src/programs/README.md b/src/programs/README.md index d936abd1a..fc4f7bcef 100644 --- a/src/programs/README.md +++ b/src/programs/README.md @@ -9,10 +9,10 @@ To run a program from code, call `runProgram`. The ## What goes in and out of `runProgram` -| Direction | What `runProgram` takes or gives | What never crosses | -| --------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ------------------------------------------- | -| In | `ProgramInput`: the run and settings built from a `ProgramConfig`, the project directory, a login and flags. `ProgramOptions`: the login, questions, approval and gate waits, flags, progress and `signal`. | A `WizardSession`, a TUI store or `getUI()` | -| Out | Progress as copies, and one `ProgramRunOutcome`. | The `ProgramStore` itself, or live objects | +- **In.** A `ProgramInput`, and callbacks for login, approval, questions and + progress. +- **Out.** One `ProgramRunOutcome` at the end. +- **Never in.** A session, a store or the UI. ## Where things live From 256d31e831a595841f2580f64223440c10d2b899 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Thu, 24 Sep 2026 20:52:03 -0400 Subject: [PATCH 21/29] docs(agent): short runner README, bindings warning up front Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- src/agent/runner/README.md | 200 ++++--------------------------------- 1 file changed, 22 insertions(+), 178 deletions(-) diff --git a/src/agent/runner/README.md b/src/agent/runner/README.md index d62313350..03f572112 100644 --- a/src/agent/runner/README.md +++ b/src/agent/runner/README.md @@ -1,187 +1,31 @@ -# agent runner +# Agent runner -How an agent run is assembled. Everything under this directory is plumbing: the -pieces that decide _how_ a program runs (which query shape, which agent SDK, -which model) and the pieces that then actually run it. +> ⚠️ **The bindings table will be gone.** The switchboard still holds the +> program bindings table. By the end of this refactor it moves to programs, and +> the runner takes a resolved route only. -``` - ┌──────────────┐ ┌─────────────┐ ┌────────────────────────────┐ - │ │ │ │────▶│ sequence (query shape) │ - │ programs │────▶│ switchboard │ │ linear | orchestrator │ - │ │ │ │ └────────────────────────────┘ - │ integration │ │ binds each │ - │ audit │ │ program to │ ┌────────────────────────────┐ - │ migration │ │ a pair │────▶│ harness (SDK adapter) │ - │ ... │ │ │ │ anthropic | pi | ... │ - └──────────────┘ └─────────────┘ └────────────────────────────┘ -``` - -## Execution policy - -Use Pi for new work and prefer the orchestrator sequence. Linear execution is -retained for very simple tasks and legacy support. The Anthropic Agent SDK is a -supported legacy fallback, deprecated as the default, retained for major Pi -vulnerabilities or gaps in support for new Anthropic models. - -Existing `DEFAULT_BINDING` is Pi + linear. Explicit program bindings and flags -determine actual behavior. Both harnesses implement `run` and `runTask`. -Composed sub-runs are clamped to linear, and linear-only post-run/outro hooks do -not automatically transfer to an orchestrated flow. - -New models require Wizard capabilities **and** mint model/effort allowlists, -gateway provider/transport support, and compatibility with required Wizard and -security-triage prompt policies. Local model constants cannot bypass admission. -See the -[development guide](../../../.claude/skills/wizard-development/SKILL.md#execution-policy-and-model-admission) -for the coordinated change checklist. - -## The pieces - -Five layers, each with its own job. Nothing crosses layers unless it has to. - -**The entry point** (`index.ts`) is the front door: -`runAgent(config, input, {onProgress?, interaction?, signal?}) → RunResult`. It -takes resolved execution data and an invocation snapshot (`shared/types.ts`), -reports through `onProgress` and asks through `interaction` (`../progress.ts`), -and returns every ending as a result. It never renders, reads a session or -exits. The caller resolves credentials, consent, flags and the route first. For -programs, `runProgram` in `src/programs/run-program.ts` does that, and the -legacy adapter in `src/programs/run-agent-legacy.ts` runs the readiness and -settings gates and maps progress back onto `getUI()`. - -The entry point also owns two per-run switches. When `config.run` sets -`collectTranscript`, it creates the transcript tail and passes it to the -sequence. After the sequence returns, it adds the tail to -`snapshot.transcriptTail`. When `config.scanReport` is `'defer'`, it skips the -end-of-run scan report flush. See -[transcript tail](../README.md#transcript-tail) and -[scan report](../README.md#scan-report). - -**Prepare** (`shared/bootstrap.ts`) is the on-ramp inside the agent: logging -targets, the gateway mint and the scan-triage classifier. Whether the run turns -out to be linear or orchestrator, anthropic or pi, the setup is the same. - -**The switchboard** (`switchboard/`) is the router. Given a program id + the -fetched flags + any CLI overrides, it returns a `ProgramBinding`: which query -shape (sequence), which agent SDK (harness), which model. Two independent -middleware chains, one per axis, apply precedence rules (CLI > flag > program -config > default). This is the only layer that makes routing decisions. - -**Sequences** (`sequence/`) are LLM query shapes. Once the switchboard has -picked one, that sequence takes over the run and owns _how the LLM's work is -shaped_. See `sequence/README.md`. - -- **linear**: one long conversation with the model, start to finish. It builds - the prompt from `run.prompt` when set, else from the project prompt plus - `run.customPrompt`. It puts the transcript tail in front of the benchmark - middleware. -- **orchestrator**: many focused conversations coordinated by a task queue, each - with its own prompt, tools, and model. - -**Harnesses** (`harness/`) are SDK adapters. Sequences don't call Anthropic's or -pi.dev's SDKs directly. They go through a harness, which knows how to translate -a run request into that SDK's shape. All harnesses drive the PostHog LLM +The runner decides how an agent run happens, then runs it. It picks a sequence +and a harness for the program, and drives the model through the PostHog LLM gateway. -- **anthropic**: wraps Anthropic's official Claude Agent SDK. See - `harness/anthropic/README.md`. Its linear run feeds every SDK message to the - middleware, so the transcript tail and its `activity` events come from here. - It passes `run.requestRemark` to the stop hook. -- **pi**: wraps pi.dev's coding-agent library. See `harness/pi/README.md`. Its - linear run always asks for the remark and takes no middleware. - -The orchestrator asks neither harness for a remark on its tasks. - -## How they connect - -- Prepare mints the gateway token and builds triage for the resolved harness. -- The switchboard knows which sequences and harnesses exist (via its two - registries), but not what they do. -- A sequence knows how to shape a conversation, but delegates the actual model - call to a harness. -- A harness adapts its SDK, gateway transport, security hooks, and tool surface. - -Each layer is replaceable. +To call it, use `runAgent`. The +[developer interfaces](../../../docs/developer-interfaces.md#runagent) cover it. -## Ownership map - -```mermaid -%%{init: {"block": {"padding": 20}}}%% -block-beta - columns 11 - hostBand["Caller: legacy adapter and UI"]:11 - runProgramAgent["runProgramAgent"]:3 space:1 wizardAbort["wizardAbort"]:3 space:4 - space:11 - programsBand["Programs"]:11 - runProgram["runProgram"]:3 space:1 programOutcome["ProgramRunOutcome"]:3 space:4 - space:11 - runnerBand["Agent runner"]:11 - runAgent["runAgent"]:3 space:1 runResult["RunResult"]:3 space:4 - space:11 - sequenceBand["Orchestrator sequence"]:11 - runOrchestrator["runOrchestrator"]:3 space:1 sequenceResult["SequenceResult"]:3 space:4 - space:11 - drainQueue["drainQueue"]:3 space:5 runAbort["AbortController"]:3 - space:11 - harnessBand["Selected harness"]:11 - agentHarness["AgentHarness"]:3 space:1 agentResult["AgentResult"]:3 space:1 signal["TaskRunInputs.signal"]:3 - space:11 - sdkBand["External model SDK"]:11 - sdk["Selected SDK"]:3 space:8 - - runProgramAgent --> runProgram - runProgram --> runAgent - runAgent --> runOrchestrator - runOrchestrator --> drainQueue - drainQueue --> agentHarness - agentHarness --> sdk - agentHarness --> agentResult - agentResult --> sequenceResult - sequenceResult --> runResult - runResult --> programOutcome - programOutcome --> wizardAbort - drainQueue --> runAbort - runAbort --> signal - - classDef owner fill:#9ca3af1f,stroke:#9ca3af,stroke-width:1.5px - classDef changed fill:#3b82f626,stroke:#3b82f6,stroke-width:2px - class hostBand,programsBand,runnerBand,sequenceBand,harnessBand,sdkBand owner - class programOutcome,runResult,agentResult,runAbort changed -``` - -Calls descend on the left, results return through the middle, and cancellation -moves down the right. Blue marks the result contracts and run-scoped abort. On -the first fatal task result, `drainQueue` stops scheduling, cancels active work -and pending asks, joins siblings, then preserves that failure for the caller to -present. +## The pieces -## Flow +| Piece | What it does | Where | +| ----------- | -------------------------------------------------------------------------------- | ------------------------------------- | +| Entry point | `runAgent` takes a run config and input, and returns one result. | [`index.ts`](index.ts) | +| Prepare | Sets up logging, the gateway token and the scan classifier. | [`bootstrap.ts`](shared/bootstrap.ts) | +| Switchboard | Picks the sequence, harness and model for a program. | [`switchboard`](switchboard/index.ts) | +| Sequence | Shapes the work: one conversation (linear), or many from a queue (orchestrator). | [`sequence`](sequence/README.md) | +| Harness | Adapts one SDK: the Anthropic Agent SDK or Pi. | [`harness`](harness/types.ts) | -1. The caller resolves credentials, consent, flags and a - `ProgramBinding { sequence, harness, model }`, and analytics tags the run. - For programs, `runProgram` does this. -2. `runAgent(config, input, options)` creates the transcript tail when the run - collects one, then prepares (mint, triage). -3. Sequence takes over. It shapes the LLM's work into one conversation (linear) - or many (orchestrator), reporting through `onProgress`. -4. Harness drives each conversation through its SDK, using the bound model, on - the PostHog LLM gateway. -5. The scan report flushes, unless `config.scanReport` is `'defer'`. `runAgent` - returns a `RunResult` whose snapshot carries the transcript tail when the run - collected one. -6. The caller applies it. `runProgram` settles it into a `ProgramRunOutcome`. - The legacy adapter sends a decided failure to `wizardAbort` with the terminal - status its outcome names, rethrows a crash for the runner's own handling, and - sends the terminal success analytics for a non-composed success. The agent - sends no terminal analytics. +## Execution policy -## Who calls `runAgent` +Use Pi for new work, and prefer the orchestrator. Linear is for very simple +tasks. The Anthropic Agent SDK is a legacy fallback. `DEFAULT_BINDING` is Pi and +linear. -- `runProgram`, for every program's agent run. -- Agentic detection, in `src/programs/detection/agentic.ts`. Each attempt is a - linear Haiku run on the Anthropic harness with its own deadline signal. It - sets `run.prompt`, `collectTranscript: true`, `requestRemark: false` and - `scanReport: 'defer'`, and reads its report from `snapshot.transcriptTail`. -- The fault probe in `scripts/a3-fault-probe.no-jest.ts`, and any standalone - caller. See the - [developer interfaces](../../../docs/developer-interfaces.md#runagent). +A new model needs gateway admission as well as wizard support. See the +[development guide](../../../.claude/skills/wizard-development/SKILL.md#execution-policy-and-model-admission). From 29758ce84cd90c8cffcee7b60a4cdc3d3ab31b08 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Thu, 24 Sep 2026 20:53:42 -0400 Subject: [PATCH 22/29] docs(programs): future work as a warning up front Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- src/agent/README.md | 10 ++++------ src/programs/README.md | 21 ++++++++------------- 2 files changed, 12 insertions(+), 19 deletions(-) diff --git a/src/agent/README.md b/src/agent/README.md index 0bf33ac32..55281ce27 100644 --- a/src/agent/README.md +++ b/src/agent/README.md @@ -1,5 +1,9 @@ # Agent +> ⚠️ **The bindings table will be gone.** The program bindings table still lives +> in the agent. By the end of this refactor it moves to programs, and the agent +> takes a resolved route only. + The agent runs one AI pipeline against a project. It takes resolved data in, reports through progress events, asks through an answerer you pass, and returns a result. It never reads a session, a store or a UI. @@ -27,9 +31,3 @@ To call it from code, use `runAgent`. The Import runtime values from `@agent` and types from `@agent/types`. Nothing outside the agent imports deeper, and lint rejects it. - -## Future work - -- **Bindings move to programs.** The program bindings table still lives in the - agent. It will move to programs, and the agent will take a resolved route - only. diff --git a/src/programs/README.md b/src/programs/README.md index fc4f7bcef..13afb934d 100644 --- a/src/programs/README.md +++ b/src/programs/README.md @@ -1,5 +1,13 @@ # Programs +> ⚠️ **This changes by the end of this refactor.** +> +> - Each program moves into its own folder with everything it owns, so a team +> can own a full program through `CODEOWNERS`. +> - The temporary adapter `runProgramAgent` in +> [`run-agent-legacy.ts`](run-agent-legacy.ts) is removed. +> - Program configs stop taking a `WizardSession` and stop calling `getUI()`. + A program is one thing the wizard does for a user, such as adding PostHog or setting up error tracking. Each program is a `ProgramConfig`: the screens the TUI walks, the agent run it performs, and its settings. @@ -36,16 +44,3 @@ Follow the skill. A framework integration uses [adding-framework-support](../../.claude/skills/adding-framework-support/SKILL.md) instead. - -## Future work - -- **One folder per program.** Each program will move into its own folder with - everything it owns, so a team can own a full program through `CODEOWNERS`. -- **No legacy adapter.** The TUI and the headless runner reach `runProgram` - through `runProgramAgent` in [`run-agent-legacy.ts`](run-agent-legacy.ts). - This is a temporary adapter, and we will remove it in the full program. -- **No session in program configs.** A program's `run` and `ciPreRun` still take - a `WizardSession`. This is temporary, and we will remove it in the full - program. -- **No direct `getUI()` calls.** Some program code still calls `getUI()`. This - is temporary, and we will remove it in the full program. From ad4ac70368342a6fe1612f9977706603b76ba8f7 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Thu, 24 Sep 2026 20:54:16 -0400 Subject: [PATCH 23/29] docs: AGENTS.md carries only the path, binding and runner-context changes Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- AGENTS.md | 42 +++++++++++++++++++++--------------------- 1 file changed, 21 insertions(+), 21 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 44b4b660c..36426b4bd 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -32,13 +32,12 @@ Each domain has a dedicated boundary: `docs/runbooks/warlock-kill-switch.md`. ONLY USE THIS IF ABSOLUTELY NECESSARY. - **Agent** → `src/agent/`, imported only through `@agent` (values) and `@agent/types` (types); see [src/agent/README.md](src/agent/README.md) -- **Shared** → `src/shared/`, stateless library code with no upward imports; see - [src/shared/README.md](src/shared/README.md) -- **Programs** → configs, detection, framework registry, task stream and - `runProgram` in `src/programs/`. The runtime and type entries are `@programs` - and `@programs/types`. See [src/programs/README.md](src/programs/README.md) - and the [developer interfaces](docs/developer-interfaces.md) -- **TUI** → screens, primitives and content decks in `src/ui/tui/` +- **Shared** → `src/shared/`, stateless library code with no upward imports; + see [src/shared/README.md](src/shared/README.md) +- **Programs** → program configs and `runProgram` in `src/programs/`; see + [src/programs/README.md](src/programs/README.md) and the + [developer interfaces](docs/developer-interfaces.md) +- **TUI** → screen components and primitives in `src/ui/tui/` Adding a new concern means finding the narrowest existing surface, not adding logic to the runner. Keep changes local to the boundary that owns them. @@ -77,8 +76,9 @@ Agent SDK is a supported legacy fallback, deprecated as the default; retain it for major Pi vulnerabilities or gaps in support for new Anthropic models. This is the contribution policy, not a claim that every existing binding has -migrated: `DEFAULT_BINDING` is Pi + linear. Set new bindings explicitly and -check sequence-specific hooks before migrating existing flows. See +migrated: `DEFAULT_BINDING` is Pi + linear. Set new bindings +explicitly and check sequence-specific hooks before migrating existing flows. +See [execution policy and model admission](.claude/skills/wizard-development/SKILL.md#execution-policy-and-model-admission) for the gateway allowlists, required system prompt, and composition constraints. @@ -103,8 +103,8 @@ aliases. | Subcommand | What it audits | | ----------------------------- | ---------------------------------------------------- | -| `wizard audit events` | event capture quality + cost | -| `wizard audit all` | comprehensive audit across every area (**default**) | +| `wizard audit events` | event capture quality + cost | +| `wizard audit all` | comprehensive audit across every area (**default**) | | `wizard audit autocapture` | autocapture setup + cost | | `wizard audit feature-flags` | feature flag usage + cost | | `wizard audit identify` | `$identify` implementation | @@ -133,8 +133,9 @@ confuse it with the top-level `wizard skill` command. ([`src/commands/factories/native-command-factory.ts`](src/commands/factories/native-command-factory.ts)). - **Family commands** (e.g. `audit`) resolve subcommands at runtime against the `cliEntries` in `skill-menu.json`. Logic lives in - [`src/programs/dispatch-family.ts`](src/programs/dispatch-family.ts). Adding a - skill-backed subcommand is a **context-mill** release, not a wizard change. + [`src/programs/dispatch-family.ts`](src/programs/dispatch-family.ts). + Adding a skill-backed subcommand is a **context-mill** release, not a wizard + change. ### Commands vs. programs (don't confuse these) @@ -175,8 +176,8 @@ nonmutating lint checks, and scope formatting fixes to edited files. Do not add tests for prose, compiler-enforced shapes, or duplicated implementation. Keep new code comments to one line; put longer explanations in linked docs. -Local `--ci`, smoke-test, and full headless runs require two separate secrets: a -PostHog personal API key and an already-issued gateway token supplied through +Local `--ci`, smoke-test, and full headless runs require two separate secrets: +a PostHog personal API key and an already-issued gateway token supplied through `WIZARD_CI_GATEWAY_TOKEN_FILE`, plus the target project ID. Follow the [credential setup](docs/local-dev.md#credentials-for-local-ci-and-headless-runs). @@ -202,13 +203,12 @@ wizard run points. Full catalog: [`docs/local-dev.md`](docs/local-dev.md). types so they satisfy `Record`. - All UI calls go through `getUI()` (returns `WizardUI` interface). Never import the store directly from business logic. A program's `run` and `ciPreRun` - callbacks use the runner context they receive (`RunnerContext`, - `CiRunnerContext`), not `getUI()`. -- Shared helpers never call `getUI()`; they take a sink or return data. - `debug()` reaches the UI through the sink `src/ui/index.ts` installs. + use the runner context they receive, not `getUI()`. +- Shared helpers never call `getUI()`; they take a sink or return data. `debug()` + reaches the UI through the sink `src/ui/index.ts` installs. - Outside `src/agent`, import the agent through `@agent` or `@agent/types`. Add - to those entry modules rather than deep-importing; lint and `pnpm test:arch` - reject `@agent/*` paths elsewhere. + to those entry modules rather than deep-importing; lint and + `pnpm test:arch` reject `@agent/*` paths elsewhere. - Session mutations go through explicit store setters that call `emitChange()`. Never mutate `session` directly — nanostore holds a shallow copy. - The router resolves the active screen from session state. No imperative From 09d91b92ac5706ea26c2db9df517ffbbe3dced15 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Fri, 25 Sep 2026 11:03:48 -0400 Subject: [PATCH 24/29] docs: fix line anchors, the cancellation sentence, DEFAULT_BINDING and the TUI line Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- .claude/skills/wizard-development/SKILL.md | 4 +-- AGENTS.md | 2 +- docs/developer-interfaces.md | 31 +++++++++++----------- 3 files changed, 19 insertions(+), 18 deletions(-) diff --git a/.claude/skills/wizard-development/SKILL.md b/.claude/skills/wizard-development/SKILL.md index 5fee5c7a5..3fa776651 100644 --- a/.claude/skills/wizard-development/SKILL.md +++ b/.claude/skills/wizard-development/SKILL.md @@ -43,8 +43,8 @@ infrastructure should consume those boundaries. new Anthropic models. Existing routing has not all migrated: -[DEFAULT_BINDING](../../../src/agent/runner/switchboard/index.ts) still -selects Anthropic + linear, with per-program and flag overrides. Set new +[DEFAULT_BINDING](../../../src/agent/runner/switchboard/index.ts) selects +Pi + linear, with per-program and flag overrides. Set new bindings explicitly. Migrating an existing program requires checking its flow, tasks, and lifecycle hooks; changing the default constant alone is insufficient. Both harnesses implement `run` and `runTask`. diff --git a/AGENTS.md b/AGENTS.md index 36426b4bd..0ecf21c7c 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -37,7 +37,7 @@ Each domain has a dedicated boundary: - **Programs** → program configs and `runProgram` in `src/programs/`; see [src/programs/README.md](src/programs/README.md) and the [developer interfaces](docs/developer-interfaces.md) -- **TUI** → screen components and primitives in `src/ui/tui/` +- **TUI** → screens, primitives and content decks in `src/ui/tui/` Adding a new concern means finding the narrowest existing surface, not adding logic to the runner. Keep changes local to the boundary that owns them. diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index ee0ff626c..6ac42f608 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -70,14 +70,14 @@ export const signature: ( ### Field definitions -| Field | Type | What it's for | -| ----------------------------------------------------------- | ------------------- | -------------------------------------------------------------------------------- | -| `programId` | `string` | Which program runs. It names analytics, the route and the gateway spend. | -| [`input`](../src/programs/run-program.ts#L66) | `ProgramInput` | What to run and where. `installDir` and `run` are required. | -| [`input.program`](../src/programs/run-program.ts#L56) | `ProgramSettings` | The program's settings from its `ProgramConfig`. | -| [`options`](../src/programs/run-program.ts#L89) | `ProgramOptions` | The login, questions, approval and gate waits, flags, progress and cancellation. | -| [`options.onProgress`](../src/programs/program-store.ts#L9) | `ProgramProgress` | Agent events and program data snapshots. Never awaited. | -| [Outcome](../src/programs/run-program.ts#L106) | `ProgramRunOutcome` | How the run ended, the agent's result, the final data and the report path. | +| Field | Type | What it's for | +| ------------------------------------------------------------ | ------------------- | -------------------------------------------------------------------------------- | +| `programId` | `string` | Which program runs. It names analytics, the route and the gateway spend. | +| [`input`](../src/programs/run-program.ts#L66) | `ProgramInput` | What to run and where. `installDir` and `run` are required. | +| [`input.program`](../src/programs/run-program.ts#L56) | `ProgramSettings` | The program's settings from its `ProgramConfig`. | +| [`options`](../src/programs/run-program.ts#L89) | `ProgramOptions` | The login, questions, approval and gate waits, flags, progress and cancellation. | +| [`options.onProgress`](../src/programs/program-store.ts#L21) | `ProgramProgress` | Agent events and program data snapshots. Never awaited. | +| [Outcome](../src/programs/run-program.ts#L106) | `ProgramRunOutcome` | How the run ended, the agent's result, the final data and the report path. | ### Example @@ -164,15 +164,16 @@ export async function runMetrics( ### Cancellation -Pass a `signal` to cancel the run. Every wait receives it, and a wait must -settle when it aborts. A cancelled run resolves to `aborted`, not a rejection. +Pass a `signal` to cancel the run. The credentials, approval and gate waits +receive it, and each must settle when it aborts. A cancelled run resolves to +`aborted`, not a rejection. ### Failures Most endings resolve to an outcome instead of throwing. Check `outcome` and read `failure`. The promise rejects only when the call itself can't run, such as an input field that can't be copied. The cases are in -[`run-program.ts`](../src/programs/run-program.ts#L160). +[`run-program.ts`](../src/programs/run-program.ts#L163). ### Program callbacks @@ -226,11 +227,11 @@ export const signature: ( | Field | Type | What it's for | | ------------------------------------------------------- | -------------------- | ------------------------------------------------------------------------------- | -| [`config`](../src/agent/runner/shared/types.ts#L153) | `RunConfig` | What the agent runs, with its route and tools. | -| [`config.run`](../src/agent/runner/shared/types.ts#L49) | `AgentRunDefinition` | The prompt and run options, such as `collectTranscript`. | -| [`input`](../src/agent/runner/shared/types.ts#L207) | `RunInput` | Where and as whom: the project, the login and the flags. | +| [`config`](../src/agent/runner/shared/types.ts#L151) | `RunConfig` | What the agent runs, with its route and tools. | +| [`config.run`](../src/agent/runner/shared/types.ts#L48) | `AgentRunDefinition` | The prompt and run options, such as `collectTranscript`. | +| [`input`](../src/agent/runner/shared/types.ts#L205) | `RunInput` | Where and as whom: the project, the login and the flags. | | `options` | | `onProgress` for agent events, `interaction` for questions, `signal` to cancel. | -| [Result](../src/agent/runner/shared/types.ts#L309) | `RunResult` | How the run ended, with a snapshot of its tasks and transcript. | +| [Result](../src/agent/runner/shared/types.ts#L307) | `RunResult` | How the run ended, with a snapshot of its tasks and transcript. | ### Callers From 3b9f54074f9b5be813d61cd991684899cf5c1db8 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Fri, 25 Sep 2026 11:23:23 -0400 Subject: [PATCH 25/29] docs(programs): say which flags skip the AI approval Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- docs/developer-interfaces.md | 3 +++ 1 file changed, 3 insertions(+) diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index 6ac42f608..8ee8a5f3a 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -79,6 +79,9 @@ export const signature: ( | [`options.onProgress`](../src/programs/program-store.ts#L21) | `ProgramProgress` | Agent events and program data snapshots. Never awaited. | | [Outcome](../src/programs/run-program.ts#L106) | `ProgramRunOutcome` | How the run ended, the agent's result, the final data and the report path. | +`flags.ci` and `flags.signup` skip the AI-processing approval. Set them only +when consent is already settled. + ### Example This runs the `metrics` program with a login you already hold. It asks your own From 574237b1fa50a2ad95fba3f9580248a1e7d6b5ad Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Fri, 25 Sep 2026 13:06:32 -0400 Subject: [PATCH 26/29] docs(programs): fix the runAgent tools example, warn about tokens in data - The runAgent example removes Write, Edit and Bash instead of adding Read and Glob, and the field table says allowedTools adds to the base tools and disallowedTools removes them. - The runProgram field table says data and program snapshots hold tokens. - The registry link points at frameworks/registry.ts, and every line anchor points at its declaration again. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- docs/developer-interfaces.md | 36 +++++++++++++++++++++--------------- src/programs/README.md | 20 ++++++++++---------- 2 files changed, 31 insertions(+), 25 deletions(-) diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index 8ee8a5f3a..841893af9 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -73,11 +73,14 @@ export const signature: ( | Field | Type | What it's for | | ------------------------------------------------------------ | ------------------- | -------------------------------------------------------------------------------- | | `programId` | `string` | Which program runs. It names analytics, the route and the gateway spend. | -| [`input`](../src/programs/run-program.ts#L66) | `ProgramInput` | What to run and where. `installDir` and `run` are required. | -| [`input.program`](../src/programs/run-program.ts#L56) | `ProgramSettings` | The program's settings from its `ProgramConfig`. | -| [`options`](../src/programs/run-program.ts#L89) | `ProgramOptions` | The login, questions, approval and gate waits, flags, progress and cancellation. | -| [`options.onProgress`](../src/programs/program-store.ts#L21) | `ProgramProgress` | Agent events and program data snapshots. Never awaited. | -| [Outcome](../src/programs/run-program.ts#L106) | `ProgramRunOutcome` | How the run ended, the agent's result, the final data and the report path. | +| [`input`](../src/programs/run-program.ts#L67) | `ProgramInput` | What to run and where. `installDir` and `run` are required. | +| [`input.program`](../src/programs/run-program.ts#L57) | `ProgramSettings` | The program's settings from its `ProgramConfig`. | +| [`options`](../src/programs/run-program.ts#L90) | `ProgramOptions` | The login, questions, approval and gate waits, flags, progress and cancellation. | +| [`options.onProgress`](../src/programs/program-store.ts#L17) | `ProgramProgress` | Agent events and program data snapshots. Never awaited. | +| [Outcome](../src/programs/run-program.ts#L107) | `ProgramRunOutcome` | How the run ended, the agent's result, the final data and the report path. | + +The outcome's `data` and the `kind: 'program'` progress snapshots hold PostHog +tokens, including the refresh token. Don't log or serialize them. `flags.ci` and `flags.signup` skip the AI-processing approval. Set them only when consent is already settled. @@ -176,7 +179,7 @@ receive it, and each must settle when it aborts. A cancelled run resolves to Most endings resolve to an outcome instead of throwing. Check `outcome` and read `failure`. The promise rejects only when the call itself can't run, such as an input field that can't be copied. The cases are in -[`run-program.ts`](../src/programs/run-program.ts#L163). +[`run-program.ts`](../src/programs/run-program.ts#L164). ### Program callbacks @@ -228,13 +231,15 @@ export const signature: ( ### Field definitions -| Field | Type | What it's for | -| ------------------------------------------------------- | -------------------- | ------------------------------------------------------------------------------- | -| [`config`](../src/agent/runner/shared/types.ts#L151) | `RunConfig` | What the agent runs, with its route and tools. | -| [`config.run`](../src/agent/runner/shared/types.ts#L48) | `AgentRunDefinition` | The prompt and run options, such as `collectTranscript`. | -| [`input`](../src/agent/runner/shared/types.ts#L205) | `RunInput` | Where and as whom: the project, the login and the flags. | -| `options` | | `onProgress` for agent events, `interaction` for questions, `signal` to cancel. | -| [Result](../src/agent/runner/shared/types.ts#L307) | `RunResult` | How the run ended, with a snapshot of its tasks and transcript. | +| Field | Type | What it's for | +| -------------------------------------------------------------------- | -------------------- | ------------------------------------------------------------------------------- | +| [`config`](../src/agent/runner/shared/types.ts#L151) | `RunConfig` | What the agent runs, with its route and tools. | +| [`config.run`](../src/agent/runner/shared/types.ts#L48) | `AgentRunDefinition` | The prompt and run options, such as `collectTranscript`. | +| [`config.allowedTools`](../src/agent/runner/shared/types.ts#L174) | `readonly string[]` | Tools added to the base tools. | +| [`config.disallowedTools`](../src/agent/runner/shared/types.ts#L176) | `readonly string[]` | Tools removed from the base tools. | +| [`input`](../src/agent/runner/shared/types.ts#L205) | `RunInput` | Where and as whom: the project, the login and the flags. | +| `options` | | `onProgress` for agent events, `interaction` for questions, `signal` to cancel. | +| [Result](../src/agent/runner/shared/types.ts#L307) | `RunResult` | How the run ended, with a snapshot of its tasks and transcript. | ### Callers @@ -246,7 +251,8 @@ export const signature: ( ### Example -This lists a project's files with a read-only agent and returns what it said: +This lists a project's files with an agent that has no Write, Edit or Bash, and +returns what it said: ```ts import { runAgent, RunOutcome } from '@agent'; @@ -288,7 +294,7 @@ export async function listProjectFiles( wizardFlags: {}, wizardFlagPayloads: {}, wizardMetadata: {}, - allowedTools: ['Read', 'Glob'], // read-only + disallowedTools: ['Write', 'Edit', 'Bash'], // no Write, Edit or Bash }; const result = await runAgent(config, input, { diff --git a/src/programs/README.md b/src/programs/README.md index 13afb934d..6a297ab32 100644 --- a/src/programs/README.md +++ b/src/programs/README.md @@ -24,16 +24,16 @@ To run a program from code, call `runProgram`. The ## Where things live -| What | Where | -| --------------------------------------------- | ----------------------------------------------------------- | -| One program's config, steps and prompt | Its own folder, such as [`metrics`](metrics/index.ts) | -| The list of every program | [`program-registry.ts`](program-registry.ts) | -| The `ProgramConfig` and step types | [`program-step.ts`](program-step.ts) | -| Running one program from explicit inputs | [`run-program.ts`](run-program.ts) | -| What a program's `run` receives from a runner | [`runner-context.ts`](runner-context.ts) | -| Framework detection and project scoping | [`detection`](detection/index.ts) | -| Framework integrations | [`frameworks`](frameworks) and [`registry.ts`](registry.ts) | -| Commands that pick a program by skill | [`dispatch-family.ts`](dispatch-family.ts) | +| What | Where | +| --------------------------------------------- | --------------------------------------------------------------------------------- | +| One program's config, steps and prompt | Its own folder, such as [`metrics`](metrics/index.ts) | +| The list of every program | [`program-registry.ts`](program-registry.ts) | +| The `ProgramConfig` and step types | [`program-step.ts`](program-step.ts) | +| Running one program from explicit inputs | [`run-program.ts`](run-program.ts) | +| What a program's `run` receives from a runner | [`runner-context.ts`](runner-context.ts) | +| Framework detection and project scoping | [`detection`](detection/index.ts) | +| Framework integrations | [`frameworks`](frameworks) and [`frameworks/registry.ts`](frameworks/registry.ts) | +| Commands that pick a program by skill | [`dispatch-family.ts`](dispatch-family.ts) | Import runtime values from `@programs` and types from `@programs/types`. From d1ed52c442955f52873c9d20eb22dd0bdcb0bcd8 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Fri, 25 Sep 2026 14:50:27 -0400 Subject: [PATCH 27/29] docs(programs): quack examples for runProgram and runAgent Two runnable scripts replace the inline examples. tsconfig includes docs/examples, so typecheck keeps them in step with the code. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- docs/developer-interfaces.md | 156 ++++------------------------- docs/examples/run-agent-quack.ts | 100 ++++++++++++++++++ docs/examples/run-program-quack.ts | 90 +++++++++++++++++ tsconfig.json | 1 + 4 files changed, 209 insertions(+), 138 deletions(-) create mode 100644 docs/examples/run-agent-quack.ts create mode 100644 docs/examples/run-program-quack.ts diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index 841893af9..6856fc639 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -85,89 +85,21 @@ tokens, including the refresh token. Don't log or serialize them. `flags.ci` and `flags.signup` skip the AI-processing approval. Set them only when consent is already settled. -### Example +### Do a quack -This runs the `metrics` program with a login you already hold. It asks your own -consent flow for approval, and logs what the program is doing: +[`run-program-quack.ts`](examples/run-program-quack.ts) is the smallest +`runProgram` call. It logs in through the wizard's browser OAuth, runs one +prompt that replies `quack`, logs status lines and prints the outcome. Each step +has a comment. -```ts -import { RunOutcome } from '@agent'; -import { getProgramConfig, runProgram } from '@programs'; -import type { - ProgramInput, - ProgramOptions, - ProgramProgress, -} from '@programs/types'; +Run it from the repository root against the [local stack](local-dev.md): -type Login = NonNullable; -type AskForApproval = NonNullable; - -// Log what the program is doing. runProgram never waits for this. -function logProgress(progress: ProgramProgress): void { - // Program data changed, for example the route resolved. - if (progress.kind === 'program') { - const route = progress.data.binding; - if (route) console.log(`route: ${route.sequence} on ${route.harness}`); - return; - } - // Otherwise it's one agent event. - const { event } = progress; - switch (event.kind) { - case 'status': - console.log(event.message); - break; - case 'tasks': { - const done = event.tasks.filter((task) => task.status === 'completed'); - console.log(`tasks: ${done.length}/${event.tasks.length} done`); - break; - } - case 'url': - console.log(`${event.which}: ${event.url}`); - break; - } -} - -export async function runMetrics( - installDir: string, - login: Login, - askForApproval: AskForApproval, - signal?: AbortSignal, -): Promise { - // Build the run from the program's own config. - const config = getProgramConfig('metrics'); - if (!config.run || typeof config.run === 'function') { - throw new Error('metrics has a static run definition'); - } - - const result = await runProgram( - config.id, - { - installDir, // the project to change - run: config.run, - credentials: login, // skip the login step - program: { - requiresAi: config.requiresAi, - agentFlow: config.agentFlow, - allowedTools: config.allowedTools, - disallowedTools: config.disallowedTools, - excludedTaskTypes: config.excludedTaskTypes, - }, - }, - { - awaitAiApproval: askForApproval, // your consent flow - onProgress: logProgress, - signal, // abort to cancel; the run resolves to aborted - }, - ); - - // Endings resolve to an outcome. Check it instead of catching. - if (result.outcome !== RunOutcome.Success) { - throw result.failure?.error ?? new Error(result.failure?.message); - } - return result.artifacts.reportFile; // where the agent wrote its report -} +```bash +QUACK_INSTALL_DIR= npx tsx --tsconfig tsconfig.json docs/examples/run-program-quack.ts ``` +It prints `reply: quack` and `outcome: success`. + ### Cancellation Pass a `signal` to cancel the run. The credentials, approval and gate waits @@ -249,67 +181,15 @@ export const signature: ( | `detectProjectsWithAgent` in [`agentic.ts`](../src/programs/detection/agentic.ts) | The agentic project scan, one call per attempt. | | [`a3-fault-probe.no-jest.ts`](../scripts/a3-fault-probe.no-jest.ts) | A fault probe against a local gateway. | -### Example +### Do a quack -This lists a project's files with an agent that has no Write, Edit or Bash, and -returns what it said: +[`run-agent-quack.ts`](examples/run-agent-quack.ts) is the smallest `runAgent` +call. It logs in the same way, builds a `RunConfig` with one prompt and no +Write, Edit or Bash, and prints the transcript tail and the outcome. Each step +has a comment. -```ts -import { runAgent, RunOutcome } from '@agent'; -import type { RunConfig, RunInput } from '@agent/types'; -import { - getSkillsBaseUrl, - Harness, - HAIKU_MODEL, - Sequence, -} from '@shared/constants'; - -export async function listProjectFiles( - programId: string, - input: RunInput, // the project, the login and the flags - signal?: AbortSignal, -): Promise { - const config: RunConfig = { - programId, // attributes the gateway spend - run: { - integrationLabel: 'list-files', - prompt: () => 'List the files in the working directory. Change nothing.', - collectTranscript: true, // keep the agent's output to return below - requestRemark: false, // no closing remark - spinnerMessage: 'Listing files...', - successMessage: 'Listed files', - estimatedDurationMinutes: 1, - reportFile: '', - docsUrl: 'https://posthog.com/docs', - }, - composed: true, // a sub-run: the caller owns the outro - // You pick the route yourself. runAgent doesn't resolve one. - binding: { - sequence: Sequence.linear, - harness: Harness.anthropic, - model: HAIKU_MODEL, - }, - switchboard: { program: programId, composed: true, flags: {} }, - skillsBaseUrl: getSkillsBaseUrl(), - wizardFlags: {}, - wizardFlagPayloads: {}, - wizardMetadata: {}, - disallowedTools: ['Write', 'Edit', 'Bash'], // no Write, Edit or Bash - }; - - const result = await runAgent(config, input, { - signal, - // Each step the agent takes, as one line. - onProgress: (event) => { - if (event.kind === 'activity') console.log(event.line); - }, - }); - if (result.outcome !== RunOutcome.Success) { - throw result.failure.error ?? new Error(result.failure.message); - } - return result.snapshot.transcriptTail ?? ''; -} +```bash +QUACK_INSTALL_DIR= npx tsx --tsconfig tsconfig.json docs/examples/run-agent-quack.ts ``` -`collectTranscript` and `requestRemark` take effect on the linear sequence with -the Anthropic harness. +It prints `transcriptTail: quack` and `outcome: success`. diff --git a/docs/examples/run-agent-quack.ts b/docs/examples/run-agent-quack.ts new file mode 100644 index 000000000..09dbde209 --- /dev/null +++ b/docs/examples/run-agent-quack.ts @@ -0,0 +1,100 @@ +/* eslint-disable no-console -- the example prints to the terminal */ +// Run only the agent with runAgent, against a local PostHog stack. +// +// npx tsx --tsconfig tsconfig.json docs/examples/run-agent-quack.ts +// +// Needs local PostHog on :8010 (with its ai-gateway) and context-mill on :8765. +// QUACK_INSTALL_DIR sets the project the agent runs in (default: the current directory). +import { runAgent, RunOutcome } from '@agent'; +import type { RunConfig, RunInput } from '@agent/types'; +import { + Harness, + HAIKU_MODEL, + Sequence, + getSkillsBaseUrl, +} from '@shared/constants'; +import { initLocalDev, POSTHOG_LOCAL_URL } from '@shared/local-dev'; +import { getOrAskForProjectData } from '@utils/setup-utils'; + +// Point PostHog, skills and MCP at the local stack, like --local-posthog --local-context-mill --local-mcp. +initLocalDev({ localPosthog: true, localContextMill: true, localMcp: true }); + +// Log in with the wizard's own browser OAuth flow. The token stays in memory. +const programId = 'posthog-integration'; // a program the local gateway admits +const login = await getOrAskForProjectData({ + signup: false, + ci: false, + baseUrl: POSTHOG_LOCAL_URL, + localMcp: true, + programId, +}); + +// What the agent runs: one prompt, a small model, no Write, Edit or Bash. +const config: RunConfig = { + programId, // pins the gateway spend + run: { + integrationLabel: 'quack', + prompt: () => 'Reply with the single word quack. Use no tools.', + collectTranscript: true, // keep the agent's output for snapshot.transcriptTail + requestRemark: false, // no closing remark + spinnerMessage: 'Quacking...', + successMessage: 'Quacked', + estimatedDurationMinutes: 1, + reportFile: '', + docsUrl: 'https://posthog.com/docs', + }, + composed: true, // a sub-run: no terminal outro + // runAgent doesn't resolve a route. Linear on the Anthropic harness keeps the transcript. + binding: { + sequence: Sequence.linear, + harness: Harness.anthropic, + model: HAIKU_MODEL, + }, + switchboard: { program: programId, composed: true, flags: {} }, + skillsBaseUrl: getSkillsBaseUrl(), + wizardFlags: {}, + wizardFlagPayloads: {}, + wizardMetadata: {}, + disallowedTools: ['Write', 'Edit', 'Bash'], +}; + +// Where and as whom: the project, the login and the flags. +const input: RunInput = { + installDir: process.env.QUACK_INSTALL_DIR ?? process.cwd(), + credentials: { + accessToken: login.accessToken, + refreshToken: login.refreshToken, + expiresAt: login.expiresAt, + projectApiKey: login.projectApiKey, + host: login.host, + projectId: login.projectId, + missingScopes: login.missingScopes, + }, + project: login.project, + apiUser: login.user, + flags: { + ci: false, + signup: false, + debug: false, + e2eAsk: false, + localMcp: true, + captureAio: false, + benchmark: false, + yaraReport: false, + }, + host: { baseUrl: POSTHOG_LOCAL_URL }, +}; + +// Run it. Each step the agent takes arrives as one activity line. +const result = await runAgent(config, input, { + onProgress: (event) => { + if (event.kind === 'activity') console.log(`activity: ${event.line}`); + }, +}); + +// Every ending resolves to a result. Print the reply and the outcome. +console.log(`transcriptTail: ${result.snapshot.transcriptTail ?? ''}`); +console.log(`outcome: ${result.outcome}`); +if (result.outcome !== RunOutcome.Success) + console.log(`failure: ${result.failure.message}`); +process.exit(result.outcome === RunOutcome.Success ? 0 : 1); diff --git a/docs/examples/run-program-quack.ts b/docs/examples/run-program-quack.ts new file mode 100644 index 000000000..c498d78cc --- /dev/null +++ b/docs/examples/run-program-quack.ts @@ -0,0 +1,90 @@ +/* eslint-disable no-console -- the example prints to the terminal */ +// Run one program with runProgram, against a local PostHog stack. +// +// npx tsx --tsconfig tsconfig.json docs/examples/run-program-quack.ts +// +// Needs local PostHog on :8010 (with its ai-gateway) and context-mill on :8765. +// QUACK_INSTALL_DIR sets the project the agent runs in (default: the current directory). +import { RunOutcome } from '@agent'; +import { runProgram } from '@programs'; +import type { ProgramProgress } from '@programs/types'; +import { Harness, HAIKU_MODEL, Sequence } from '@shared/constants'; +import { initLocalDev, POSTHOG_LOCAL_URL } from '@shared/local-dev'; +import { getOrAskForProjectData } from '@utils/setup-utils'; + +// Point PostHog, skills and MCP at the local stack, like --local-posthog --local-context-mill --local-mcp. +initLocalDev({ localPosthog: true, localContextMill: true, localMcp: true }); + +// Log in with the wizard's own browser OAuth flow. The token stays in memory. +const programId = 'posthog-integration'; // a program the local gateway admits +const login = await getOrAskForProjectData({ + signup: false, + ci: false, + baseUrl: POSTHOG_LOCAL_URL, + localMcp: true, + programId, +}); + +// Log status lines as the program reports them. runProgram never waits for this. +function logProgress(progress: ProgramProgress): void { + if (progress.kind === 'program') return; // a data snapshot; it holds tokens, don't log it + const { event } = progress; + if (event.kind === 'status') console.log(`status: ${event.message}`); + if (event.kind === 'lifecycle') console.log(`lifecycle: ${event.phase}`); +} + +const result = await runProgram( + programId, + { + installDir: process.env.QUACK_INSTALL_DIR ?? process.cwd(), + // A caller-built run in place of the program's own: one prompt, and keep the reply. + run: { + integrationLabel: 'quack', + prompt: () => 'Reply with the single word quack. Use no tools.', + collectTranscript: true, // keep the agent's output for the reply below + requestRemark: false, // no closing remark + spinnerMessage: 'Quacking...', + successMessage: 'Quacked', + estimatedDurationMinutes: 1, + reportFile: '', + docsUrl: 'https://posthog.com/docs', + }, + program: { disallowedTools: ['Write', 'Edit', 'Bash'] }, + // The login from above, so runProgram skips its own login step. + credentials: { + posthog: { + accessToken: login.accessToken, + refreshToken: login.refreshToken, + expiresAt: login.expiresAt, + projectApiKey: login.projectApiKey, + host: login.host, + projectId: login.projectId, + missingScopes: login.missingScopes, + }, + project: login.project, + apiUser: login.user, + }, + composed: true, // a sub-run: no terminal outro + // A small model on the linear Anthropic route, where the transcript is kept. + overrides: { + sequence: Sequence.linear, + harness: Harness.anthropic, + model: HAIKU_MODEL, + }, + flags: { localMcp: true }, + host: { baseUrl: POSTHOG_LOCAL_URL }, + wizardFlags: {}, // no flag snapshot to load + }, + { + // You approved AI data processing for this local test user. + awaitAiApproval: () => Promise.resolve(true), + onProgress: logProgress, + }, +); + +// Endings resolve to an outcome. The agent's reply is in its settled run's transcript. +const reply = result.settledRuns[0]?.result.snapshot.transcriptTail ?? ''; +console.log(`reply: ${reply}`); +console.log(`outcome: ${result.outcome}`); +if (result.failure) console.log(`failure: ${result.failure.message}`); +process.exit(result.outcome === RunOutcome.Success ? 0 : 1); diff --git a/tsconfig.json b/tsconfig.json index b3b6d448f..9947307f6 100644 --- a/tsconfig.json +++ b/tsconfig.json @@ -19,6 +19,7 @@ "src/**/*", "test/**/*", "e2e-harness/**/*", + "docs/examples/**/*", "types/**/*" ], "exclude": [ From 3844eb8e165403416d9b8c6a4e42301ad054f997 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Fri, 25 Sep 2026 15:32:25 -0400 Subject: [PATCH 28/29] docs(programs): quack examples log in with keys Both quack examples log in with a personal API key the way --ci does and read the gateway token from WIZARD_CI_GATEWAY_TOKEN_FILE, so a run needs no browser and spends no mint. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- docs/developer-interfaces.md | 20 +++++++++++++++----- docs/examples/run-agent-quack.ts | 16 +++++++++++++--- docs/examples/run-program-quack.ts | 12 +++++++++--- 3 files changed, 37 insertions(+), 11 deletions(-) diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index 6856fc639..07aa07d73 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -88,14 +88,22 @@ when consent is already settled. ### Do a quack [`run-program-quack.ts`](examples/run-program-quack.ts) is the smallest -`runProgram` call. It logs in through the wizard's browser OAuth, runs one +`runProgram` call. It logs in with keys, the same way `--ci` does, runs one prompt that replies `quack`, logs status lines and prints the outcome. Each step has a comment. -Run it from the repository root against the [local stack](local-dev.md): +It takes the two keys from +[local credentials](local-dev.md#credentials-for-local-ci-and-headless-runs). On +the [local stack](local-dev.md), the personal API key is the one PostHog's +`setup_local_api_key` command creates, and the gateway token is the `phs_` key +PostHog's `setup-gateway-e2e` provisions. Run it from the repository root: ```bash -QUACK_INSTALL_DIR= npx tsx --tsconfig tsconfig.json docs/examples/run-program-quack.ts +export POSTHOG_PERSONAL_API_KEY= +export WIZARD_CI_GATEWAY_TOKEN_FILE= +export WIZARD_CI_GATEWAY_URL=http://localhost:8080 +export QUACK_INSTALL_DIR= +npx tsx --tsconfig tsconfig.json docs/examples/run-program-quack.ts ``` It prints `reply: quack` and `outcome: success`. @@ -184,12 +192,14 @@ export const signature: ( ### Do a quack [`run-agent-quack.ts`](examples/run-agent-quack.ts) is the smallest `runAgent` -call. It logs in the same way, builds a `RunConfig` with one prompt and no +call. It logs in with the same keys, builds a `RunConfig` with one prompt and no Write, Edit or Bash, and prints the transcript tail and the outcome. Each step has a comment. +With the same variables set: + ```bash -QUACK_INSTALL_DIR= npx tsx --tsconfig tsconfig.json docs/examples/run-agent-quack.ts +npx tsx --tsconfig tsconfig.json docs/examples/run-agent-quack.ts ``` It prints `transcriptTail: quack` and `outcome: success`. diff --git a/docs/examples/run-agent-quack.ts b/docs/examples/run-agent-quack.ts index 09dbde209..4d1f73aff 100644 --- a/docs/examples/run-agent-quack.ts +++ b/docs/examples/run-agent-quack.ts @@ -4,8 +4,13 @@ // npx tsx --tsconfig tsconfig.json docs/examples/run-agent-quack.ts // // Needs local PostHog on :8010 (with its ai-gateway) and context-mill on :8765. +// POSTHOG_PERSONAL_API_KEY logs in. WIZARD_CI_GATEWAY_TOKEN_FILE holds the gateway token. // QUACK_INSTALL_DIR sets the project the agent runs in (default: the current directory). -import { runAgent, RunOutcome } from '@agent'; +import { + configureGatewayFromCIEnvironment, + runAgent, + RunOutcome, +} from '@agent'; import type { RunConfig, RunInput } from '@agent/types'; import { Harness, @@ -19,15 +24,20 @@ import { getOrAskForProjectData } from '@utils/setup-utils'; // Point PostHog, skills and MCP at the local stack, like --local-posthog --local-context-mill --local-mcp. initLocalDev({ localPosthog: true, localContextMill: true, localMcp: true }); -// Log in with the wizard's own browser OAuth flow. The token stays in memory. +// Log in with keys instead of the browser, the same way --ci does. +const apiKey = process.env.POSTHOG_PERSONAL_API_KEY; +if (!apiKey) throw new Error('Set POSTHOG_PERSONAL_API_KEY'); const programId = 'posthog-integration'; // a program the local gateway admits const login = await getOrAskForProjectData({ signup: false, - ci: false, + ci: true, // with apiKey, this skips OAuth + apiKey, baseUrl: POSTHOG_LOCAL_URL, localMcp: true, programId, }); +// Use the token in WIZARD_CI_GATEWAY_TOKEN_FILE at WIZARD_CI_GATEWAY_URL instead of minting one. +configureGatewayFromCIEnvironment(login.projectId, 'us'); // What the agent runs: one prompt, a small model, no Write, Edit or Bash. const config: RunConfig = { diff --git a/docs/examples/run-program-quack.ts b/docs/examples/run-program-quack.ts index c498d78cc..587ce3cea 100644 --- a/docs/examples/run-program-quack.ts +++ b/docs/examples/run-program-quack.ts @@ -4,8 +4,9 @@ // npx tsx --tsconfig tsconfig.json docs/examples/run-program-quack.ts // // Needs local PostHog on :8010 (with its ai-gateway) and context-mill on :8765. +// POSTHOG_PERSONAL_API_KEY logs in. WIZARD_CI_GATEWAY_TOKEN_FILE holds the gateway token. // QUACK_INSTALL_DIR sets the project the agent runs in (default: the current directory). -import { RunOutcome } from '@agent'; +import { configureGatewayFromCIEnvironment, RunOutcome } from '@agent'; import { runProgram } from '@programs'; import type { ProgramProgress } from '@programs/types'; import { Harness, HAIKU_MODEL, Sequence } from '@shared/constants'; @@ -15,15 +16,20 @@ import { getOrAskForProjectData } from '@utils/setup-utils'; // Point PostHog, skills and MCP at the local stack, like --local-posthog --local-context-mill --local-mcp. initLocalDev({ localPosthog: true, localContextMill: true, localMcp: true }); -// Log in with the wizard's own browser OAuth flow. The token stays in memory. +// Log in with keys instead of the browser, the same way --ci does. +const apiKey = process.env.POSTHOG_PERSONAL_API_KEY; +if (!apiKey) throw new Error('Set POSTHOG_PERSONAL_API_KEY'); const programId = 'posthog-integration'; // a program the local gateway admits const login = await getOrAskForProjectData({ signup: false, - ci: false, + ci: true, // with apiKey, this skips OAuth + apiKey, baseUrl: POSTHOG_LOCAL_URL, localMcp: true, programId, }); +// Use the token in WIZARD_CI_GATEWAY_TOKEN_FILE at WIZARD_CI_GATEWAY_URL instead of minting one. +configureGatewayFromCIEnvironment(login.projectId, 'us'); // Log status lines as the program reports them. runProgram never waits for this. function logProgress(progress: ProgramProgress): void { From 65c7d7230a1408e27129308698fb780e23f305a0 Mon Sep 17 00:00:00 2001 From: "Vincent (Wen Yu) Ge" Date: Fri, 25 Sep 2026 15:40:59 -0400 Subject: [PATCH 29/29] docs(programs): trim the quack sections to the command Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589 --- docs/developer-interfaces.md | 22 +++++----------------- 1 file changed, 5 insertions(+), 17 deletions(-) diff --git a/docs/developer-interfaces.md b/docs/developer-interfaces.md index 07aa07d73..4f4247304 100644 --- a/docs/developer-interfaces.md +++ b/docs/developer-interfaces.md @@ -88,21 +88,12 @@ when consent is already settled. ### Do a quack [`run-program-quack.ts`](examples/run-program-quack.ts) is the smallest -`runProgram` call. It logs in with keys, the same way `--ci` does, runs one -prompt that replies `quack`, logs status lines and prints the outcome. Each step -has a comment. +`runProgram` call. It runs one prompt that replies `quack`, logs status lines +and prints the outcome. Each step has a comment. -It takes the two keys from -[local credentials](local-dev.md#credentials-for-local-ci-and-headless-runs). On -the [local stack](local-dev.md), the personal API key is the one PostHog's -`setup_local_api_key` command creates, and the gateway token is the `phs_` key -PostHog's `setup-gateway-e2e` provisions. Run it from the repository root: +Run it from the repository root against the [local stack](local-dev.md): ```bash -export POSTHOG_PERSONAL_API_KEY= -export WIZARD_CI_GATEWAY_TOKEN_FILE= -export WIZARD_CI_GATEWAY_URL=http://localhost:8080 -export QUACK_INSTALL_DIR= npx tsx --tsconfig tsconfig.json docs/examples/run-program-quack.ts ``` @@ -192,11 +183,8 @@ export const signature: ( ### Do a quack [`run-agent-quack.ts`](examples/run-agent-quack.ts) is the smallest `runAgent` -call. It logs in with the same keys, builds a `RunConfig` with one prompt and no -Write, Edit or Bash, and prints the transcript tail and the outcome. Each step -has a comment. - -With the same variables set: +call. It builds a `RunConfig` with one prompt and no Write, Edit or Bash, and +prints the transcript tail and the outcome. Each step has a comment. ```bash npx tsx --tsconfig tsconfig.json docs/examples/run-agent-quack.ts