Repository navigation
refactor(agent): WIP publish the agent through entry modules - #1303
Conversation
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
`@agent` exports the runtime values callers use and `@agent/types` the types; everything outside `src/agent` imports one of the two. The MCP prompt streaming export loads on first call so the startup chunk does not grow. The architecture scanner enforces the entries on resolved paths and ESLint rejects deep `@agent/*` specifiers in the editor. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
`debug()` reports through an injected sink that the UI module installs, so shared no longer looks the UI up. The progress tag helpers move from `src/telemetry.ts` to `@utils/telemetry` and the detect error map moves to programs, which own the detect error kinds. `src/shared` is now its own surface in the architecture test; the remaining upward edges are listed with their owners in the stack plan. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
🧙 Wizard CIRun the Wizard CI and test your changes against wizard-workbench example apps by replying with a GitHub comment using one of the following commands: Test all apps:
Test all apps in a directory:
Test an individual app:
Show more apps
Test against a Context Mill branch:
Add Results will be posted here when complete. |
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
| export type * from './types'; | ||
| export { runAgent, RunOutcome } from './runner'; | ||
| export { AgentSignals } from './agent-interface'; | ||
| export { WIZARD_TOOL_NAMES } from './tools'; |
There was a problem hiding this comment.
This will be the eventual clean surface.
There was a problem hiding this comment.
The "Stays" group is that surface, and each group below it names the stage that removes it:
Lines 11 to 19 in 38f57e8
|
/wizard-ci self-driving/nuxt |
🧙 Wizard CI ResultsTrigger ID:
Configuration
|
| import { | ||
| initializeAgent, | ||
| runAgent as executeAgent, | ||
| executeAgent, |
There was a problem hiding this comment.
fable says we need to update line 433 now that executeAgent returns failure instead of error
if (result.failure) throw result.failure.error ?? new WizardError(result.failure.message, {}, result.failure.code)
There was a problem hiding this comment.
Detection now throws on every non-success result, and a decided failure throws its own Error or one with its message:
wizard/src/lib/detection/agentic.ts
Lines 467 to 474 in 200961f
Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Brings main's eight fixes and A1's per-request ask signal onto A3's cancellation contract. Each ask and task notice owns one AbortController, which aborts on its own timeout, on the run signal, and on a sibling's fatal failure through the orchestrator's internal signal. cancelQuestion and cancelTaskNotice are gone, and the catch around host dismissal lives in the answerer. On timeout or cancel, the bridge and the seeded-task offer now settle their own result before aborting the request, so a host that rejects on dismissal can't win the race. Agentic detection keeps main's two attempts and A3's tagged results; its timeout returns a tagged AGENTIC_DETECTION_TIMEOUT failure after the host-abort check. AgentErrorType joins the agent entry for detection. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
…n a run ends Terminal analytics. Since A3's tagged results, a decided failure with no Error object (MCP missing, YARA violation, no progress, API errors read from output, an SDK result failure, the agent's own [ABORT]) reached wizardAbort without an Error. So the run was labelled 'cancelled' and skipped error tracking. wizardAbort now takes an explicit status. The legacy host passes 'cancelled' only for a host-cancelled run and 'error' for everything else, and captures a WizardError built from the failure's code and message. The agent's own [ABORT] returns Failed; Aborted now means only that the host's signal cancelled the run. What the user sees is unchanged. Orchestrator funnel. When a sibling's fatal result ended the drain, the accepted steps it stopped sent no 'orchestrator task blocked' event. The fatal path now sends them too, through one helper that catches analytics errors, so a failed capture can't turn a decided result into Crashed. Linear cancellation. With no siblings, an ask left open when the harness ended terminally stayed pending until its own timeout. The linear sequence now owns a run controller, as the orchestrator does: the host signal forwards to it, and it aborts when the run ends. The runner and agent READMEs say terminal analytics are the host's. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Carries the guard that keeps a successful run successful when its terminal analytics flush fails, next to this PR's explicit terminal status on failure. One import conflict in the adapter test. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Release A: real-TUI sweep, nine programs on 200961fEach program ran once through the real Ink TUI in a PTY, from a detached worktree pinned to 200961f, against a temporary copy of its workbench app, on project 228144 (US), with its All nine passed: exit 0,
Known and not from A: warehouse-source says "Data warehouse source connected!" but creates 0 of 1 sources, and replay-vision drafts its scanners but doesn't create them. In both cases the test key lacks the MCP scopes, and main does the same. Other notes:
Rendered from the real TUI frames ( ✅ posthog-integration · express-todo · 5m35s · 7/8 tasks · 25 framesFull journey including MCP and Slack follow-ups. Screen path: intro → auth → run → outro Not completed: Add user identification (skipped) 01-intro 02-auth 03-run 04-run 07-run 10-run 13-run 16-run 19-run 20-run 21-outro 22-outro 23-mcp 24-slack-connect 25-keep-skills ✅ error-tracking · express-todo · 6m06s · 7/8 tasks · 24 framesDetect screen, then the orchestrator run. Screen path: error-tracking-intro → auth → error-tracking-detect → run → outro Not completed: Collect upload credentials (skipped) 01-error-tracking-intro 02-auth 03-run 04-run 07-run 10-run 14-run 17-run 20-run 21-run 22-outro 23-outro 24-keep-skills ✅ metrics · express-todo · 3m40s · 3/3 tasks · 13 framesPlatform flow on a plain Node API. Screen path: metrics-intro → auth → run → outro 01-metrics-intro 02-run 03-run 04-run 05-run 07-run 08-run 09-run 10-run 11-outro 12-outro 13-keep-skills ✅ ai-observability · node-weather · 4m24s · 7/7 tasks · 27 framesThe agent picks the OpenAI provider variant, instruments calls, writes env vars and a report. Screen path: ai-observability-intro → auth → run → outro 01-ai-observability-intro 02-run 03-run 07-run 11-run 16-run 20-run 24-run 25-run 26-outro 27-keep-skills ✅ audit · posthog-demo-3000 · 5m39s · checklist · 13 framesRead-only audit. Progress lives in the audit checklist, not the task list. Screen path: audit-intro → auth → audit-run → audit-outro → keep-skills 01-audit-intro 02-auth 03-audit-run 04-audit-run 05-audit-run 06-audit-run 08-audit-run 09-audit-run 10-audit-run 11-audit-run 12-audit-outro 13-keep-skills ✅ error-tracking-upload-source-maps · cicd-github-actions-node-raw · 3m46s · 8/8 tasks · 38 framesDetect picks the project; wizard_ask overlays answered by the driver. Screen path: source-maps-intro → auth → source-maps-detect → run → wizard-ask → run → wizard-ask → run → source-maps-outro → keep-skills 01-source-maps-intro 02-auth 03-source-maps-detect 04-run 05-run 06-run 07-run 08-wizard-ask 09-run 10-run 14-run 18-run 21-run 25-run 29-run 30-run 31-wizard-ask 32-run 33-run 34-run 35-run 36-source-maps-outro 37-source-maps-outro 38-keep-skills ✅ replay-vision · court-booking · 6m16s · 5/7 tasks · 25 framesAgainst court-booking, the app inside replay-vision/react-vite. Screen path: agent-skill-intro → auth → run → outro Not completed: Scan for user frustration (skipped), Summarize sessions (skipped) 01-agent-skill-intro 02-auth 03-auth 04-auth 05-auth 06-auth 07-run 08-run 11-run 13-run 16-run 18-run 21-run 22-run 23-outro 24-outro 25-keep-skills ✅ self-driving · expense-splitter · 8m31s · 9/9 tasks · 68 framesIntegration-first: composed integration run, handoff, GitHub gate, terminal outro. Screen path: self-driving-intro → self-driving-integration-check → auth → self-driving-integration-detect → run → self-driving-handoff → self-driving-github → run → wizard-ask → run → wizard-ask → run → outro 01-self-driving-intro 02-self-driving-integration-check 03-self-driving-integration-check 04-auth 05-self-driving-integration-detect 06-run 07-run 12-run 17-run 21-run 26-run 31-run 32-run 33-self-driving-handoff 34-run 35-run 39-run 42-run 46-run 49-run 53-run 54-run 55-wizard-ask 56-run 57-run 58-run 59-run 60-run 61-wizard-ask 62-run 63-run 64-run 65-run 66-run 67-run 68-outro ✅ warehouse-source · stripe-saas-demo · 2m28s · 6/6 tasks · 31 framesStripe detected from package.json. Screen path: warehouse-intro → auth → run → wizard-ask → run → outro 01-warehouse-intro 02-run 03-run 06-run 09-run 13-run 16-run 19-run 20-run 21-wizard-ask 22-wizard-ask 23-wizard-ask 24-run 25-run 26-run 27-run 28-run 29-run 30-outro 31-keep-skills |
gewenyu99
left a comment
There was a problem hiding this comment.
NotVincent — automated review. Not written or checked by a person. Verify before acting on any of it.
| if (from === 'tui' && target !== AGENT_TYPES_ENTRY) { | ||
| return `matrix:${from}->${to}`; | ||
| } | ||
| return null; |
There was a problem hiding this comment.
Here's the potential issue: The known entry shortcut lets shared and environment code bypass their runtime import restrictions.
shared helper imports runAgent from @agent -> checks pass -> forbidden upward dependency goes unflagged
Suggested fix: Address caller restrictions and regression cases in planned C2's boundary-enforcement pass. This is an accepted follow-up, not a blocker for A3.
There was a problem hiding this comment.
C2d replaces this checker with per-layer TypeScript configs, and at the C3 head skill-map.ts no longer imports the agent:
wizard/src/shared/errors/skill-map.ts
Line 2 in 548ef5f
There was a problem hiding this comment.
Disregarding this one with no change, since the checker is only a sanity ledger and C2d replaces it with compiler-enforced layer configs, where src/shared can resolve only @env, @shared and @utils:
wizard/src/shared/tsconfig.layer.json
Lines 7 to 16 in 548ef5f
Carries the fence's directory-import patterns and the handoff, benchmark and scan-summary wiring tests. No textual conflicts. The scan-summary table's abort case now uses A3's tagged result and expects Failed, since here an agent's own [ABORT] fails the run and only the host's signal aborts it. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
… starts runAgent's pre-aborted return flushed the scan report but dropped the summary line it returned, so the report file could be written with no line in the terminal. Every termination path now flushes through one helper that emits the line as a log event. The scan-summary test table gains the two host-cancel cases, mid-run and before start; the second failed before this change. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Brings in main (#1235, 2.77.0) through A2b. One conflict: the legacy adapter takes A3's @agent and @agent/types imports and adds TASK_OUTCOMES_KEY. The key and its TaskOutcome type join the public entry in the B2 group, and the e2e harness reads them from there instead of the orchestrator's queue. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Brings in main (#1319). No conflicts. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589
Release A landed on main as squash commits (#1293, #1297, #1299, #1303). B1 already carries that content through the A3 branch, so the merge keeps B1's tree and adds #1334, the one change main has beyond A3, with B1 import paths. Generated-By: PostHog Desktop Task-Id: d14e92bb-6ee1-49b5-8502-39cb80079589



























































































































The agent now exposes its runtime API through
@agentand its types through@agent/types. Later stack layers can change the implementation behind those entries without changing callers.Terminal SDK failures and aborts now have explicit result states. Agentic detection stops before a partial transcript can be read as success, preserving an attached
Errorwhen one exists.Implementation and checks
The runtime entry and type entry group exports by the stack layer that will own them. The MCP prompt wrapper loads its streaming module on first use. Skill menu fetching lives in shared code, and
debug()writes through a sink installed by the UI.Lint and architecture checks guard deep agent imports. The architecture list has 38 recorded exceptions at this head. Three production deep imports remain documented for later moves. The runner classifies terminal SDK results and cancels active sibling work after a fatal result. The detector checks the result state before parsing JSON, preserving an original
Erroror reporting the failure message.The updated A3 behavior passed in the B1 integration: 3,205 unit tests, typecheck, architecture, lint, and bundle. A3's current remote head passed its configured CI checks. No credentialed or snapshot rerun was made for these fixes.
The frames below were captured before the terminal-result fixes. They have not been rerun at
1443587c. The prior harness run completed 7/7 steps, but two assertions also failed on #1299.Real-TUI snapshots: express-todo, 26 frames, earlier head
Captured by the wizard-workbench snapshot route (
pnpm wizard-ci-snapshots, realstartTUIin a PTY) against this branch, project 228144, US. The run completed 7/7 steps with a dashboard and a notebook. Frames are on the image-only branchposthog/a3-snapshots; the.anssources and a colored HTML report sit next to them.01-intro
02-run
03-run
04-run
05-run
06-run
07-run
08-run
09-run
10-run
11-run
12-run
13-run
14-run
15-run
16-run
17-run
18-run
19-run
20-run
21-run
22-outro
23-outro
24-mcp
25-slack-connect
26-keep-skills