You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
The snapshot route, where scripts/tui-snapshots.no-jest.ts drives scripts/tui-host.no-jest.ts in a PTY, has seven gaps. A failed run sits until the outer timeout. Some frames get the wrong label, and some show a blank screen. result.json can leave off the end of the flow. A missing report, or a host crash, still exits 0. The profile's skills: "delete" deletes nothing.
The timeout kill does not reach the host. node-pty starts the host in a new session. It uses POSIX_SPAWN_SETSID on macOS and forkpty on Linux. A process-group kill of the driver therefore never signals the host. tui-snapshots installs no signal handler and never calls cap.kill(). Cleanup depends on the SIGHUP from the PTY hangup, and any descendant that survives SIGHUP is left running.
A fast screen gets the next screen's label.snap() adds the screen to screenPath when it is called. It then sleeps 500 ms and writes store.currentScreen as the frame label. If the screen changed during that sleep, the frame shows the next screen and carries its name, so the first screen has no frame. With CI-mode auth this happened to auth on every B2 run, and on 3 of 9 Release A runs.
A failed run waits on the mint-failure screen until the timeout.decideE2eAction has no case for mint-failure ("The Wizard's a little busy"). It returns { wait: true }, and the driver waits in 600 s turns until the outer timeout kills it. The action registry already has dismiss_outro for this screen. Choosing it is not enough on its own: ExitScreen does not exit while mintHandoff is set, because run-wizard normally does that exit, and the host does not use run-wizard.
result.json stops at outro. The outro subscription writes the result and sets resultWritten, so the final write after keep-skills returns early. For every program whose outro screen id is outro, screenPath ends at outro and skillsComplete is false, even when mcp, slack-connect and keep-skills ran. feat(programs): B3 plumbing and tests — runProgram, adapter wiring, detection via runAgent #1308 fixes this on the stack, but main still has the bug.
No check reads the report outcome or the host's exit code.result.json includes reportFile.exists, but no code in this repo reads it. It also does not record whether the agent published a handoff, which is the report channel the publish_handoff description now asks for (feat: export handoffs to host paths #1141). TuiCapture.exited also drops the host's exit code, and tui-snapshots then always exits 0.
The capture can run after the driver has acted. The host keeps the screen for 300 ms after it writes the label. The capturer polls every 150 ms and then waits 200 ms before it reads, so the read can come up to 350 ms after the label. term.write(d) is asynchronous, and frameAnsi() does not wait for pending writes to finish. When the machine is loaded, the timers slip further.
skills: "delete" deletes nothing. Every test/e2e.json sets it, but the keep-skills action only records kept: false in the store. No harness run exercises Remove, and runs leave their skills in the app directory.
Evidence
The sweep ran all nine programs through tui-snapshots.no-jest.ts on the refactor stack's B2 head, d938a43 (#1307): nine runs in parallel, then a sequential rerun of source-maps and self-driving. It also ran on the Release A top, 200961f. Apart from import paths, the harness files match main, so the line numbers below are for main (d8486dc).
1. Timeout kill
e2e-harness/tui-capture.ts:108 starts the host with pty.spawn.
scripts/tui-snapshots.no-jest.ts:29-61 installs no signal handler and exits only after the child exits (:55-61). Nothing calls cap.kill() (e2e-harness/tui-capture.ts:155-161).
I reproduced this locally in the sweep driver's shape: Popen(start_new_session=True), then killpg(SIGTERM). The host's session and group id did not match the driver's group. In that minimal case the PTY hangup still ended the host and a plain grandchild. A process is left behind only when a descendant outlives SIGHUP.
2. Labels
scripts/tui-host.no-jest.ts:471-472 adds screen to screenPath. Lines :474-475 sleep 500 ms and then append store.currentScreen to the control file.
On B2, all nine screenPaths list auth, but none of the nine runs has an auth frame. In posthog-integration, frame 02 is labelled run and shows the run screen with "Using provided API key (CI mode - OAuth bypassed)".
On the Release A top, auth was slower, and 6 of the 9 runs have an auth frame.
3. Mint failure
e2e-harness/e2e-profile.ts:322-325 handles the outro screens. There is no case for ScreenId.MintFailure, so it falls to :389-390 and returns { wait: true }.
scripts/tui-host.no-jest.ts:619 waits up to 600 s per turn.
e2e-harness/action-registry.ts:220-231 already defines continue_setup and dismiss_outro for mint-failure.
src/ui/tui/screens/ExitScreen.tsx:14 exits only when mintHandoff is unset. Otherwise src/lib/runners/run-wizard.ts:284-297 does the exit.
In an earlier sweep of the C stack, replay-vision (replay-vision/react-vite) reached mint-failure at frame 04. It captured frames 04 to 08 and then waited until the 2400 s timeout.
On B2, posthog-integration captured frames 28-mcp, 29-slack-connect and 30-keep-skills. Its result.json still has screenPath intro → auth → run → outro and skillsComplete: false. Metrics, error-tracking, ai-observability, replay-vision and warehouse-source show the same pattern: the last frame is keep-skills, but screenPath ends at outro. Audit and source-maps have their own outro screen ids, so the bug does not affect them.
The workbench wizard-ci --e2e checks "full interactive flow reached keep-skills" and "skillsComplete" read these fields (PostHog/wizard-workbench services/wizard-ci/e2e.ts:470-471).
e2e-harness/e2e-result.ts:170-193 defines readReportFile, and scripts/tui-host.no-jest.ts:696 stores its result. No code in the repo reads the field.
On B2, metrics (posthog-metrics-report.md), audit (posthog-audit-report.md) and warehouse-source (posthog-warehouse-report.md) all have reportFile.exists: false. All three still have runPhase: completed and exit 0. The files are also absent from the app copies.
The workbench's "report written" check (services/wizard-ci/warehouse-checks.ts:806-831) runs only for apps that have apps/<app>/.wizard-ci/expect.json. revenue/stripe/stripe-saas-demo has no such file. Where it does run, it passes only when the file exists (:813-820).
A missing file is not always a failed run on main. The publish_handoff description says "do not write the report to a file yourself" (src/agent/tools/handoff.ts:33), while skillPrompt still asks for the file (src/agent/agent-prompt.ts:54). In the B2 audit and warehouse runs the agent published the handoff and wrote no file. So a check on reportFile.exists alone would fail runs that followed the tool description. The report-file issue covers picking one channel.
e2e-harness/tui-capture.ts:121-123 resolves exited without the exit code, and scripts/tui-snapshots.no-jest.ts:60 exits 0. The workbench pass condition includes run.status === 0 (e2e.ts:399), so it never sees a host failure.
6. Capture race
scripts/tui-host.no-jest.ts:476 keeps the screen for 300 ms after the label is written. scripts/tui-snapshots.no-jest.ts:54 polls every 150 ms, and :45 waits 200 ms before it calls frameAnsi().
e2e-harness/tui-capture.ts:118 calls term.write(d) with no callback. frameAnsi() (:136-151) reads the buffer in whatever state it is in.
On B2 with nine parallel runs, source-maps frame 31 (wizard-ask) shows only the header and the "↑↓ navigate enter select" footer. In the sequential rerun, the same frame shows the full test-affordance question.
On the Release A top, self-driving frames 55 and 61 (wizard-ask) are blank in the parallel sweep and complete in a solo rerun (frames 64 and 70).
On B2, self-driving frames 58 and 64 are blank, and they were still blank in the sequential rerun (frames 63 and 69). Load is not the only cause.
7. skills: "delete"
e2e-harness/e2e-profile.ts:343-351: the keep-skills decision returns kept: profile.skills === 'keep' and skillsPolicy. The comment at :149 says "the orchestrator does the fs deletion", and the doc comment at :265 says the caller "handles skillsPolicy itself".
e2e-harness/action-registry.ts:326-330: keep_skills only calls store.setSkillsComplete(...).
Nothing in the repo reads skillsPolicy. E2E_KEEP_SKILLS, which PostHog/wizard-workbench services/wizard-ci/e2e.ts:300 sets, is not read either.
e2e-harness/ARCHITECTURE.md:92-94 says keep-skills outcomes only commit store state. So "delete" in every test/e2e.json is a label that nothing acts on.
The keep-skills screen issue shows the effect on the product side: a run's skills stay in the directory, and a later run lists them as its own.
Suggested fix
In tui-snapshots, handle SIGTERM, SIGINT and SIGHUP. Send SIGTERM to the PTY child's process group, send SIGKILL after a grace period, then exit. Expose pid on TuiCapture so this is possible. Add a wall-clock deadline in the harness so a stuck run does not rely on the outer driver.
Label each frame with the screen that snap() recorded. If the screen changed before the capture, write a marker such as NN-auth.missed, so that screenPath and the frames agree. A better option is for the capturer to wait until the output goes quiet, instead of sleeping a fixed time.
In decideE2eAction, add case ScreenId.MintFailure: return { action: { id: 'dismiss_outro' }, done: true }. The host then writes the result and exits 1 when mintHandoff === 'exit'. chore(harness): WIP drive the e2e routes over the control socket #1278 made the profile change on a route that was later closed.
Record the handoff in result.json (for example handoff: { published, chars } from the store's handoffText). Declare the expected report channel in each program's test/e2e.json, and exit non-zero when neither the handoff nor the named file exists. Once the product keeps one channel, check only that one. Pass the host's exit code through TuiCapture.exited and use it as the exit code of tui-snapshots. The e2e-key issue adds a scope-blocked result on top of this.
Replace the timed handshake with an acknowledgement. The capturer writes the frame and then appends an ack, and the host waits for the ack before it acts. Use term.write(d, cb) and capture only after all pending writes have finished.
Make skills: "delete" do something, or remove it. Either move the removal into a shared function that KeepSkillsScreen's handleRemove and the keep_skills action both call when kept is false, or delete skillsPolicy, the profile field and the E2E_KEEP_SKILLS env, and say the profile only records the choice.
Problem
The snapshot route, where
scripts/tui-snapshots.no-jest.tsdrivesscripts/tui-host.no-jest.tsin a PTY, has seven gaps. A failed run sits until the outer timeout. Some frames get the wrong label, and some show a blank screen.result.jsoncan leave off the end of the flow. A missing report, or a host crash, still exits 0. The profile'sskills: "delete"deletes nothing.POSIX_SPAWN_SETSIDon macOS andforkptyon Linux. A process-group kill of the driver therefore never signals the host.tui-snapshotsinstalls no signal handler and never callscap.kill(). Cleanup depends on the SIGHUP from the PTY hangup, and any descendant that survives SIGHUP is left running.snap()adds the screen toscreenPathwhen it is called. It then sleeps 500 ms and writesstore.currentScreenas the frame label. If the screen changed during that sleep, the frame shows the next screen and carries its name, so the first screen has no frame. With CI-mode auth this happened toauthon every B2 run, and on 3 of 9 Release A runs.decideE2eActionhas no case formint-failure("The Wizard's a little busy"). It returns{ wait: true }, and the driver waits in 600 s turns until the outer timeout kills it. The action registry already hasdismiss_outrofor this screen. Choosing it is not enough on its own:ExitScreendoes not exit whilemintHandoffis set, becauserun-wizardnormally does that exit, and the host does not userun-wizard.result.jsonstops atoutro. Theoutrosubscription writes the result and setsresultWritten, so the final write after keep-skills returns early. For every program whose outro screen id isoutro,screenPathends atoutroandskillsCompleteis false, even when mcp, slack-connect and keep-skills ran. feat(programs): B3 plumbing and tests — runProgram, adapter wiring, detection via runAgent #1308 fixes this on the stack, but main still has the bug.result.jsonincludesreportFile.exists, but no code in this repo reads it. It also does not record whether the agent published a handoff, which is the report channel thepublish_handoffdescription now asks for (feat: export handoffs to host paths #1141).TuiCapture.exitedalso drops the host's exit code, andtui-snapshotsthen always exits 0.term.write(d)is asynchronous, andframeAnsi()does not wait for pending writes to finish. When the machine is loaded, the timers slip further.skills: "delete"deletes nothing. Everytest/e2e.jsonsets it, but the keep-skills action only recordskept: falsein the store. No harness run exercises Remove, and runs leave their skills in the app directory.Evidence
The sweep ran all nine programs through
tui-snapshots.no-jest.tson the refactor stack's B2 head, d938a43 (#1307): nine runs in parallel, then a sequential rerun of source-maps and self-driving. It also ran on the Release A top, 200961f. Apart from import paths, the harness files match main, so the line numbers below are for main (d8486dc).1. Timeout kill
e2e-harness/tui-capture.ts:108starts the host withpty.spawn.scripts/tui-snapshots.no-jest.ts:29-61installs no signal handler and exits only after the child exits (:55-61). Nothing callscap.kill()(e2e-harness/tui-capture.ts:155-161).Popen(start_new_session=True), thenkillpg(SIGTERM). The host's session and group id did not match the driver's group. In that minimal case the PTY hangup still ended the host and a plain grandchild. A process is left behind only when a descendant outlives SIGHUP.2. Labels
scripts/tui-host.no-jest.ts:471-472addsscreentoscreenPath. Lines:474-475sleep 500 ms and then appendstore.currentScreento the control file.screenPaths listauth, but none of the nine runs has anauthframe. In posthog-integration, frame 02 is labelledrunand shows the run screen with "Using provided API key (CI mode - OAuth bypassed)".3. Mint failure
e2e-harness/e2e-profile.ts:322-325handles the outro screens. There is no case forScreenId.MintFailure, so it falls to:389-390and returns{ wait: true }.scripts/tui-host.no-jest.ts:619waits up to 600 s per turn.e2e-harness/action-registry.ts:220-231already definescontinue_setupanddismiss_outrofor mint-failure.src/ui/tui/screens/ExitScreen.tsx:14exits only whenmintHandoffis unset. Otherwisesrc/lib/runners/run-wizard.ts:284-297does the exit.replay-vision/react-vite) reached mint-failure at frame 04. It captured frames 04 to 08 and then waited until the 2400 s timeout.4.
result.jsontailscripts/tui-host.no-jest.ts:631-632, theoutrowrite at:703-705, and the final write at:720. The final write returns early. The comment at:624-629says integration rewrites the result after keep-skills. The guard came in with feat(e2e): let the e2e harness drive the warehouse flow #1152, and theoutrowrite came in with 6f0c731 (chore(wizard): bind source maps to sol medium #1151).result.jsonstill hasscreenPathintro → auth → run → outro andskillsComplete: false. Metrics, error-tracking, ai-observability, replay-vision and warehouse-source show the same pattern: the last frame is keep-skills, butscreenPathends at outro. Audit and source-maps have their own outro screen ids, so the bug does not affect them.wizard-ci --e2echecks "full interactive flow reached keep-skills" and "skillsComplete" read these fields (PostHog/wizard-workbenchservices/wizard-ci/e2e.ts:470-471).createE2eResultWritertoe2e-harness/e2e-result.tsand callswriteResult(true)in the host.5. Missing report and exit code
e2e-harness/e2e-result.ts:170-193definesreadReportFile, andscripts/tui-host.no-jest.ts:696stores its result. No code in the repo reads the field.posthog-metrics-report.md), audit (posthog-audit-report.md) and warehouse-source (posthog-warehouse-report.md) all havereportFile.exists: false. All three still haverunPhase: completedand exit 0. The files are also absent from the app copies.services/wizard-ci/warehouse-checks.ts:806-831) runs only for apps that haveapps/<app>/.wizard-ci/expect.json.revenue/stripe/stripe-saas-demohas no such file. Where it does run, it passes only when the file exists (:813-820).publish_handoffdescription says "do not write the report to a file yourself" (src/agent/tools/handoff.ts:33), whileskillPromptstill asks for the file (src/agent/agent-prompt.ts:54). In the B2 audit and warehouse runs the agent published the handoff and wrote no file. So a check onreportFile.existsalone would fail runs that followed the tool description. The report-file issue covers picking one channel.e2e-harness/tui-capture.ts:121-123resolvesexitedwithout the exit code, andscripts/tui-snapshots.no-jest.ts:60exits 0. The workbench pass condition includesrun.status === 0(e2e.ts:399), so it never sees a host failure.6. Capture race
scripts/tui-host.no-jest.ts:476keeps the screen for 300 ms after the label is written.scripts/tui-snapshots.no-jest.ts:54polls every 150 ms, and:45waits 200 ms before it callsframeAnsi().e2e-harness/tui-capture.ts:118callsterm.write(d)with no callback.frameAnsi()(:136-151) reads the buffer in whatever state it is in.7.
skills: "delete"e2e-harness/e2e-profile.ts:343-351: the keep-skills decision returnskept: profile.skills === 'keep'andskillsPolicy. The comment at:149says "the orchestrator does the fs deletion", and the doc comment at:265says the caller "handlesskillsPolicyitself".e2e-harness/action-registry.ts:326-330:keep_skillsonly callsstore.setSkillsComplete(...).skillsPolicy.E2E_KEEP_SKILLS, which PostHog/wizard-workbenchservices/wizard-ci/e2e.ts:300sets, is not read either.e2e-harness/ARCHITECTURE.md:92-94says keep-skills outcomes only commit store state. So "delete" in everytest/e2e.jsonis a label that nothing acts on.Suggested fix
tui-snapshots, handle SIGTERM, SIGINT and SIGHUP. Send SIGTERM to the PTY child's process group, send SIGKILL after a grace period, then exit. ExposepidonTuiCaptureso this is possible. Add a wall-clock deadline in the harness so a stuck run does not rely on the outer driver.snap()recorded. If the screen changed before the capture, write a marker such asNN-auth.missed, so thatscreenPathand the frames agree. A better option is for the capturer to wait until the output goes quiet, instead of sleeping a fixed time.decideE2eAction, addcase ScreenId.MintFailure: return { action: { id: 'dismiss_outro' }, done: true }. The host then writes the result and exits 1 whenmintHandoff === 'exit'. chore(harness): WIP drive the e2e routes over the control socket #1278 made the profile change on a route that was later closed.createE2eResultWriterfrom feat(programs): B3 plumbing and tests — runProgram, adapter wiring, detection via runAgent #1308, or cherry-pick it to main.result.json(for examplehandoff: { published, chars }from the store'shandoffText). Declare the expected report channel in each program'stest/e2e.json, and exit non-zero when neither the handoff nor the named file exists. Once the product keeps one channel, check only that one. Pass the host's exit code throughTuiCapture.exitedand use it as the exit code oftui-snapshots. The e2e-key issue adds a scope-blocked result on top of this.term.write(d, cb)and capture only after all pending writes have finished.skills: "delete"do something, or remove it. Either move the removal into a shared function thatKeepSkillsScreen'shandleRemoveand thekeep_skillsaction both call whenkeptis false, or deleteskillsPolicy, the profile field and theE2E_KEEP_SKILLSenv, and say the profile only records the choice.