Skip to content

Real-TUI snapshot runs hang on a failed run, mislabel frames, and never fail on a missing report or a host exit code #1333

Description

@gewenyu99

Problem

The snapshot route, where scripts/tui-snapshots.no-jest.ts drives scripts/tui-host.no-jest.ts in a PTY, has seven gaps. A failed run sits until the outer timeout. Some frames get the wrong label, and some show a blank screen. result.json can leave off the end of the flow. A missing report, or a host crash, still exits 0. The profile's skills: "delete" deletes nothing.

  • The timeout kill does not reach the host. node-pty starts the host in a new session. It uses POSIX_SPAWN_SETSID on macOS and forkpty on Linux. A process-group kill of the driver therefore never signals the host. tui-snapshots installs no signal handler and never calls cap.kill(). Cleanup depends on the SIGHUP from the PTY hangup, and any descendant that survives SIGHUP is left running.
  • A fast screen gets the next screen's label. snap() adds the screen to screenPath when it is called. It then sleeps 500 ms and writes store.currentScreen as the frame label. If the screen changed during that sleep, the frame shows the next screen and carries its name, so the first screen has no frame. With CI-mode auth this happened to auth on every B2 run, and on 3 of 9 Release A runs.
  • A failed run waits on the mint-failure screen until the timeout. decideE2eAction has no case for mint-failure ("The Wizard's a little busy"). It returns { wait: true }, and the driver waits in 600 s turns until the outer timeout kills it. The action registry already has dismiss_outro for this screen. Choosing it is not enough on its own: ExitScreen does not exit while mintHandoff is set, because run-wizard normally does that exit, and the host does not use run-wizard.
  • result.json stops at outro. The outro subscription writes the result and sets resultWritten, so the final write after keep-skills returns early. For every program whose outro screen id is outro, screenPath ends at outro and skillsComplete is false, even when mcp, slack-connect and keep-skills ran. feat(programs): B3 plumbing and tests — runProgram, adapter wiring, detection via runAgent #1308 fixes this on the stack, but main still has the bug.
  • No check reads the report outcome or the host's exit code. result.json includes reportFile.exists, but no code in this repo reads it. It also does not record whether the agent published a handoff, which is the report channel the publish_handoff description now asks for (feat: export handoffs to host paths #1141). TuiCapture.exited also drops the host's exit code, and tui-snapshots then always exits 0.
  • The capture can run after the driver has acted. The host keeps the screen for 300 ms after it writes the label. The capturer polls every 150 ms and then waits 200 ms before it reads, so the read can come up to 350 ms after the label. term.write(d) is asynchronous, and frameAnsi() does not wait for pending writes to finish. When the machine is loaded, the timers slip further.
  • skills: "delete" deletes nothing. Every test/e2e.json sets it, but the keep-skills action only records kept: false in the store. No harness run exercises Remove, and runs leave their skills in the app directory.

Evidence

The sweep ran all nine programs through tui-snapshots.no-jest.ts on the refactor stack's B2 head, d938a43 (#1307): nine runs in parallel, then a sequential rerun of source-maps and self-driving. It also ran on the Release A top, 200961f. Apart from import paths, the harness files match main, so the line numbers below are for main (d8486dc).

1. Timeout kill

  • e2e-harness/tui-capture.ts:108 starts the host with pty.spawn.
  • scripts/tui-snapshots.no-jest.ts:29-61 installs no signal handler and exits only after the child exits (:55-61). Nothing calls cap.kill() (e2e-harness/tui-capture.ts:155-161).
  • I reproduced this locally in the sweep driver's shape: Popen(start_new_session=True), then killpg(SIGTERM). The host's session and group id did not match the driver's group. In that minimal case the PTY hangup still ended the host and a plain grandchild. A process is left behind only when a descendant outlives SIGHUP.

2. Labels

  • scripts/tui-host.no-jest.ts:471-472 adds screen to screenPath. Lines :474-475 sleep 500 ms and then append store.currentScreen to the control file.
  • On B2, all nine screenPaths list auth, but none of the nine runs has an auth frame. In posthog-integration, frame 02 is labelled run and shows the run screen with "Using provided API key (CI mode - OAuth bypassed)".
  • On the Release A top, auth was slower, and 6 of the 9 runs have an auth frame.

3. Mint failure

  • e2e-harness/e2e-profile.ts:322-325 handles the outro screens. There is no case for ScreenId.MintFailure, so it falls to :389-390 and returns { wait: true }.
  • scripts/tui-host.no-jest.ts:619 waits up to 600 s per turn.
  • e2e-harness/action-registry.ts:220-231 already defines continue_setup and dismiss_outro for mint-failure.
  • src/ui/tui/screens/ExitScreen.tsx:14 exits only when mintHandoff is unset. Otherwise src/lib/runners/run-wizard.ts:284-297 does the exit.
  • In an earlier sweep of the C stack, replay-vision (replay-vision/react-vite) reached mint-failure at frame 04. It captured frames 04 to 08 and then waited until the 2400 s timeout.

4. result.json tail

  • The guard is at scripts/tui-host.no-jest.ts:631-632, the outro write at :703-705, and the final write at :720. The final write returns early. The comment at :624-629 says integration rewrites the result after keep-skills. The guard came in with feat(e2e): let the e2e harness drive the warehouse flow #1152, and the outro write came in with 6f0c731 (chore(wizard): bind source maps to sol medium #1151).
  • On B2, posthog-integration captured frames 28-mcp, 29-slack-connect and 30-keep-skills. Its result.json still has screenPath intro → auth → run → outro and skillsComplete: false. Metrics, error-tracking, ai-observability, replay-vision and warehouse-source show the same pattern: the last frame is keep-skills, but screenPath ends at outro. Audit and source-maps have their own outro screen ids, so the bug does not affect them.
  • The workbench wizard-ci --e2e checks "full interactive flow reached keep-skills" and "skillsComplete" read these fields (PostHog/wizard-workbench services/wizard-ci/e2e.ts:470-471).
  • The feat(programs): B3 plumbing and tests — runProgram, adapter wiring, detection via runAgent #1308 head, 17c02e4, adds createE2eResultWriter to e2e-harness/e2e-result.ts and calls writeResult(true) in the host.

5. Missing report and exit code

  • e2e-harness/e2e-result.ts:170-193 defines readReportFile, and scripts/tui-host.no-jest.ts:696 stores its result. No code in the repo reads the field.
  • On B2, metrics (posthog-metrics-report.md), audit (posthog-audit-report.md) and warehouse-source (posthog-warehouse-report.md) all have reportFile.exists: false. All three still have runPhase: completed and exit 0. The files are also absent from the app copies.
  • The workbench's "report written" check (services/wizard-ci/warehouse-checks.ts:806-831) runs only for apps that have apps/<app>/.wizard-ci/expect.json. revenue/stripe/stripe-saas-demo has no such file. Where it does run, it passes only when the file exists (:813-820).
  • A missing file is not always a failed run on main. The publish_handoff description says "do not write the report to a file yourself" (src/agent/tools/handoff.ts:33), while skillPrompt still asks for the file (src/agent/agent-prompt.ts:54). In the B2 audit and warehouse runs the agent published the handoff and wrote no file. So a check on reportFile.exists alone would fail runs that followed the tool description. The report-file issue covers picking one channel.
  • e2e-harness/tui-capture.ts:121-123 resolves exited without the exit code, and scripts/tui-snapshots.no-jest.ts:60 exits 0. The workbench pass condition includes run.status === 0 (e2e.ts:399), so it never sees a host failure.

6. Capture race

  • scripts/tui-host.no-jest.ts:476 keeps the screen for 300 ms after the label is written. scripts/tui-snapshots.no-jest.ts:54 polls every 150 ms, and :45 waits 200 ms before it calls frameAnsi().
  • e2e-harness/tui-capture.ts:118 calls term.write(d) with no callback. frameAnsi() (:136-151) reads the buffer in whatever state it is in.
  • On B2 with nine parallel runs, source-maps frame 31 (wizard-ask) shows only the header and the "↑↓ navigate enter select" footer. In the sequential rerun, the same frame shows the full test-affordance question.
  • On the Release A top, self-driving frames 55 and 61 (wizard-ask) are blank in the parallel sweep and complete in a solo rerun (frames 64 and 70).
  • On B2, self-driving frames 58 and 64 are blank, and they were still blank in the sequential rerun (frames 63 and 69). Load is not the only cause.

7. skills: "delete"

  • e2e-harness/e2e-profile.ts:343-351: the keep-skills decision returns kept: profile.skills === 'keep' and skillsPolicy. The comment at :149 says "the orchestrator does the fs deletion", and the doc comment at :265 says the caller "handles skillsPolicy itself".
  • e2e-harness/action-registry.ts:326-330: keep_skills only calls store.setSkillsComplete(...).
  • Nothing in the repo reads skillsPolicy. E2E_KEEP_SKILLS, which PostHog/wizard-workbench services/wizard-ci/e2e.ts:300 sets, is not read either.
  • e2e-harness/ARCHITECTURE.md:92-94 says keep-skills outcomes only commit store state. So "delete" in every test/e2e.json is a label that nothing acts on.
  • The keep-skills screen issue shows the effect on the product side: a run's skills stay in the directory, and a later run lists them as its own.

Suggested fix

  1. In tui-snapshots, handle SIGTERM, SIGINT and SIGHUP. Send SIGTERM to the PTY child's process group, send SIGKILL after a grace period, then exit. Expose pid on TuiCapture so this is possible. Add a wall-clock deadline in the harness so a stuck run does not rely on the outer driver.
  2. Label each frame with the screen that snap() recorded. If the screen changed before the capture, write a marker such as NN-auth.missed, so that screenPath and the frames agree. A better option is for the capturer to wait until the output goes quiet, instead of sleeping a fixed time.
  3. In decideE2eAction, add case ScreenId.MintFailure: return { action: { id: 'dismiss_outro' }, done: true }. The host then writes the result and exits 1 when mintHandoff === 'exit'. chore(harness): WIP drive the e2e routes over the control socket #1278 made the profile change on a route that was later closed.
  4. Land createE2eResultWriter from feat(programs): B3 plumbing and tests — runProgram, adapter wiring, detection via runAgent #1308, or cherry-pick it to main.
  5. Record the handoff in result.json (for example handoff: { published, chars } from the store's handoffText). Declare the expected report channel in each program's test/e2e.json, and exit non-zero when neither the handoff nor the named file exists. Once the product keeps one channel, check only that one. Pass the host's exit code through TuiCapture.exited and use it as the exit code of tui-snapshots. The e2e-key issue adds a scope-blocked result on top of this.
  6. Replace the timed handshake with an acknowledgement. The capturer writes the frame and then appends an ack, and the host waits for the ack before it acts. Use term.write(d, cb) and capture only after all pending writes have finished.
  7. Make skills: "delete" do something, or remove it. Either move the removal into a shared function that KeepSkillsScreen's handleRemove and the keep_skills action both call when kept is false, or delete skillsPolicy, the profile field and the E2E_KEEP_SKILLS env, and say the profile only records the choice.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions