Skip to content

fix(desktop): recover when a wsl-only backend never becomes ready - #84

Closed
macodev00 wants to merge 1 commit into
mainfrom
cursor/wsl-startup-failure-fallback-8565
Closed

macodev00 wants to merge 1 commit into
mainfrom
cursor/wsl-startup-failure-fallback-8565

Conversation

@macodev00

Copy link
Copy Markdown
Owner

Fixes pingdotgg#14393. Supersedes pingdotgg#14740.

Problem

WSL-only mode can sit on "Connecting to WSL..." forever when the backend keeps crashing, or stays up but never answers, after preflight has already passed. The splash has no controls. Exits before ready restart without a cap, and a live unreachable process is probed again every minute.

pingdotgg#14393

Change

After three consecutive pre-ready failures, the primary stops scheduling restarts and asks the pool to recover. That covers both loops: a child that exits before it is ready, and a live child whose readiness budget expires three times (the process is killed so the same exit path owns recovery).

The pool shows a fixed error box and applies the existing in-memory Windows fallback, so the next configResolve starts Windows for this launch and the next app start tries WSL again. A Windows primary has no distro and keeps the normal restart loop.

Restarts are suppressed for the whole time a startup-failure hook is pending, including inside scheduleRestart and in the restart fiber. stop() interrupts the hook and a recovery start() no-ops if quit already moved the stop generation. Dialogs and log annotations only carry a bounded category (exited or unreachable), an optional exit code 0–255, and the distro name. The raw cause stays on the log cause.

Scope and approval

Accepted bug: pingdotgg#14393 (apps/desktop). Triage asked for the same in-memory Windows fallback used for a non-fatal preflight failure, for both the crash loop and the live-but-unreachable loop.

This replaces pingdotgg#14740. Recovery is one hook after a fixed cap, not a concurrent replace-by-run-id lifecycle. A process exit does not schedule a restart while that hook is pending.

Verification

Platform: Linux 6.12.94+, Node v24.13.1, Cursor cloud VM. No Windows host and no WSL distro were available, so this was not exercised in a live Electron session. The new UI is a native error box with fixed copy, not a layout change, so there is no before/after screenshot.

export PATH="$HOME/.nvm/versions/node/v24.13.1/bin:$PATH"
cd /workspace
./node_modules/.bin/vp test run apps/desktop/src/backend/DesktopBackendManager.test.ts apps/desktop/src/backend/DesktopBackendPool.test.ts --testTimeout 20000

Observed:

Test Files  2 passed (2)
     Tests  35 passed (35)
  Duration  2.10s
./node_modules/.bin/vp lint apps/desktop/src/backend/DesktopBackendManager.ts apps/desktop/src/backend/DesktopBackendManager.test.ts apps/desktop/src/backend/DesktopBackendPool.ts apps/desktop/src/backend/DesktopBackendPool.test.ts

Observed: exit 0, no findings.

cd /workspace/apps/desktop && /workspace/node_modules/.bin/tsc --noEmit --pretty false

Observed: exit 0. One pre-existing suggestion in src/app/DesktopClerk.test.ts (preferSucceedSomeOrNone). No errors in the changed files.

The new tests check: the hook runs after three pre-ready exits and a 30s clock advance does not spawn again; a declined hook resumes the restart loop once; exits after ready do not count; stop() during the hook does not apply recovery; three readiness budgets kill the live child and fall back without an early restart; the pool dialog text is the fixed template (no URL, no MODULE_NOT_FOUND) and the in-memory settings switch to Windows while keeping the distro.

Limitations

  • The cap is three consecutive pre-ready failures. With the default one-minute readiness budget, a live unreachable backend waits about three minutes before the dialog.
  • The fallback is in memory for this launch. wslOnly and the distro stay on disk, so the next launch tries WSL again.
  • A primary that already resolved to Windows keeps restarting.
  • libatomic.so.1 and ffi-rs MODULE_NOT_FOUND are separate and not handled here.

Model: Grok 4.7. Harness: Cursor cloud agent.

Open in Web Open in Cursor 

After three consecutive pre-ready crashes or readiness timeouts, stop the
restart loop and use the Windows backend for this launch. Restarts stay
suppressed while that recovery is pending, and the dialog only reports a
bounded failure category.

Co-authored-by: maco <macodev00@users.noreply.github.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L labels Oct 2, 2026
@macodev00

Copy link
Copy Markdown
Owner Author

Opened upstream.

@macodev00 macodev00 closed this Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: WSL-only mode stuck on "Connecting to WSL..." forever when the WSL backend keeps crashing on startup

2 participants