Skip to content

fix(desktop): recover when WSL backend crashes during connect - #82

Closed
macodev00 wants to merge 1 commit into
mainfrom
cursor/wsl-connecting-crash-startup-3c9f
Closed

macodev00 wants to merge 1 commit into
mainfrom
cursor/wsl-connecting-crash-startup-3c9f

Conversation

@macodev00

Copy link
Copy Markdown
Owner

Problem

In WSL-only mode the desktop app shows a "Connecting to WSL…" splash until the primary backend reports ready. That splash has no controls. If the WSL backend passes preflight and then never becomes ready — it exits on startup (for example code=1), or it stays up but never answers — the restart loop is uncapped. The splash stays up and the only way out is Task Manager plus editing desktop-settings.json.

Preflight failures already fall back to Windows. Exits before ready and repeated readiness timeouts did not.

Fixes pingdotgg#14393

Why

The accepted triage on pingdotgg#14393 asks for the same recovery as a bounded preflight failure: after repeated startup failures, show an error and use the in-memory Windows fallback for this launch, so wslOnly stays on disk and the next launch tries WSL again. Both loops need the cap: exits before the backend was ever ready, and a live process whose readiness probes keep failing.

Change

DesktopBackendManager counts those failures after a clean preflight and resets the count once the backend is ready. After three in a row it calls a new onStartupFailed hook. The hook runs off the start mutex so stop() can interrupt it. If the hook asks for a replacement, the failed run is stopped and started again only when that stop() was the only one (stopGeneration), so a quit during the swap is not undone. On the crash path the next restart waits until the hook finishes, so the fallback is visible to the next config resolve.

DesktopBackendPool wires that hook for the primary. When the failed run has a WSL distro, it logs the failure, shows the existing native error box ("WSL backend isn't responding"), and calls applyWslWindowsFallbackInMemory. A Windows primary returns false and keeps the current restart loop. Exits after a successful ready do not count, so a backend that was healthy and later crashes is not switched to Windows.

The connecting splash is unchanged. It closes the same way it does after any successful primary ready: handleBackendReady opens the main window, which dismisses the splash.

Scope and approval

pingdotgg#14393 is labeled accepted / via-triage. Julius's triage comment confirms the failure on main and names this in-memory Windows fallback, covering both the exit loop and the unreachable-process loop, as the recovery. This change does that and nothing else. It does not change how the staged WSL runtime or the mounted server tree is built.

Verification

Focused desktop tests, after vp fmt on the edited files. Node v24.13.1.

export PATH="$HOME/.nvm/versions/node/v24.13.1/bin:$PATH"
cd /workspace
./node_modules/.bin/vp test run apps/desktop/src/backend/DesktopBackendManager.test.ts apps/desktop/src/backend/DesktopBackendPool.test.ts

Observed: Test Files 2 passed (2), Tests 38 passed (38), duration 2.19s.

The new cases:

  • A live process that keeps returning 503 is replaced with the fallback config on the third readiness budget, then becomes ready.
  • Repeated exits before ready call the hook on the third exit. While the hook is blocked, advancing the clock 30s does not spawn again. After the hook applies the fallback, the next spawn is the Windows URL and becomes ready.
  • If the hook returns false, the restart loop continues.
  • Exits after the backend was ready never call the hook (five crashes).
  • stop() during a blocked hook does not apply the fallback and does not stay stuck, including after pre-ready exits.
  • stop() while a replacement is closing the old run does not start another process.
  • A backend that becomes ready while the hook is still pending is kept.
  • Pool: a wsl-only primary that exits code=1 three times shows the error box titled "WSL backend isn't responding" (body includes the distro, code=1, and the Windows-for-this-launch sentence), clears wslOnly and wslBackendEnabled in memory, keeps wslDistro, spawns the Windows config next, and fires handleBackendReady.
cd /workspace/apps/desktop
../../node_modules/.bin/tsc --noEmit --pretty false

Observed: exit 0. The only diagnostic is a pre-existing suggestion in src/app/DesktopClerk.test.ts (preferSucceedSomeOrNone), unrelated to this change.

cd /workspace
./node_modules/.bin/vp lint apps/desktop/src/backend/DesktopBackendManager.ts apps/desktop/src/backend/DesktopBackendManager.test.ts apps/desktop/src/backend/DesktopBackendPool.ts apps/desktop/src/backend/DesktopBackendPool.test.ts

Observed: exit 0, no findings.

Platform checks and limitations

Checked on Linux (cloud agent VM) with the desktop unit tests and tsc. This machine has no Windows desktop session and no WSL distro, so the native error box and the splash-to-window transition were not captured. The splash markup did not change. The dialog is Electron.dialog.showErrorBox, the same primitive the bounded preflight path already uses, and the pool test asserts the title and body it is called with.

Not covered here, and not changed:

  • The staged runtime still fails t3 --version when the distro is missing libatomic.so.1. That linker error is still reported as a bad archive, and the install still falls through to the mounted server tree.
  • That mounted tree is the Windows server sidecar, so Linux Node can still exit with Cannot find module '@yuuang/ffi-rs-linux-x64-gnu'.
  • An unreachable-but-alive backend still gets three full readiness rounds (one minute each in production) before the fallback, so a slow cold boot is not cut off early. A crash loop falls back on the third exit.
  • The persisted wslOnly setting is not cleared. The next launch tries WSL again.
  • A Windows primary that never becomes ready keeps restarting. Only a run that resolved a WSL distro falls back.
Open in Web Open in Cursor 

WSL-only mode kept the connecting splash up when the primary backend
exited before it was ready, or stayed up and never answered. After three
consecutive startup failures, show an error and use the in-memory Windows
fallback for this launch.

Fixes pingdotgg#14393

Co-authored-by: maco <macodev00@users.noreply.github.com>
@github-actions github-actions Bot added vouch:trusted PR author is trusted by repo permissions or the VOUCHED list. size:L labels Oct 2, 2026
@macodev00

Copy link
Copy Markdown
Owner Author

Opened upstream.

@macodev00 macodev00 closed this Oct 2, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size:L vouch:trusted PR author is trusted by repo permissions or the VOUCHED list.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: WSL-only mode stuck on "Connecting to WSL..." forever when the WSL backend keeps crashing on startup

1 participant