Skip to content

fix(web): close wedged websockets with an app.version heartbeat - #1771

Merged
chuks-qua merged 1 commit into
mainfrom
fix/ws-liveness-heartbeat
Sep 29, 2026
Merged

chuks-qua merged 1 commit into
mainfrom
fix/ws-liveness-heartbeat

Conversation

@chuks-qua

@chuks-qua chuks-qua commented Sep 27, 2026 •

Copy link
Copy Markdown
Contributor

What

Previously every recovery mechanism in the client transport (reconnect scheduling, running-thread hydration, subscription resync, thread-list refresh) hung off ws.onclose. A half-open or wedged socket never fires it, so pushes silently stopped while runningThreadIds kept stale state and the UI stayed in "streaming" forever.

The transport now runs a liveness probe. Every 30s on an open socket it issues the existing app.version RPC with a 10s deadline. An unanswered probe closes the socket locally, so onclose fires and the established recovery ladder takes over. agent.send also gains a 20s timeout so a composer submit errors instead of parking in pending forever.

Why

Composer and project tree could report a thread as active/streaming while the socket had silently died. No ping/pong or application heartbeat existed anywhere in the transport, so nothing could trigger recovery for a dead-but-open connection.

Evidence

  • src/__tests__/connection.test.ts: silent socket is closed and reports reconnecting after one interval plus the timeout; an answered heartbeat leaves the socket open.
  • src/transport suite: the wedged-socket premise had to acknowledge the probe, which is itself the fix working. 69 transport tests pass; tsc --noEmit clean for apps/web.

Review notes

This is PR 1 of a stacked series finishing the agent.event to agent.canonical runtime migration. The plan lives at docs/plans/canonical-runtime-state-migration.md (gitignored working doc). Live lanes and perf probes per the plan are still outstanding.


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.

A half-open socket answers nothing and never fires close, so every
recovery mechanism (reconnect, hydration, subscription replay) waited
forever while the UI showed stale streaming state. A 30s interval now
issues the existing app.version RPC with a 10s deadline; an unanswered
probe closes the socket locally and the established onclose ladder owns
recovery. agent.send gains a 20s timeout so a composer submit cannot
park forever on a dead socket.

The wedged-socket test now answers the liveness probe, since a socket
that ignores it is exactly what the watchdog closes.
@chuks-qua
chuks-qua merged commit 8b60925 into main Sep 29, 2026
6 checks passed
@chuks-qua
chuks-qua deleted the fix/ws-liveness-heartbeat branch September 29, 2026 20:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant