Skip to content

[compass] Fix: a request queued behind a long seat operation loses its reply and detaches the seat - #24

Merged
skulitom merged 4 commits into
mainfrom
compass/anode-fix-2026-10-06-seat-queue-time
Oct 8, 2026
Merged

skulitom merged 4 commits into
mainfrom
compass/anode-fix-2026-10-06-seat-queue-time

Conversation

@skulitom

@skulitom skulitom commented Oct 6, 2026 •

Copy link
Copy Markdown
Owner

The report

Compass's Anode inbox (ANODE-BUGS.md), B-017: long seat operations count their time from when the seat host starts them, not from the caller's deadline. Filed during #18's reviews. compass-merge's Codex review of #18 added a concrete case:

  • Two concurrent 3,000-character seat_type calls. The second waits about 45 s for the daemon's one seat connection, then gets a fresh 60 s clock in the seat host.
  • The daemon times it out at 60 s. Since the request was already sent, JsonPipeClient closes the seat connection, and the daemon marks the seat detached.
  • The seat host types on until its 52 s stop, and the reply is lost.

Root cause

MCP and the CLI give a seat request 60 s (or timeoutMs). The daemon forwards it with the same timeoutMs on its single connection to the seat host, and that connection carries one request at a time.

The seat host counted its time from when the request arrived:

  • DesktopLease.HandleAsync started a full timeoutMs clock;
  • typing planned for 3/4 of it and stopped at 7/8;
  • emulator start and Steam launch kept reserves of 10 s and 15 s.

So time spent queued in the daemon was counted twice. So was time in the seat host before the operation began: the same agent's earlier action, or the previous lease's display restore.

There was a second problem. An operation cut off by the lease's deadline, rather than sized by timeoutMs, was stopped at the same moment the daemon gave up. Its reply usually lost the race. Then:

  • the reply was lost;
  • the connection closed;
  • the seat stayed detached until someone reconnected it.

The fix

Each hop hands on the time left, in the existing timeoutMs field, and the seat stops in time to answer.

  • Daemon: JsonPipeClient.RequestAsync gains sendTimeLeft, which sets timeoutMs to the time left as the request goes out, after any queue. A request left with almost no time (under 100 ms, or an eighth of a shorter time) isn't sent: its reply would come too late, and a late reply costs the connection. The daemon's forwards use it through AnodeDaemon.SendOnAsync.
  • Seat host:
    • DesktopLease takes its wait for an earlier action off, through AgentAccess.Spend.
    • SeatDisplay.RestoredAsync takes off the wait for a display restore.
    • DesktopLease.StopAfterMs stops a desktop action a second (or an eighth of a shorter time) before its caller gives up. So even an action cut short, such as a queued seat_wait, answers in time, and the connection survives. Typing's own stop at 7/8 comes earlier and is unchanged.
  • Tick noise: waits under 100 ms (a sixteenth of a shorter time) aren't counted (JsonPipeClient.TimeLeft). Otherwise a clock tick would shrink the time and randomly refuse a 3,001-character seat_type that fits.
  • Refusal message: a long text refused because its call waited now says the call had less than the usual 60 s.

In the reported case: the second call reaches the seat with about 15 s left. It is refused before anything is typed, with a message that says why, and the seat stays attached.

Why not the absolute deadline the inbox suggested. It needs a new envelope field. Argument checks are strict, so an older running daemon would refuse every request that carries it. timeoutMs already exists on both hops: it is validated as 1 to 180000, and every seat host since 0.5.0 accepts and strips it.

What's left. The client-to-daemon hop still starts its clock a few milliseconds before the daemon. MCP serializes tool calls before its clock starts, and each CLI process has its own connection, so there is no queue there.

Docs: PROTOCOL.md (the envelope paragraph), the CHANGELOG, and comments in TextTyping, SeatEmulators and AgentAccess.

How I tested it

Second review

Codex is at its usage limit until 9 October, so fresh read-only Claude subagents reviewed.

First pass, on dc7cc07: problems found. It confirmed the reported case is fixed end to end. It also confirmed that every seat-side reader of timeoutMs reads it after the deduction, that "no new envelope field" holds, and that leaving the client-to-daemon gap is honest. Its findings, all fixed in 88df5aa and 9d1ccde:

  • P1: most operations don't size their work by timeoutMs. The lease's deadline cut them off just as the daemon gave up, so the docs overclaimed. Fixed with the stop before the caller's deadline, and the docs now say exactly that.
  • P2: a queued long text got advice without the reason. It now says the call had less time.
  • P3: a request with no time left was still sent (now it isn't). Clock ticks shrank the typing limit (waits under 100 ms are now ignored). Also "until it reconnected", the untested restore wait, and stale comments.

Second pass, on 88df5aa and 9d1ccde: all P3, nothing blocking. It confirmed that all five first-pass findings are fixed. It checked that StopAfterMs cuts no operation's own limit short in normal use: typing, seat_wait, audio, exec.start, emulator start, display.set and Steam launch. It also found RestoredAsync's accounting and the new checks sound. Its P3s:

  • Fixed in 7158c1e: a request with almost no time left (under 100 ms, or an eighth of a shorter time) was still sent; now it isn't, and a check covers it. The live stop check's slack is widened to 240 ms.
  • Left as a follow-up (Compass's B-023): an action cut short by its deadline replies with a bare "TaskCanceledException: A task was canceled.". It now arrives in time, but doesn't say whether anything ran. A timed_out error code, with "nothing was done" when the action was never admitted, would be clearer.

CI: green on 88df5aa, 9d1ccde and 7158c1e.

All four open fix PRs merged together on main (#20, #21, #23, #24, locally): 95/95 quick checks, documentation consistent.

Still to check in a real seat

Run two anode type calls of about 2,500 characters each, into Notepad, from two terminals about a second apart, with the same agent's lease. The second should come back refused inside 60 s, saying it had less time, and anode status should stay ready.

🤖 Generated with Claude Code

skulitom and others added 2 commits October 6, 2026 23:30
MCP and the CLI give a seat request 60 s, and the daemon forwards it on its
one connection to the seat host with the same timeoutMs. The seat host started
a fresh clock when the request arrived, so time spent queued behind another
long operation, such as typing, was counted twice. A queued request could
still be running when the daemon gave up on it: JsonPipeClient then closed the
connection, the daemon marked the seat detached, and the reply was lost.

JsonPipeClient.RequestAsync can now send the time left as timeoutMs when the
request goes out (sendTimeLeft), and the daemon's forwards use it through
AnodeDaemon.SendOnAsync. In the seat host, DesktopLease takes its own wait for
an earlier action off before dispatch, and HandleOwnedAsync takes off a wait
for the previous lease's display restore. Every operation that bounds its work
by timeoutMs (typing, emulator start, Steam launch) then answers in time.

No new envelope field: timeoutMs already exists on both hops, so an older
peer is unaffected.

Reported in Compass's Anode inbox as B-017.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…overs it

SeatDisplay.RestoredAsync takes the request and spends the time it waited for
the previous lease's restore, instead of SeatHost measuring around the call.
"display restored after a lease" now checks that an action which waited for
the restore keeps less than its 10,000 ms.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.
To continue using code reviews, add credits to your account and enable them for code reviews in your settings.

…ed waits

Review follow-up.
- The lease's deadline now falls a second (or an eighth of a shorter time)
  before the request's timeoutMs, so an action cut short by it, such as a
  queued seat_wait, still answers before the daemon gives up, and the seat
  connection survives. Typing's own stop (7/8) comes before it, unchanged.
- Waits under 100 ms aren't counted (JsonPipeClient.TimeLeft), so the time
  forwarded to the seat, and limits derived from it such as the 3,001
  characters one call may type, no longer depend on clock ticks.
- A request whose time ran out while it queued isn't sent.
- A refusal of text too long for a call that waited says the call had less
  than the usual 60 s.
- PROTOCOL.md, the CHANGELOG and comments say exactly which operations size
  their work by timeoutMs and that the rest are stopped before it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Second review follow-up (P3s): a forwarded request left with less than
100 ms (an eighth of a shorter time), or whose deadline already fired, isn't
sent, so its late reply can't cost the connection; waits under a sixteenth of
a short time aren't counted either. "pipe forwarded time left" now checks the
unsent path, and "agent queue time" gives its live stop 240 ms of slack.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@skulitom

skulitom commented Oct 7, 2026

Copy link
Copy Markdown
Owner Author

Compass review: ready to merge

Tested by compass-merge on 2026-10-07 on head 7158c1e, with the current main (7f14875, after tonight's merges of #20, #21 and #22) merged in locally:

  • FOCUS test command (scripts\build.ps1 -QuickTest): 95/95 quick checks pass, including "pipe forwarded time left" and "agent queue time"; documentation check consistent. With [compass] Fix: a seat_type call Windows refuses partway loses the count of what was typed #23 merged in as well: 95/95.
  • CI: build green on this head.
  • Review (the whole diff): RequestAsync(sendTimeLeft) stamps the time left once the request holds the one seat connection. Its early "nothing was sent" refusal returns before sent, so the finally releases the queue and the connection stays open. DesktopLease and SeatDisplay.RestoredAsync each take only their own wait off timeoutMs, so no wait is counted twice. StopAfterMs stops an action 1 s (or an eighth) before the caller's deadline, which also covers the waits under 100 ms that aren't counted. The daemon's timeout was already request.timeoutMs ?? 60000, so forwarding an explicit value changes nothing for an unqueued request.
  • Second review: compass-anode-fixer's two fresh read-only Claude subagent passes (Codex is at its usage limit). The P1 and P2 are fixed in 88df5aa/9d1ccde, and two P3s in 7158c1e; one more is B-023.
  • Note, not blocking: the "held" step of "agent queue time" must stop between 1,750 and 1,990 ms, so a runner more than about 240 ms late would fail it, as B-012's runners were. If CI flakes there, that is the place to look.

Not merged tonight only because of the five-merges-per-run cap. The next compass-merge run merges it if nothing changes.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant