Summary
On the in-process shared daemon, concurrent open_session requests are processed one at a time. A burst of 32 opens on an idle machine completes in ~14 s with the fastest one taking 10.5 s — every open waits for nearly all the others. Under real load (a parent fanning out 16 children, twice) the serialized queue crosses the 30 s open deadline and every open behind the crossing point times out with an empty stderr.
This is distinct from #1719 (a 0 ms timer from a construction-time deadline — fixed in #1843) and from #1492 (a sync mkdir blocking the main thread). Nothing blocks synchronously here; the no-sync session-path audit is green. It is async per-open work on the single session loop that does not interleave across concurrent callers.
Reproduction
Start an in-process daemon with a mock provider, then from 32 independent RpcClients fire open_session (kind: "worker", retain_on_disconnect: true) simultaneously via Promise.allSettled.
Measured (idle sandbox, engine at the #1843 merge commit)
single open (extension via -e) OPENED 710 ms
single open (extension via launch spec) OPENED 738 ms
150 opens, SEQUENTIAL all opened, refusal null
32 opens, CONCURRENT:
opened 32 / failed 0
wall 13,973 ms
ok_ms min 10,550 ms <- if opens ran in parallel this would be ~700 ms
ok_ms median 12,796 ms
ok_ms max 13,971 ms
32 × ~430 ms of serialized loop work. Sequential opens never expose it (each waits only its own ~700 ms), which is why capacity tests at 150 sessions pass while burst topologies fail.
Actual (under load)
A live QA matrix — two parents × 16 children on a loaded host — reports per scenario:
childStart: { total: 14, running: 3, completed: 0, errored: 11 }
rootCause: "Timeout waiting for response to open_session. Stderr: "
A few opens land before the 30 s deadline, most do not, and nothing crashed — so stderr is empty. Applying the #1719 fix changed nothing (identical counts, identical reason strings), which is what separated the two defects.
Expected
A machine-wide daemon whose purpose is one host serving several parents and their children must serve exactly that topology. One of these has to hold:
open_session work interleaves across concurrent callers, so burst latency is ~O(1) rather than O(n) per caller; or
- the open deadline scales with observed queue depth instead of being a fixed 30 s wall clock; or
- the protocol tells a client it is queued (position / ETA) so the client can wait or stagger instead of timing out blind.
(1) is the real fix. (2) and (3) are mitigations that keep the current serialization but stop it presenting as silent failure.
Acceptance
- 32 concurrent opens on an idle host:
ok_ms min within ~2× of a single open, not 15×.
- The burst scenario above completes with
errored: 0 on a loaded host.
- A queued open that will miss its deadline reports why (queue depth / position), never
Stderr: with nothing after it.
Related
Evidence held locally (per-open latency table, matrix logs, cleanup receipts).
Summary
On the in-process shared daemon, concurrent
open_sessionrequests are processed one at a time. A burst of 32 opens on an idle machine completes in ~14 s with the fastest one taking 10.5 s — every open waits for nearly all the others. Under real load (a parent fanning out 16 children, twice) the serialized queue crosses the 30 s open deadline and every open behind the crossing point times out with an empty stderr.This is distinct from #1719 (a 0 ms timer from a construction-time deadline — fixed in #1843) and from #1492 (a sync
mkdirblocking the main thread). Nothing blocks synchronously here; the no-sync session-path audit is green. It is async per-open work on the single session loop that does not interleave across concurrent callers.Reproduction
Start an in-process daemon with a mock provider, then from 32 independent
RpcClients fireopen_session(kind: "worker",retain_on_disconnect: true) simultaneously viaPromise.allSettled.Measured (idle sandbox, engine at the #1843 merge commit)
32 × ~430 ms of serialized loop work. Sequential opens never expose it (each waits only its own ~700 ms), which is why capacity tests at 150 sessions pass while burst topologies fail.
Actual (under load)
A live QA matrix — two parents × 16 children on a loaded host — reports per scenario:
A few opens land before the 30 s deadline, most do not, and nothing crashed — so stderr is empty. Applying the #1719 fix changed nothing (identical counts, identical reason strings), which is what separated the two defects.
Expected
A machine-wide daemon whose purpose is one host serving several parents and their children must serve exactly that topology. One of these has to hold:
open_sessionwork interleaves across concurrent callers, so burst latency is ~O(1) rather than O(n) per caller; or(1) is the real fix. (2) and (3) are mitigations that keep the current serialization but stop it presenting as silent failure.
Acceptance
ok_ms minwithin ~2× of a single open, not 15×.errored: 0on a loaded host.Stderr:with nothing after it.Related
senpi hostCLI #1782 — the shared-daemon tracking issue; this is a gap in its burst behaviourmkdiron main thread)Evidence held locally (per-open latency table, matrix logs, cleanup receipts).