Skip to content

rpc: in-process daemon serializes concurrent open_session, so a parent fan-out crosses the 30s open deadline #1844

Description

@code-yeongyu

Summary

On the in-process shared daemon, concurrent open_session requests are processed one at a time. A burst of 32 opens on an idle machine completes in ~14 s with the fastest one taking 10.5 s — every open waits for nearly all the others. Under real load (a parent fanning out 16 children, twice) the serialized queue crosses the 30 s open deadline and every open behind the crossing point times out with an empty stderr.

This is distinct from #1719 (a 0 ms timer from a construction-time deadline — fixed in #1843) and from #1492 (a sync mkdir blocking the main thread). Nothing blocks synchronously here; the no-sync session-path audit is green. It is async per-open work on the single session loop that does not interleave across concurrent callers.

Reproduction

Start an in-process daemon with a mock provider, then from 32 independent RpcClients fire open_session (kind: "worker", retain_on_disconnect: true) simultaneously via Promise.allSettled.

Measured (idle sandbox, engine at the #1843 merge commit)

single open (extension via -e)            OPENED   710 ms
single open (extension via launch spec)   OPENED   738 ms
150 opens, SEQUENTIAL                     all opened, refusal null

32 opens, CONCURRENT:
  opened 32 / failed 0
  wall            13,973 ms
  ok_ms min       10,550 ms     <- if opens ran in parallel this would be ~700 ms
  ok_ms median    12,796 ms
  ok_ms max       13,971 ms

32 × ~430 ms of serialized loop work. Sequential opens never expose it (each waits only its own ~700 ms), which is why capacity tests at 150 sessions pass while burst topologies fail.

Actual (under load)

A live QA matrix — two parents × 16 children on a loaded host — reports per scenario:

childStart: { total: 14, running: 3, completed: 0, errored: 11 }
rootCause:  "Timeout waiting for response to open_session. Stderr: "

A few opens land before the 30 s deadline, most do not, and nothing crashed — so stderr is empty. Applying the #1719 fix changed nothing (identical counts, identical reason strings), which is what separated the two defects.

Expected

A machine-wide daemon whose purpose is one host serving several parents and their children must serve exactly that topology. One of these has to hold:

  1. open_session work interleaves across concurrent callers, so burst latency is ~O(1) rather than O(n) per caller; or
  2. the open deadline scales with observed queue depth instead of being a fixed 30 s wall clock; or
  3. the protocol tells a client it is queued (position / ETA) so the client can wait or stagger instead of timing out blind.

(1) is the real fix. (2) and (3) are mitigations that keep the current serialization but stop it presenting as silent failure.

Acceptance

  • 32 concurrent opens on an idle host: ok_ms min within ~2× of a single open, not 15×.
  • The burst scenario above completes with errored: 0 on a loaded host.
  • A queued open that will miss its deadline reports why (queue depth / position), never Stderr: with nothing after it.

Related

Evidence held locally (per-open latency table, matrix logs, cleanup receipts).

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions