You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
Repository navigation
[Feature request] Shared background renderer for all Claude Code sessions #14
Add an opt-in shared-process mode: all Claude Code sessions belonging to one OS user send status-line requests to one long-lived ccstatusline process. Preserve the existing widgets, formatting, per-session data, CLI and TUI.
The intended outcome is to stop loading Node and the rendering module graph on every repaint, share reusable work at the correct scope, and keep a slow session from delaying every other status line. This is a feature proposal; the daemon has not been implemented.
The motivation comes from a local macOS diagnosis with 36 open Claude Code sessions. Snapshots showed 18–25 simultaneous ccstatusline processes, plus Git/custom-command/Keychain subprocesses. A separate local 30-second negative-cache mitigation for unsuccessful Keychain discovery reduced CPU time in repeated test renders from approximately 1.10 s to 0.50 s, including child commands. Those measurements were taken on a heavily loaded machine: they are not a daemon benchmark, not proof of a fixed 55% improvement for all workloads, and not evidence that ccstatusline caused all system load. They show why sharing work is worth investigating. Related fork work: #6 and #7; runtime compatibility is separately tracked in #4.
Claude Code invokes the configured command again on status-line updates, supplies JSON on stdin, and uses stdout as the displayed text. It can cancel an in-flight invocation when a newer update arrives. We must keep that contract; a daemon cannot push directly into Claude's terminal. See the official status-line protocol.
The proposed flow is:
flowchart LR
A["Claude Code sessions"] -->|"stdin JSON"| B["Short-lived shell + curl clients"]
B -->|"HTTP over private Unix socket"| C
subgraph D["One shared Node process"]
C["Validate request and resolve context"] --> E["Async providers and scoped caches"]
E --> F["Existing status-line formatter"]
end
F -->|"ANSI text response"| B
B -->|"stdout"| A
Loading
One shared process means one persistent renderer per OS user and compatible installation/build, not one per session, worktree or account. Lightweight clients still start for each invocation. Git and configured shell commands can still create bounded subprocesses; reducing those remains part of the measurement.
Keep the client genuinely lightweight. Start with the installed shell and curl, without a Node/Bun/Python/jq bootstrap on the healthy repaint path. curl already supports Unix-domain sockets; the server can use Node's built-in HTTP API, without a web framework.
The client forwards the original stdin bytes, adds the invocation context, and prints only the successful response body. Diagnostics go to stderr. Disable user curl configuration, proxy routing and redirects for this IPC call; use explicit connection/total deadlines and HTTP error handling. A partially received/error response must not become a malformed status line. Include all client helper processes in the performance accounting.
Define a small, versioned IPC contract. Use POST /v1/render with the original StatusJSON as the JSON body, and GET /v1/health for readiness/build compatibility. A successful render returns complete UTF-8 ANSI text with the same line breaks, reset codes, non-breaking spaces and OSC links as the existing command.
Transport metadata must include protocol/build identity, request identity/deadline, the caller's actual working directory, ccstatusline config selection and caller environment. A request context is conceptually:
Keep context encoding independent of the status JSON. The first client prototype should use a bounded, encoded metadata header block; a base64-encoded NUL-delimited environment snapshot is one practical candidate that avoids writing a JSON serializer in shell. Supply sensitive header material through an inherited pipe/file descriptor, not command-line arguments, URLs or temporary credential files. Test the actual macOS/Linux tooling and header-size limits before freezing this encoding.
Preserve distinctions between absent and empty environment values, including CLAUDE_CONFIG_DIR and CLAUDE_SECURESTORAGE_CONFIG_DIR. Capture terminal dimensions and overrides for each invocation, including COLUMNS and CCSTATUSLINE_WIDTH; the daemon's own terminal ancestry is irrelevant. Paths and environment values must survive spaces, Unicode, quotes and newlines.
Arbitrary custom commands currently inherit the caller's environment. Full fidelity is the default requirement: pass that snapshot to the relevant child process, keep it transient, and never log it. An allowlist-only mode would need an explicit documented compatibility choice. If the shell/curl prototype cannot preserve these semantics cheaply and safely, benchmark a tiny native transport client before adding another runtime dependency.
Separate request context from process globals. Extract the orchestration in renderMultipleLines() into a function that accepts validated input and explicit invocation/settings context and returns text. Keep stdout/process-exit behavior in the command entry point.
config.ts currently has mutable module-level settingsPath and lastLoadError; terminal.ts memoizes width for the process; usage-fetch.ts has a single in-memory usage result; ccstatusline.ts changes global Chalk/color state. custom-command.ts reads process.cwd() and inherits process.env. These assumptions are valid for short-lived isolated renders but unsafe when unrelated requests overlap.
Resolve config, config errors, profile/account identity, cwd, terminal width and subprocess environment per request. Never switch process.chdir() or mutate process.env to impersonate another caller. Make styling request-local, or serialize only the small final synchronous formatting section with the appropriate style state; do not keep mutable style state across asynchronous waits. Preserve existing config validation, atomic migrations, warning badges and malformed-file recovery.
Gather data asynchronously, then render using the existing widgets. Reuse the widget registry, formatter, transcript parser and existing provider logic. Prefetch the data required by configured widgets into RenderContext so synchronous widget formatting can consume resolved results.
In particular, custom-command capture currently creates an extra Node process with spawnSync(process.execPath, ['-e', ...]) and its helper writes to stdout/exits. Adapt its capture logic to an asynchronous result in the daemon, retaining output limits, timeouts, EPIPE handling and descendant/process-group cleanup. Do not call that helper unchanged inside the server, or preserve a hidden Node-per-widget hot path.
The Git review cache also has a detached refresh path that re-enters the Node executable. In shared mode, route those refreshes through the same asynchronous scheduler instead of spawning another ccstatusline runtime. Audit the other providers for self-spawns and caller-global state as part of the one-process invariant; keep the legacy execution path available for one-shot mode.
Share caches at the scope of their inputs. Reuse existing TTL settings and on-disk formats where compatible. Add bounded in-memory caches and reuse an in-flight refresh for the same key, so a burst of 36 callers does not launch 36 identical lookups. Account/profile separation must happen before any memory-cache fast return.
Data
Scope and invalidation
Settings
Absolute config path plus file revision; revalidate after change and keep errors local to that config.
Usage/credentials
Auth profile, resolved account/credential generation and effective provider configuration. Respect TTL/backoff and logout/account changes. Missing credentials must never select another profile's cached usage.
Service health
Endpoint/query requirements and relevant network configuration; reuse the existing freshness rules.
Git
Actual worktree/cwd, command, relevant environment and existing repository metadata/TTL rules. Linked worktrees must not share branch/index state accidentally.
Transcript analysis
Canonical file identity, size/mtime/replacement indicators and requested analysis options. Account for referenced subagent changes when included.
Custom commands
Existing command/session/width/timeout/TTL semantics, extended with explicit cwd and effective environment identity. TTL zero continues to execute for each eligible request.
Last complete rendered line
Exact display context: session, config revision, cwd, auth profile, terminal dimensions and formatting inputs; short bounded lifetime.
Avoid unconditional complete transcript scans for unchanged files. Start with validated whole-analysis reuse using the existing parser; any dependency that cannot be checked reliably remains uncached. File truncation, replacement, compaction, partial trailing records and subagent updates need invalidation tests. Incremental tail parsing is a follow-up only if profiling still shows a meaningful bottleneck.
Bound cache memory, active jobs and idle-session retention. Do not retain full environments, tokens or transcript contents in diagnostics or an unbounded map. No periodic whole-machine scan is needed: refresh on requests and expire idle entries.
Handle overlapping requests and cancellation deliberately. Use bounded queues and provider concurrency, with at most one active and one newest pending render per exact display context. Resolve superseded requests explicitly; do not leave their connections waiting indefinitely. Different accounts/configs/terminal widths are distinct contexts, even when a conversation ID is reused.
When Claude cancels the client, detach that consumer. Cancel provider work when it has no remaining consumers; cancellation of one request must not kill a shared refresh needed by others. Slow network calls, custom commands and large transcript reads must not prevent cached responses for unrelated sessions. Chunk/yield long scans and limit concurrent expensive scans before introducing a worker pool.
Make lifecycle and failure behavior predictable. Proposed user commands are ccstatusline daemon start|stop|status, plus an explicit installation choice for shared mode. Installation writes a client-wrapper command to Claude settings; returning to one-shot mode restores the previous command.
Coordinate cold startup using the platform's per-user service manager where available, or an atomic per-user startup lock with a readiness handshake. Concurrent first requests must converge on one server. A health response reports protocol and build identity; old clients must not silently talk to an incompatible server after an upgrade. Handle stale sockets, crashed owners, PID reuse and interrupted upgrades without unlinking a live server's endpoint or killing an unrelated process.
Disconnect/restart loops need bounded retry and backoff. On daemon failure, prefer a correctly scoped recent completed render or a short unavailable indicator while one restart attempt proceeds. Do not silently start a full legacy Node render for every failed client: that would recreate the original process storm. Keep explicit one-shot mode available for compatibility and debugging. Do not automatically replay an ambiguous failed render containing a potentially side-effecting custom command.
Keep IPC private and bounded. Store the filesystem socket under a short per-user runtime path with a private directory (0700) and socket permissions restricted to the owner (0600). Verify ownership/type and avoid following attacker-controlled symlinks during stale-file cleanup. Account for Unix socket path-length limits.
Use no TCP listener in the initial implementation. Validate JSON/context, request/header size, cwd/config paths, queue limits and deadlines at the boundary. Same-user clients retain only the capabilities of the existing local command; another OS user must not be able to reach the renderer or induce configured shell commands. Health/status diagnostics should show aggregate counters, versions and timings, never credentials, full environment snapshots, prompts or transcript content.
Extract request-scoped orchestration and prove output parity while retaining the current one-shot entry point.
Add the private IPC server/client and coordinated startup; establish the one-process invariant before adding caching.
Move blocking providers off the event loop, adapt custom-command capture, and add scoped caches, refresh deduplication and cancellation.
Add opt-in installation/status/recovery and benchmarks. Start on macOS/Linux; keep the existing path on Windows until an equivalent private transport/client is tested. Preserve Bun/Node behavior on genuinely supported runtimes and resolve/document Node 14+ distribution requirement is already broken on main (pre-existing blocker) #4 rather than assuming an untested Node 14 claim.
Evaluate incremental transcript parsing or a worker pool only if the resulting profile warrants them.
Acceptance criteria:
With 36 concurrent sessions, exactly one compatible shared renderer remains after startup; a healthy repaint launches no Node/Bun process, including through custom-command capture.
Fixed input, fixed provider results and fixed time produce byte-equivalent output in one-shot/shared modes, including multiline layouts, colors, Powerline, custom commands and warning badges.
Simultaneous different configs, accounts, worktrees, environments and terminal widths cannot exchange data or formatting state; account switches/logout invalidate the relevant cache.
One slow/hung command or usage fetch does not stall cached/simple requests from another session; queues, cancellation and timeout behavior remain bounded.
Cold-start races, crashes, stale sockets, incompatible versions and upgrades recover without duplicate persistent servers or restart/fallback storms.
Truncation/replacement/compaction/partial JSONL records and subagent changes produce correct transcript results.
Cache/session memory stays bounded during a sustained multi-session run and is reclaimed after sessions become idle.
IPC rejects other users and invalid/oversized requests; secrets are absent from logs, URLs, argv and persisted transport metadata.
Benchmarks report end-to-end cost for the daemon and all clients/children, with unchanged output and equivalent data freshness.
Shared mode remains opt-in, with a tested return to the existing one-shot command and no TUI/CLI regression.
Extend scripts/benchmark-render.py rather than introducing a benchmark framework. Compare one-shot and shared mode using 1/10/36/50 sessions, cold and warm starts, idle timer refreshes, active updates, several repositories, different configs/accounts, growing transcripts and slow-command/recovery cases. Record CPU seconds per completed render and per minute, aggregate RSS/peak RSS, p50/p95/p99 latency, subprocess counts, provider/cache hits and output equality. Count persistent daemon CPU explicitly; child-process-only accounting misses it.
A suggested performance gate is at least 50% lower aggregate CPU in the warmed 36-session representative workload, with lower aggregate memory and no p95 latency regression for equivalent work/freshness. This is a target to validate, not a claimed result. Keep the feature opt-in until repeatable measurements support enabling it more broadly.
Add an opt-in shared-process mode: all Claude Code sessions belonging to one OS user send status-line requests to one long-lived ccstatusline process. Preserve the existing widgets, formatting, per-session data, CLI and TUI.
The intended outcome is to stop loading Node and the rendering module graph on every repaint, share reusable work at the correct scope, and keep a slow session from delaying every other status line. This is a feature proposal; the daemon has not been implemented.
The motivation comes from a local macOS diagnosis with 36 open Claude Code sessions. Snapshots showed 18–25 simultaneous ccstatusline processes, plus Git/custom-command/Keychain subprocesses. A separate local 30-second negative-cache mitigation for unsuccessful Keychain discovery reduced CPU time in repeated test renders from approximately 1.10 s to 0.50 s, including child commands. Those measurements were taken on a heavily loaded machine: they are not a daemon benchmark, not proof of a fixed 55% improvement for all workloads, and not evidence that ccstatusline caused all system load. They show why sharing work is worth investigating. Related fork work: #6 and #7; runtime compatibility is separately tracked in #4.
Claude Code invokes the configured command again on status-line updates, supplies JSON on stdin, and uses stdout as the displayed text. It can cancel an in-flight invocation when a newer update arrives. We must keep that contract; a daemon cannot push directly into Claude's terminal. See the official status-line protocol.
The proposed flow is:
flowchart LR A["Claude Code sessions"] -->|"stdin JSON"| B["Short-lived shell + curl clients"] B -->|"HTTP over private Unix socket"| C subgraph D["One shared Node process"] C["Validate request and resolve context"] --> E["Async providers and scoped caches"] E --> F["Existing status-line formatter"] end F -->|"ANSI text response"| B B -->|"stdout"| AOne shared process means one persistent renderer per OS user and compatible installation/build, not one per session, worktree or account. Lightweight clients still start for each invocation. Git and configured shell commands can still create bounded subprocesses; reducing those remains part of the measurement.
Keep the client genuinely lightweight. Start with the installed shell and curl, without a Node/Bun/Python/jq bootstrap on the healthy repaint path. curl already supports Unix-domain sockets; the server can use Node's built-in HTTP API, without a web framework.
The client forwards the original stdin bytes, adds the invocation context, and prints only the successful response body. Diagnostics go to stderr. Disable user curl configuration, proxy routing and redirects for this IPC call; use explicit connection/total deadlines and HTTP error handling. A partially received/error response must not become a malformed status line. Include all client helper processes in the performance accounting.
Define a small, versioned IPC contract. Use
POST /v1/renderwith the original StatusJSON as the JSON body, andGET /v1/healthfor readiness/build compatibility. A successful render returns complete UTF-8 ANSI text with the same line breaks, reset codes, non-breaking spaces and OSC links as the existing command.Transport metadata must include protocol/build identity, request identity/deadline, the caller's actual working directory, ccstatusline config selection and caller environment. A request context is conceptually:
Keep context encoding independent of the status JSON. The first client prototype should use a bounded, encoded metadata header block; a base64-encoded NUL-delimited environment snapshot is one practical candidate that avoids writing a JSON serializer in shell. Supply sensitive header material through an inherited pipe/file descriptor, not command-line arguments, URLs or temporary credential files. Test the actual macOS/Linux tooling and header-size limits before freezing this encoding.
Preserve distinctions between absent and empty environment values, including
CLAUDE_CONFIG_DIRandCLAUDE_SECURESTORAGE_CONFIG_DIR. Capture terminal dimensions and overrides for each invocation, includingCOLUMNSandCCSTATUSLINE_WIDTH; the daemon's own terminal ancestry is irrelevant. Paths and environment values must survive spaces, Unicode, quotes and newlines.Arbitrary custom commands currently inherit the caller's environment. Full fidelity is the default requirement: pass that snapshot to the relevant child process, keep it transient, and never log it. An allowlist-only mode would need an explicit documented compatibility choice. If the shell/curl prototype cannot preserve these semantics cheaply and safely, benchmark a tiny native transport client before adding another runtime dependency.
Separate request context from process globals. Extract the orchestration in
renderMultipleLines()into a function that accepts validated input and explicit invocation/settings context and returns text. Keep stdout/process-exit behavior in the command entry point.config.tscurrently has mutable module-levelsettingsPathandlastLoadError;terminal.tsmemoizes width for the process;usage-fetch.tshas a single in-memory usage result;ccstatusline.tschanges global Chalk/color state.custom-command.tsreadsprocess.cwd()and inheritsprocess.env. These assumptions are valid for short-lived isolated renders but unsafe when unrelated requests overlap.Resolve config, config errors, profile/account identity, cwd, terminal width and subprocess environment per request. Never switch
process.chdir()or mutateprocess.envto impersonate another caller. Make styling request-local, or serialize only the small final synchronous formatting section with the appropriate style state; do not keep mutable style state across asynchronous waits. Preserve existing config validation, atomic migrations, warning badges and malformed-file recovery.Gather data asynchronously, then render using the existing widgets. Reuse the widget registry, formatter, transcript parser and existing provider logic. Prefetch the data required by configured widgets into RenderContext so synchronous widget formatting can consume resolved results.
Git, Keychain and custom-command work must use asynchronous subprocess execution in shared mode. Node documents that the synchronous child-process APIs block its event loop; putting the current render path behind an HTTP handler would otherwise make one slow command stall every session.
In particular, custom-command capture currently creates an extra Node process with
spawnSync(process.execPath, ['-e', ...])and its helper writes to stdout/exits. Adapt its capture logic to an asynchronous result in the daemon, retaining output limits, timeouts, EPIPE handling and descendant/process-group cleanup. Do not call that helper unchanged inside the server, or preserve a hidden Node-per-widget hot path.The Git review cache also has a detached refresh path that re-enters the Node executable. In shared mode, route those refreshes through the same asynchronous scheduler instead of spawning another ccstatusline runtime. Audit the other providers for self-spawns and caller-global state as part of the one-process invariant; keep the legacy execution path available for one-shot mode.
Share caches at the scope of their inputs. Reuse existing TTL settings and on-disk formats where compatible. Add bounded in-memory caches and reuse an in-flight refresh for the same key, so a burst of 36 callers does not launch 36 identical lookups. Account/profile separation must happen before any memory-cache fast return.
Avoid unconditional complete transcript scans for unchanged files. Start with validated whole-analysis reuse using the existing parser; any dependency that cannot be checked reliably remains uncached. File truncation, replacement, compaction, partial trailing records and subagent updates need invalidation tests. Incremental tail parsing is a follow-up only if profiling still shows a meaningful bottleneck.
Bound cache memory, active jobs and idle-session retention. Do not retain full environments, tokens or transcript contents in diagnostics or an unbounded map. No periodic whole-machine scan is needed: refresh on requests and expire idle entries.
Handle overlapping requests and cancellation deliberately. Use bounded queues and provider concurrency, with at most one active and one newest pending render per exact display context. Resolve superseded requests explicitly; do not leave their connections waiting indefinitely. Different accounts/configs/terminal widths are distinct contexts, even when a conversation ID is reused.
When Claude cancels the client, detach that consumer. Cancel provider work when it has no remaining consumers; cancellation of one request must not kill a shared refresh needed by others. Slow network calls, custom commands and large transcript reads must not prevent cached responses for unrelated sessions. Chunk/yield long scans and limit concurrent expensive scans before introducing a worker pool.
Make lifecycle and failure behavior predictable. Proposed user commands are
ccstatusline daemon start|stop|status, plus an explicit installation choice for shared mode. Installation writes a client-wrapper command to Claude settings; returning to one-shot mode restores the previous command.Coordinate cold startup using the platform's per-user service manager where available, or an atomic per-user startup lock with a readiness handshake. Concurrent first requests must converge on one server. A health response reports protocol and build identity; old clients must not silently talk to an incompatible server after an upgrade. Handle stale sockets, crashed owners, PID reuse and interrupted upgrades without unlinking a live server's endpoint or killing an unrelated process.
Disconnect/restart loops need bounded retry and backoff. On daemon failure, prefer a correctly scoped recent completed render or a short unavailable indicator while one restart attempt proceeds. Do not silently start a full legacy Node render for every failed client: that would recreate the original process storm. Keep explicit one-shot mode available for compatibility and debugging. Do not automatically replay an ambiguous failed render containing a potentially side-effecting custom command.
Keep IPC private and bounded. Store the filesystem socket under a short per-user runtime path with a private directory (0700) and socket permissions restricted to the owner (0600). Verify ownership/type and avoid following attacker-controlled symlinks during stale-file cleanup. Account for Unix socket path-length limits.
Use no TCP listener in the initial implementation. Validate JSON/context, request/header size, cwd/config paths, queue limits and deadlines at the boundary. Same-user clients retain only the capabilities of the existing local command; another OS user must not be able to reach the renderer or induce configured shell commands. Health/status diagnostics should show aggregate counters, versions and timings, never credentials, full environment snapshots, prompts or transcript content.
The main implementation touchpoints are src/ccstatusline.ts, RenderContext, config, terminal, usage fetching, Git, Git review refresh, custom-command execution, capture and transcript analysis. Keep added IPC/client/lifecycle code separate from the reusable rendering pipeline, without building a general daemon framework.
Deliver this in reviewable steps:
Acceptance criteria:
Extend scripts/benchmark-render.py rather than introducing a benchmark framework. Compare one-shot and shared mode using 1/10/36/50 sessions, cold and warm starts, idle timer refreshes, active updates, several repositories, different configs/accounts, growing transcripts and slow-command/recovery cases. Record CPU seconds per completed render and per minute, aggregate RSS/peak RSS, p50/p95/p99 latency, subprocess counts, provider/cache hits and output equality. Count persistent daemon CPU explicitly; child-process-only accounting misses it.
A suggested performance gate is at least 50% lower aggregate CPU in the warmed 36-session representative workload, with lower aggregate memory and no p95 latency regression for equivalent work/freshness. This is a target to validate, not a claimed result. Keep the feature opt-in until repeatable measurements support enabling it more broadly.