Skip to content

Should HarnessRouter support persistent harness runtimes instead of spawning a CLI process per turn? #25

Description

@yunhao-tech

Summary

HarnessRouter currently uses the harness CLI as the primary execution interface and generally starts a new process for each turn.

This is a pragmatic way to support multiple harnesses behind one protocol: the CLI is often the lowest common interface, it already encapsulates the complete agent behavior, and a subprocess provides a useful isolation boundary.

However, using a one-process-per-turn execution model introduces several limitations:

  • process startup overhead on every turn;
  • session continuity depends on each vendor CLI's resume implementation;
  • interactive capabilities are reduced to what a one-shot CLI command can expose;
  • the UHP session lifecycle is not backed by a persistent harness runtime;
  • harness-specific features such as mid-turn input, interactive tool approval, pause/resume, and runtime-local state become difficult to preserve.

Would it make sense to support persistent, capability-aware harness adapters while retaining CLI execution as a fallback?

Why CLI-first makes sense

There are good reasons for the current design.

1. The CLI is the lowest common interface

Not every supported harness exposes a stable SDK or service protocol. A process with stdin/stdout is therefore the broadest integration surface and makes it possible to support more harnesses without requiring a vendor-specific in-process integration for each one.

2. The CLI already contains the full harness behavior

Using a model SDK directly would require HarnessRouter to reimplement substantial parts of each harness, including:

  • the agent loop;
  • tool execution;
  • MCP integration;
  • project instruction and skill discovery;
  • context management and compaction;
  • provider configuration;
  • permission handling;
  • persistence and recovery.

The CLI allows HarnessRouter to reuse the actual harness rather than reproduce it.

3. Subprocesses provide isolation

A subprocess boundary makes cancellation, timeout handling, environment isolation, working-directory control, and process-group termination relatively straightforward.

For these reasons, CLI support should probably remain an important compatibility path.

The problem with one process per turn

The concern is not primarily that a CLI is used. The concern is that the CLI invocation also defines the runtime lifecycle.

Today, a logical HarnessRouter session does not necessarily correspond to a live harness session. Each turn starts a new process and attempts to reconstruct continuity through vendor-specific state.

For example, continuity may depend on mechanisms such as:

  • Claude Code --resume;
  • Codex resume --last;
  • Pi session identifiers;
  • Hermes state stored in SQLite;
  • DeepSeek resume-or-create behavior.

These mechanisms are not equivalent. Some resume a transcript, some recover vendor state, and some only approximate continuity. As a result, the UHP session abstraction can be stronger than the execution guarantees provided by the underlying adapter.

This creates several concrete problems.

Startup overhead

Every turn pays for process initialization, configuration discovery, authentication setup, MCP initialization, workspace scanning, and potentially other harness-specific startup work.

For a multi-turn interactive session, this work is repeated even though the logical session remains active.

Vendor-dependent session semantics

Session continuity is only as reliable as the CLI's resume functionality. A change in vendor behavior, session file format, or CLI flags can affect HarnessRouter's session guarantees.

A workspace checkpoint plus a vendor session ID is useful for recovery, but it is not equivalent to a persistent runtime.

Limited interactivity

A one-shot command model makes it difficult to support capabilities such as:

  • appending user input while a turn is running;
  • interactive tool approval;
  • pausing and resuming a turn;
  • changing execution parameters during a turn;
  • maintaining runtime-local caches and context;
  • coordinating multiple active agents.

These are runtime operations rather than ordinary request/response operations.

Lowest-common-denominator behavior

If all adapters are forced into the same one-shot execution model, harnesses with richer interfaces cannot expose their full capabilities.

For example, a persistent Codex app-server integration should not need to behave as if it were another codex exec invocation.

Proposed direction

Could HarnessRouter separate the protocol-level session lifecycle from the process lifecycle and introduce capability-aware harness adapters?

A possible interface could look like this:

create_session
send_turn
append_input
approve_tool
cancel_turn
checkpoint
restore
close_session
capabilities

Each adapter would declare the operations it actually supports:

{
  "persistent_runtime": true,
  "resume": true,
  "append_input": true,
  "interactive_approval": false,
  "checkpoint": true
}

This would allow HarnessRouter to use the strongest available integration for each harness without pretending that all harnesses have identical capabilities.

Possible adapter strategies could include:

  • Codex: a persistent app-server adapter;
  • Claude Code: an Agent SDK or another managed persistent-runtime adapter;
  • DeepSeek Harness: an SDK-backed persistent JSON-RPC runtime;
  • other harnesses: a CLI adapter when no stable persistent interface exists.
    The CLI adapter would remain valuable as a compatibility fallback rather than being the lifecycle model imposed on every backend.

Runtime lifecycle

Persistent runtimes would not need to stay alive indefinitely.

HarnessRouter could manage them with an idle policy such as:

  1. create or restore a runtime when a session becomes active;
  2. keep the runtime alive across turns;
  3. checkpoint it after an idle timeout;
  4. terminate the process;
  5. restore it from the checkpoint when the next turn arrives.

This preserves most of the isolation and resource-control benefits of subprocess execution while avoiding startup on every turn.

Session ownership

A useful architectural principle may be:

HarnessRouter owns the logical session lifecycle; vendor resume mechanisms are recovery tools, not the definition of the session.

Under this model:

  • UHP defines the durable session and response objects;
  • the adapter manages the live harness runtime;
  • workspace checkpoints provide portable recovery state;
  • vendor session IDs are adapter-specific metadata;
  • CLI resume is used when necessary but does not determine the protocol semantics.

Questions

  1. Was the one-process-per-turn model chosen mainly for implementation simplicity and backend compatibility, or is it an intentional long-term isolation boundary?
  2. Are persistent harness runtimes already part of the roadmap?
  3. Would a capability negotiation model fit the current UHP abstraction?
  4. Could the existing Codex app-server support become the first persistent adapter rather than being started again for every turn?
  5. Would the maintainers accept an adapter lifecycle that supports both persistent runtimes and one-shot CLI fallbacks?

I think the current CLI-first approach is a reasonable bootstrap strategy. The main concern is making it the permanent execution model for every harness, including those that expose richer and more persistent interfaces.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions