Repository navigation
feat(agents): experimental ContainerHarness: Claude Code or Codex in a Container, driven from a Durable Object - #2514
Merged
Conversation
🦋 Changeset detectedLatest commit: b28edcd The changes in this PR will be included in the next version bump. This PR includes changesets to release 2 packages
Not sure what this means? Click here to learn what changesets are. Click here if you're a maintainer who wants to add another changeset to this PR |
Contributor
🟢 agents import sizes: 4 entry points changed, no growth
Changed exports (30)
How this worksEach runtime export is bundled on its own, minified, and gzipped. Changes smaller than 100 B, or smaller than 1% and 1 KiB, are ignored. Growth over 10% or 5 KiB is marked 🔴. This report is informational and does not fail CI. The workflow artifact contains every measurement. Compared |
agents
@cloudflare/ai-chat
@cloudflare/codemode
hono-agents
@cloudflare/shell
@cloudflare/think
@cloudflare/voice
@cloudflare/worker-bundler
commit: |
…a Container, driven from a Durable Object
mattzcarey
force-pushed
the
feat/container-harness-pi-shape
branch
from
October 7, 2026 04:16
0be9dae to
cd23a9d
Compare
The platform does not say why a start failed, so a failure no longer deletes the snapshot it came from. Snapshots are skipped once they are too old to restore (their 30-day lifetime runs from the last restore), and otherwise forgotten only after failing three times over at least ten minutes. Failures after the container started (inactivity timeout, egress intercepts) say nothing about the snapshot and are not counted.
A snapshot that fails two starts in a row is skipped (not deleted) in favour of the next source, so a prompt's five attempts run workspace, workspace, setup snapshot, setup snapshot, fresh setup. The fallback's result decides who was at fault: if it starts, the skipped snapshot is broken and dropped; if it fails too, the snapshot is not to blame and its count is cleared. A wake that gives up clears every count, so the next prompt tries the workspace first again.
…dropping it A fallback that starts proves only that the platform can start containers again, not that the skipped snapshot is broken. The harness now stops that container and tries the skipped snapshot once more: if it starts, the workspace is kept; if it fails, it is broken and dropped, and the fallback starts again. The probe does not count against the start budget.
Merged
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Adds experimental
agents/harness/container:ContainerHarnessruns an agent CLI (Claude Code or Codex) in a Cloudflare Container and drives it from a Durable Object. Its API matchesPiHarness.You don't build an image or write a Dockerfile, and the container never holds a credential.
How it works
flowchart LR subgraph DO[Durable Object] H[ContainerHarness] --- S[(HarnessStore<br/>sessions · operations<br/>transcript · CLI session files)] end subgraph W[Worker] E[ContainerEgress] end subgraph C[Container: cloudflare/debian-trixie] D[daemon] -->|one process per turn, as uid 10001| CLI[claude -p / codex exec] end H <-->|WebSocket per session via getTcpPort| D CLI -->|http://anthropic.harness.internal<br/>placeholder key| E E -->|real key| GW[AI Gateway / provider]No image. On first start the harness:
cloudflare/debian-trixie;setupsteps withexec()(an unprivileged user,npm install -gof the CLI, your own steps);agentsat build time);Later starts restore a snapshot. Changing the setup sets up again.
Customise with
setuponly. A step withuser: "agent"runs in the agent's home, and every session home starts as a copy of it. Installing plugins and mods, settings, skills or a fork works as it does on a laptop:Credentials stay outside the container. The CLI calls a placeholder host over plain HTTP with a placeholder key.
interceptOutboundHttphands that request toContainerEgress, which swaps inapiKey/headersand forwards tobaseUrl(https only).The daemon (
agents/harness/container/runtime, no dependencies):~/.claude/projects,~/.codex/sessions) to the object.For any other CLI,
cliAdapter({ command, parser, stateDirs })is the extension point.Durable, like
PiHarness.singleflight,recoveryLoop, a heartbeat) that starts or reattaches, reconciles and waits.Session store (
agents/harness/store): sessions, operations with idempotency keys, and logs of opaque JSON. It doesn't depend on any one harness.Recovery
idleTimeoutMs(5 min) the harness snapshots the container, workspace included, and stops it. The next prompt starts from that snapshot; the CLI resumes its own session (--resume,codex exec resume).container_lost(or reruns withonContainerLost: "retry"). The next container starts from the last snapshot plus the CLI's session files, and the session continues.Example
examples/next/harnesses/containerruns Claude Code and Codex side by side: two Durable Objects that differ only in their preset.All new entries are marked
@experimental.