Skip to content

Raise transactor throughput and add a Cloudflare write benchmark - #641

Draft
tvanhens wants to merge 12 commits into
masterfrom
claude/transactor-throughput-optimization-opr95i
Draft

tvanhens wants to merge 12 commits into
masterfrom
claude/transactor-throughput-optimization-opr95i

Conversation

@tvanhens

@tvanhens tvanhens commented Sep 1, 2026 •

Copy link
Copy Markdown
Owner

Primary outcome

Per-transaction work in the transactor no longer grows with the size of unindexed novelty, operations group-commit alongside transactions, and the operation path costs one transactor request per Worker batch instead of two per operation. Every existing commit guarantee is unchanged: one ordered writer, dense t, all-or-nothing batches, receipt claim before body execution, authorization freshness checked at apply and again at commit, ack only after the durable log row, ordered subscriber frames, and the client-tx replay cache.

Transactor loop

  • SortedNovelty keeps log-structured sorted runs instead of rebuilding one array on every flush. A staged transaction's post-commit view layers only its own datoms over the live novelty; the full copy remains only for transactions that define schema attributes.
  • takeBatch no longer isolates an operation into a batch of one. A resolve-time budget (RAMOSE_BATCH_BUDGET_MS, default 20) bounds how long a batch runs before the commit-time freshness recheck. A repeated invocation whose receipt was claimed earlier in the same batch is deferred until after that batch commits.
  • The indexer merges the four index trees concurrently and remembers the armed alarm instead of reading it on every commit.
  • expandTx probes each unique value once per transaction.
  • Log rows commit in multi-row inserts sized under the Durable Object parameter limit.

Measured with bench:transactor at 64 concurrent writers over fsync'd SQLite:

before after
Group commit tx/s 1239 ~7000
One tx per write tx/s 141 499
Ack p50 (group commit) 48 ms 8.5 ms
Core tx cost at 48k unindexed datoms 4.06 ms 0.07 ms

Operation path

Catalog preparation and per-unit policy tables are memoized by descriptor identity, so admission no longer revalidates the catalog on every invocation. Receipt claims reuse the inspected row and complete with one guarded update. Catalog resolution and receipt preparation run without Effect wrappers, with the principal scope digest cached. The replay fence keeps the admission witness when an operation's datoms touch nothing that witness observed, and recomputes otherwise. bun run bench:op drives the bench catalog through the in-process transactor; locally these take operations from 3,577/s to 4,907/s.

On Cloudflare a no-op route into the transactor Durable Object tops out near 1,000 requests/s, and each operation cost two of them: a catalog provisioning call, then the invoke. The transactor now provisions a routed catalog once per object lifetime when the invoke request names its route root, and exposes /invoke-batch, which settles each entry with the same error mapping as a single invoke. Each Worker isolate groups operations for the same database that arrive within RAMOSE_OP_COALESCE_MS (default 2 ms, 0 disables) into one transactor request; every waiting request keeps its own timer so the Workers runtime never cancels it as hung.

Cloudflare benchmark

bun run bench:cf deploys a throwaway stage with test hooks, generates load from inside Cloudflare through a service binding to the stage itself, and drives the raw transaction path and the /op path with lanes × parallel concurrent writers. It reports client-side throughput and latency, the peer's warn and error telemetry captured inside the stage, and the transactor's batch counters as wall-clock cadence. BENCH_PING=1 adds a no-op request phase that measures the per-object request ceiling, and BENCH_SOCKET=1 adds ping and write phases over a WebSocket into the same object. bun run bench:cf:compare runs it against master and the current branch. Both Cloudflare scripts accept an Alchemy profile login in place of explicit credentials.

Measured at 256 concurrent writers, with the no-op info lane as the request ceiling on the same stage:

stage raw tx/s info req/s ops/s
before, typical 600–890 ~980 240–370
after, stage A 482 628 552
after, stage B 345 535 566

Operations now exceed both the single-request write path and the per-request ceiling on their stage, with zero errors. Stage-to-stage variance is large, so compare phases within one run.

The socket phases show where the next ceiling is. On one stage: 667 no-op requests/s with overload errors against 2,004 no-op frames/s with none, and 534 raw writes/s over requests against 878 over frames. The production write protocol is unchanged here; the direction is tracked in #642.

Public API additions

None. Two new environment variables, RAMOSE_BATCH_BUDGET_MS and RAMOSE_OP_COALESCE_MS, forwarded through alchemy.run.ts.

Test lanes

  • Unit: test/internal/core/novelty-runs.test.ts covers sorted runs, the overlay view, and validator visibility; test/worker/operation-coalescer.test.ts covers batching, the size cap, and failure propagation; operations-runtime.test.ts gains clean and dirty replay fence cases. Typecheck, unit, browser, and build pass.
  • Local: not run in the authoring environment (no Docker).
  • Cloudflare: the e2e suite passes on a live stage, and the benchmark numbers above come from live stages.

🤖 Generated with Claude Code

Keep per-transaction work independent of unindexed novelty: sorted novelty
now holds log-structured sorted runs instead of rebuilding one array on
every flush, and a staged transaction's post-commit view layers only its
own datoms over the live novelty instead of copying it. The copy remains
only for transactions that define schema.

Group-commit operations alongside transactions, bounded by a resolve-time
budget (RAMOSE_BATCH_BUDGET_MS) so the commit-time authorization freshness
recheck stays tight, and defer a repeated invocation whose receipt was
claimed earlier in the same batch until after that batch commits.

Merge the four index trees concurrently, remember the armed alarm instead
of reading it on every commit, and probe each unique value once per
transaction.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1hDDT7sST15978317T6HE
bun run bench:cf deploys a throwaway stage with test hooks and a small
operation catalog, drives the raw transaction path and the /op path at a
chosen concurrency, prints throughput, latency, and transactor batch
statistics per phase, and destroys the stage. Transactor tuning variables
are forwarded to the stage so runs can be compared.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01D1hDDT7sST15978317T6HE
@tvanhens
tvanhens deployed to Development September 1, 2026 23:55 — with GitHub Actions Active
@tvanhens
tvanhens deployed to Development September 1, 2026 23:55 — with GitHub Actions Active
@github-actions

github-actions Bot commented Sep 1, 2026 •

Copy link
Copy Markdown

Reef preview: https://ramose-reef-api-pr-641-ftjn6v6zlcs2sfzw.tvanhens.workers.dev

Deployed from 253d4dd as stage pr-641. Torn down when this PR closes.

tvanhens and others added 2 commits September 1, 2026 17:54
R2 refuses to delete a bucket that still holds objects, and the benchmark
writes index segments, so teardown failed with BucketNotEmpty on every run
that reached the load phase and left the bucket behind.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
Client-side throughput only measures the transactor once enough requests are
in flight to keep the commit queue non-empty; below that it measures the round
trip to the edge. Reporting the transactor's own committed-transaction count
and rate alongside the client numbers shows which of the two is being measured.

The per-batch timing counters are not usable here: a Worker isolate does not
advance its clock during synchronous execution, so loopMs, resolveMs, commitMs
and fenceMs all read zero on Cloudflare regardless of RAMOSE_TIMING_YIELDS.
metrics.txPerSec is driven by a clock that advances on request I/O and survives.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@tvanhens
tvanhens deployed to Development September 2, 2026 00:54 — with GitHub Actions Active
@tvanhens
tvanhens deployed to Development September 2, 2026 00:54 — with GitHub Actions Active
The bench stage now exposes a lane route that calls the peer through a
service binding to itself. The local script only orchestrates lanes, so
the transactor sees lanes × parallel concurrent writers instead of being
bounded by the round trip from the developer machine.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@tvanhens
tvanhens deployed to Development September 2, 2026 01:05 — with GitHub Actions Active
@tvanhens
tvanhens deployed to Development September 2, 2026 01:05 — with GitHub Actions Active
The bench Worker tees warn and error telemetry into each lane report so
the phase output names the server-side cause of failed requests, such as
Cloudflare's Durable Object overload back-pressure. Setup calls retry
while a fresh deployment propagates, batch cadence is reported as wall
time because Workers freeze timers during synchronous work, transactor
restarts are detected, and the capability-gated test-admin route carries
the upstream error detail on 500s. The script accepts an Alchemy profile
login in place of explicit Cloudflare variables.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@tvanhens
tvanhens deployed to Development September 2, 2026 01:24 — with GitHub Actions Active
@tvanhens
tvanhens deployed to Development September 2, 2026 01:24 — with GitHub Actions Active
tvanhens and others added 3 commits September 1, 2026 18:56
Catalog preparation and per-unit policy tables are memoized by descriptor
identity, so admission no longer revalidates the catalog on every
invocation. Receipt claims reuse the inspected row and complete with a
single guarded update, catalog resolution and receipt preparation run
without Effect wrappers with the principal scope digest cached, and the
replay fence keeps the admission witness when an operation's datoms touch
nothing that witness observed. A local operation benchmark drives the
bench catalog through the in-process transactor.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…object

A no-op route into the transactor Durable Object tops out near a thousand
requests per second, and every operation cost two of them: a catalog
provisioning call and the invoke. The transactor now provisions a routed
catalog once per object lifetime when the invoke request names its route
root, and accepts an invocation batch that it settles per entry with the
same error mapping as a single invoke. Each Worker isolate groups the
operations that arrive within RAMOSE_OP_COALESCE_MS (default 2) for the
same database into one transactor request, with every waiting request
keeping its own timer so the runtime never sees it as hung. Log rows are
inserted in multi-row statements sized under the Durable Object parameter
limit. The bench gains an info lane that measures the request ceiling, and
both Cloudflare scripts accept an Alchemy profile login.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@tvanhens
tvanhens deployed to Development September 2, 2026 02:38 — with GitHub Actions Active
@tvanhens
tvanhens deployed to Development September 2, 2026 02:38 — with GitHub Actions Active
With test hooks enabled the transactor accepts a writer socket that is not
a subscriber and answers write frames with acks, and the bench gains
socket-ping and socket-write lanes so frames on an open socket can be
compared with requests into the same object.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

This branch was successfully deployed

1 active deployment
Development — 253d4dd9 Deployed Sep 2, 2026 by tvanhens via e2e-cloudflare #1438
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants