What
Every JSON-RPC response goes through jrpc2's jhttp.Bridge, which routes the call through an in-memory server.Local client/server pair. On the way out, the handler result is marshaled by Server.invoke, re-marshaled into the envelope by the bridge (the stdlib validates and re-compacts every embedded json.RawMessage), and re-parsed on the internal client hop. For a 2.8 MB pubnet getLatestLedger response that is about 35 ms and 33 MB of allocation per call, with the handler itself under 6%. The bridge, not the handler, bounds throughput on every large response (getLedgers, getTransactions and getEvents pages too), and it makes a pre-rendered json.RawMessage result 3x slower than a struct, which is exactly what the memo in #989 returns.
Part III of epic #1055: rework the server side so nothing is marshaled or decoded beyond what the wire needs.
json.RawMessage results reach the socket verbatim.
- Requests are dispatched from the HTTP handler straight to the handler map, with no in-memory client hop.
- The envelope is written without copying the result.
jsonrpc.NewHandler sets ServerOptions.Concurrency: math.MaxInt, so jrpc2's handler semaphore (default runtime.NumCPU()) never blocks. In-flight work is bounded by stellar-rpc's own per-method backlog and request-duration limiters, which wrap every handler.
Done in #1037 through a fork, stellar-experimental/jrpc2 branch raw-message-passthrough (on upstream main, see stellar-experimental/jrpc2#1), rather than a parallel dispatcher inside stellar-rpc. jrpc2 keeps owning the wire shape. The SDK's rpcclient, which the integration tests use as their client, moves to the fork too (stellar/go-stellar-sdk#6020), so stellar-rpc carries one jrpc2 module and one jrpc2.Error type.
Supersedes PR #969
There is no prior issue for this. The prior work is @tamirms's #969, which writes the response envelope in-repo and deletes the bridge. Both remove the bridge round trip and only one should land. Measured on a 2.8 MB result: equal on struct results (bridge ~32 ms, ~1.5 ms either way); on json.RawMessage results #969's wire.invoke still runs the bytes through json.Marshal (6.8 ms vs 87 µs when passed through verbatim). #969 also documents seven wire-visible deltas and moves the handler context to server scope; #1037 keeps jrpc2's semantics. Two things from #969 worth carrying over once #1037 lands: the Handler.Shutdown drain, and registering the six limiter metric families that were never registered.
Acceptance criteria
Measurements
Local, through jsonrpc.NewHandler with httptest on real pubnet ledgers (the dispatcher A/B that motivated the fork; the fork's passthrough does the same work):
|
bridge |
passthrough |
getLatestLedger (2.8 MB) |
35.2 ms / 35 MB |
3.8 ms / 15.6 MB |
getLatestLedger with the rendered memo from #989 |
38.9 ms |
0.34 ms / 13 KB |
getLedgers limit=5 (14 MB) |
160 ms / 191 MB |
16.5 ms / 100 MB |
getLedgers limit=20 (56 MB) |
694 ms / 753 MB |
72 ms / 421 MB |
The remaining cost is the handler's own base64 and struct marshal. #969 measured the same bottleneck on the raw wire: getLedgers limit=1 (16.9 MB), p50 410 ms → 69 ms.
Out of scope
What
Every JSON-RPC response goes through jrpc2's
jhttp.Bridge, which routes the call through an in-memoryserver.Localclient/server pair. On the way out, the handler result is marshaled byServer.invoke, re-marshaled into the envelope by the bridge (the stdlib validates and re-compacts every embeddedjson.RawMessage), and re-parsed on the internal client hop. For a 2.8 MB pubnetgetLatestLedgerresponse that is about 35 ms and 33 MB of allocation per call, with the handler itself under 6%. The bridge, not the handler, bounds throughput on every large response (getLedgers,getTransactionsandgetEventspages too), and it makes a pre-renderedjson.RawMessageresult 3x slower than a struct, which is exactly what the memo in #989 returns.Part III of epic #1055: rework the server side so nothing is marshaled or decoded beyond what the wire needs.
json.RawMessageresults reach the socket verbatim.jsonrpc.NewHandlersetsServerOptions.Concurrency: math.MaxInt, so jrpc2's handler semaphore (defaultruntime.NumCPU()) never blocks. In-flight work is bounded by stellar-rpc's own per-method backlog and request-duration limiters, which wrap every handler.Done in #1037 through a fork,
stellar-experimental/jrpc2branchraw-message-passthrough(on upstreammain, see stellar-experimental/jrpc2#1), rather than a parallel dispatcher inside stellar-rpc. jrpc2 keeps owning the wire shape. The SDK'srpcclient, which the integration tests use as their client, moves to the fork too (stellar/go-stellar-sdk#6020), so stellar-rpc carries one jrpc2 module and onejrpc2.Errortype.Supersedes PR #969
There is no prior issue for this. The prior work is @tamirms's #969, which writes the response envelope in-repo and deletes the bridge. Both remove the bridge round trip and only one should land. Measured on a 2.8 MB result: equal on struct results (bridge ~32 ms, ~1.5 ms either way); on
json.RawMessageresults #969'swire.invokestill runs the bytes throughjson.Marshal(6.8 ms vs 87 µs when passed through verbatim). #969 also documents seven wire-visible deltas and moves the handler context to server scope; #1037 keeps jrpc2's semantics. Two things from #969 worth carrying over once #1037 lands: theHandler.Shutdowndrain, and registering the six limiter metric families that were never registered.Acceptance criteria
json.RawMessagehandler results are written to the socket without re-marshal or re-parse, and the bridge's in-memory client hop is gone (Bypass jrpc2 marshal to enable raw endpoint response passthrough #1037, via the fork)rpcclienton the fork (clients/rpcclient: switch to the stellar-experimental/jrpc2 fork go-stellar-sdk#6020, pinned by Bypass jrpc2 marshal to enable raw endpoint response passthrough #1037 until it lands)getLatestLedger,getLedgers) recorded with the perf-eval harness; baseline in [Benchmarking] revert query-path backends to pre-XDR-views to measure speedup #1018, runs on [DRAFT] Optimize query-path caching/memoization #1003's sticky commentMeasurements
Local, through
jsonrpc.NewHandlerwith httptest on real pubnet ledgers (the dispatcher A/B that motivated the fork; the fork's passthrough does the same work):getLatestLedger(2.8 MB)getLatestLedgerwith the rendered memo from #989getLedgerslimit=5 (14 MB)getLedgerslimit=20 (56 MB)The remaining cost is the handler's own base64 and struct marshal. #969 measured the same bottleneck on the raw wire:
getLedgerslimit=1 (16.9 MB), p50 410 ms → 69 ms.Out of scope