Support bounded parallel chunk transfers - #29341
Conversation
|
FYI @tjgq I think this was a follow-up from the original PR |
|
@bazel-io fork 9.2.0 |
sluongng
left a comment
There was a problem hiding this comment.
I'm a bit skeptical of the current approach, so I will skip on reading the window filling implementation right now. (though Codex does suggest there is a problem there)
Each invocation may have multiple actions/spawns running in parallel, each creates multiple blob uploads/downloads, and some of those blobs are chunked blobs. Adding parallelism on the blob level feels like a local optimization, and the new flag does not offer strong control over the total parallelism of the invocation.
I wonder if we need something higher-level that lets us effectively enforce a global parallelism for uploads and downloads🤔
8969051 to
33689c3
Compare
There was a problem hiding this comment.
Pull request overview
Enables bounded parallelism for content-defined chunk (CDC) blob uploads/downloads, improving throughput while avoiding unbounded per-blob fanout.
Changes:
- Implement sliding-window style, per-blob bounded concurrency for chunk uploads in
ChunkedBlobUploader. - Implement sliding-window style, per-blob bounded concurrency for chunk downloads (including in-flight dedup) in
ChunkedBlobDownloader. - Add/expand unit tests for window refill, cancellation, and failure propagation; add a JMH benchmark binary target.
Reviewed changes
Copilot reviewed 6 out of 6 changed files in this pull request and generated 5 comments.
Show a summary per file
| File | Description |
|---|---|
src/main/java/com/google/devtools/build/lib/remote/ChunkedBlobUploader.java |
Adds bounded in-flight chunk upload window and cancellation on failure. |
src/main/java/com/google/devtools/build/lib/remote/ChunkedBlobDownloader.java |
Adds bounded in-flight chunk download window with reassembly and in-flight dedup. |
src/test/java/com/google/devtools/build/lib/remote/ChunkedBlobUploaderTest.java |
Adds tests for window refill, cancellation, and failure handling for parallel uploads. |
src/test/java/com/google/devtools/build/lib/remote/ChunkedBlobDownloaderTest.java |
Updates tests for new download API and adds parallel-window behavior tests. |
src/test/java/com/google/devtools/build/lib/remote/ChunkedTransferBenchmark.java |
Introduces a JMH benchmark for chunked upload/download with latency + jitter. |
src/test/java/com/google/devtools/build/lib/remote/BUILD |
Adds a java_opt_binary target to run the new benchmark. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
Updated this in the follow-up direction you suggested: I removed the flag/plumbing and kept only a small hardcoded per-blob window of 32 as a guard against huge single-blob fanout. The actual global active RPC limit is still the shared gRPC channel pool ( |
33689c3 to
8363b37
Compare
8363b37 to
976a25f
Compare
c64cbdb to
d846665
Compare
There was a problem hiding this comment.
My main point of contention with this approach is that files are buffered in memory in their entirety while being transferred, which could easily cause us to run out of heap. I'm concerned that this might make chunked transfers useless in practice for builds with lots of parallelism.
I think this is easy to fix for the upload side: instead of passing a byte[] around, pass a (path, offset) and delay the opening/reading of the InputStream until the upload starts. (That does increase the number of open descriptors, but as long as they're still bounded by MAX_IN_FLIGHT_CHUNK_UPLOADS, it might be ok.)
The download side unfortunately isn't as easy. We could try to write a chunk out as soon as it completes instead of waiting for all of them the next one in order, but that doesn't help much if the downloads complete roughly at the same time (which is what we'd expect if the transport is completely fair).
Any thoughts?
d846665 to
2a2987d
Compare
I like this. I adjusted the upload path to lazily matierialize the chunk. But for download we can't really do this. I adjusted the lookahead to match the concurrency, but realistically this would be rare that these take up a lot of memory, since we should immediately consume the memory once the next chunk in order is available. It's a tad messier |
2a2987d to
cc1c408
Compare
tjgq
left a comment
There was a problem hiding this comment.
Thanks - I'm happy with the current state; we can revisit the downloader if anyone reports problems. I only have a couple of nits, will approve once they've been addressed.
cc1c408 to
5e2051f
Compare
|
@bazel-io fork 9.2.0 |
|
Hmm... During import, certain tests are somehow failing for RBE (https://buildkite.com/bazel/google-bazel-presubmit/builds/102948/list?jid=019e4143-51bb-4115-9a99-73ded8407a81&tab=artifacts). We recently upgraded RBE VMs to Ubuntu 24.04 from 20.04. Could you rebase on top of HEAD and see if you can trigger the same failures? |
5e2051f to
16f48f9
Compare
@Wyverald I think there's some underlying permissions issues with 24.04 because of unprivileged user namespace restrictions. These test failure do not seem like they could be caused by anything in this PR, probably just because they were previously cache hits which are now needing to be re-run |
|
EDIT: seems to still be failing 😞 |
|
Thanks for the investigation! @coeuvre to take a look as well. |
|
It might be related to a recent upgrade of the rbe worker pool. Still investigating. In the meanwhile, the affected tests on RBE have been disabled by bca5e06. |
16f48f9 to
cb1a9c2
Compare
|
@tyler-french Sure, this will be cherry-picked as soon as it's merged to master. |
…bazelbui… (bazelbuild#29614) …ld/bazel/pull/29341) For `--experimental_remote_cache_chunking` implemented in bazelbuild#28437 This PR enables parallel uploads and downloads for chunked files to improve performance. Transport-level concurrency is still bounded by the existing remote cache gRPC connection/concurrency limits. The chunk transfer managers add a fixed per-blob window of 16 to prevent a single large blob from fanning out too aggressively. To avoid issues with batches, uploads and downloads use simple sliding-window style transfer managers. RELNOTES: CDC chunk uploads and downloads can now happen in parallel within a large blob. Benchmarks were rerun on 2026-05-12 with the JMH target `//src/test/java/com/google/devtools/build/lib/remote:ChunkedTransferBenchmark`. The parent baseline was measured by running the same benchmark harness against the parent commit. With the synthetic benchmark of network delays and simulated jitter, the current branch is about 11-13x faster than the parent baseline for these cases. As usual, this synthetic benchmark is not a substitute for real remote-cache measurements. After Change: ``` Benchmark (avgChunkSizeBytes) (chunkCount) (chunkSizeBytes) (delayMillis) (fileSizeBytes) (jitterMillis) (schedulerThreads) Mode Cnt Score Error Units ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 1 avgt 3 67.963 ± 2.039 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 2 avgt 3 67.925 ± 2.257 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 4 avgt 3 67.915 ± 2.497 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 8 avgt 3 67.946 ± 2.863 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 1 avgt 3 61.357 ± 7.385 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 2 avgt 3 61.373 ± 7.920 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 4 avgt 3 61.326 ± 8.220 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 8 avgt 3 61.393 ± 8.376 ms/op ``` Before Change: ``` Benchmark (avgChunkSizeBytes) (chunkCount) (chunkSizeBytes) (delayMillis) (fileSizeBytes) (jitterMillis) (schedulerThreads) Mode Cnt Score Error Units ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 1 avgt 3 812.081 ± 463.666 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 2 avgt 3 812.570 ± 442.021 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 4 avgt 3 811.883 ± 459.534 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 8 avgt 3 812.371 ± 461.231 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 1 avgt 3 740.734 ± 389.653 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 2 avgt 3 742.434 ± 412.117 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 4 avgt 3 742.483 ± 395.466 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 8 avgt 3 742.509 ± 397.122 ms/op ``` Big File: ``` CURRENT BRANCH (512 MiB) Benchmark (avgChunkSizeBytes) (chunkCount) (chunkSizeBytes) (delayMillis) (fileSizeBytes) (jitterMillis) (schedulerThreads) Mode Cnt Score Error Units ChunkedTransferBenchmark.downloadChunked N/A 512 1048576 25 N/A 10 8 avgt 3 991.101 ± 173.890 ms/op ChunkedTransferBenchmark.uploadChunked 1048576 N/A N/A 25 536870912 10 8 avgt 3 1087.504 ± 123.857 ms/op ``` ``` PARENT BASELINE (512 MiB) Benchmark (avgChunkSizeBytes) (chunkCount) (chunkSizeBytes) (delayMillis) (fileSizeBytes) (jitterMillis) (schedulerThreads) Mode Cnt Score Error Units ChunkedTransferBenchmark.downloadChunked N/A 512 1048576 25 N/A 10 8 avgt 3 12849.136 ± 1733.063 ms/op ChunkedTransferBenchmark.uploadChunked 1048576 N/A N/A 25 536870912 10 8 avgt 3 11888.109 ± 2591.883 ms/op ``` Closes bazelbuild#29341. PiperOrigin-RevId: 919144974 Change-Id: Iaa0ca8971bd21c879f21c708327b5ddd837ecf1f <!-- Thank you for contributing to Bazel! Please read the contribution guidelines: https://bazel.build/contribute.html --> ### Description <!-- Please provide a brief summary of the changes in this PR. --> ### Motivation <!-- Why is this change important? Does it fix a specific bug or add a new feature? If this PR fixes an existing issue, please link it here (e.g. "Fixes bazelbuild#1234"). --> ### Build API Changes <!-- Does this PR affect the Build API? (e.g. Starlark API, providers, command-line flags, native rules) If yes, please answer the following: 1. Has this been discussed in a design doc or issue? (Please link it) 2. Is the change backward compatible? 3. If it's a breaking change, what is the migration plan? --> No ### Checklist - [ ] I have added tests for the new use cases (if any). - [ ] I have updated the documentation (if applicable). ### Release Notes <!-- If this is a new feature, please add 'RELNOTES[NEW]: <description>' here. If this is a breaking change, please add 'RELNOTES[INC]: <reason>' here. If this change should be mentioned in release notes, please add 'RELNOTES: <reason>' here. --> RELNOTES: None Commit bazelbuild@084958b Co-authored-by: Tyler French <french.tyler.d@gmail.com>
### Description For `--experimental_remote_cache_chunking` implemented in bazelbuild#28437 This PR enables parallel uploads and downloads for chunked files to improve performance. Transport-level concurrency is still bounded by the existing remote cache gRPC connection/concurrency limits. The chunk transfer managers add a fixed per-blob window of 16 to prevent a single large blob from fanning out too aggressively. To avoid issues with batches, uploads and downloads use simple sliding-window style transfer managers. RELNOTES: CDC chunk uploads and downloads can now happen in parallel within a large blob. ### Benchmarking Benchmarks were rerun on 2026-05-12 with the JMH target `//src/test/java/com/google/devtools/build/lib/remote:ChunkedTransferBenchmark`. The parent baseline was measured by running the same benchmark harness against the parent commit. With the synthetic benchmark of network delays and simulated jitter, the current branch is about 11-13x faster than the parent baseline for these cases. As usual, this synthetic benchmark is not a substitute for real remote-cache measurements. After Change: ``` Benchmark (avgChunkSizeBytes) (chunkCount) (chunkSizeBytes) (delayMillis) (fileSizeBytes) (jitterMillis) (schedulerThreads) Mode Cnt Score Error Units ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 1 avgt 3 67.963 ± 2.039 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 2 avgt 3 67.925 ± 2.257 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 4 avgt 3 67.915 ± 2.497 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 8 avgt 3 67.946 ± 2.863 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 1 avgt 3 61.357 ± 7.385 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 2 avgt 3 61.373 ± 7.920 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 4 avgt 3 61.326 ± 8.220 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 8 avgt 3 61.393 ± 8.376 ms/op ``` Before Change: ``` Benchmark (avgChunkSizeBytes) (chunkCount) (chunkSizeBytes) (delayMillis) (fileSizeBytes) (jitterMillis) (schedulerThreads) Mode Cnt Score Error Units ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 1 avgt 3 812.081 ± 463.666 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 2 avgt 3 812.570 ± 442.021 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 4 avgt 3 811.883 ± 459.534 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 8 avgt 3 812.371 ± 461.231 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 1 avgt 3 740.734 ± 389.653 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 2 avgt 3 742.434 ± 412.117 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 4 avgt 3 742.483 ± 395.466 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 8 avgt 3 742.509 ± 397.122 ms/op ``` Big File: ``` CURRENT BRANCH (512 MiB) Benchmark (avgChunkSizeBytes) (chunkCount) (chunkSizeBytes) (delayMillis) (fileSizeBytes) (jitterMillis) (schedulerThreads) Mode Cnt Score Error Units ChunkedTransferBenchmark.downloadChunked N/A 512 1048576 25 N/A 10 8 avgt 3 991.101 ± 173.890 ms/op ChunkedTransferBenchmark.uploadChunked 1048576 N/A N/A 25 536870912 10 8 avgt 3 1087.504 ± 123.857 ms/op ``` ``` PARENT BASELINE (512 MiB) Benchmark (avgChunkSizeBytes) (chunkCount) (chunkSizeBytes) (delayMillis) (fileSizeBytes) (jitterMillis) (schedulerThreads) Mode Cnt Score Error Units ChunkedTransferBenchmark.downloadChunked N/A 512 1048576 25 N/A 10 8 avgt 3 12849.136 ± 1733.063 ms/op ChunkedTransferBenchmark.uploadChunked 1048576 N/A N/A 25 536870912 10 8 avgt 3 11888.109 ± 2591.883 ms/op ``` Closes bazelbuild#29341. PiperOrigin-RevId: 919144974 Change-Id: Iaa0ca8971bd21c879f21c708327b5ddd837ecf1f
…azelbuild#29796) For `--experimental_remote_cache_chunking` implemented in bazelbuild#28437 This PR enables parallel uploads and downloads for chunked files to improve performance. Transport-level concurrency is still bounded by the existing remote cache gRPC connection/concurrency limits. The chunk transfer managers add a fixed per-blob window of 16 to prevent a single large blob from fanning out too aggressively. To avoid issues with batches, uploads and downloads use simple sliding-window style transfer managers. RELNOTES: CDC chunk uploads and downloads can now happen in parallel within a large blob. Benchmarks were rerun on 2026-05-12 with the JMH target `//src/test/java/com/google/devtools/build/lib/remote:ChunkedTransferBenchmark`. The parent baseline was measured by running the same benchmark harness against the parent commit. With the synthetic benchmark of network delays and simulated jitter, the current branch is about 11-13x faster than the parent baseline for these cases. As usual, this synthetic benchmark is not a substitute for real remote-cache measurements. After Change: ``` Benchmark (avgChunkSizeBytes) (chunkCount) (chunkSizeBytes) (delayMillis) (fileSizeBytes) (jitterMillis) (schedulerThreads) Mode Cnt Score Error Units ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 1 avgt 3 67.963 ± 2.039 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 2 avgt 3 67.925 ± 2.257 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 4 avgt 3 67.915 ± 2.497 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 8 avgt 3 67.946 ± 2.863 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 1 avgt 3 61.357 ± 7.385 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 2 avgt 3 61.373 ± 7.920 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 4 avgt 3 61.326 ± 8.220 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 8 avgt 3 61.393 ± 8.376 ms/op ``` Before Change: ``` Benchmark (avgChunkSizeBytes) (chunkCount) (chunkSizeBytes) (delayMillis) (fileSizeBytes) (jitterMillis) (schedulerThreads) Mode Cnt Score Error Units ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 1 avgt 3 812.081 ± 463.666 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 2 avgt 3 812.570 ± 442.021 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 4 avgt 3 811.883 ± 459.534 ms/op ChunkedTransferBenchmark.downloadChunked N/A 32 1024 25 N/A 10 8 avgt 3 812.371 ± 461.231 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 1 avgt 3 740.734 ± 389.653 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 2 avgt 3 742.434 ± 412.117 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 4 avgt 3 742.483 ± 395.466 ms/op ChunkedTransferBenchmark.uploadChunked 1024 N/A N/A 25 32768 10 8 avgt 3 742.509 ± 397.122 ms/op ``` Big File: ``` CURRENT BRANCH (512 MiB) Benchmark (avgChunkSizeBytes) (chunkCount) (chunkSizeBytes) (delayMillis) (fileSizeBytes) (jitterMillis) (schedulerThreads) Mode Cnt Score Error Units ChunkedTransferBenchmark.downloadChunked N/A 512 1048576 25 N/A 10 8 avgt 3 991.101 ± 173.890 ms/op ChunkedTransferBenchmark.uploadChunked 1048576 N/A N/A 25 536870912 10 8 avgt 3 1087.504 ± 123.857 ms/op ``` ``` PARENT BASELINE (512 MiB) Benchmark (avgChunkSizeBytes) (chunkCount) (chunkSizeBytes) (delayMillis) (fileSizeBytes) (jitterMillis) (schedulerThreads) Mode Cnt Score Error Units ChunkedTransferBenchmark.downloadChunked N/A 512 1048576 25 N/A 10 8 avgt 3 12849.136 ± 1733.063 ms/op ChunkedTransferBenchmark.uploadChunked 1048576 N/A N/A 25 536870912 10 8 avgt 3 11888.109 ± 2591.883 ms/op ``` Closes bazelbuild#29341. PiperOrigin-RevId: 919144974 Change-Id: Iaa0ca8971bd21c879f21c708327b5ddd837ecf1f <!-- Thank you for contributing to Bazel! Please read the contribution guidelines: https://bazel.build/contribute.html --> ### Description <!-- Please provide a brief summary of the changes in this PR. --> ### Motivation <!-- Why is this change important? Does it fix a specific bug or add a new feature? If this PR fixes an existing issue, please link it here (e.g. "Fixes bazelbuild#1234"). --> ### Build API Changes <!-- Does this PR affect the Build API? (e.g. Starlark API, providers, command-line flags, native rules) If yes, please answer the following: 1. Has this been discussed in a design doc or issue? (Please link it) 2. Is the change backward compatible? 3. If it's a breaking change, what is the migration plan? --> No ### Checklist - [ ] I have added tests for the new use cases (if any). - [ ] I have updated the documentation (if applicable). ### Release Notes <!-- If this is a new feature, please add 'RELNOTES[NEW]: <description>' here. If this is a breaking change, please add 'RELNOTES[INC]: <reason>' here. If this change should be mentioned in release notes, please add 'RELNOTES: <reason>' here. --> RELNOTES: None Co-authored-by: Tyler French <french.tyler.d@gmail.com>
|
we have a user on 8.7 complaining about upload slowness when enabled chunking cc: @fmeum |
|
@bazel-io fork 8.8.0 |
1 similar comment
|
@bazel-io fork 8.8.0 |
The 8.8.0 cherry-pick tracked by bazelbuild#29870 is needed because users enabling `--experimental_remote_cache_chunking` can still see large CDC blobs transfer chunks serially on this release branch. Backport bazelbuild#29341 so chunked uploads and downloads use a bounded per-blob sliding window while existing remote cache connection limits continue to bound global RPC concurrency. The 8.8.0 branch does not have the newer `GrpcCacheClient.shouldVerifyDownloads()` helper or the split remote Java BUILD targets. Wire whole-blob verification through `RemoteOptions.remoteVerifyDownloads` and keep the benchmark on this branch's coarse remote target layout. Refs bazelbuild#29870. PiperOrigin-RevId: 919144974 Change-Id: Iaa0ca8971bd21c879f21c708327b5ddd837ecf1f (cherry picked from commit 084958b)
The 8.8.0 cherry-pick tracked by bazelbuild#29870 is needed because users enabling `--experimental_remote_cache_chunking` can still see large CDC blobs transfer chunks serially on this release branch. Backport bazelbuild#29341 so chunked uploads and downloads use a bounded per-blob sliding window while existing remote cache connection limits continue to bound global RPC concurrency. The 8.8.0 branch does not have the newer `GrpcCacheClient.shouldVerifyDownloads()` helper or the split remote Java BUILD targets. Wire whole-blob verification through `RemoteOptions.remoteVerifyDownloads` and keep the benchmark on this branch's coarse remote target layout. Refs bazelbuild#29870. PiperOrigin-RevId: 919144974 Change-Id: Iaa0ca8971bd21c879f21c708327b5ddd837ecf1f (cherry picked from commit 084958b)
Fixes bazelbuild#29870. This cherry-picks bazelbuild#29341 to release-8.8.0 because users enabling `--experimental_remote_cache_chunking` can still see large CDC blobs transfer chunks serially on this release branch. The backport lets chunked uploads and downloads use a bounded per-blob sliding window while existing remote cache connection limits continue to bound global RPC concurrency. The 8.8.0 branch does not have the newer `GrpcCacheClient.shouldVerifyDownloads()` helper or the split remote Java BUILD targets. This backport wires whole-blob verification through `RemoteOptions.remoteVerifyDownloads` and keeps the benchmark on this branch's coarse remote target layout. Co-authored-by: Tyler French <french.tyler.d@gmail.com>
Description
For
--experimental_remote_cache_chunkingimplemented in #28437This PR enables parallel uploads and downloads for chunked files to improve performance. Transport-level concurrency is still bounded by the existing remote cache gRPC connection/concurrency limits. The chunk transfer managers add a fixed per-blob window of 16 to prevent a single large blob from fanning out too aggressively.
To avoid issues with batches, uploads and downloads use simple sliding-window style transfer managers. This PR also fixes #29544 by properly copying the data to
outRELNOTES: CDC chunk uploads and downloads can now happen in parallel within a large blob.
Benchmarking
Benchmarks were rerun on 2026-05-12 with the JMH target
//src/test/java/com/google/devtools/build/lib/remote:ChunkedTransferBenchmark. The parent baseline was measured by running the same benchmark harness against the parent commit.With the synthetic benchmark of network delays and simulated jitter, the current branch is about 11-13x faster than the parent baseline for these cases. As usual, this synthetic benchmark is not a substitute for real remote-cache measurements.
After Change:
Before Change:
Big File:
Resolves: #29544