Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
64 commits
Select commit Hold shift + click to select a range
b521be0
kv-stream: add block-granular residency planner
RaymondHuang210129 Aug 21, 2026
40d1772
kv-stream: add extent reserve and shrink hysteresis
RaymondHuang210129 Aug 21, 2026
47709b4
kv-stream: build block-level cache regions
RaymondHuang210129 Aug 21, 2026
63c1f8d
kv-stream: define exact block attention merge
RaymondHuang210129 Aug 21, 2026
45e4629
kv-stream: add guarded CUDA transfer harness
RaymondHuang210129 Aug 21, 2026
9b9a197
kv-stream: add pinned host CUDA runtime
RaymondHuang210129 Aug 21, 2026
775b63a
kv-stream: register CUDA stream buffers safely
RaymondHuang210129 Aug 21, 2026
aaff66b
kv-stream: add bounded CUDA stage transfers
RaymondHuang210129 Aug 21, 2026
9d050a9
kv-stream: execute one staged attention block
RaymondHuang210129 Aug 21, 2026
4093b17
kv-stream: merge exact streamed attention blocks
RaymondHuang210129 Aug 21, 2026
290c183
kv-stream: pack real cache views by stride
RaymondHuang210129 Aug 21, 2026
b91e010
kv-stream: commit cache rows to pinned storage
RaymondHuang210129 Aug 21, 2026
bc37a3b
server: wire experimental block KV streaming
RaymondHuang210129 Aug 21, 2026
832ce53
kv-stream: partition resident pages per layer
RaymondHuang210129 Aug 21, 2026
07302b6
kv-stream: retain stable cache pages on CUDA
RaymondHuang210129 Aug 21, 2026
9eac316
kv-stream: preserve ordinary attention for one-page caches
RaymondHuang210129 Aug 21, 2026
0378b6f
kv-stream: add deadline-driven prefetch scheduler
RaymondHuang210129 Aug 21, 2026
ef81bdc
kv-stream: pipeline staged attention pages asynchronously
RaymondHuang210129 Aug 21, 2026
5d81908
kv-stream: adapt resident and ring partition from runtime pressure
RaymondHuang210129 Aug 21, 2026
9e17817
kv-stream: move resident boundary inside fixed CUDA pool
RaymondHuang210129 Aug 21, 2026
f2c055a
kv-stream: prefetch pages across attention layers
RaymondHuang210129 Aug 21, 2026
6ed33be
kv-stream: adapt shared ring from runtime feedback
RaymondHuang210129 Aug 21, 2026
f96d8bf
kv-stream: bound attention reduction workspace
RaymondHuang210129 Aug 21, 2026
22814d0
kv-stream: fuse resident attention spans
RaymondHuang210129 Aug 22, 2026
748392d
kv-stream: coalesce adjacent streamed pages
RaymondHuang210129 Aug 22, 2026
61c6b04
kv-stream: defer decode attention reduction
RaymondHuang210129 Aug 22, 2026
0b87564
Revert "kv-stream: defer decode attention reduction"
RaymondHuang210129 Aug 22, 2026
a8b3c0f
kv-stream: coalesce adjacent page uploads
RaymondHuang210129 Aug 22, 2026
2888aa6
Revert "kv-stream: coalesce adjacent page uploads"
RaymondHuang210129 Aug 22, 2026
6d5350a
kv-stream: track resident dirty pages precisely
RaymondHuang210129 Aug 22, 2026
0432759
Revert "kv-stream: track resident dirty pages precisely"
RaymondHuang210129 Aug 22, 2026
7f8d149
kv-stream: rate-limit residency repartition
RaymondHuang210129 Aug 22, 2026
4ba9c78
kv-stream: use pool remainder for prefetch slots
RaymondHuang210129 Aug 22, 2026
63fba29
kv-stream: preserve token-major staging locality
RaymondHuang210129 Aug 22, 2026
cb71402
kv-stream: overlap copies with bounded attention spans
RaymondHuang210129 Aug 22, 2026
666c786
Revert "kv-stream: overlap copies with bounded attention spans"
RaymondHuang210129 Aug 22, 2026
9a70ebb
kv-stream: use write-combined host cache
RaymondHuang210129 Aug 22, 2026
91653a5
cuda: accelerate adaptive KV prefill staging
RaymondHuang210129 Aug 22, 2026
4500335
cuda: accelerate streamed KV prefill attention
RaymondHuang210129 Aug 22, 2026
0c9cadf
kv stream: size transfer ring to layer working set
RaymondHuang210129 Aug 22, 2026
f3473db
kv stream: bound decode residency by transfer ring
RaymondHuang210129 Aug 23, 2026
0cec723
kv stream: migrate resident pages across decode layouts
RaymondHuang210129 Aug 23, 2026
aace8e9
kv stream: support serial prompts and cache restore
RaymondHuang210129 Aug 23, 2026
388264c
kv stream: bound decode layout by layer capacity
RaymondHuang210129 Aug 23, 2026
dfd0899
kv stream: surface repartition events at default verbosity
RaymondHuang210129 Aug 23, 2026
bf27cde
kv stream: allow multi-wave small-ring layouts
RaymondHuang210129 Aug 23, 2026
553ef79
kv stream: batch page uploads and bound feedback
RaymondHuang210129 Aug 24, 2026
cf20cc7
cuda: keep adaptive KV pool device-resident under UVM
RaymondHuang210129 Aug 25, 2026
afc6b4a
bench: publish adaptive KV streaming workflow
RaymondHuang210129 Aug 27, 2026
58cd05f
bench: automate adaptive KV context sweep
RaymondHuang210129 Aug 28, 2026
400b0ff
docs: link adaptive KV streaming article
RaymondHuang210129 Aug 28, 2026
4660373
fix(cuda): avoid write-combined KV storage on Windows
RaymondHuang210129 Aug 29, 2026
5a50cbe
kv stream: generalize cache types and tune prefetch
RaymondHuang210129 Aug 29, 2026
90c91f7
cuda: allow wide adaptive KV prefill batches
RaymondHuang210129 Aug 29, 2026
c260fe5
cuda: scale KV stream scratch to attention mode
RaymondHuang210129 Aug 29, 2026
d1f53af
bench: make adaptive KV batch sizes configurable
RaymondHuang210129 Aug 29, 2026
6cf5610
cuda: bound vector KV stream query workspace
RaymondHuang210129 Aug 29, 2026
b742479
docs: describe adaptive KV ubatch support
RaymondHuang210129 Aug 29, 2026
5c8dee6
make benchmark_kv_stream.py compatible with Windows (#1)
linuxtextadventurer Sep 6, 2026
017a512
baseline: mirror production CUDA unified-memory build
RaymondHuang210129 Aug 20, 2026
5539185
feat(server): LLAMA_KV_STREAM_DEVICE device filter + opt-in parallel-…
Sep 11, 2026
20a949b
test(arm117): rig scripts — boot/cell-A parity/cell-C long-context/sn…
Sep 11, 2026
65644fb
fix(arm117): parity script reads reasoning_content; rpc -d CUDA1; sna…
Sep 11, 2026
ef3b119
fix(cuda): gate kv-stream FA dispatch on full fits() predicate — mult…
Sep 11, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
16 changes: 1 addition & 15 deletions .devops/cuda.Dockerfile
Original file line number Diff line number Diff line change
Expand Up @@ -13,18 +13,6 @@ ARG APP_REVISION=N/A

ARG NODE_VERSION=24

FROM docker.io/node:$NODE_VERSION AS web

ARG APP_VERSION

WORKDIR /app/tools/ui

COPY tools/ui/package.json tools/ui/package-lock.json ./
RUN npm ci

COPY tools/ui/ ./
RUN LLAMA_BUILD_NUMBER="$APP_VERSION" npm run build

FROM ${BASE_CUDA_DEV_CONTAINER} AS build

ARG GCC_VERSION
Expand All @@ -40,12 +28,10 @@ WORKDIR /app

COPY . .

COPY --from=web /app/tools/ui/dist tools/ui/dist

RUN if [ "${CUDA_DOCKER_ARCH}" != "default" ]; then \
export CMAKE_ARGS="-DCMAKE_CUDA_ARCHITECTURES=${CUDA_DOCKER_ARCH}"; \
fi && \
cmake -B build -DGGML_NATIVE=OFF -DGGML_CUDA=ON -DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON -DLLAMA_BUILD_TESTS=OFF ${CMAKE_ARGS} -DCMAKE_EXE_LINKER_FLAGS=-Wl,--allow-shlib-undefined . && \
cmake -B build -DGGML_NATIVE=OFF -DGGML_CUDA=ON -DGGML_CUDA_NO_VMM=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DGGML_BACKEND_DL=ON -DGGML_CPU_ALL_VARIANTS=ON -DLLAMA_BUILD_TESTS=OFF -DLLAMA_BUILD_UI=OFF ${CMAKE_ARGS} -DCMAKE_EXE_LINKER_FLAGS=-Wl,--allow-shlib-undefined . && \
cmake --build build --config Release -j$(nproc)

RUN mkdir -p /app/lib && \
Expand Down
1 change: 1 addition & 0 deletions .gitignore
Original file line number Diff line number Diff line change
Expand Up @@ -130,6 +130,7 @@ flake.lock
/hellaswag_val_full.txt
/winogrande-debiased-eval.csv
/wikitext-2-raw/
/benchmarks/results/

# Test models for lora adapters

Expand Down
90 changes: 90 additions & 0 deletions README.md
Original file line number Diff line number Diff line change
@@ -1,3 +1,93 @@
# Adaptive KV Streaming for llama.cpp

This branch adds an experimental, block-granular KV cache streaming path to the CUDA `llama-server`. It is intended for running long contexts when model weights leave too little VRAM for the complete KV cache.

With `--kv-stream-stage-mib N`, the authoritative KV tensors are stored in pinned host memory while a bounded CUDA pool is shared by resident KV pages and a transfer ring. The runtime adapts that split as the context grows: it keeps as many pages resident as the budget allows, reclaims resident space for staging when more streaming is required, and prefetches later layers while the current layer computes. This avoids relying on uncontrolled Unified Memory page thrashing and preserves exact attention over the full context.

Detailed project story, design, implementation, and benchmark results are in
[Running Qwen 27B on 16G VRAM with Full Context Length: Building Adaptive KV Cache Streaming for llama.cpp](https://medium.com/@raymond860909/running-qwen-27b-on-16g-vram-with-full-context-length-building-adaptive-kv-cache-streaming-for-bf1e819116e9).

> [!WARNING]
> This is research code optimized and production-validated primarily for an RTX 5070 Ti with 16 GB VRAM, `unsloth/Qwen3.8-27B-GGUF` `UD-Q3_K_XL`, a 262144-token context, Flash Attention, a Q8_0 K cache, a Q4_0 V cache, and one server slot.
> CUDA correctness tests cover every KV type currently accepted by the CLI, including native and F16-conversion fallback paths. Production performance for other models, KV combinations, parallel slots, and non-CUDA backends is not yet broadly characterized.

## Build the modified server

Install a C++ compiler, CMake, and the CUDA toolkit, then run this command from the repository root:

```bash
cmake -S . -B build -DGGML_CUDA=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DCMAKE_BUILD_TYPE=Release && cmake --build build --config Release --target llama-server -j
```

The executable is created at `build/bin/llama-server`.

Example using the tested cache configuration:

```bash
./build/bin/llama-server \
--model /path/to/model.gguf \
--ctx-size 262144 \
-fa on \
-ctk q8_0 \
-ctv q4_0 \
-ngl all \
-b 512 \
-ub 512 \
-np 1 \
--kv-stream-stage-mib 2304
```

The best value for `--kv-stream-stage-mib` depends on the model, context capacity, GPU, and other VRAM consumers. Start conservatively and increase it while checking startup and peak VRAM use.

### Batch and micro-batch sizes

`-b` sets the logical prompt batch size and `-ub` sets the largest physical batch submitted to one graph. This branch no longer requires `256/256`; `-ub` may be any positive value no larger than `-b`.

The Qwen3.8 MMA prefill path processes the full physical batch and allocates only the partial workspace that kernel actually emits. Generic vector and F16-conversion fallback paths use a bounded 256-query workspace: each staged KV span is consumed by all query tiles before its ring slot is released, so wider micro-batches do not multiply KV host-to-device transfers.

The Q8_0/Q4_0 Qwen configuration has been exercised with `b/ub` values `256/256`, `512/512`, `768/512`, and `1024/1024`, including non-divisible final micro-batches. A 122880-token production-shaped run at `512/512` completed with adaptive streaming active. Wider values can require more graph and accumulator memory, so validate them on the target GPU.

### Optional Unified Memory for model weights

Adaptive KV streaming works with or without Unified Memory. Leave `GGML_CUDA_ENABLE_UNIFIED_MEMORY` unset for ordinary CUDA device allocations. To make GPU-offloaded model buffers CUDA managed allocations, launch the same server with the environment variable enabled:

```bash
GGML_CUDA_ENABLE_UNIFIED_MEMORY=1 \
./build/bin/llama-server \
--model /path/to/model.gguf \
--ctx-size 262144 \
-fa on \
-ctk q8_0 \
-ctv q4_0 \
-ngl all \
-b 512 \
-ub 512 \
-np 1 \
--kv-stream-stage-mib 2304
```

With this flag, CUDA-backed model buffers, including GPU-offloaded weights, are allocated with `cudaMallocManaged` and their pages can migrate between VRAM and host memory. The adaptive resident-page and transfer-ring pool is intentionally different: it is still allocated with `cudaMalloc`, so that fixed-size pool remains physically allocated in VRAM instead of becoming managed memory. UVM is therefore optional for this branch and does not change the KV streaming pool into pageable storage.

## Recreate the benchmark graph

The benchmark driver automatically selects the largest practical adaptive KV pool for each configured context capacity, sweeps from 8K through the requested maximum, and generates the CSV, PNG, and SVG results:

```bash
python3 -m pip install matplotlib

python3 benchmarks/benchmark_kv_stream.py \
--model /path/to/model.gguf \
--max-context 192K \
--batch-size 512 \
--ubatch-size 512
```

The only required arguments are the model GGUF and maximum context. See [benchmarks/README.md](benchmarks/README.md) for the pool-probing algorithm, generated files, optional settings, and resumable output directories.

---

## Upstream llama.cpp README

# llama.cpp

![llama](https://raw.githubusercontent.com/ggml-org/llama.brand/refs/heads/master/cover/llama-cpp/cover-llama-cpp-dark.svg)
Expand Down
2 changes: 2 additions & 0 deletions arm117-artifacts/Aoff-np1/gpu-mem.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
0, 14149 MiB, 16311 MiB
1, 9080 MiB, 12288 MiB
3 changes: 3 additions & 0 deletions arm117-artifacts/Aoff-np1/gpu-procs.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
1311054 /tmp/opencode/arm117/port/build/bin/llama-server -m /mnt/SSD/Qwen3.8-27B-UD-Q5_K_M.gguf --rpc 127.0.0.1:50052 -ts 27,38 -ngl 99 --rope-scaling yarn --rope-scale 5 --yarn-orig-ctx 32768 -fa on -ctk q8_0 -ctv q5_1 -ctkd q8_0 -ctvd q5_1 --no-kv-unified --cache-prompt --cache-reuse 64 --cache-idle-slots --cache-ram 1024 --ubatch-size 512 --cont-batching -np 1 -c 65536 --parallel-ctx-threshold 100000 --spec-type draft-mtp --prio-batch 1 --kv-stream-stage-mib 0 --jinja --host 0.0.0.0 --port 8080 --metrics --slots --log-verbosity 4
1311032 /tmp/opencode/arm117/port/build/bin/ggml-rpc-server --host 127.0.0.1 --port 50052 -d CUDA1
1311054 /tmp/opencode/arm117/port/build/bin/llama-server -m /mnt/SSD/Qwen3.8-27B-UD-Q5_K_M.gguf --rpc 127.0.0.1:50052 -ts 27,38 -ngl 99 --rope-scaling yarn --rope-scale 5 --yarn-orig-ctx 32768 -fa on -ctk q8_0 -ctv q5_1 -ctkd q8_0 -ctvd q5_1 --no-kv-unified --cache-prompt --cache-reuse 64 --cache-idle-slots --cache-ram 1024 --ubatch-size 512 --cont-batching -np 1 -c 65536 --parallel-ctx-threshold 100000 --spec-type draft-mtp --prio-batch 1 --kv-stream-stage-mib 0 --jinja --host 0.0.0.0 --port 8080 --metrics --slots --log-verbosity 4
1 change: 1 addition & 0 deletions arm117-artifacts/Aoff-np1/health.json
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
{"status":"ok"}
Empty file.
1 change: 1 addition & 0 deletions arm117-artifacts/Aoff-np1/slots.json
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
[{"id":0,"n_ctx":65536,"speculative":true,"is_processing":false,"id_task":33,"n_prompt_tokens":2260,"n_prompt_tokens_processed":0,"n_prompt_tokens_cache":0,"params":{"seed":4294967295,"temperature":0.0,"dynatemp_range":0.0,"dynatemp_exponent":1.0,"top_k":20,"top_p":0.949999988079071,"min_p":0.05000000074505806,"top_n_sigma":-1.0,"xtc_probability":0.0,"xtc_threshold":0.10000000149011612,"typical_p":1.0,"repeat_last_n":64,"repeat_penalty":1.0,"presence_penalty":0.0,"frequency_penalty":0.0,"dry_multiplier":0.0,"dry_base":1.75,"dry_allowed_length":2,"dry_penalty_last_n":64,"mirostat":0,"mirostat_tau":5.0,"mirostat_eta":0.10000000149011612,"adaptive_target":-1.0,"adaptive_decay":0.8999999761581421,"max_tokens":96,"n_predict":96,"n_keep":0,"n_discard":0,"ignore_eos":false,"stream":false,"n_probs":0,"min_keep":0,"chat_format":"peg-native","reasoning_format":"deepseek","reasoning_in_content":false,"generation_prompt":"<|im_start|>assistant\n<think>\n","samplers":["penalties","dry","top_n_sigma","top_k","typ_p","top_p","min_p","xtc","temperature"],"speculative.types":"none,draft-mtp","timings_per_token":false,"post_sampling_probs":false,"backend_sampling":false,"lora":[]},"next_token":[{"has_next_token":false,"has_new_line":false,"n_remain":-1,"n_decoded":0}]}]
2 changes: 2 additions & 0 deletions arm117-artifacts/Aon-np1/gpu-mem.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,2 @@
0, 15184 MiB, 16311 MiB
1, 9142 MiB, 12288 MiB
3 changes: 3 additions & 0 deletions arm117-artifacts/Aon-np1/gpu-procs.txt
Original file line number Diff line number Diff line change
@@ -0,0 +1,3 @@
1312993 /tmp/opencode/arm117/port/build/bin/llama-server -m /mnt/SSD/Qwen3.8-27B-UD-Q5_K_M.gguf --rpc 127.0.0.1:50052 -ts 27,38 -ngl 99 --rope-scaling yarn --rope-scale 5 --yarn-orig-ctx 32768 -fa on -ctk q8_0 -ctv q5_1 -ctkd q8_0 -ctvd q5_1 --no-kv-unified --cache-prompt --cache-reuse 64 --cache-idle-slots --cache-ram 1024 --ubatch-size 512 --cont-batching -np 1 -c 65536 --parallel-ctx-threshold 100000 --spec-type draft-mtp --prio-batch 1 --kv-stream-stage-mib 2048 --jinja --host 0.0.0.0 --port 8080 --metrics --slots --log-verbosity 4
1312976 /tmp/opencode/arm117/port/build/bin/ggml-rpc-server --host 127.0.0.1 --port 50052 -d CUDA1
1312993 /tmp/opencode/arm117/port/build/bin/llama-server -m /mnt/SSD/Qwen3.8-27B-UD-Q5_K_M.gguf --rpc 127.0.0.1:50052 -ts 27,38 -ngl 99 --rope-scaling yarn --rope-scale 5 --yarn-orig-ctx 32768 -fa on -ctk q8_0 -ctv q5_1 -ctkd q8_0 -ctvd q5_1 --no-kv-unified --cache-prompt --cache-reuse 64 --cache-idle-slots --cache-ram 1024 --ubatch-size 512 --cont-batching -np 1 -c 65536 --parallel-ctx-threshold 100000 --spec-type draft-mtp --prio-batch 1 --kv-stream-stage-mib 2048 --jinja --host 0.0.0.0 --port 8080 --metrics --slots --log-verbosity 4
1 change: 1 addition & 0 deletions arm117-artifacts/Aon-np1/health.json
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
{"status":"ok"}
Empty file.
1 change: 1 addition & 0 deletions arm117-artifacts/Aon-np1/slots.json
Original file line number Diff line number Diff line change
@@ -0,0 +1 @@
[{"id":0,"n_ctx":65536,"speculative":true,"is_processing":false,"id_task":20,"n_prompt_tokens":2260,"n_prompt_tokens_processed":0,"n_prompt_tokens_cache":0,"params":{"seed":4294967295,"temperature":0.0,"dynatemp_range":0.0,"dynatemp_exponent":1.0,"top_k":20,"top_p":0.949999988079071,"min_p":0.05000000074505806,"top_n_sigma":-1.0,"xtc_probability":0.0,"xtc_threshold":0.10000000149011612,"typical_p":1.0,"repeat_last_n":64,"repeat_penalty":1.0,"presence_penalty":0.0,"frequency_penalty":0.0,"dry_multiplier":0.0,"dry_base":1.75,"dry_allowed_length":2,"dry_penalty_last_n":64,"mirostat":0,"mirostat_tau":5.0,"mirostat_eta":0.10000000149011612,"adaptive_target":-1.0,"adaptive_decay":0.8999999761581421,"max_tokens":96,"n_predict":96,"n_keep":0,"n_discard":0,"ignore_eos":false,"stream":false,"n_probs":0,"min_keep":0,"chat_format":"peg-native","reasoning_format":"deepseek","reasoning_in_content":false,"generation_prompt":"<|im_start|>assistant\n<think>\n","samplers":["penalties","dry","top_n_sigma","top_k","typ_p","top_p","min_p","xtc","temperature"],"speculative.types":"none,draft-mtp","timings_per_token":false,"post_sampling_probs":false,"backend_sampling":false,"lora":[]},"next_token":[{"has_next_token":false,"has_new_line":false,"n_remain":-1,"n_decoded":0}]}]
Loading
Loading