Measure whether query-selected KV paging can work for your model and your workloads, before you build it.
We measured it for ours and the answer was no. The harness, the method and the full negative result are here so the finding can be checked and the measurement repeated elsewhere.
The plan was to run a 262k-token context on a single 24 GB GPU by keeping a small query-selected working set of KV resident and paging the rest from host RAM. Landmark scoring, page table, LRU pool and gather kernel all rest on one assumption: that a small budget of query-selected pages captures most of the attention mass.
Measured on Qwen3.6-27B (a hybrid SSM/attention model; 16 of 65 blocks carry KV) over 9 real long-context work sessions, 450,983 prefill tokens:
| Budget B (% of context) | oracle ceiling | landmark selector | needed |
|---|---|---|---|
| 1% | 0.404 | 0.314 | |
| 2% | 0.507 | 0.409 | ≥ 0.95 |
| 4% | 0.619 | 0.516 | ≥ 0.85 |
The oracle column ranks pages by their true attention mass. It is a perfect selector that cannot actually be built, and at a 2% budget it reaches 0.507 against the 0.95 we need.
The landmark selector tracks that oracle to within about 0.10, so the scoring is working roughly as well as scoring can. The ceiling is the problem. No amount of tuning the score, the budget, the per-layer split or the quota scheme closes a 0.54 gap when perfect selection only reaches 0.51.
Attention in this model is diffuse, and page granularity makes it worse:
- Effective attention support runs 473–4,062 tokens, and the hottest single token holds only 2.4–6.9% of the mass. There is no small hot set to page in.
- The same 2% budget captures 0.42–0.89 of the mass at token granularity but only 0.16–0.76 at page granularity. High-mass tokens are scattered rather than clustered, so each 32-token page drags in ~31 near-irrelevant neighbours.
Our guess at the cause: in a hybrid SSM/attention model the recurrent backbone handles local structure while the sparse attention layers do global mixing, so diffuse attention may be what those layers are there for. Most published KV-sparsity work is measured on dense transformers, and this result should not be generalised to them. Run a different model and the numbers may look completely different, which is why the harness is published.
Full results: results/PROFILE_REPORT.md. Post-mortem: CLOSEOUT.md.
A dense-model control run is planned, to establish that the harness reports high coverage where high coverage is expected. It has not been run yet; the numbers will go here when it has.
For each sampled decode step and every attention layer, the harness computes the exact attention distribution of the current query over all cached keys, softmax(kq_scale · q·Kᵀ). It is a single-vector matmul, cheap even at 128k. Token-level mass is then aggregated into fixed-size pages, and the harness reports:
| Metric | Meaning |
|---|---|
coverage@B |
mass captured by the top-B pages ranked by true mass — the oracle ceiling |
landmark-coverage@B |
same, ranked by landmark score (q · page-mean-of-keys, max over the GQA group) — the realistic number |
| selection overlap | Jaccard of the selected page set across decode steps, which predicts LRU hit rate and miss traffic |
| per-layer profile | which layers are dense, and so candidates to keep fully resident |
| head starvation | per-kv-head quota vs global top-B |
| dim ablation | full head-dim scoring vs unrotated dims only, for partial-RoPE models |
Sinks (first S tokens) and the recent window (last R) are masked before scoring. They are free under any paging scheme and would otherwise inflate the result.
Requires Python 3.10+, NumPy, and a CUDA build of llama.cpp.
# 1. patch and build the capture tool (see patches/README.md for detail)
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp
git checkout fb92d8f1873c96ec63f9c59721d58a55bf46d441
git apply /path/to/kv-sparsity-profiler/patches/kvpage-capture.patch
cmake -B build -DGGML_CUDA=ON && cmake --build build -j --target llama-kvpage-capture# 2. capture one trace (prompt.txt = any long prompt; see the ChatML note below)
export KVPROF_ROOT=$PWD LLAMA_BUILD=$PWD/llama.cpp/build
./llama.cpp/build/bin/llama-kvpage-capture \
-m /path/to/model.gguf -c 40960 -ngl 99 \
--prompt-file prompt.txt -n 256 \
--sampled-steps 0,8,16,24,32,40,48,56,64,72,80,88,96,104,112,120,128,136,144,152,160,168,176,184,192,200,208,216,224,232,240,248 \
--out-dir profiling-out/mytrace \
--cache-type-k q8_0 --cache-type-v q8_0 --flash-attn on --ignore-eos# 3. convert to the analysis format, then compute metrics and print the tables
python3 harness/capture.py convert --out-dir profiling-out/mytrace --log capture.log
python3 harness/run_profile.py
python3 harness/make_report.pyChatML note. If your prompt is a completed conversation the model emits EOS immediately and you get no usable decode steps. Drop the trailing turn and append your model's assistant-turn cue (for ChatML: <|im_end|>\n<|im_start|>assistant\n). --ignore-eos is a safety net rather than a substitute; forcing generation past a natural EOS produces filler whose attention is unrepresentative.
Verify your capture before trusting any numbers:
python3 harness/verify_reference.py --out-dir profiling-out/mytrace --log capture.logThis recomputes attention from the captured keys and compares it against the model's own softmax, dumped from a flash-attention-off reference run. It should agree to ~1e-6; the gate is 1e-3. It caught a wrong KV-cache read path in our own harness, which disagreed at 0.231.
- Use a per-kv-head quota rather than global top-B. 0.2254 vs 0.2035. Global top-B lets one head monopolise the budget and starve the others.
- Score on all head dims. With partial RoPE it is tempting to score only the position-invariant dims, since that is cheaper and avoids the rotation question. Measured, it is worse: 0.3740 vs 0.4037. The rotated dims carry real ranking signal.
- Mask sinks and the recent window before scoring, and express the recent window as a fraction of context rather than an absolute token count. An absolute R tuned for a long deployment context hands a shorter profiling run a far larger budget relative to the eligible population, which inflates coverage. In our case by roughly 5×, enough to turn a NO-GO into a false GO.
Whether attention is concentrated enough for query-selected paging is model-dependent, and we have one data point on an unusual architecture. If you run this on a dense transformer, an MoE, another hybrid, or at a different context length, please open a PR or issue with:
- model name, quantisation,
n_head_q/n_head_kv/head_dim, and how many layers carry KV; - the
coverage@Bandlandmark-coverage@Btable at B ∈ {0.5, 1, 2, 4}%, with your S and R; - context lengths measured and roughly what the workload was (document QA, agentic tool output, chat, code…);
- the per-layer table if you have it.
Aggregate tables only, please: no prompts, traces or captured tensors. Ours are not published for the same reason.
A dense model showing coverage well above 0.5 at 2% would be a useful counterweight to this result. It would suggest the paging design is sound and that we picked the wrong architecture for it.
harness/ Python: capture driver, metrics, landmark scoring, report generation,
independent cross-check, unit tests (55, all passing)
patches/ llama.cpp instrumentation as a .patch against a pinned base commit
results/ full measured results — aggregate tables only, no trace content
(dense-model control run pending — see note above)
docs/ method and the decision log from the original project
CLOSEOUT.md project post-mortem: thesis, verdict, what was ruled out
The paging implementation this profiling was meant to justify was never built. The gate failed first, which is why the gate went before the build. What is here is the measurement apparatus and the result.
The harness is research code. It was written to answer one question on one model and is published as-is rather than polished into a library. The metrics module has unit tests; the capture instrumentation is throwaway and pinned to one upstream commit.
MIT. See LICENSE.