Name and Version
$ /app/llama cli --version
version: 0.4.0-dev (build 10909, commit 44b47fd77)
built with GNU 12.3.0 for Linux x86_64
Operating systems
Linux
GGML backends
CUDA
Hardware
Ryzen 7950x on x670 MB 192MB RAM 5x RTX3090 (3 of them via USB4, but this detail is not relevant to the bug)
Models
unsloth/Qwen3.8-Flash-Next-GGUF UD-Q4_K_XL -ngl 99, layer split across 5 GPUs,
-ot per_layer_token_embd.weight=CPU -c 524288 --parallel 2
Problem description & steps to reproduce
This is a likely a duplicate of a previous report erroneusly closed by reporter #27856, I can't reopen it so I'm submitting this one with all my detail. The issue #27856 is not a duplicate of #27856, that one appears to be fixed (it reports much more severe deterioration).
When runnig the PR #27742 branch I initaly noticed substantial performance degradation about halving the decode speed starting at 54 tok/s every 50k context size or so.
Using the latest master df03399 this is much improved to:
- 37.2tok/s to 50.8tok/s at 42k ctx (+37%)
- 22.5tok/s 48.0tok/s at 111k (+113%),
- 11tok/s to 41.8 tok/s at 250k
but still deteriorating to "nighly speed" on my hardware at the bottom of the context.
Please note both generation is affected (severely) as well aas prefill (a lot less). Prefill almost halves the speed at max 260k context so far less impactful.
I let my LLM run over the code to see what improvements it finds and by the morning indeed it managed to make changes that helped a lot to:
- 37.2tok/s to 50.8tok/s at 42k ctx (+37%)
- 22.5tok/s to 48.0tok/s at 111k (+113%)
- 11tok/s to 41.8tok/s at 250k (+280%)
It performed some correctness validation that suggests we didn't brake anything, but I'm far from claiming a "fix" at this stage (have yet to validate properly).
The diff is almost 2k lines over 7 patches so it is too big to include here. I will provide it upon request.
First Bad Commit
it was built from pull/27742/head
Relevant log output
I also include a brief LLM written sumamry of changes made below:
qwen4exp decode depth-decay — LLM generated "fix" summary
Arch: qwen4exp (Qwen3.8-Flash-Next, 125B-A6B; 12 QSA sparse-attn + 36 Gated-DeltaNet layers, indexer.top_k=2048, compress_ratio=4).
Problem: single-stream decode slows ~linearly with context depth even though QSA is supposed to bound attention to a fixed top-k budget. Root cause (nsys-profiled): the top-k path was running as a dense attention scan over all n_kv, plus per-token re-pooling of the whole cache and MBs of host→device + P2P input traffic.
Base: llama.cpp master df03399b8, branch levers. Build/run: CUDA 12.4.1, sm_86, image llama.cpp:lever2b-cuda124; KV_TYPE=f16 is mandatory (quantized KV can't reach the sparse MMA kernel at batch 1 on Ampere).
Results (single-stream decode, t/s, 5×RTX 3090)
37.2tok/s to 50.8tok/s at 42k ctx (+37%)
22.5tok/s to 48.0tok/s at 111k (+113%)
11tok/s to 41.8tok/s at 250k (+280%)
(shallow essentially unchanged: 54.4 → 56.5. Prefill improves as a side effect: 609 → 716 t/s @111k.)
Changes (7 patches, llama-levers/000*.patch)
lever1 src/models/qwen4exp.cpp | 9 +++++---- (+ incidental Dockerfile pip fix)
lever3 src/llama-memory-hybrid-idx.{cpp,h} | 99 ++++++ block-granular top-k
src/models/qwen4exp.cpp | 72 ++++
sparse256 ggml/src/ggml-cuda/fattn-mma-f16.cuh | 3 + whitelist DKQ=DV=256 MMA
lever2 src/llama-memory-hybrid-idx.{cpp,h} | 590 ++++++ persistent pooled block-key cache
src/llama-kv-cache.{cpp,h}, qwen4exp.cpp
lever2b src/llama-memory-hybrid-idx.{cpp,h} | 59 ++++ device-resident block-member cells
src/llama-kv-cache.{cpp,h}, qwen4exp.cpp
lever1 — enable the sparse-FA path (the big one). The enablement PR #27742 shipped top-k as a dense mask with a // TODO: enable sparse attention. Un-TODO it: pass top_k->ne[0] as n_kv_max so flash-attn treats the finite mask entries as a sparse K/V set — O(top_k) instead of O(n_kv). Drops attention from ~11 ms to 0.22 ms/token @250K.
- // TODO: enable sparse attention when we are ready
- ggml_tensor * cur = build_attn_mha(q, k, v, nullptr, kq_mask_top_k, nullptr, nullptr, 0, kq_scale, il);
+ // LEVER1: passing top_k->ne[0] as n_kv_max -> sparse K/V set (deepseek4/GLM-DSA pattern, #27970)
+ ggml_tensor * cur = build_attn_mha(q, k, v, nullptr, kq_mask_top_k, nullptr, nullptr, top_k->ne[0], kq_scale, il);
sparse256 — instantiate the kernel for this head shape. may_use_sparse only whitelisted the MLA shapes; qwen4exp is DKQ=DV=256, ncols2=8 and that MMA template instance already exists. One clause:
- return (DKQ == 512 && DV == 512 && ncols1 == 1 && ncols2 == 8) ||
+ return (DKQ == 256 && DV == 256 && ncols1 == 1 && ncols2 == 8) ||
+ (DKQ == 512 && DV == 512 && ncols1 == 1 && ncols2 == 8) ||
(DKQ == 576 && DV == 512 && ncols1 == 1 && ncols2 == 16);
lever3 — block-granular top-k. Scores are block-constant and the budget is whole blocks, so select blocks first and expand only the ~513 selected blocks to cell indices (via blk_cells), instead of expanding block scores to all n_kv cells and sorting those. Removes the O(n_kv) expansion get_rows + one O(n_kv) mask add and shrinks the sort by compress_ratio. Fully-future blocks get −inf bias; the incomplete tail block's cells ride a small per-token side input.
lever2 — persistent pooled block-key cache. The indexer re-gathered, mean-pooled, RMS-normed and re-RoPEd all cached blocks every token though only the newest changed. Add a third per-QSA-layer buffer of pooled+normed+roped block keys; each ubatch recomputes only blocks it dirtied (host-side structural compare of member cells — also catches defrag/rollback/seq-ops with no invalidation hooks) and ggml_set_rows them in. Dirty count is pow2-bucketed (≥4) for CUDA-graph topology stability.
lever2b — device-resident block-member cells. After 1–3 the residual depth cost was input traffic (~3.1 MB/token HtoD + ~13.2 MB/token P2P for blk_cells/mask). Store each block's member cells in the block cache's V tensor (f32; ggml_set_rows takes no i32) and gather the top-k expansion on-device — removes the blk_cells term entirely.
Validity
Every patch was A/B'd against the immediately-prior build and kept only if it passed all gates, all at temperature 0:
- Exact needle-in-haystack retrieval at 63k and 111k (a random code string embedded mid-context must be reproduced verbatim) — guards against the sparse path silently dropping the right K/V.
- Multi-turn prefix reuse + rewind and 2 concurrent slots — guards against the persistent block-key cache going stale across defrag/rollback/seq ops.
- Throughput read from
timings.predicted_per_second on POST /completion, cache_prompt:false, n_predict:48, ignore_eos:true.
A comparable dense model on the same rig (arch=qwen35, plain GQA, no indexer) stays flat at these depths, confirming the decay was QSA/indexer-specific, not KV bandwidth.
Status / upstream
Upstream #28012 (this exact CUDA decay) is closed. Patches are self-contained and PR-able in order: lever1+sparse256 (smallest), then lever3, then lever2+lever2b (shared machinery). Remaining depth losses: host-side O(n_kv) grouping scan in set_input_qsa; per-GPU f16 mask fan-out; flash_attn_mask_to_sparse_indices re-scanning the mask per layer (an indices-in FA API would remove both). Full detail + nsys attribution in new-issue.md / issue-speed.md §9.
Name and Version
$ /app/llama cli --version
version: 0.4.0-dev (build 10909, commit 44b47fd77)
built with GNU 12.3.0 for Linux x86_64
Operating systems
Linux
GGML backends
CUDA
Hardware
Ryzen 7950x on x670 MB 192MB RAM 5x RTX3090 (3 of them via USB4, but this detail is not relevant to the bug)
Models
unsloth/Qwen3.8-Flash-Next-GGUF UD-Q4_K_XL -ngl 99, layer split across 5 GPUs,
-ot per_layer_token_embd.weight=CPU -c 524288 --parallel 2
Problem description & steps to reproduce
This is a likely a duplicate of a previous report erroneusly closed by reporter #27856, I can't reopen it so I'm submitting this one with all my detail. The issue #27856 is not a duplicate of #27856, that one appears to be fixed (it reports much more severe deterioration).
When runnig the PR #27742 branch I initaly noticed substantial performance degradation about halving the decode speed starting at 54 tok/s every 50k context size or so.
Using the latest master df03399 this is much improved to:
but still deteriorating to "nighly speed" on my hardware at the bottom of the context.
Please note both generation is affected (severely) as well aas prefill (a lot less). Prefill almost halves the speed at max 260k context so far less impactful.
I let my LLM run over the code to see what improvements it finds and by the morning indeed it managed to make changes that helped a lot to:
It performed some correctness validation that suggests we didn't brake anything, but I'm far from claiming a "fix" at this stage (have yet to validate properly).
The diff is almost 2k lines over 7 patches so it is too big to include here. I will provide it upon request.
First Bad Commit
it was built from pull/27742/head
Relevant log output
I also include a brief LLM written sumamry of changes made below:
qwen4exp decode depth-decay — LLM generated "fix" summary
Arch:
qwen4exp(Qwen3.8-Flash-Next, 125B-A6B; 12 QSA sparse-attn + 36 Gated-DeltaNet layers,indexer.top_k=2048,compress_ratio=4).Problem: single-stream decode slows ~linearly with context depth even though QSA is supposed to bound attention to a fixed top-k budget. Root cause (nsys-profiled): the top-k path was running as a dense attention scan over all
n_kv, plus per-token re-pooling of the whole cache and MBs of host→device + P2P input traffic.Base: llama.cpp master
df03399b8, branchlevers. Build/run: CUDA 12.4.1,sm_86, imagellama.cpp:lever2b-cuda124;KV_TYPE=f16is mandatory (quantized KV can't reach the sparse MMA kernel at batch 1 on Ampere).Results (single-stream decode, t/s, 5×RTX 3090)
37.2tok/s to 50.8tok/s at 42k ctx (+37%)
22.5tok/s to 48.0tok/s at 111k (+113%)
11tok/s to 41.8tok/s at 250k (+280%)
(shallow essentially unchanged: 54.4 → 56.5. Prefill improves as a side effect: 609 → 716 t/s @111k.)
Changes (7 patches,
llama-levers/000*.patch)lever1 — enable the sparse-FA path (the big one). The enablement PR #27742 shipped top-k as a dense mask with a
// TODO: enable sparse attention. Un-TODO it: passtop_k->ne[0]asn_kv_maxso flash-attn treats the finite mask entries as a sparse K/V set — O(top_k) instead of O(n_kv). Drops attention from ~11 ms to 0.22 ms/token @250K.sparse256 — instantiate the kernel for this head shape.
may_use_sparseonly whitelisted the MLA shapes; qwen4exp isDKQ=DV=256, ncols2=8and that MMA template instance already exists. One clause:lever3 — block-granular top-k. Scores are block-constant and the budget is whole blocks, so select blocks first and expand only the ~513 selected blocks to cell indices (via
blk_cells), instead of expanding block scores to alln_kvcells and sorting those. Removes the O(n_kv) expansionget_rows+ one O(n_kv) mask add and shrinks the sort bycompress_ratio. Fully-future blocks get −inf bias; the incomplete tail block's cells ride a small per-token side input.lever2 — persistent pooled block-key cache. The indexer re-gathered, mean-pooled, RMS-normed and re-RoPEd all cached blocks every token though only the newest changed. Add a third per-QSA-layer buffer of pooled+normed+roped block keys; each ubatch recomputes only blocks it dirtied (host-side structural compare of member cells — also catches defrag/rollback/seq-ops with no invalidation hooks) and
ggml_set_rowsthem in. Dirty count is pow2-bucketed (≥4) for CUDA-graph topology stability.lever2b — device-resident block-member cells. After 1–3 the residual depth cost was input traffic (~3.1 MB/token HtoD + ~13.2 MB/token P2P for
blk_cells/mask). Store each block's member cells in the block cache's V tensor (f32;ggml_set_rowstakes no i32) and gather the top-k expansion on-device — removes theblk_cellsterm entirely.Validity
Every patch was A/B'd against the immediately-prior build and kept only if it passed all gates, all at temperature 0:
timings.predicted_per_secondonPOST /completion,cache_prompt:false,n_predict:48,ignore_eos:true.A comparable dense model on the same rig (
arch=qwen35, plain GQA, no indexer) stays flat at these depths, confirming the decay was QSA/indexer-specific, not KV bandwidth.Status / upstream
Upstream #28012 (this exact CUDA decay) is closed. Patches are self-contained and PR-able in order: lever1+sparse256 (smallest), then lever3, then lever2+lever2b (shared machinery). Remaining depth losses: host-side O(n_kv) grouping scan in
set_input_qsa; per-GPU f16 mask fan-out;flash_attn_mask_to_sparse_indicesre-scanning the mask per layer (an indices-in FA API would remove both). Full detail + nsys attribution innew-issue.md/issue-speed.md §9.