Skip to content

Eval bug: qwen4exp (Qwen3.8-Flash-Next) CUDA: decode slows linearly with context #28734

Description

@lukolszewski

Name and Version

$ /app/llama cli --version
version: 0.4.0-dev (build 10909, commit 44b47fd77)
built with GNU 12.3.0 for Linux x86_64

Operating systems

Linux

GGML backends

CUDA

Hardware

Ryzen 7950x on x670 MB 192MB RAM 5x RTX3090 (3 of them via USB4, but this detail is not relevant to the bug)

Models

unsloth/Qwen3.8-Flash-Next-GGUF UD-Q4_K_XL -ngl 99, layer split across 5 GPUs,
-ot per_layer_token_embd.weight=CPU -c 524288 --parallel 2

Problem description & steps to reproduce

This is a likely a duplicate of a previous report erroneusly closed by reporter #27856, I can't reopen it so I'm submitting this one with all my detail. The issue #27856 is not a duplicate of #27856, that one appears to be fixed (it reports much more severe deterioration).

When runnig the PR #27742 branch I initaly noticed substantial performance degradation about halving the decode speed starting at 54 tok/s every 50k context size or so.

Using the latest master df03399 this is much improved to:

  • 37.2tok/s to 50.8tok/s at 42k ctx (+37%)
  • 22.5tok/s 48.0tok/s at 111k (+113%),
  • 11tok/s to 41.8 tok/s at 250k

but still deteriorating to "nighly speed" on my hardware at the bottom of the context.

Please note both generation is affected (severely) as well aas prefill (a lot less). Prefill almost halves the speed at max 260k context so far less impactful.

I let my LLM run over the code to see what improvements it finds and by the morning indeed it managed to make changes that helped a lot to:

  • 37.2tok/s to 50.8tok/s at 42k ctx (+37%)
  • 22.5tok/s to 48.0tok/s at 111k (+113%)
  • 11tok/s to 41.8tok/s at 250k (+280%)

It performed some correctness validation that suggests we didn't brake anything, but I'm far from claiming a "fix" at this stage (have yet to validate properly).

The diff is almost 2k lines over 7 patches so it is too big to include here. I will provide it upon request.

First Bad Commit

it was built from pull/27742/head

Relevant log output

I also include a brief LLM written sumamry of changes made below:

qwen4exp decode depth-decay — LLM generated "fix" summary

Arch: qwen4exp (Qwen3.8-Flash-Next, 125B-A6B; 12 QSA sparse-attn + 36 Gated-DeltaNet layers, indexer.top_k=2048, compress_ratio=4).
Problem: single-stream decode slows ~linearly with context depth even though QSA is supposed to bound attention to a fixed top-k budget. Root cause (nsys-profiled): the top-k path was running as a dense attention scan over all n_kv, plus per-token re-pooling of the whole cache and MBs of host→device + P2P input traffic.
Base: llama.cpp master df03399b8, branch levers. Build/run: CUDA 12.4.1, sm_86, image llama.cpp:lever2b-cuda124; KV_TYPE=f16 is mandatory (quantized KV can't reach the sparse MMA kernel at batch 1 on Ampere).

Results (single-stream decode, t/s, 5×RTX 3090)

37.2tok/s to 50.8tok/s at 42k ctx (+37%)
22.5tok/s to 48.0tok/s at 111k (+113%)
11tok/s to 41.8tok/s at 250k (+280%)

(shallow essentially unchanged: 54.4 → 56.5. Prefill improves as a side effect: 609 → 716 t/s @111k.)

Changes (7 patches, llama-levers/000*.patch)

 lever1     src/models/qwen4exp.cpp                 |  9 +++++----   (+ incidental Dockerfile pip fix)
 lever3     src/llama-memory-hybrid-idx.{cpp,h}     |  99 ++++++      block-granular top-k
            src/models/qwen4exp.cpp                 |  72 ++++
 sparse256  ggml/src/ggml-cuda/fattn-mma-f16.cuh    |  3 +           whitelist DKQ=DV=256 MMA
 lever2     src/llama-memory-hybrid-idx.{cpp,h}     | 590 ++++++     persistent pooled block-key cache
            src/llama-kv-cache.{cpp,h}, qwen4exp.cpp
 lever2b    src/llama-memory-hybrid-idx.{cpp,h}     |  59 ++++       device-resident block-member cells
            src/llama-kv-cache.{cpp,h}, qwen4exp.cpp

lever1 — enable the sparse-FA path (the big one). The enablement PR #27742 shipped top-k as a dense mask with a // TODO: enable sparse attention. Un-TODO it: pass top_k->ne[0] as n_kv_max so flash-attn treats the finite mask entries as a sparse K/V set — O(top_k) instead of O(n_kv). Drops attention from ~11 ms to 0.22 ms/token @250K.

-    // TODO: enable sparse attention when we are ready
-    ggml_tensor * cur = build_attn_mha(q, k, v, nullptr, kq_mask_top_k, nullptr, nullptr, 0, kq_scale, il);
+    // LEVER1: passing top_k->ne[0] as n_kv_max -> sparse K/V set (deepseek4/GLM-DSA pattern, #27970)
+    ggml_tensor * cur = build_attn_mha(q, k, v, nullptr, kq_mask_top_k, nullptr, nullptr, top_k->ne[0], kq_scale, il);

sparse256 — instantiate the kernel for this head shape. may_use_sparse only whitelisted the MLA shapes; qwen4exp is DKQ=DV=256, ncols2=8 and that MMA template instance already exists. One clause:

-    return (DKQ == 512 && DV == 512 && ncols1 == 1 && ncols2 == 8) ||
+    return (DKQ == 256 && DV == 256 && ncols1 == 1 && ncols2 == 8) ||
+           (DKQ == 512 && DV == 512 && ncols1 == 1 && ncols2 == 8) ||
            (DKQ == 576 && DV == 512 && ncols1 == 1 && ncols2 == 16);

lever3 — block-granular top-k. Scores are block-constant and the budget is whole blocks, so select blocks first and expand only the ~513 selected blocks to cell indices (via blk_cells), instead of expanding block scores to all n_kv cells and sorting those. Removes the O(n_kv) expansion get_rows + one O(n_kv) mask add and shrinks the sort by compress_ratio. Fully-future blocks get −inf bias; the incomplete tail block's cells ride a small per-token side input.

lever2 — persistent pooled block-key cache. The indexer re-gathered, mean-pooled, RMS-normed and re-RoPEd all cached blocks every token though only the newest changed. Add a third per-QSA-layer buffer of pooled+normed+roped block keys; each ubatch recomputes only blocks it dirtied (host-side structural compare of member cells — also catches defrag/rollback/seq-ops with no invalidation hooks) and ggml_set_rows them in. Dirty count is pow2-bucketed (≥4) for CUDA-graph topology stability.

lever2b — device-resident block-member cells. After 1–3 the residual depth cost was input traffic (~3.1 MB/token HtoD + ~13.2 MB/token P2P for blk_cells/mask). Store each block's member cells in the block cache's V tensor (f32; ggml_set_rows takes no i32) and gather the top-k expansion on-device — removes the blk_cells term entirely.

Validity

Every patch was A/B'd against the immediately-prior build and kept only if it passed all gates, all at temperature 0:

  • Exact needle-in-haystack retrieval at 63k and 111k (a random code string embedded mid-context must be reproduced verbatim) — guards against the sparse path silently dropping the right K/V.
  • Multi-turn prefix reuse + rewind and 2 concurrent slots — guards against the persistent block-key cache going stale across defrag/rollback/seq ops.
  • Throughput read from timings.predicted_per_second on POST /completion, cache_prompt:false, n_predict:48, ignore_eos:true.

A comparable dense model on the same rig (arch=qwen35, plain GQA, no indexer) stays flat at these depths, confirming the decay was QSA/indexer-specific, not KV bandwidth.

Status / upstream

Upstream #28012 (this exact CUDA decay) is closed. Patches are self-contained and PR-able in order: lever1+sparse256 (smallest), then lever3, then lever2+lever2b (shared machinery). Remaining depth losses: host-side O(n_kv) grouping scan in set_input_qsa; per-GPU f16 mask fan-out; flash_attn_mask_to_sparse_indices re-scanning the mask per layer (an indices-in FA API would remove both). Full detail + nsys attribution in new-issue.md / issue-speed.md §9.

No activity

Activity on this issue will appear here.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions