Skip to content

Misc. bug: Performance regression: CUDA sparse flash attention decode 1.6x slower (b11047 -> b11062) #29281

Description

@goodbadwolf

Name and Version

version: 0.4.1-dev (build 11062, commit 3cf03257f)
built with GNU 13.3.0 for Linux x86_64

Compared against build 11047, commit b23701f77.

Operating systems

Linux

Which llama.cpp modules do you know to be affected?

ggml (CUDA backend)

Command line

test-backend-ops perf -b CUDA0 -o FLASH_ATTN_EXT

Built with:

-DGGML_CUDA=ON -DGGML_CUDA_FA=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DGGML_CUDA_GRAPHS=ON -DCMAKE_CUDA_ARCHITECTURES=120

Problem description & steps to reproduce

FLASH_ATTN_EXT on the sparse path costs 64 percent more at batch one after
#28770, median over the 11 sparse configurations
test-backend-ops generates. The worst is 113 percent. Dense shapes are untouched. RTX PRO 6000 Blackwell
Workstation Edition, sm_120, driver 615.71.09, CUDA 13.4.

Run the command above at 3cf03257f and at b23701f77 and compare the configurations whose n_kv_max is not 0.
Use console output, since --output csv drops the timings.

Every sparse configuration, us/run:

hsk nh kv n_kv_max b11047 b11062 delta
512 1 49152 2048 25.80 55.01 +113%
512 1 32768 512 18.29 37.58 +106%
576 1 49152 2048 29.17 58.83 +102%
576 1 32768 512 21.78 41.60 +91%
512 1 16384 512 13.58 23.27 +71%
256 2 32768 2048 22.55 36.97 +64%
576 1 16384 512 17.12 27.14 +59%
256 2 16384 2048 15.03 22.53 +50%
512 1 4096 512 10.21 12.64 +24%
576 1 4096 512 13.59 16.18 +19%
256 2 4096 2048 16.45 11.81 -28%

Dense shapes over the same range moved 0.1 percent median on 58 decode configurations and 0.3 percent on 29 batched
ones. Nothing outside the sparse path changed.

Where the time goes. Dividing the added time by the mask width gives a constant, within each head size and across
a twelve-fold range of kv:

hsk nh kv n_kv_max added us ns per mask column
512 1 4096 512 +2.43 0.593
512 1 16384 512 +9.69 0.591
512 1 32768 512 +19.29 0.589
512 1 49152 2048 +29.21 0.594
576 1 4096 512 +2.59 0.632
576 1 16384 512 +10.02 0.612
576 1 32768 512 +19.82 0.605
576 1 49152 2048 +29.66 0.603

The cost tracks kv and ignores n_kv_max, including across the pair at n_kv_max 512 and 2048 that read 0.589 and
0.594. The attention kernel's work is bounded by n_kv_max, so it cannot produce this. The mask scan reads exactly
kv columns per list, so it can. I have not profiled the two kernels separately, so this is inference from the
scaling rather than a direct attribution.

It points at the new inner loop in flash_attn_mask_to_sparse_indices:

for (int q = 0; q < q1 - q0 && !selected; ++q) {
    selected = i < ne30 && isfinite(__half2float(mask[q*s31 + i]));
}

where it used to be a single test. The trip count is a runtime value and the exit is data dependent, so the
#pragma unroll over values_per_lane around it no longer yields straight line code. At ncols1 == 1 the body
always runs once and the union reduces to the old single-query test, so specialising that case, or templating the
kernel on ncols1, would recover it without touching the batched path the PR was written for.

The one configuration that improved, kv=4096 at n_kv_max=2048, fits the same reading. The new per-list count
bounding kb0_stop is a real win, and at that shape it is larger than the scan cost.

Scope. test-backend-ops only generates sparse configurations at nb=1, so this says nothing about the batched
sparse path. The models reaching it are the ones calling ggml_flash_attn_ext_set_n_kv_max, currently qwen4exp and
deepseek4. A server decoding at batch one, meaning one slot and no draft model, sits on the regressed shape for
every token.

First Bad Commit

3cf03257f (#28770, "CUDA: enable sparse fa for qwen4").

15 commits separate b23701f77 from 3cf03257f, and only 3cf03257f touches ggml/src/ggml-cuda/ at all. The rest
are Metal, Hexagon, server, chat and UI changes.

I measured #28536 separately, as the other CUDA flash attention
change in the neighbourhood. It costs 0.7 percent on these same configurations, so it is not involved.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions