Name and Version
version: 0.4.1-dev (build 11062, commit 3cf03257f)
built with GNU 13.3.0 for Linux x86_64
Compared against build 11047, commit b23701f77.
Operating systems
Linux
Which llama.cpp modules do you know to be affected?
ggml (CUDA backend)
Command line
test-backend-ops perf -b CUDA0 -o FLASH_ATTN_EXT
Built with:
-DGGML_CUDA=ON -DGGML_CUDA_FA=ON -DGGML_CUDA_FA_ALL_QUANTS=ON -DGGML_CUDA_GRAPHS=ON -DCMAKE_CUDA_ARCHITECTURES=120
Problem description & steps to reproduce
FLASH_ATTN_EXT on the sparse path costs 64 percent more at batch one after
#28770, median over the 11 sparse configurations
test-backend-ops generates. The worst is 113 percent. Dense shapes are untouched. RTX PRO 6000 Blackwell
Workstation Edition, sm_120, driver 615.71.09, CUDA 13.4.
Run the command above at 3cf03257f and at b23701f77 and compare the configurations whose n_kv_max is not 0.
Use console output, since --output csv drops the timings.
Every sparse configuration, us/run:
| hsk |
nh |
kv |
n_kv_max |
b11047 |
b11062 |
delta |
| 512 |
1 |
49152 |
2048 |
25.80 |
55.01 |
+113% |
| 512 |
1 |
32768 |
512 |
18.29 |
37.58 |
+106% |
| 576 |
1 |
49152 |
2048 |
29.17 |
58.83 |
+102% |
| 576 |
1 |
32768 |
512 |
21.78 |
41.60 |
+91% |
| 512 |
1 |
16384 |
512 |
13.58 |
23.27 |
+71% |
| 256 |
2 |
32768 |
2048 |
22.55 |
36.97 |
+64% |
| 576 |
1 |
16384 |
512 |
17.12 |
27.14 |
+59% |
| 256 |
2 |
16384 |
2048 |
15.03 |
22.53 |
+50% |
| 512 |
1 |
4096 |
512 |
10.21 |
12.64 |
+24% |
| 576 |
1 |
4096 |
512 |
13.59 |
16.18 |
+19% |
| 256 |
2 |
4096 |
2048 |
16.45 |
11.81 |
-28% |
Dense shapes over the same range moved 0.1 percent median on 58 decode configurations and 0.3 percent on 29 batched
ones. Nothing outside the sparse path changed.
Where the time goes. Dividing the added time by the mask width gives a constant, within each head size and across
a twelve-fold range of kv:
| hsk |
nh |
kv |
n_kv_max |
added us |
ns per mask column |
| 512 |
1 |
4096 |
512 |
+2.43 |
0.593 |
| 512 |
1 |
16384 |
512 |
+9.69 |
0.591 |
| 512 |
1 |
32768 |
512 |
+19.29 |
0.589 |
| 512 |
1 |
49152 |
2048 |
+29.21 |
0.594 |
| 576 |
1 |
4096 |
512 |
+2.59 |
0.632 |
| 576 |
1 |
16384 |
512 |
+10.02 |
0.612 |
| 576 |
1 |
32768 |
512 |
+19.82 |
0.605 |
| 576 |
1 |
49152 |
2048 |
+29.66 |
0.603 |
The cost tracks kv and ignores n_kv_max, including across the pair at n_kv_max 512 and 2048 that read 0.589 and
0.594. The attention kernel's work is bounded by n_kv_max, so it cannot produce this. The mask scan reads exactly
kv columns per list, so it can. I have not profiled the two kernels separately, so this is inference from the
scaling rather than a direct attribution.
It points at the new inner loop in flash_attn_mask_to_sparse_indices:
for (int q = 0; q < q1 - q0 && !selected; ++q) {
selected = i < ne30 && isfinite(__half2float(mask[q*s31 + i]));
}
where it used to be a single test. The trip count is a runtime value and the exit is data dependent, so the
#pragma unroll over values_per_lane around it no longer yields straight line code. At ncols1 == 1 the body
always runs once and the union reduces to the old single-query test, so specialising that case, or templating the
kernel on ncols1, would recover it without touching the batched path the PR was written for.
The one configuration that improved, kv=4096 at n_kv_max=2048, fits the same reading. The new per-list count
bounding kb0_stop is a real win, and at that shape it is larger than the scan cost.
Scope. test-backend-ops only generates sparse configurations at nb=1, so this says nothing about the batched
sparse path. The models reaching it are the ones calling ggml_flash_attn_ext_set_n_kv_max, currently qwen4exp and
deepseek4. A server decoding at batch one, meaning one slot and no draft model, sits on the regressed shape for
every token.
First Bad Commit
3cf03257f (#28770, "CUDA: enable sparse fa for qwen4").
15 commits separate b23701f77 from 3cf03257f, and only 3cf03257f touches ggml/src/ggml-cuda/ at all. The rest
are Metal, Hexagon, server, chat and UI changes.
I measured #28536 separately, as the other CUDA flash attention
change in the neighbourhood. It costs 0.7 percent on these same configurations, so it is not involved.
Name and Version
Compared against build 11047, commit
b23701f77.Operating systems
Linux
Which llama.cpp modules do you know to be affected?
ggml (CUDA backend)
Command line
Built with:
Problem description & steps to reproduce
FLASH_ATTN_EXTon the sparse path costs 64 percent more at batch one after#28770, median over the 11 sparse configurations
test-backend-opsgenerates. The worst is 113 percent. Dense shapes are untouched. RTX PRO 6000 BlackwellWorkstation Edition, sm_120, driver 615.71.09, CUDA 13.4.
Run the command above at
3cf03257fand atb23701f77and compare the configurations whosen_kv_maxis not 0.Use console output, since
--output csvdrops the timings.Every sparse configuration,
us/run:Dense shapes over the same range moved 0.1 percent median on 58 decode configurations and 0.3 percent on 29 batched
ones. Nothing outside the sparse path changed.
Where the time goes. Dividing the added time by the mask width gives a constant, within each head size and across
a twelve-fold range of
kv:The cost tracks
kvand ignoresn_kv_max, including across the pair atn_kv_max512 and 2048 that read 0.589 and0.594. The attention kernel's work is bounded by
n_kv_max, so it cannot produce this. The mask scan reads exactlykvcolumns per list, so it can. I have not profiled the two kernels separately, so this is inference from thescaling rather than a direct attribution.
It points at the new inner loop in
flash_attn_mask_to_sparse_indices:where it used to be a single test. The trip count is a runtime value and the exit is data dependent, so the
#pragma unrollovervalues_per_lanearound it no longer yields straight line code. Atncols1 == 1the bodyalways runs once and the union reduces to the old single-query test, so specialising that case, or templating the
kernel on
ncols1, would recover it without touching the batched path the PR was written for.The one configuration that improved,
kv=4096atn_kv_max=2048, fits the same reading. The new per-list countbounding
kb0_stopis a real win, and at that shape it is larger than the scan cost.Scope.
test-backend-opsonly generates sparse configurations atnb=1, so this says nothing about the batchedsparse path. The models reaching it are the ones calling
ggml_flash_attn_ext_set_n_kv_max, currentlyqwen4expanddeepseek4. A server decoding at batch one, meaning one slot and no draft model, sits on the regressed shape forevery token.
First Bad Commit
3cf03257f(#28770, "CUDA: enable sparse fa for qwen4").15 commits separate
b23701f77from3cf03257f, and only3cf03257ftouchesggml/src/ggml-cuda/at all. The restare Metal, Hexagon, server, chat and UI changes.
I measured #28536 separately, as the other CUDA flash attention
change in the neighbourhood. It costs 0.7 percent on these same configurations, so it is not involved.