Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
69 commits
Select commit Hold shift + click to select a range
b1f86ee
experiment: dequant-once FA scratch for all KV quant types (q4_0/q4_1…
Nathanw1014 Jul 13, 2026
6dc9116
vulkan : gate FA dequant-once scratch on device-local capacity
Nathanw1014 Jul 14, 2026
e52cc7c
vulkan: contiguize strided f16 KV for FA prefill (GGML_VK_FA_KV_CONTI…
Nathanw1014 Jul 28, 2026
d61a7fd
tests: dense-permuted K/V option + Strix FA prefill perf/probe cases
Nathanw1014 Jul 28, 2026
6d056bf
vulkan: enable the f16 KV contiguize pass by default (GGML_VK_FA_KV_C…
Nathanw1014 Jul 28, 2026
82c3fe9
vulkan: single source of truth for native FA K/V types + non-native h…
Nathanw1014 Jul 28, 2026
fa117d1
tests: dense-permuted iq4_nl FA cases (dequant-once route coverage)
Nathanw1014 Jul 28, 2026
5d5759a
vulkan: bill the FA K/V contiguize pass on its own perf-logger line
Nathanw1014 Jul 31, 2026
8f178b8
vulkan: scale the FA MMQ dot product in fp32 before narrowing
Nathanw1014 Aug 9, 2026
e72ffec
vulkan: DeepSeek V4 lightning indexer kernels + indexed sparse FA
gaetan-puleo Aug 1, 2026
56a1843
vulkan: harden the sparse-FA shader and document the top-k API
Nathanw1014 Aug 1, 2026
79149d1
tests: sparse top-k FA parity + perf coverage (V4 CSA shape)
Nathanw1014 Aug 1, 2026
5d1180a
vulkan: gather-to-compact sparse decode FA for DeepSeek V4 top-k sele…
Nathanw1014 Aug 1, 2026
554519b
vulkan: fused DeepSeek V4 hyper-connection ops (HC pre / comb / post)
Nathanw1014 Aug 2, 2026
b8d2503
llama: keep DeepSeek lightning-indexer key cache f16 under quantized …
Nathanw1014 Aug 2, 2026
9cb8c28
llama: contiguize grouped o-proj input for small multi-token batches …
Nathanw1014 Aug 3, 2026
78e31af
vulkan: accelerate DeepSeek V4 sparse prefill FA
Mushoz Aug 12, 2026
d4cf907
vulkan: split sparse prefill attention
Mushoz Aug 13, 2026
57a64cb
vulkan: tile sparse prefill scratch
Mushoz Aug 13, 2026
989c21b
vulkan: reuse sparse FA probability fragments
Mushoz Aug 13, 2026
94ecd38
vulkan: cache sparse FA masks per key block
Mushoz Aug 13, 2026
5f60911
vulkan: fix DeepSeek V4 sparse split attention with multiple sequences
Nathanw1014 Aug 13, 2026
76b18c3
vulkan: harden the sparse FA split path and drop its debug scaffolding
Nathanw1014 Aug 13, 2026
75e195f
test-backend-ops: cover sparse top-k FA with more than one sequence
Nathanw1014 Aug 13, 2026
04311a0
vulkan: let sparse FA query tiling and multiple sequences coexist
Nathanw1014 Aug 13, 2026
feff28e
vulkan: parallelize DSV4 Lightning Indexer prefill
Mushoz Aug 13, 2026
2a6d1de
vulkan: extend DeepSeek V4 gather-to-compact to small batches
Nathanw1014 Aug 13, 2026
6a31bf0
vulkan: record the resource limits behind the two DSV4 prefill kernels
Nathanw1014 Aug 13, 2026
964a218
vulkan: deduplicated union for DeepSeek V4 small-batch decode
Nathanw1014 Aug 14, 2026
628788d
vulkan: gate the DeepSeek V4 small-batch union on the measured union …
Nathanw1014 Aug 14, 2026
857f81c
vulkan: measure the V4 union on real draft tokens, and correct the fi…
Nathanw1014 Aug 14, 2026
fa82576
vulkan: default the DeepSeek V4 small-batch union on
Nathanw1014 Aug 14, 2026
1481bbe
vulkan: let the DeepSeek V4 small-batch gather serve quantised K/V
Nathanw1014 Aug 14, 2026
35b7578
vulkan: dequantise q8_0 K/V inside the DeepSeek V4 small-batch gather
Nathanw1014 Aug 14, 2026
987f662
vulkan: decode q4_0 in the V4 gather too, with the real element mapping
Nathanw1014 Aug 14, 2026
fa5f251
vulkan : restore the unrolled row copy in the DSV4 gather shaders
Nathanw1014 Aug 17, 2026
7ceebb6
vulkan: decode quantised K/V inside the DeepSeek V4 per-token gather
Nathanw1014 Aug 20, 2026
f9471c6
vulkan: dequantise the cache for the DeepSeek V4 sparse prefill
Nathanw1014 Aug 20, 2026
f7ffc13
vulkan: use small Lightning Indexer CM for batches 4-15
pepuscz Aug 24, 2026
ef56966
vulkan: hoist the Lightning Indexer K fragments out of the head loop
Nathanw1014 Aug 26, 2026
24d18f7
vulkan: route the whole small-batch Lightning Indexer window to the d…
Nathanw1014 Aug 26, 2026
4d83452
vulkan: hoist the Lightning Indexer K fragments out of the CM head loop
Nathanw1014 Aug 26, 2026
883b4d8
vulkan: support arbitrary Lightning Indexer head counts via specializ…
Nathanw1014 Aug 26, 2026
808ce1c
vulkan: N_HEAD spec constant for the scalar-64 lightning indexer
Nathanw1014 Aug 30, 2026
1007fdc
vulkan: hoist the coopmat1 FA P-fragment load out of the hsv_tile loop
Nathanw1014 Jul 30, 2026
11eaefa
vulkan: store coopmat1 FA Psh query-major so the GEMM2 A load vectorizes
Nathanw1014 Jul 30, 2026
bc2b9ad
vulkan: pin a 32-wide subgroup for coopmat1 FA where narrowing is free
Nathanw1014 Jul 30, 2026
1556cd1
vulkan: enable the coopmat1 FA wave32 narrowing rule by default
Nathanw1014 Aug 30, 2026
e75a22b
vulkan: mul_mat_id per-expert-n tile selection (Stage 2a, env-gated)
Nathanw1014 Jul 14, 2026
33a08e8
vulkan: mul_mat_id small-tile shape probes (env-gated)
Nathanw1014 Jul 14, 2026
fe10c7f
vulkan: mul_mat_id taller medium tile probe (GGML_VK_MMID_M128, env-g…
Nathanw1014 Jul 14, 2026
07e35fe
vulkan: mmid wave32 probe (GGML_VK_MMID_WAVE32, env-gated)
Nathanw1014 Jul 14, 2026
18569a9
vulkan: mul_mat_id f16-B probe (GGML_VK_MMID_F16B, env-gated)
Nathanw1014 Jul 14, 2026
8ad46e1
vulkan: guard mmid f16-B path on pipeline existence (Q2_0 fallback)
Nathanw1014 Jul 26, 2026
68df071
tests: cover MMQ tile boundaries in MUL_MAT and MUL_MAT_ID
Nathanw1014 Aug 6, 2026
77214b5
vulkan: create coopmat2 mul_mat_id pipelines with the real param count
Nathanw1014 Aug 8, 2026
6192a05
vulkan: enable the Strix mmid tile gates by default
Nathanw1014 Aug 30, 2026
891cfff
vulkan: run the quantised dense coopmat pipelines at wave32
Nathanw1014 Aug 18, 2026
f4d6f9e
vulkan: four env-gated Strix Halo prefill fixes for delta-net MoE
Nathanw1014 Aug 6, 2026
874b904
vulkan: reject ne[3] > 1 in the mul_mat_id scale epilogue
Nathanw1014 Aug 6, 2026
e249248
vulkan : default the transposed-concat path on
Nathanw1014 Aug 17, 2026
7db0111
vulkan: flush pending compute ctx before perf logger timestamps
Nathanw1014 Aug 2, 2026
6ab7cb6
vulkan: bound command buffers by memory traffic, not just flops
Nathanw1014 Aug 2, 2026
e387252
ggml: cut backend splits on the input constant, not the grown capacity
Nathanw1014 Aug 9, 2026
492e443
vulkan : optional f16 B operand for quantized MUL_MAT on coopmat1
Nathanw1014 Aug 16, 2026
d737bd5
vulkan : add auto mode to GGML_VK_DENSE_F16B
Nathanw1014 Aug 16, 2026
9f5ece3
vulkan: derive the mul_mat_id tiles from the wave64 dense tiles
Nathanw1014 Sep 8, 2026
1debd52
vulkan: keep the coopmat1 FA wave32 pin to multi-row dispatches
Nathanw1014 Sep 8, 2026
2f311ba
Merge branch 'master' into strix/vulkan-stack-only
Nathanw1014 Sep 8, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
158 changes: 158 additions & 0 deletions docs/development/DSV4-vulkan-lightning-indexer-progress.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,158 @@
# DeepSeek V4 Vulkan Lightning Indexer progress

This is a restart note for the Strix Halo Lightning Indexer optimization. It is a development scratch pad and can be removed before the final PR.

## Repository state

- Main repository: `/home/jaap/Projects/git/llama.cpp`
- Optimization worktree: `/tmp/llama-strix-beta-bench`
- Branch: `strix-halo-vulkan-lightning-indexer`
- Base commit: `316c72ee9eab590f5891089d3b6bfc0d01d00d19`
- Base branch: Nathan's `strix-halo-vulkan-beta`
- Decode microbench work is stored in the main worktree as `stash@{0}: On strix-halo-vulkan: wip: DSV4 decode microbench depth matrix`.
- The Indexer changes are uncommitted. Do not commit without explicit user approval. An assisted commit needs an `Assisted-by:` trailer.
- Do not run builds and GPU benchmarks together. The APU shares its power and memory-bandwidth budget.
- GPU commands need sandbox escalation.

## Objective and result

After sparse prefill attention was flattened, the context-dependent Lightning Indexer became the next prefill bottleneck. The old cooperative-matrix shader used one wave64 subgroup per workgroup, processed one 16-key tile, and loaded one query head at a time.

The new wide pipeline uses eight wave64 subgroups per workgroup. Each subgroup processes a separate 16-key tile, so one workgroup covers 128 keys. It stages four query heads and their weights together, reuses them across all eight subgroups, and uses subgroup-scoped synchronization between cooperative-matrix result stores. A workgroup barrier remains between four-head groups because all subgroups reuse the shared query storage.

The optimized shader requires 512 workgroup invocations and 64 KiB shared memory. Pipeline creation is capability-based. Devices without those limits use a one-wave, one-head cooperative-matrix specialization. The scalar implementation remains the fallback when cooperative matrices are unavailable. The decode-specific cooperative-matrix pipeline is unchanged.

At the 32k-equivalent prefill microbench shape:

| Version | Time per layer | Throughput |
| --- | ---: | ---: |
| Baseline | 50.51 ms | 5.83 TFLOPS |
| Optimized | 31.81 ms | 9.25 TFLOPS |

This is a 37.0% reduction in Lightning Indexer kernel time.

The canonical 32k llama-bench improved from 209.45 to 216.32 tokens/s. Total Vulkan time fell from 9.73729 to 9.42448 seconds. Total Lightning Indexer time fell from 1.11860 to 0.697276 seconds. Sparse attention and top-K were effectively unchanged.

## Changed files

- `ggml/src/ggml-vulkan/vulkan-shaders/lightning_indexer_cm.comp`: parameterizes the shader, adds the eight-wave four-head implementation, and remains usable for the small fallback.
- `ggml/src/ggml-vulkan/vulkan-shaders/vulkan-shaders-gen.cpp`: generates wide `N_WAVES=8`, `HEADS_PER_TILE=4` and small `N_WAVES=1`, `HEADS_PER_TILE=1` variants.
- `ggml/src/ggml-vulkan/ggml-vulkan.cpp`: creates and selects the capability-gated wide pipeline and the small cooperative-matrix fallback.
- `tests/test-backend-ops.cpp`: adds 127, 128, and 129-key correctness boundaries and PP2048 performance shapes through 512k simulated source context.

## Performance data

The performance rows model the actual PP2048 Indexer shapes after the source context is filled in 2048-token batches. `kv=8704` is the measured shape near 32k source context. It differs from 32768 because the Indexer compresses source tokens into rows.

| Source depth | `kv` | Baseline | Optimized | Reduction |
| ---: | ---: | ---: | ---: | ---: |
| 0 | 512 | 4.23 ms | 2.14 ms | 49.5% |
| 8k | 2560 | 17.10 ms | 10.09 ms | 41.0% |
| 16k | 4608 | 29.97 ms | 17.62 ms | 41.2% |
| 32k | 8704 | 50.51 ms | 32.94 ms | 34.8% |
| 64k | 16896 | 94.36 ms | 64.87 ms | 31.3% |
| 128k | 33280 | 174.96 ms | 128.76 ms | 26.4% |
| 256k | 66048 | 352.73 ms | 250.88 ms | 28.9% |
| 512k | 131584 | 696.79 ms | 495.90 ms | 28.8% |

The final isolated 32k run after cleanup measured 31.81159 ms. Small matrix differences are normal laptop GPU clock variation.

| Canonical 32k metric | Baseline | Optimized |
| --- | ---: | ---: |
| PP2048 | 209.45 tokens/s | 216.32 tokens/s |
| Total Vulkan | 9.73729 s | 9.42448 s |
| Lightning Indexer | 1.11860 s | 0.697276 s |
| Sparse FA raw | 0.198438 s | 0.195367 s |
| Sparse FA selected | 0.835110 s | 0.843385 s |
| Sparse FA reduce | 0.088463 s | 0.086678 s |
| TOP_K | 0.075473 s | 0.075350 s |

Logs:

- Baseline matrix: `/tmp/dsv4-lightning-prefill-baseline.log`
- Optimized matrix: `/tmp/dsv4-lightning-four-head-matrix.log`
- Final selected-pipeline 32k microbench: `/tmp/dsv4-lightning-final-selected-32k.log`
- Baseline canonical 32k llama-bench: `/tmp/dsv4-nathan-beta-32k-rerun-new-first.log`
- Optimized canonical 32k llama-bench: `/tmp/dsv4-lightning-final-32k-llama-bench.log`

## Correctness and resources

- Wide pipeline: all 20 focused F16 cases passed, including 127, 128, and 129-key boundaries.
- Small cooperative-matrix fallback: temporarily forced and all the same 20 cases passed.
- Wide shader on gfx1151: 168 VGPRs, 63,488 bytes LDS, no spills, eight subgroups per SIMD.
- The final pipeline-statistics run confirmed `lightning_indexer_cm_f16` was selected.
- No NaN, Inf, or comparison failures were reported.

Correctness logs:

- Wide: `/tmp/dsv4-lightning-consolidated-wide-correctness.log`
- Small fallback: `/tmp/dsv4-lightning-consolidated-small-correctness.log`

## Experiments and decisions

- Four waves improved the 32k shape about 3% and became worse at deep simulated contexts.
- Eight waves improved it about 8% before the other changes.
- Subgroup-scoped synchronization after cooperative-matrix stores improved the eight-wave version.
- Staging four query heads and weights produced the large gain by reducing redundant loads and barriers.
- Using only a subgroup barrier between head groups failed four boundary tests. One subgroup could overwrite shared query data while another still read it. A workgroup barrier is required there.
- A separate fallback shader source was avoided. Generator definitions create both variants from one file.

## Commands

Build only, with no GPU benchmark running:

```sh
cd /tmp/llama-strix-beta-bench
git diff --check
cmake --build build --config Release --target test-backend-ops llama-bench -j "$(nproc)"
```

Focused correctness:

```sh
cd /tmp/llama-strix-beta-bench
./build/bin/test-backend-ops test -b Vulkan0 -o LIGHTNING_INDEXER -p 'type_K=f16' > /tmp/dsv4-lightning-correctness.log 2>&1
tail -n 30 /tmp/dsv4-lightning-correctness.log
```

Final 32k-equivalent microbench and pipeline selection:

```sh
cd /tmp/llama-strix-beta-bench
GGML_VK_PIPELINE_STATS=lightning_indexer_cm_f16 ./build/bin/test-backend-ops perf -b Vulkan0 -o LIGHTNING_INDEXER -p 'kv=8704' > /tmp/dsv4-lightning-final-selected-32k.log 2>&1
tail -n 16 /tmp/dsv4-lightning-final-selected-32k.log
```

Full Indexer depth matrix:

```sh
cd /tmp/llama-strix-beta-bench
./build/bin/test-backend-ops perf -b Vulkan0 -o LIGHTNING_INDEXER -p 'nb=2048,nh=64,ns=1,nm=1,type_K=f16' > /tmp/dsv4-lightning-matrix.log 2>&1
rg 'kv=(512|2560|4608|8704|16896|33280|66048|131584),nb=2048' /tmp/dsv4-lightning-matrix.log
```

Canonical 32k llama-bench. Run it only when needed, never while compiling, and inspect only the final block:

```sh
cd /tmp/llama-strix-beta-bench
GGML_VK_PERF_LOGGER=1 ./build/bin/llama-bench -m /home/jaap/Projects/docker/localLLaMA/models/models--unsloth--DeepSeek-V4-Flash-0731-GGUF/snapshots/109848da2469efe1f1aab9e11acea08a065ccd4f/UD-IQ3_XXS/DeepSeek-V4-Flash-0731-UD-IQ3_XXS-00001-of-00004.gguf -r 1 -d 32768 -p 2048 -ub 2048 -fa 1 -n 0 > /tmp/dsv4-lightning-32k-llama-bench.log 2>&1
last=$(grep -n 'Vulkan Timings:' /tmp/dsv4-lightning-32k-llama-bench.log | tail -n 1 | cut -d: -f1)
sed -n "${last},\$p" /tmp/dsv4-lightning-32k-llama-bench.log | tail -n 180
```

Patch inspection:

```sh
cd /tmp/llama-strix-beta-bench
git diff --check
git diff --stat
git diff
git status --short
```

## Next actions

1. The user reviews and understands the four-file implementation and result summary.
2. Commit only after explicit user approval for that commit action.
3. Remove this scratch pad before a PR if it is not useful as permanent documentation.
4. Restore the separate decode microbench stash from the main worktree only if that work resumes.
Loading
Loading