docs: record why upstream Vulkan sparse FA was deferred - #65
Merged
Merged
Conversation
PR #64 reverted ggml/src/ggml-vulkan/ wholesale rather than merge ggml-org#28105: upstream's USE_SPARSE and strix's DYNAMIC_KV claim the same Flags bit, the two declare incompatible flash-attn push-constant layouts, and they are rival implementations of overlapping functionality. Records the collision, what integration requires (renumber a bit, merge the push-constant structs and their C++ writer, decide which mechanism survives), and the test-backend-ops runs that gate it -- including that a mismatched push-constant layout yields silently wrong tensors rather than a crash, so it has to be validated numerically. Docs only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
dzannotti
force-pushed
the
docs/deferred-strix-verification
branch
from
September 17, 2026 18:41
3731c05 to
05e2e39
Compare
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Companion to #64, trimmed to the part that is worth carrying as a document.
#64 reverted
ggml/src/ggml-vulkan/wholesale instead of merging upstream's sparseFlash Attention (
ggml-org#28105). This records why, so the next person does notrediscover it: upstream's
USE_SPARSEand strix'sDYNAMIC_KVclaim the sameFlagsbit, the two declare incompatible flash-attn push-constant layouts, and they are rival
implementations of overlapping functionality — strix already has its gather/compaction
family and a DeepSeek-V4 sparse split. Taking either side alone leaves
flash_attn{,_cm1,_cm2}.compcalling an undefinedfa_kv_index().Also records what integration requires (renumber one feature's bit, merge the
push-constant structs and their single C++ writer, decide whether the two mechanisms
coexist), and the
test-backend-ops -o FLASH_ATTN_EXTruns that gate it. The importantwarning: a mismatched push-constant layout does not crash, it emits plausible-looking
wrong tensors, so it has to be validated numerically.
Includes a measured baseline: FLASH_ATTN_EXT 5339/5339 on Vulkan0 (Radeon 8060S, RADV
STRIX_HALO) at
dffbb7888.Docs only, one new file. The PLE-on-disk and unverified-claims material that was here
earlier is dropped: #64's description is now the permanent record of what changed there,
and the prompt-processing regression reported on #63 belongs in an issue rather than a
document.