Repository navigation
Merge halo-box/llama.cpp master into strix (2026-09-29) - #111
Merged
Merged
Conversation
* vulkan: support qwen4exp hc ops * fix stale comment [no-ci]
…8996) Both im2col.comp and im2col_3d.comp declare D_ptr without an explicit buffer_reference_align, so glslang emits writes through it as Aligned 16. The shaders advance the pointer by D_SIZE, a per-variant define set to 4 for float and 2 for float16_t, so most write addresses are not 16-byte aligned. This triggers VUID-RuntimeSpirv-PhysicalStorageBuffer64-06315 under GPU-AV. Declaring buffer_reference_align = D_SIZE matches the alignment to the actual write stride and takes validation hits from 20 to 0 for both IM2COL and IM2COL_3D. Fixes ggml-org#28960
* opencl: fix warnings * opencl: fix warnings for non adreno
ggml-org#28993) * gguf : align the data section relative to the GGUF start, not the file gguf_init_from_file_ptr reads a GGUF from the current file position, but padded the data section from file offset 0, so a GGUF embedded at an offset that is not a multiple of the alignment loaded without error and returned wrong tensor data. Also adds llama_adapter_lora_init_from_file_ptr, and disables mmap with a warning when an embedded data section is not aligned, instead of asserting in ggml. Assisted-by: Claude Opus 5 * llama : load lora from path through the FILE* variant The test now checks that mmap is disabled only for an unaligned offset. Assisted-by: Claude Fable 5.1 * Update ggml/src/gguf.cpp Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * Update include/llama.h Co-authored-by: Johannes Gäßler <johannesg@5d6.de> * llama : error on unaligned mmap of an embedded GGUF, drop test-load-file-ptr --------- Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
…g#29008) * chat : add message delimiters to the DeepSeek V3.2/V4 parser Assisted-by: Claude Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
* ci : add API/ABI check to make-release workflow [no ci] This commit adds an API/ABI compatibility check to the make-release workflow. The motivation for this to allow us to detect any potential breaking changes in API/ABI compatibility between releases and fail the the release if there are any. The workflow can be triggered manually as before and this check can be skipped if needed as it does take some time which might be useful when doing a dry-run and not specifically interested in the API/ABI check. By default this will check the current release against the latest release, but this can also be configured in the workflow, or in the script run on the command line, to check a different tag. * add check for minor version bumps [no ci] This commit also changes the build type to be RelWithDebInfo so that the reported information is more useful.
…org#27985) * ui: fix accidentally removed reasoning menu in single model mode on desktop * ui: formatting task run to fix storybook test * ui: mount the add menu reasoning submenu outside router mode only The models selector already owns the reasoning submenu in router mode, so the add menu only mounts it in single model mode. The first enabled item of the add menu is now the reasoning submenu, the accessibility story expects it. --------- Co-authored-by: Ben Babik <work@benjaminbabik.com> Co-authored-by: Pascal <admin@serveurperso.com>
…org#29009) * Update to openvino-2026.4 * Update OV docs * ggml-openvino : fix clangd and MSVC warnings * fix int to ptr cast, more internal linkage enforcement, and avoiding duplicate switch case --------- Co-authored-by: Mostafa Faheem <mostafaaafaheem@gmail.com>
* first fix * removed unnecessary declarations
required for qwen35moe if MTP tensors are fused but not loaded
… experts (ggml-org#28501) * vulkan: raise the hoisted row-id limit for mul_mat_id to 512 experts The expert-count shader (count_experts.comp) sizes its shared arrays with BLOCK_SIZE, which is 256. Because of that, row-id hoisting is switched off for any model with more than 256 experts, and every mul_mat_id workgroup has to rescan the whole ids tensor on its own. Qwen3.8-Flash-Next has 512 experts and was quietly running on that slow path. This change sizes the arrays with a separate MAX_EXPERTS constant (512), clears them in a loop instead of one entry per thread, and raises the matching limit on the host side. On Strix Halo at batch 2048 the expert matmuls drop from 12.5 to 9.5 ms (iq3_s) and from 14.0 to 7.5 ms (iq4_nl) per op, and prompt processing gets about 19 % faster at 8k tokens. test-backend-ops MUL_MAT_ID passes (891/891) with new 512-expert test cases. Assisted-by: Claude Fable 5.1 * vulkan: raise the hoisted row-id limit for mul_mat_id to 1024 experts Follow-up to review feedback: 1024 matches LLAMA_MAX_EXPERTS instead of stopping at 512. The three shared arrays in count_experts.comp grow to 3 * 1024 * 4 = 12 KiB, which fits the 16 KiB that Vulkan guarantees for maxComputeSharedMemorySize. Adds mul_mat_id test cases at 1024 experts alongside the existing 512 ones. test-backend-ops MUL_MAT_ID passes on Vulkan (RADV, Strix Halo, Radeon 8060S): 889/889.
…9036) * gguf-py: fix Q8_1 block size in GGML_QUANT_SIZES * --whitespace --------- Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
* fix: build fails when GGML_CPU=OFF and GGML_CUDA=ON * fix: eol in examples/convert-llama2c-to-ggml/CMakeLists.txt file
* vocab : add ufakzeka pre-tokenizer * vocab : move ufakzeka to the models list and regenerate the hash mapping
* vulkan: add IQ3_S MMQ matmul kernels * Make block_a_to_shmem do 2-byte loads (110 bytes is divisible by 2) * Align the check, IQ3_S is also using K tile size
* fix get_rows vec4 handling * Add src strides checking to vec4_aligned of get_rows and the new test case.
…gml-org#29042) * llama: read the SWA pattern as a period or a per-layer array Add llama_model_base::load_swa_pattern(), which reads sliding_window_pattern either as one flag per layer or as a period expanded by set_swa_pattern(), and use it in every loader that reads the key as a period. These loaders silently ignored an array and applied their default period, although the converters of olmo2, gemma3n and exaone4 write arrays. The published GGUFs match the defaults, so their outputs do not change. The loaders that already accepted both forms lose their duplicated scalar-then-array block, and use their declared default period when the key is absent. * model-saver: write the SWA pattern and the MLA SWA geometry Write sliding_window_pattern as one flag per layer, nextn layers included, for every model using SWA. The array is never collapsed to a scalar, since the loaders read a scalar as a period. Also write the MLA key/value lengths and KV LoRA rank of the SWA layers, required by dots3note. This enables the saver for plamo3, gemma3, cohere2, cohere2moe, olmo2, exaone-moe, afmoe, mimo2, spark2_5, muse-glimmer, mellum, laguna, granite_swa, dots3note and maple, all passing the bit-exact roundtrip of test-llama-archs.
* ggml : check for allocation failures to prevent crashes * wording
tests/test-backend-ops.cpp: keep both new test cases, upstream's test_mul_mat_id_w4a8/w4a4 from the sync and test_moe_prefill from #91. Assisted-by: Claude Opus 5.5 Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
gaetan-puleo
approved these changes
Sep 29, 2026
pwilkin
approved these changes
Sep 30, 2026
Keep ROCmFP4 FAST and MXFP4/NVFP4 W4A4 MMQ declarations. Keep explicit rollback registrations, including Qwen4exp, without duplicate directory runs. Combine the status/all-model harness with scaled-weight seq_cp/graph-reuse coverage. Assisted-by: Codex
…ndency pass Pass the inclusive final node index instead of the exclusive group end. This keeps allocation dependencies in the graph when the group is terminal. Assisted-by: Codex
This was referenced Oct 1, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Overview
Merges
halo-box/llama.cppmaster (e57160642, 215 commits, including halo's 2026-09-29 upstream sync) into strix. This is a real merge commit, do not squash it (seeSYNC.mdin halo-box). After it lands, strix is 0 commits behind halo.21 files conflicted. Each hunk was traced to the commit and PR on each side. Where the two sides competed, the decision was made by measuring on this hardware, not by picking a side.
Vulkan: neither side wholesale
Syncs #47 and #64 had reverted upstream's matmul framework (ggml-org#25773), and every upstream Vulkan matmul change since then builds on it. The two stacks were measured head to head: strix
origin/masteragainst halomaster, same session, ABBA order.GGML_VK_PERF_LOGGER, pp2048):So the resolution is upstream's matmul framework plus strix's non-matmul Vulkan stack. That stack covers FA (dynamic KV, dequant-once, contiguize, gather/union, DeepSeek V4 sparse), transposed CONCAT, the lightning indexer, mat-vec chunking and submit bounding, moved into upstream's new
ggml-vulkan-{types,push-constants,common}.hsplit.docs/development/vulkan-sparse-fa-deferred.mdis updated with this sync's decisions.Retired: strix's env-gated mul_mat_id experiments (
GGML_VK_MMID_*,GGML_VK_DENSE_F16B,GGML_VK_MMID_SCALE_EPILOGUE). Two measured strix wins are lost and not yet ported onto the new framework (see follow-ups):HIP
prec_src1template parameter.switch_typedefaultsprec_src1to Q8, so the paired-MMQ call sites are unchanged.ggml-cuda.cualloc deps: strix's fusion loops are kept. Oneif (op != GGML_OP_MUL) continue;is dropped, because it would have made upstream's top-k MoE alloc deps unreachable.DSV4_HC_PRE/HC_POSTreplaced strix'shc_mix/hc_combinefusions and cost -8.6% pp2048. All four combinations were measured (see below). The HIP backend now declines the gated PRE and comb-less POST on RDNA3.5. NewLLAMA_FUSED_DSV4_HC_PRE/_POST=0env overrides follow the existingLLAMA_FUSED_GDN_CHpattern.Breakage that merged without a conflict (fixed)
llama-context.cpp: the PLE prefetch readbatch_inp.token, which the newllama_batch_extdoes not have.llama-memory-hybrid-idx: upstream'scausal_attnbranch landed inside strix'sset_input_qsa_scan.causal_attnis now passed through, and the causal-only prefix fast path, block selection andqsa_cacheare gated on it.ggml-cuda.cu:hinthad become undeclared after upstream's FWHT refactor.norm.cu: strix'sncols == 128RMS_NORM launch was missing upstream's new scale argument.test-backend-ops.cpp: theinit_mul_mat_id_tensorsforward declaration no longer matched the definition.test-save-load-state.cpp: both sides claimed "Test 9". Strix's test becomes Test 10.build-self-hosted.ymlinto 7ci-self-hosted-*.ymlfiles. All 7 get strix'sbranches-ignoreonpull_request, so the 24h runner timeouts seen on Merge halo-box/llama.cpp master into strix (2026-09-15) #64 don't come back.Measurements
All runs:
-ngl 99 -fa 1 -b 2048 -ub 2048, ABBA order, 2 reps per run, so n=4 per arm. Mean ± sd in t/s.*means the difference is larger than the combined sd. Background GPU workloads were stopped for the whole window.Vulkan, strix main vs this PR:
Vulkan, halo master vs this PR (what taking halo's Vulkan wholesale would have cost):
HIP, strix main vs this PR:
AGENTS.md prefill protocol, Qwen3.8-Flash-Next:
llama-bench -p 2048 -d 0,12000,32000,64000 -b 2048 -ub 2048 -n 0 -r 4 -ngl 99 -fa on -ctk f16 -ctv f16 --load-mode noneqwen4exp HIP hyper-connection decision (strix main = 910.6 pp2048 / 30.1 tg128):
Isolated A/Bs of the other contested choices (toggle build, not part of this PR):
ncols_opton RDNA3.5: gpt-oss pp512 +8.5%*, neutral elsewhere. Kept.convert.cu4-wide vs strix 2-wide: identical within noise on 3 models; kernel geomean 1.004. Kept upstream's.Correctness:
[CONCAT]MUL_MAT cases and sixTOP_K extreme=1cases, all pre-existing.llama-perplexity --kl-divergence,-c 512):Setup: 4
tests/corpusfiles, each repeated to ≥12 KB. Hashes in the build log: code112a1f7f, numericdf11b694, prose56ec4dc3, structuredff746217.HIP: qwen4exp, gpt-oss and Qwen3.6 × 4 corpora: mean KLD ≤ 0.000001 everywhere, top-1 ≥ 99.94% (100% in 9 of 12).
Vulkan: mean KLD 0.004-0.061, top-1 90-98%. Upstream's int8 path quantizes activations to q8_1, so this is not math-preserving, and byte-identical output does not apply.
To check whether it is a quality loss, both builds were compared against a CPU reference. The PR is closer to the reference than strix main in 6 of 8 runs and equal within error in 2:
qwen4exp Vulkan PPL: 5.058 vs 5.090 on base (prose, -0.6%, within error).
Additional information
Follow-ups, not in this PR:
tc_mmqidconfigs.TOP_K extreme=1, and the[CONCAT]MUL_MAT cases.Requirements
-DGGML_VULKAN_RUN_TESTS/CHECK_RESULTS(debug.cpp was taken from upstream unchanged)batch_extmigration touched it.