Skip to content

Merge halo-box/llama.cpp master into strix (2026-09-29) - #111

Merged
pwilkin merged 271 commits into
masterfrom
sync/halo-master-2026-09-29
Sep 30, 2026
Merged

pwilkin merged 271 commits into
masterfrom
sync/halo-master-2026-09-29

Conversation

@dzannotti

@dzannotti dzannotti commented Sep 29, 2026 •

Copy link
Copy Markdown
Collaborator

Overview

Merges halo-box/llama.cpp master (e57160642, 215 commits, including halo's 2026-09-29 upstream sync) into strix. This is a real merge commit, do not squash it (see SYNC.md in halo-box). After it lands, strix is 0 commits behind halo.

21 files conflicted. Each hunk was traced to the commit and PR on each side. Where the two sides competed, the decision was made by measuring on this hardware, not by picking a side.

Vulkan: neither side wholesale

Syncs #47 and #64 had reverted upstream's matmul framework (ggml-org#25773), and every upstream Vulkan matmul change since then builds on it. The two stacks were measured head to head: strix origin/master against halo master, same session, ABBA order.

  • Matmul: upstream is faster.
    • test-backend-ops perf: MUL_MAT geomean 1.10x, MUL_MAT_ID geomean 1.14x in upstream's favour.
    • q6_K and q8_0 up to 2x faster.
    • IQ4_XS: strix hit a slow path, up to 13.9x slower.
  • Everything else: strix is faster. Per-op timings (GGML_VK_PERF_LOGGER, pp2048):
    • Upstream CONCAT is 12-25x slower on the delta-net models: 335 ms vs 28 ms per graph on Qwen3.6-35B, 1448 ms vs 58 ms on Signal-3.8-27B.
    • On Signal-3.8-27B, upstream FA, SSM_CONV and GATED_DELTA_NET are 2-3x slower.

So the resolution is upstream's matmul framework plus strix's non-matmul Vulkan stack. That stack covers FA (dynamic KV, dequant-once, contiguize, gather/union, DeepSeek V4 sparse), transposed CONCAT, the lightning indexer, mat-vec chunking and submit bounding, moved into upstream's new ggml-vulkan-{types,push-constants,common}.h split.

  • ROCmFPx now uses upstream's per-type (LUT) matmul path and q8_1 MMQ list.
  • Hyper-connection ops are upstream's. The merged qwen4exp graph emits gated and identity-comb forms that the strix kernels could not run. Strix's supports_op would dereference a null comb.
  • Upstream sparse FA (vulkan: support sparse Flash Attention ggml-org/llama.cpp#28105) stays deferred. docs/development/vulkan-sparse-fa-deferred.md is updated with this sync's decisions.

Retired: strix's env-gated mul_mat_id experiments (GGML_VK_MMID_*, GGML_VK_DENSE_F16B, GGML_VK_MMID_SCALE_EPILOGUE). Two measured strix wins are lost and not yet ported onto the new framework (see follow-ups):

  • f16-weight MoE at n=128-256
  • iq3_xxs MoE, about -22 ms/graph on Qwen3.6

HIP

  • MMQ: strix's RDNA3.5 tuning (prefetch whitelist, routed-compact MoE, per-expert J tables, Ling J32 paths) is kept around upstream's new prec_src1 template parameter. switch_type defaults prec_src1 to Q8, so the paired-MMQ call sites are unchanged.
  • ggml-cuda.cu alloc deps: strix's fusion loops are kept. One if (op != GGML_OP_MUL) continue; is dropped, because it would have made upstream's top-k MoE alloc deps unreachable.
  • qwen4exp hyper-connection ops: upstream's fused DSV4_HC_PRE/HC_POST replaced strix's hc_mix/hc_combine fusions and cost -8.6% pp2048. All four combinations were measured (see below). The HIP backend now declines the gated PRE and comb-less POST on RDNA3.5. New LLAMA_FUSED_DSV4_HC_PRE/_POST=0 env overrides follow the existing LLAMA_FUSED_GDN_CH pattern.

Breakage that merged without a conflict (fixed)

  • llama-context.cpp: the PLE prefetch read batch_inp.token, which the new llama_batch_ext does not have.
  • llama-memory-hybrid-idx: upstream's causal_attn branch landed inside strix's set_input_qsa_scan. causal_attn is now passed through, and the causal-only prefix fast path, block selection and qsa_cache are gated on it.
  • ggml-cuda.cu: hint had become undeclared after upstream's FWHT refactor.
  • norm.cu: strix's ncols == 128 RMS_NORM launch was missing upstream's new scale argument.
  • test-backend-ops.cpp: the init_mul_mat_id_tensors forward declaration no longer matched the definition.
  • test-save-load-state.cpp: both sides claimed "Test 9". Strix's test becomes Test 10.
  • CI: upstream split build-self-hosted.yml into 7 ci-self-hosted-*.yml files. All 7 get strix's branches-ignore on pull_request, so the 24h runner timeouts seen on Merge halo-box/llama.cpp master into strix (2026-09-15) #64 don't come back.

Measurements

Device:     Ryzen AI Max+ 395 / Minisforum MS-S1 MAX
Memory:     128 GB LPDDR5X-8000
Power:      governor performance, no platform_profile exposed
BIOS:       UMA split 1 GB (GTT via ttm.pages_limit=30408704)
Kernel:     7.0.0-34-generic, amd_iommu=off ttm.pages_limit=30408704
Backend:    Vulkan RADV, Mesa 26.0.8 (Ubuntu 26.04 container); ROCm HIP 7.15 (rocm10.0 toolchain container)
Build:      -DCMAKE_BUILD_TYPE=Release -DGGML_NATIVE=ON -DLLAMA_BUILD_TESTS=ON; +-DGGML_VULKAN=ON | +-DGGML_HIP=ON -DGPU_TARGETS=gfx1151
Baseline:   strix origin/master 7a9196dad (built and run in this session); halo master e57160642 as the second parent
Change:     a4a4bfa93
Models:     Qwen3.8-Flash-Next GSQ-RCO Q2_0 (qwen4exp), agentionai/Signal-3.8-27B AP-Q4_K_XL, unsloth/Qwen3.6-35B-A3B UD-Q3_K_XL,
            gpt-oss-20b MXFP4, google gemma-4-26B-A4B q4_0 + unsloth UD-Q4_K_XL, Qwen3-1.7B Q8_0

All runs: -ngl 99 -fa 1 -b 2048 -ub 2048, ABBA order, 2 reps per run, so n=4 per arm. Mean ± sd in t/s. * means the difference is larger than the combined sd. Background GPU workloads were stopped for the whole window.

Vulkan, strix main vs this PR:

model pp512 pp2048 tg128
Qwen3.8-Flash-Next Q2_0 561.9±3.7 → 596.5±4.8 +6.2%* 623.0±0.4 → 767.1±0.6 +23.1%* 34.0±0.0 → 35.7±0.0 +5.2%*
Signal-3.8-27B Q4_K_XL 405.7±0.9 → 449.6±0.5 +10.8%* 387.9±1.6 → 438.4±2.0 +13.0%* 12.6 → 12.7 +0.5%*
Qwen3.6-35B-A3B Q3_K_XL 1466.8±4.2 → 1582.8±4.6 +7.9%* 1776.4±2.3 → 2065.0±1.8 +16.2%* 66.5 → 66.7 +0.2%
gpt-oss-20b MXFP4 1920.4±2.8 → 2387.5±29.1 +24.3%* 2236.9±3.3 → 2561.2±4.6 +14.5%* 79.8 → 80.1 +0.3%*
gemma-4-26B-A4B q4_0 1774.7±4.4 → 2012.9±9.4 +13.4%* 2170.3±1.4 → 2268.5±0.8 +4.5%* 73.8 → 73.9 +0.2%*
gemma-4-26B-A4B UD-Q4_K_XL 1778.9±4.7 → 2019.7±4.0 +13.5%* 2177.8±0.6 → 2279.9±0.9 +4.7%* 78.7 → 78.6 -0.1%
Qwen3-1.7B Q8_0 6743.8±5.4 → 6324.4±45.7 -6.2%* 5315.1±78.0 → 6415.9±54.4 +20.7%* 114.6 → 114.5 -0.1%

Vulkan, halo master vs this PR (what taking halo's Vulkan wholesale would have cost):

model pp512 pp2048 tg128
Qwen3.8-Flash-Next 560.7 → 598.1 +6.7%* 524.8±7.9 → 767.2±0.8 +46.2%* 34.0 → 35.7 +5.1%*
Signal-3.8-27B 423.8 → 449.1 +6.0%* 337.0±4.5 → 438.9±1.2 +30.3%* =
Qwen3.6-35B 1520.5 → 1581.9 +4.0%* 1526.0±2.7 → 2065.4±1.7 +35.4%* 67.1 → 66.6 -0.6%*
gpt-oss-20b = = =

HIP, strix main vs this PR:

model pp512 pp2048 tg128
Qwen3.8-Flash-Next 646.6±44.5 → 649.5±47.3 +0.5% 902.3±21.0 → 908.4±20.1 +0.7% 29.9±0.2 → 30.2±0.1 +1.1%*
Qwen3.6-35B 1719.1±149.7 → 1925.1±176.8 +12.0% 2150.3±54.7 → 2338.7±57.3 +8.8%* 59.1 → 60.8 +2.7%*
gpt-oss-20b 1944.4±98.3 → 2117.6±72.4 +8.9%* 2492.3 → 2479.7 -0.5% =
Signal-3.8-27B = = 12.1 → 12.1 +0.3%*

AGENTS.md prefill protocol, Qwen3.8-Flash-Next:

  • Command: llama-bench -p 2048 -d 0,12000,32000,64000 -b 2048 -ub 2048 -n 0 -r 4 -ngl 99 -fa on -ctk f16 -ctv f16 --load-mode none
  • Order: base/cand/cand/base, so n=8 per arm.
depth HIP base HIP PR gain Vulkan base Vulkan PR gain
0 915.9±15.5 916.7±17.6 +0.1% 621.3±1.0 765.9±1.5 +23.3%
12000 863.3±8.1 862.8±7.9 -0.1% 489.8±1.3 594.8±0.3 +21.4%
32000 832.2±7.0 834.6±6.7 +0.3% 393.4±1.2 454.2±1.0 +15.4%
64000 793.9±7.6 793.6±6.1 -0.0% 276.9±3.0 308.9±2.8 +11.5%

qwen4exp HIP hyper-connection decision (strix main = 910.6 pp2048 / 30.1 tg128):

HC ops pp512 pp2048 tg128
both fused (upstream default) -7.2% -8.6%* +2.2%*
pre fused, post unfused -7.0% -4.7%* +4.5%*
post fused, pre unfused -4.4% -6.4%* -1.6%*
both unfused (this PR on RDNA3.5) +0.9% -0.2% +0.5%*

Isolated A/Bs of the other contested choices (toggle build, not part of this PR):

  • Upstream ncols_opt on RDNA3.5: gpt-oss pp512 +8.5%*, neutral elsewhere. Kept.
  • convert.cu 4-wide vs strix 2-wide: identical within noise on 3 models; kernel geomean 1.004. Kept upstream's.
  • Vulkan q8_0, int8 vs f16: Qwen3-1.7B pp512 -5.0%, pp2048 +15.9%; Signal +1.2%. Kept int8.
  • Vulkan MXFP4, int8 vs f16: pp512 +15.7%, pp2048 +9.9%. Kept int8.

Correctness:

  • test-backend-ops ROCm0 (full): 29804/29804 on base, 30197/30197 on this PR.
  • test-backend-ops Vulkan0 (full): 33745/33799 on base, 33875/33929 on this PR.
    • Outside the QSA probes the failure sets are identical: three [CONCAT] MUL_MAT cases and six TOP_K extreme=1 cases, all pre-existing.
    • The QSA_* probes fail on base too. Run alone, base fails 0/25 QSA_DECODE_MASKLESS, and full runs fail a varying subset. Pre-existing, see follow-ups.
  • Targeted Vulkan0 on this PR:
    • MUL_MAT_ID 7530/7530, FLASH_ATTN_EXT 5344/5344
    • DSV4_HC_PRE 6/6, HC_POST 9/9, HC_COMB 11/11
    • GET_ROWS 591/591
  • Logits vs strix main (llama-perplexity --kl-divergence, -c 512):
    • Setup: 4 tests/corpus files, each repeated to ≥12 KB. Hashes in the build log: code 112a1f7f, numeric df11b694, prose 56ec4dc3, structured ff746217.

    • HIP: qwen4exp, gpt-oss and Qwen3.6 × 4 corpora: mean KLD ≤ 0.000001 everywhere, top-1 ≥ 99.94% (100% in 9 of 12).

    • Vulkan: mean KLD 0.004-0.061, top-1 90-98%. Upstream's int8 path quantizes activations to q8_1, so this is not math-preserving, and byte-identical output does not apply.

    • To check whether it is a quality loss, both builds were compared against a CPU reference. The PR is closer to the reference than strix main in 6 of 8 runs and equal within error in 2:

      vs CPU reference strix main this PR
      gpt-oss prose 0.0226 / 90.8% 0.0177 / 92.1%
      gpt-oss code 0.0065 / 97.3% 0.0031 / 98.8%
      gpt-oss structured 0.0497 / 93.2% 0.0445 / 94.0%
      gpt-oss numeric 0.0468 / 89.2% 0.0360 / 91.2%
      Qwen3.6 prose 0.0066 / 96.2% 0.0051 / 96.4%
      Qwen3.6 code 0.0041 / 98.3% 0.0037 / 98.0%
      Qwen3.6 structured 0.0373 / 96.6% 0.0404 / 96.4% (±0.0085)
      Qwen3.6 numeric 0.0047 / 98.1% 0.0044 / 98.2%
    • qwen4exp Vulkan PPL: 5.058 vs 5.090 on base (prose, -0.6%, within error).

Additional information

Follow-ups, not in this PR:

  1. HIP qwen4exp: choose fused HC_PRE per batch size. The "pre fused" configuration is +4.5% decode, but it costs prefill as long as the choice is per context.
  2. Vulkan: port strix's f16-expert MoE tiles and iq3_xxs MoE tile choice onto upstream's tc_mmqid configs.
  3. Vulkan q8_0 at small n: the int8 path loses 5% at pp512 on Qwen3-1.7B. A dense tile/threshold tweak may recover it.
  4. Pre-existing Vulkan test failures: the QSA_* probes, TOP_K extreme=1, and the [CONCAT] MUL_MAT cases.
  5. Upstream sparse FA (vulkan: support sparse Flash Attention ggml-org/llama.cpp#28105) remains deferred, per the doc.

Requirements

  • I have read and agree with the contributing guidelines
  • This change is Strix Halo specific, or justified by measurements on Strix Halo. General llama.cpp improvements belong in halo-box/llama.cpp instead
  • AI usage disclosure: AGENT-AUTHORED. Claude (Opus 5.5) resolved the merge, ported the Vulkan stack, and ran every build, test and benchmark above on this machine. It also wrote this description. @dzannotti directed the work and reviews.
  • What was NOT verified:
    • ROCmFPx: no ROCmFPx GGUF was available, so the new upstream-framework path is covered only by the build and generic tests, not by a model run or a perf measurement.
    • Models not run: DeepSeek-V4 / GLM-DSA models, so the DSV4 sparse FA and indexer paths are covered only by test-backend-ops.
    • Builds not run:
      • -DGGML_VULKAN_RUN_TESTS / CHECK_RESULTS (debug.cpp was taken from upstream unchanged)
      • coopmat2 (NVIDIA) and non-RDNA3.5 HIP paths
      • CPU-only and other backends beyond compiling
    • Speculative decoding (llama-benchy): not run. No speculative code changed on the strix side, but upstream's batch_ext migration touched it.

mctylr-gh and others added 30 commits September 16, 2026 22:15
* vulkan: support qwen4exp hc ops

* fix stale comment [no-ci]
…8996)

Both im2col.comp and im2col_3d.comp declare D_ptr without an explicit
  buffer_reference_align, so glslang emits writes through it as Aligned
  16. The shaders advance the pointer by D_SIZE, a per-variant define
  set to 4 for float and 2 for float16_t, so most write addresses are
  not 16-byte aligned. This triggers
  VUID-RuntimeSpirv-PhysicalStorageBuffer64-06315 under GPU-AV.

  Declaring buffer_reference_align = D_SIZE matches the alignment to the
  actual write stride and takes validation hits from 20 to 0 for both
  IM2COL and IM2COL_3D.

  Fixes ggml-org#28960
* opencl: fix warnings

* opencl: fix warnings for non adreno
ggml-org#28993)

* gguf : align the data section relative to the GGUF start, not the file

gguf_init_from_file_ptr reads a GGUF from the current file position, but padded
the data section from file offset 0, so a GGUF embedded at an offset that is not
a multiple of the alignment loaded without error and returned wrong tensor data.

Also adds llama_adapter_lora_init_from_file_ptr, and disables mmap with a warning
when an embedded data section is not aligned, instead of asserting in ggml.

Assisted-by: Claude Opus 5

* llama : load lora from path through the FILE* variant

The test now checks that mmap is disabled only for an unaligned offset.

Assisted-by: Claude Fable 5.1

* Update ggml/src/gguf.cpp

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* Update include/llama.h

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>

* llama : error on unaligned mmap of an embedded GGUF, drop test-load-file-ptr

---------

Co-authored-by: Johannes Gäßler <johannesg@5d6.de>
…g#29008)

* chat : add message delimiters to the DeepSeek V3.2/V4 parser

Assisted-by: Claude
Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
* ci : add API/ABI check to make-release workflow [no ci]

This commit adds an API/ABI compatibility check to the make-release
workflow.

The motivation for this to allow us to detect any potential breaking
changes in API/ABI compatibility between releases and fail the the
release if there are any.

The workflow can be triggered manually as before and this check can be
skipped if needed as it does take some time which might be useful when
doing a dry-run and not specifically interested in the API/ABI check.

By default this will check the current release against the latest
release, but this can also be configured in the workflow, or in the
script run on the command line, to check a different tag.

* add check for minor version bumps [no ci]

This commit also changes the build type to be RelWithDebInfo so that the
reported information is more useful.
…org#27985)

* ui: fix accidentally removed reasoning menu in single model mode on desktop

* ui: formatting task run to fix storybook test

* ui: mount the add menu reasoning submenu outside router mode only

The models selector already owns the reasoning submenu in router mode,
so the add menu only mounts it in single model mode. The first enabled
item of the add menu is now the reasoning submenu, the accessibility
story expects it.

---------

Co-authored-by: Ben Babik <work@benjaminbabik.com>
Co-authored-by: Pascal <admin@serveurperso.com>
…org#29009)

* Update to openvino-2026.4

* Update OV docs

* ggml-openvino : fix clangd and MSVC warnings

* fix int to ptr cast, more internal linkage enforcement, and avoiding duplicate switch case

---------

Co-authored-by: Mostafa Faheem <mostafaaafaheem@gmail.com>
* first fix

* removed unnecessary declarations
required for qwen35moe if MTP tensors are fused but not loaded
… experts (ggml-org#28501)

* vulkan: raise the hoisted row-id limit for mul_mat_id to 512 experts

The expert-count shader (count_experts.comp) sizes its shared arrays
with BLOCK_SIZE, which is 256. Because of that, row-id hoisting is
switched off for any model with more than 256 experts, and every
mul_mat_id workgroup has to rescan the whole ids tensor on its own.
Qwen3.8-Flash-Next has 512 experts and was quietly running on that
slow path.

This change sizes the arrays with a separate MAX_EXPERTS constant (512),
clears them in a loop instead of one entry per thread, and raises the
matching limit on the host side.

On Strix Halo at batch 2048 the expert matmuls drop from 12.5 to 9.5 ms
(iq3_s) and from 14.0 to 7.5 ms (iq4_nl) per op, and prompt processing
gets about 19 % faster at 8k tokens. test-backend-ops MUL_MAT_ID passes
(891/891) with new 512-expert test cases.

Assisted-by: Claude Fable 5.1

* vulkan: raise the hoisted row-id limit for mul_mat_id to 1024 experts

Follow-up to review feedback: 1024 matches LLAMA_MAX_EXPERTS instead of
stopping at 512. The three shared arrays in count_experts.comp grow to
3 * 1024 * 4 = 12 KiB, which fits the 16 KiB that Vulkan guarantees for
maxComputeSharedMemorySize.

Adds mul_mat_id test cases at 1024 experts alongside the existing 512
ones. test-backend-ops MUL_MAT_ID passes on Vulkan (RADV, Strix Halo,
Radeon 8060S): 889/889.
…9036)

* gguf-py: fix Q8_1 block size in GGML_QUANT_SIZES

* --whitespace

---------

Co-authored-by: Sigbjørn Skjæret <sigbjorn.skjaeret@huggingface.co>
* fix: build fails when GGML_CPU=OFF and GGML_CUDA=ON

* fix: eol in examples/convert-llama2c-to-ggml/CMakeLists.txt file
* vocab : add ufakzeka pre-tokenizer

* vocab : move ufakzeka to the models list and regenerate the hash mapping
* vulkan: add IQ3_S MMQ matmul kernels

* Make block_a_to_shmem do 2-byte loads (110 bytes is divisible by 2)

* Align the check, IQ3_S is also using K tile size
* fix get_rows vec4 handling

* Add src strides checking to vec4_aligned of get_rows and the new test case.
…gml-org#29042)

* llama: read the SWA pattern as a period or a per-layer array

Add llama_model_base::load_swa_pattern(), which reads
sliding_window_pattern either as one flag per layer or as a period
expanded by set_swa_pattern(), and use it in every loader that reads
the key as a period.

These loaders silently ignored an array and applied their default
period, although the converters of olmo2, gemma3n and exaone4 write
arrays. The published GGUFs match the defaults, so their outputs do
not change. The loaders that already accepted both forms lose their
duplicated scalar-then-array block, and use their declared default
period when the key is absent.

* model-saver: write the SWA pattern and the MLA SWA geometry

Write sliding_window_pattern as one flag per layer, nextn layers
included, for every model using SWA. The array is never collapsed to
a scalar, since the loaders read a scalar as a period.

Also write the MLA key/value lengths and KV LoRA rank of the SWA
layers, required by dots3note.

This enables the saver for plamo3, gemma3, cohere2, cohere2moe,
olmo2, exaone-moe, afmoe, mimo2, spark2_5, muse-glimmer, mellum,
laguna, granite_swa, dots3note and maple, all passing the bit-exact
roundtrip of test-llama-archs.
* ggml : check for allocation failures to prevent crashes

* wording
tests/test-backend-ops.cpp: keep both new test cases, upstream's
test_mul_mat_id_w4a8/w4a4 from the sync and test_moe_prefill from #91.

Assisted-by: Claude Opus 5.5
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Keep ROCmFP4 FAST and MXFP4/NVFP4 W4A4 MMQ declarations.
Keep explicit rollback registrations, including Qwen4exp, without duplicate directory runs.
Combine the status/all-model harness with scaled-weight seq_cp/graph-reuse coverage.

Assisted-by: Codex
…ndency pass

Pass the inclusive final node index instead of the exclusive group end.
This keeps allocation dependencies in the graph when the group is terminal.

Assisted-by: Codex
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.