Skip to content

Merge halo-box/llama.cpp master into strix (2026-09-15) - #64

Merged
dzannotti merged 81 commits into
masterfrom
sync/halo-master-2026-09-15
Sep 17, 2026
Merged

dzannotti merged 81 commits into
masterfrom
sync/halo-master-2026-09-15

Conversation

@dzannotti

@dzannotti dzannotti commented Sep 15, 2026 •

Copy link
Copy Markdown
Collaborator

Brings strix-llama.cpp up to date with halo-box/llama.cpp master (was 77 behind /
164 ahead), and brings in strix master's #63 (pwilkin/strix-halo), which landed after
this branch was cut. This is a real merge commit — do not squash it. Squashing an upstream
sync is what left halo-box stuck at "72 commits behind" and silently dropped two
upstream commits; see SYNC.md in halo-box.

Conflict resolutions (5 files)

CONTRIBUTING.md — kept strix's policy. Agents may author code, commit messages
and PR descriptions, and open PRs, per AGENTS.md. The incoming side was upstream's
stricter policy (mandatory disclosure, ~1h review per 200-400 LOC, no AI-written
posts); explicitly rejected.

tests/test-backend-ops.cpp — union, both sides kept. strix's Strix Halo probes
(MALL-spill, L2-residence, KV-cache-layout) and upstream's sparse-FA cases are
independent list entries. Verified the merged test_flash_attn_ext constructor
accepts both call signatures: it combines strix's permute/kv_view with upstream's
v_is_view_of_k/n_kv_max, and VARS_TO_STR17 exists to match the 17 vars.

src/models/qwen4exp.cpp — resolved per hunk, not wholesale:

  • Kept strix — the rms_norm+mul fusion (gamma reshaped to [n_embd, hc] and
    force-expanded so RMS_NORM and MUL stay graph-adjacent for the backend fuser).
    Traced to 6130b7262 hip: optimize RDNA3.5 MoE inference paths. Upstream's
    incoming 41abbfd59 does the same thing less carefully as a one-liner.
  • Took halo-box — the PLE-disk-entangled hunks, per halo-box Merge: from halobox 2026 09 06 uma ring #29.

ggml/src/ggml-vulkan/ — reverted wholesale to strix (5 files). See below.

Vulkan: upstream sparse FA deliberately NOT integrated

Upstream fc82583e6 vulkan: support sparse Flash Attention (ggml-org#28105) cannot
be merged mechanically with strix's Vulkan stack:

  • Both claim the same specialization-constant bit:
    strix DYNAMIC_KV = (Flags & 16) vs upstream USE_SPARSE = (Flags & 16).
  • They declare incompatible flash-attn push-constant layouts
    (strix: m_row_len, gqa_ratio, split_k_num, output_k_num, partial_output;
    upstream: sparse_base). A mismatched layout produces silent garbage, not a crash.
  • They are rival implementations of overlapping functionality — strix already has
    its own gather/compaction family (flash_attn_gather{,_dq,_union,_union_dq}.comp)
    and a DeepSeek-V4 sparse split path keyed on bit 31 of gqa_ratio.

Taking either side alone breaks the build: upstream's sparse code merged cleanly
into flash_attn.comp, flash_attn_cm1.comp and flash_attn_cm2.comp (strix had
not touched those regions), so keeping strix's flash_attn_base.glsl leaves those
three files calling a now-undefined fa_kv_index() / USE_SPARSE / data_sparse.

Reverting the whole directory to strix restores a self-consistent, known-good state.
Orphaned upstream shader flash_attn_sparse_compact.comp removed (the shader
generator uses an explicit list, so nothing referenced it).

Also deferred by this: f1e44dcc1 vulkan: workaround NV queuesubmit driver bug.

Follow-up PR needed to integrate sparse FA properly: renumber one feature's flag
bit, reconcile the push-constant structs and their C++ writer, and decide whether
DYNAMIC_KV and USE_SPARSE coexist or one supersedes. That work needs a GPU.

Update 2026-09-17: #63 merged in (82aed6017)

pwilkin's #63 landed on master after this PR was opened and conflicted in 4 files, all
around the qwen4exp PLE n-gram table. #63 builds on the PLE disk reader this sync had
deleted
:

  • --lazy-mode on-direct (LLAMA_LAZY_MODE_DIRECT) reads PLE rows through the reader
    with page-cached pread()s
  • llama_ple_disk::prefetch() / page_cached(), and qwen4exp_ple_prefetch() called
    from llama_context for batches of 4096+ tokens
  • the host_gather path and mmap madvise prefetch, kept and extended
  • tests/test-ple-disk.cpp

Taking this branch's deletion would have silently dropped all of it. Resolution:

halo-box #29's user-facing cleanup stays. --ngram-on-disk is an alias for
--lazy-mode on, and --ngram-io-threads/-cache/-direct-io and the ple_*
llama_model_params fields are removed. Net effect versus pre-sync strix: only
O_DIRECT reads and the tuning knobs are gone. The prefetch and host_gather paths
the original description listed as collateral loss are back.

Audit: every other file #63 touched is byte-identical to master. Where the merge differs
from master, the difference is exactly this branch's own delta from the merge base.

Also merged: halo-box #31 (SYNC.md, docs only). Now 0 behind halo-box, 0 behind strix
master.

Verification

Built and tested on a Strix Halo (Ryzen AI MAX+ 395 / Radeon 8060S, gfx1151) in
containers: rocm/dev-ubuntu-24.04:7.2.1-complete for HIP, ubuntu:26.04 for Vulkan.

result
HIP build (GGML_HIP=ON, GPU_TARGETS=gfx1151, tests on) ✅ 602/602 targets
Vulkan + CPU build (GGML_VULKAN=ON, tests on) ✅
test-backend-ops -b ROCm0 (full) ✅ 29695/29695
test-backend-ops -b Vulkan0 FLASH_ATTN_EXT / MUL_MAT_ID / RMS_NORM / GET_ROWS ✅ 5339 / 7498 / 51 / 590
test-llama-archs (ROCm) ✅ no failures; qwen4exp OK (9.97e-14) on GPU
test-ple-disk ✅ 960 gathers byte-identical to CPU dequant, all configs
test-qsa-prefix ✅
test-arg-parser ✅

PP/TG: not measured yet. An A/B against master 0636c9aee ran, but the box was
serving inference at the same time, so the numbers are discarded.

Not verified: no qwen4exp GGUF was available, so the on-direct / prefetch /
host_gather paths are covered only by test-ple-disk and test-llama-archs, not by
a real model run. CPU-only test-backend-ops was not run. The PP regression reported
on #63 (HIP 700→400, Qwen3.8-27B IQ4_XS halved) predates this PR, and I didn't
reproduce it. Tracked in #65's doc.

CI: the red checks on this PR are jobs that never got a runner and timed out after
exactly 24h with no logs. They aren't build failures.

Original description: PLE on disk (partly superseded by the update above)

Deliberately dropped: PLE on disk

halo-box #29 replaced the bespoke --ngram-on-disk with an alias to upstream
--lazy-mode, removing the custom Qwen4Exp PLE disk reader, O_DIRECT handling, row
cache, worker-thread controls, C API params and build wiring. strix still carried the
full implementation; this merge follows halo-box.

Note this was not visible as a conflict. strix had not modified those files since
the merge-base, so git applied halo-box's deletions silently. Only qwen4exp.cpp
conflicted — 1 of 12 affected files. What went:

src/llama-ple-disk.cpp, src/llama-ple-disk.h   (deleted)
src/CMakeLists.txt, src/llama-model.cpp        (wiring reverted)
common/common.{cpp,h}, include/llama.h         (C API params reverted)
src/models/models.h, README.md, docs/          (reverted)

Collateral loss, both from 6130b7262 and both referencing ple_disk, so neither
survives the feature's removal as written:

  • madvise(MADV_WILLNEED) batch prefetch for random mmap PLE rows (each row is
    otherwise a serial page fault, ~0.2 ms from NVMe).
  • the host_gather path, which avoided a CPU get_rows node splitting the graph
    into GPU→CPU→GPU with an extra launch and device sync per ubatch.

The host_gather optimization has a non-disk half (host-resident buffers) that could
be preserved independently. Not attempted here — it would be new hybrid code, and it
needs benchmarking to prove it is worth it.

ngxson and others added 30 commits September 12, 2026 00:53
* server: refactor subproc handling

* fix Windows build

* download: keep concurrent downloads of one blob apart

Every process writes the same path + .downloadInProgress, so a second
download of the same blob finds that file, takes it for its own partial
transfer and asks for the bytes after it, which produces a corrupt
result. The in-progress file now carries the pid of the process writing
it.

std::rename also replaces an existing destination on POSIX but fails on
Windows, so a download whose blob appeared in the meantime is dropped
after every retry and an etag rewrite silently keeps the old value.
std::filesystem::rename has the POSIX behaviour everywhere, and the
error now carries the reason reported by the system.

* Revert "download: keep concurrent downloads of one blob apart"

This reverts commit 917b83f.

* tests: serialize the router tests that download the same model

Parallel workers share one cache, so the two tests fetch the same blob
into the same in-progress file and race to rename it. They now take a
file lock around the download, like the session fixture does for the
preset models.

* Revert "tests: serialize the router tests that download the same model"

This reverts commit c368a4a.

---------

Co-authored-by: Pascal <admin@serveurperso.com>
* ggml-webgpu: Update to a recent version of Dawn

* No module scanning

* Accept review suggestion to update comment

Co-authored-by: Masashi Yoshimura <yoshimura.masashi.frbs@gmail.com>

---------

Co-authored-by: Masashi Yoshimura <yoshimura.masashi.frbs@gmail.com>
…rg#28589)

* hex-row-split: add support for multi-device row spliting

Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com>

* hex-mdev: add work splitting to fused kernels

* hex-mdev: use mdev_ prefix for all multi-device state

* hex-mdev: make device configuration more expressive to support device groups

* hex-mdev: fix mdev session init

* hex-mdev: fused nx (2x,3x) matmuls must update row counts for each w/o

* hex-mdev: fix MUL_MAT work partitioning bugs introduced by mdev

* hex-cont: fix crashes with new tests due to wrong striding

* hex-mdev: move fences after l2flushes

* hex-cont: fix work splitting for mnpu -- align chunks to cachelines

* hex-mdev: fix CPY tests with multi-dev

* hex-mmid: fix work partitioning with mnpu

* hex-mm: fix test failures with mdev

* hex-binary: fix work partitioning for mdev

* hex-argsort: fix mdev partitioning

* hex-mdev: fix work partitioning and general updates for all simple ops

* hex-fa: fix mdev work splitting issues

* hex-mdev: fixing more failing ops test

* hex-mdev: update the rest of the ops

* hex-mdev: refactor all mdev splitting logic to be contained within if (mdev_count > 1) {...}

* hex-mdev: fix macros

* hex-mdev: simplify session flush logic

* hex-sync: fix recursion in session flush

* hex-mdev: factor out fence buffer and allocator

* hex-fence: make fence allocation more robust with reserved slots for mdev

* hex-mdev: keep all mdev state in htp_mdev_group

* hex-mdev: further cleanup mdev group handling at the host

* hex-mdev: update group idx in the opbatch before serializing

* hex-batch: remove separate op_pending and use batch_req/rsp_seq

* hex-async: workaround another missing tensor_init in ggml-meta

* hex-fence: cleanup and robustify fences and error handling in multi-device scenarios

* hex-ar: improve ALLREDUCE error handling

* hex-async: robust error handling for op_cpy_fence

* hex-async: use seq0 from allreduce context to allocate fence_seq

* hex-mdev: fix remaining issues with fence and barrier clearing in CPY_FENCE

* hex-misc: realign macros and fix misplaces trace events

* hex-misc: align macros

* hex-mdev: fix unclone buffer re-entrancy

* hex-glu: fix mdev partitioning logic

* hex-mdev: make buffer uncloning/cleanup work with tensor-split scenarios

* hex-mdev: tighten up the can_split check in act-ops

* hex-mdev: factor out common bits of the partitioning logic

* hex-mm: minor realignment of the macros

* hex-bufs: fix incorrectly placed assert for MAX_BUFS

* hex-pad: tighten up gating checks for PAD

* hex-kparams: make sure all kernels properly use kparams->n_threads

* hex-docs: update user and developer docs with new features and detailed guide for ops development

* hex-scripts: update run script to properly parse dev groups

* hex-misc: formatting

* hex-sess: minor cleanup for session init

* hex-ar: fix vtcm size calc in allreduce kparams

* hex-scripts: fix flake8 warnings

* hex-rope: update ROPE to support mdev work split

* hex-ops: remove redunant checks and minor reformat

* hex-dev-guide: update dev-guide to avoid redundant null checks

* hex-async: improve event_wait, event_sync and fence implementations

* hex-async: remove synchronous flush from event_sync

* hex-async: symplify fence recovery protocol and make sync more robust

* hex-async: futher simplify error recovery for fences

* hex-err: return status instead of just -1

* hex-async: print all seq nums in hex

* hex-async: make sure fences flush dirty ranges

* hex-async: add dirty ranges merging to reduce fence flushes

* hex-async: properly sync before freeing the event

* hex-async: make sure fence owner session is not overriden

* hex-async: more fence write order more robust

* hex-async: make sure not to fuse ALLREDUCE+ADD if their dsts overlap

* hex-fusion: cleanup redundant checks

---------

Co-authored-by: Alexander Lu <alexlu@qti.qualcomm.com>
Walk the binding offset back until the distance to the tensor is a
whole number of blocks, so block quantized views get a valid element
offset in the shader.
…a8_bin` (ggml-org#28677)

* opencl: add A8 Q4_K non-MoE binary kernel

* opencl: fix layout compatibility

* opencl: rename binary kernel selection helpers

---------

Co-authored-by: Li He <lih@qti.qualcomm.com>
…g#28747)

The child writes its state commands on stdout while the logger writes
on stderr, and both share a single pipe. The logger emits the trailing
color reset after the newline of a debug, warn or error entry, so that
escape sequence has no newline of its own and the router reads it glued
in front of the next command. The line prefix check then fails and the
command is forwarded as a log line instead of being handled, which
leaves a finished download stuck in the downloading state.

Writing the command with a leading newline closes the pending line so
it always starts at a line boundary.
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
…org#28816)

Clang stores the modification time of the precompiled header sources
inside the header and refuses the header when they differ. A cached
header restored from another checkout carries the timestamps of that
checkout, so the build fails. The option covers the compilers ccache
treats as MSVC while they are clang underneath, clang-cl and the Intel
LLVM drivers.
…emas (ggml-org#28736)

* common : implement common_schema types

* common : implement a json schema optimizer

* common : reduce optimizations

* common : refactor json-schema-to-grammar to use common_schema

* common : use common_trie

* common/schema : implement type/kind resolution

* cont : cleanup

* cont : remove common_chat_tool_parameters

* cont : simplify schema resolution

* cont : pass common_schema through the json-schema-to-grammar builder

* cont : cleanup

* cont : move enums under common_schema and add type enum

* cont : reduce test cases

* cont : clean up

* cont : clean up

* refactor : rename common_schema_parse to common_schema_from_json

* tests : fix gcc dangling-reference warning in test-json-schema

* tests : take the schema label as const char * to satisfy gcc dangling-reference

* refactor : rename common_schema_builder parse_* methods to build_*

* cont : fix may_be_string

* cont : properly handle empty tool parameters

* cont : add tests for empty $ref

* cont : remove dead code

* cont : update docs

* cont : make "{}" mean any object for json_object as well

* cont : restore (min|max)Length to imply string type

* cont : rename common_schema to common_chat_schema
* add LOG_JSON macro

* fit: add demo LOG_JSON
* chat : improve schema support in qwen3 parser

* cont : clean up grammar a bit
There is a driver bug where two queues on the same VkDevice simultaneously
submitting can break some internal synchronization. Until it's fixed, add a
mutex around queuesubmit.
…gml-org#28833)

- Clamp the -j parallelism to min(nproc, 2) so a single-core runner
  uses -j 1 and multi-core runners use at most -j 2, instead of
  unconditionally using $(nproc).
- Add a 3600s timeout to both test-backend-ops runs (the high-perf CPU
  path and the default path) so a hung test cannot stall CI indefinitely.
- Note a TODO to reduce the timeout to 1800s in the future.

Assisted-by: pi:llama.cpp/Qwen3.8-27B
…28854)

Move the EditorConfig Checker and Code Style Checker workflows from the
`[self-hosted, fast]` runners to `ubuntu-slim`, which is an established
runner label in the repo.

Assisted-by: pi:llama.cpp/Qwen3.8-27B
* fix for unsupport zes API

* optimize the code

* adjust the log level

* rm unused head files

* Update docs/backend/SYCL.md

Co-authored-by: Titaniumtown <titaniumtown@proton.me>

* fix the error to detect level zero SDK/dev package, stop build after detect the error

* update the message

* fix the build error when missed to install level zero dev package

* rm GGML_SYCL_DEV_DEBUG, mv read env vars in all entry functions

---------

Co-authored-by: Neo Zhang Jianyu <jianyu.zhang@intel.com>
Co-authored-by: Titaniumtown <titaniumtown@proton.me>
Co-authored-by: Neo Zhang <NA>
…ml-org#28835)

Corrects a typo in `tests/test-quant-type-selection` for the
Nvidia Nemotron 3 Nano 30B A3B model, which was referred to as
*nvidia-nemotron-nano-3-30b-a3b*.

The error made the test skip that test case, rather than failing
the test.

[no release]
…ero divisor (ggml-org#28779)

The NextN/MTP tail loop derives the expert FFN size as n_ff/n_expert_used
when expert_feed_forward_length gives nothing for the layer. Both values come
from per-layer arrays that legitimately hold 0 on layers that are not MoE, so
a checkpoint whose predict layers hold 0 in both divides by zero and dies with
SIGFPE at load time, with no error message. Report the malformed metadata
instead.
Conflict resolutions:

- CONTRIBUTING.md: kept strix's policy (agents may author code, commit
  messages, PR descriptions, and open PRs per AGENTS.md).

- tests/test-backend-ops.cpp: union. strix's Strix Halo probes (MALL-spill,
  L2-residence, KV-layout) and upstream's sparse-FA cases are independent
  entries; the merged test_flash_attn_ext ctor supports both signatures.

- src/models/qwen4exp.cpp: split by hunk. Kept strix's rms_norm+mul fusion
  (graph-adjacent gamma view, from 6130b72 "hip: optimize RDNA3.5 MoE
  inference paths"). Took halo-box's side for the PLE-disk-entangled hunks,
  following halo-box #29, which deliberately replaced --ngram-on-disk with
  an alias to upstream --lazy-mode.

- ggml/src/ggml-vulkan/: reverted wholesale to strix. Upstream's sparse
  Flash Attention (ggml-org#28105) collides with strix's DYNAMIC_KV on
  Flags bit 16 and on the flash-attn push-constant layout. The two are
  rival implementations and cannot be merged mechanically. Deferred to a
  separate PR that can be compiled and validated on hardware.

Deliberately dropped with halo-box #29 (PLE on disk): src/llama-ple-disk.{h,cpp},
the CMake wiring, C API params, and the madvise(MADV_WILLNEED) row prefetch
plus host_gather graph-split avoidance, both of which referenced ple_disk.

NOT BUILT. Requires CPU, ROCm/HIP and Vulkan builds before merge.
dzannotti and others added 2 commits September 17, 2026 17:23
Brings pwilkin's strix-halo merge (#63) into the 2026-09-15 halo-box sync
(#64). #63 landed after the sync branch was cut and conflicted with it in
four files, all around the qwen4exp PLE n-gram table.

#63 builds on the llama_ple_disk reader that the sync removed (following
#29): it adds --lazy-mode on-direct (LLAMA_LAZY_MODE_DIRECT), which
routes PLE rows through that reader page-cached, plus prefetch()/
page_cached(), qwen4exp_ple_prefetch() called from llama_context, the
host_gather path and test-ple-disk. Taking the sync's deletion would have
silently dropped all of that, so the reader is kept as the backend of
on-direct only:

- src/llama-ple-disk.{cpp,h}: kept as in #63, re-added to src/CMakeLists.txt
- src/models/qwen4exp.cpp: #63's file, with the reader gated on
  LAZY_MODE_DIRECT alone (defaults: 64 readers, 256 MiB row cache, no
  O_DIRECT, which is what #63 chose for on-direct) and upstream's
  [n_embd, hc] TENSOR_ALLOW_RESHAPE gamma shapes from the sync re-applied,
  including #63's new nextn.hc_head_norm. The graph reshapes the gamma
  either way, so the load shape does not change the fused RMS_NORM+MUL.
- src/models/models.h: ple_disk member restored
- src/llama-model.cpp: init_mappings populates unless on-direct

#29's user-facing cleanup stays: --ngram-on-disk is an alias for
--lazy-mode on, and the --ngram-io-threads/-cache/-direct-io switches and
the ple_* llama_model_params fields remain removed.

Every other file #63 touched is byte-identical to strix master.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Docs only (SYNC.md). Leaves strix 0 commits behind halo-box.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
dzannotti added a commit that referenced this pull request Sep 17, 2026
#63 landed on master after #64 was cut and built on the PLE disk reader the
sync had removed. Merging it into the sync kept the reader as the backend of
--lazy-mode on-direct, so section 2 no longer describes madvise prefetch and
host_gather as dropped. Rewritten to say what actually changed (O_DIRECT and
the --ngram-* knobs are gone, on-direct and the prefetch paths stay) and to
A/B the lazy modes instead.

Adds a section for the prompt-processing regression reported on #63
(HIP gfx1151 ~700 -> ~400 t/s; Qwen3.8-27B IQ4_XS PP halved), unreproduced,
with a bench recipe and the shared HIP files to bisect first.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
@dzannotti
dzannotti merged commit 2ee5fe2 into master Sep 17, 2026
12 of 23 checks passed
@dzannotti
dzannotti deleted the sync/halo-master-2026-09-15 branch September 17, 2026 18:36
dzannotti added a commit that referenced this pull request Sep 17, 2026
PR #64 reverted ggml/src/ggml-vulkan/ wholesale rather than merge
ggml-org#28105: upstream's USE_SPARSE and strix's DYNAMIC_KV claim the same
Flags bit, the two declare incompatible flash-attn push-constant layouts, and
they are rival implementations of overlapping functionality.

Records the collision, what integration requires (renumber a bit, merge the
push-constant structs and their C++ writer, decide which mechanism survives),
and the test-backend-ops runs that gate it -- including that a mismatched
push-constant layout yields silently wrong tensors rather than a crash, so it
has to be validated numerically.

Docs only.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.