Merge halo-box/llama.cpp master into strix (2026-09-15) - #64
Merged
Merged
Conversation
* server: refactor subproc handling * fix Windows build * download: keep concurrent downloads of one blob apart Every process writes the same path + .downloadInProgress, so a second download of the same blob finds that file, takes it for its own partial transfer and asks for the bytes after it, which produces a corrupt result. The in-progress file now carries the pid of the process writing it. std::rename also replaces an existing destination on POSIX but fails on Windows, so a download whose blob appeared in the meantime is dropped after every retry and an etag rewrite silently keeps the old value. std::filesystem::rename has the POSIX behaviour everywhere, and the error now carries the reason reported by the system. * Revert "download: keep concurrent downloads of one blob apart" This reverts commit 917b83f. * tests: serialize the router tests that download the same model Parallel workers share one cache, so the two tests fetch the same blob into the same in-progress file and race to rename it. They now take a file lock around the download, like the session fixture does for the preset models. * Revert "tests: serialize the router tests that download the same model" This reverts commit c368a4a. --------- Co-authored-by: Pascal <admin@serveurperso.com>
* ggml-webgpu: Update to a recent version of Dawn * No module scanning * Accept review suggestion to update comment Co-authored-by: Masashi Yoshimura <yoshimura.masashi.frbs@gmail.com> --------- Co-authored-by: Masashi Yoshimura <yoshimura.masashi.frbs@gmail.com>
…rg#28589) * hex-row-split: add support for multi-device row spliting Co-authored-by: Max Krasnyansky <maxk@qti.qualcomm.com> * hex-mdev: add work splitting to fused kernels * hex-mdev: use mdev_ prefix for all multi-device state * hex-mdev: make device configuration more expressive to support device groups * hex-mdev: fix mdev session init * hex-mdev: fused nx (2x,3x) matmuls must update row counts for each w/o * hex-mdev: fix MUL_MAT work partitioning bugs introduced by mdev * hex-cont: fix crashes with new tests due to wrong striding * hex-mdev: move fences after l2flushes * hex-cont: fix work splitting for mnpu -- align chunks to cachelines * hex-mdev: fix CPY tests with multi-dev * hex-mmid: fix work partitioning with mnpu * hex-mm: fix test failures with mdev * hex-binary: fix work partitioning for mdev * hex-argsort: fix mdev partitioning * hex-mdev: fix work partitioning and general updates for all simple ops * hex-fa: fix mdev work splitting issues * hex-mdev: fixing more failing ops test * hex-mdev: update the rest of the ops * hex-mdev: refactor all mdev splitting logic to be contained within if (mdev_count > 1) {...} * hex-mdev: fix macros * hex-mdev: simplify session flush logic * hex-sync: fix recursion in session flush * hex-mdev: factor out fence buffer and allocator * hex-fence: make fence allocation more robust with reserved slots for mdev * hex-mdev: keep all mdev state in htp_mdev_group * hex-mdev: further cleanup mdev group handling at the host * hex-mdev: update group idx in the opbatch before serializing * hex-batch: remove separate op_pending and use batch_req/rsp_seq * hex-async: workaround another missing tensor_init in ggml-meta * hex-fence: cleanup and robustify fences and error handling in multi-device scenarios * hex-ar: improve ALLREDUCE error handling * hex-async: robust error handling for op_cpy_fence * hex-async: use seq0 from allreduce context to allocate fence_seq * hex-mdev: fix remaining issues with fence and barrier clearing in CPY_FENCE * hex-misc: realign macros and fix misplaces trace events * hex-misc: align macros * hex-mdev: fix unclone buffer re-entrancy * hex-glu: fix mdev partitioning logic * hex-mdev: make buffer uncloning/cleanup work with tensor-split scenarios * hex-mdev: tighten up the can_split check in act-ops * hex-mdev: factor out common bits of the partitioning logic * hex-mm: minor realignment of the macros * hex-bufs: fix incorrectly placed assert for MAX_BUFS * hex-pad: tighten up gating checks for PAD * hex-kparams: make sure all kernels properly use kparams->n_threads * hex-docs: update user and developer docs with new features and detailed guide for ops development * hex-scripts: update run script to properly parse dev groups * hex-misc: formatting * hex-sess: minor cleanup for session init * hex-ar: fix vtcm size calc in allreduce kparams * hex-scripts: fix flake8 warnings * hex-rope: update ROPE to support mdev work split * hex-ops: remove redunant checks and minor reformat * hex-dev-guide: update dev-guide to avoid redundant null checks * hex-async: improve event_wait, event_sync and fence implementations * hex-async: remove synchronous flush from event_sync * hex-async: symplify fence recovery protocol and make sync more robust * hex-async: futher simplify error recovery for fences * hex-err: return status instead of just -1 * hex-async: print all seq nums in hex * hex-async: make sure fences flush dirty ranges * hex-async: add dirty ranges merging to reduce fence flushes * hex-async: properly sync before freeing the event * hex-async: make sure fence owner session is not overriden * hex-async: more fence write order more robust * hex-async: make sure not to fuse ALLREDUCE+ADD if their dsts overlap * hex-fusion: cleanup redundant checks --------- Co-authored-by: Alexander Lu <alexlu@qti.qualcomm.com>
Walk the binding offset back until the distance to the tensor is a whole number of blocks, so block quantized views get a valid element offset in the shader.
…a8_bin` (ggml-org#28677) * opencl: add A8 Q4_K non-MoE binary kernel * opencl: fix layout compatibility * opencl: rename binary kernel selection helpers --------- Co-authored-by: Li He <lih@qti.qualcomm.com>
…g#28747) The child writes its state commands on stdout while the logger writes on stderr, and both share a single pipe. The logger emits the trailing color reset after the newline of a debug, warn or error entry, so that escape sequence has no newline of its own and the router reads it glued in front of the next command. The line prefix check then fails and the command is forwarded as a log line instead of being handled, which leaves a finished download stuck in the downloading state. Writing the command with a leading newline closes the pending line so it always starts at a line boundary.
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
Signed-off-by: Adrien Gallouët <angt@huggingface.co>
…org#28816) Clang stores the modification time of the precompiled header sources inside the header and refuses the header when they differ. A cached header restored from another checkout carries the timestamps of that checkout, so the build fails. The option covers the compilers ccache treats as MSVC while they are clang underneath, clang-cl and the Intel LLVM drivers.
…emas (ggml-org#28736) * common : implement common_schema types * common : implement a json schema optimizer * common : reduce optimizations * common : refactor json-schema-to-grammar to use common_schema * common : use common_trie * common/schema : implement type/kind resolution * cont : cleanup * cont : remove common_chat_tool_parameters * cont : simplify schema resolution * cont : pass common_schema through the json-schema-to-grammar builder * cont : cleanup * cont : move enums under common_schema and add type enum * cont : reduce test cases * cont : clean up * cont : clean up * refactor : rename common_schema_parse to common_schema_from_json * tests : fix gcc dangling-reference warning in test-json-schema * tests : take the schema label as const char * to satisfy gcc dangling-reference * refactor : rename common_schema_builder parse_* methods to build_* * cont : fix may_be_string * cont : properly handle empty tool parameters * cont : add tests for empty $ref * cont : remove dead code * cont : update docs * cont : make "{}" mean any object for json_object as well * cont : restore (min|max)Length to imply string type * cont : rename common_schema to common_chat_schema
* add LOG_JSON macro * fit: add demo LOG_JSON
* chat : improve schema support in qwen3 parser * cont : clean up grammar a bit
There is a driver bug where two queues on the same VkDevice simultaneously submitting can break some internal synchronization. Until it's fixed, add a mutex around queuesubmit.
…gml-org#28833) - Clamp the -j parallelism to min(nproc, 2) so a single-core runner uses -j 1 and multi-core runners use at most -j 2, instead of unconditionally using $(nproc). - Add a 3600s timeout to both test-backend-ops runs (the high-perf CPU path and the default path) so a hung test cannot stall CI indefinitely. - Note a TODO to reduce the timeout to 1800s in the future. Assisted-by: pi:llama.cpp/Qwen3.8-27B
Assisted-by: pi:llama.cpp/Qwen3.8-27B
…28854) Move the EditorConfig Checker and Code Style Checker workflows from the `[self-hosted, fast]` runners to `ubuntu-slim`, which is an established runner label in the repo. Assisted-by: pi:llama.cpp/Qwen3.8-27B
* fix for unsupport zes API * optimize the code * adjust the log level * rm unused head files * Update docs/backend/SYCL.md Co-authored-by: Titaniumtown <titaniumtown@proton.me> * fix the error to detect level zero SDK/dev package, stop build after detect the error * update the message * fix the build error when missed to install level zero dev package * rm GGML_SYCL_DEV_DEBUG, mv read env vars in all entry functions --------- Co-authored-by: Neo Zhang Jianyu <jianyu.zhang@intel.com> Co-authored-by: Titaniumtown <titaniumtown@proton.me> Co-authored-by: Neo Zhang <NA>
…ml-org#28835) Corrects a typo in `tests/test-quant-type-selection` for the Nvidia Nemotron 3 Nano 30B A3B model, which was referred to as *nvidia-nemotron-nano-3-30b-a3b*. The error made the test skip that test case, rather than failing the test. [no release]
…ero divisor (ggml-org#28779) The NextN/MTP tail loop derives the expert FFN size as n_ff/n_expert_used when expert_feed_forward_length gives nothing for the layer. Both values come from per-layer arrays that legitimately hold 0 on layers that are not MoE, so a checkpoint whose predict layers hold 0 in both divides by zero and dies with SIGFPE at load time, with no error message. Report the malformed metadata instead.
Conflict resolutions: - CONTRIBUTING.md: kept strix's policy (agents may author code, commit messages, PR descriptions, and open PRs per AGENTS.md). - tests/test-backend-ops.cpp: union. strix's Strix Halo probes (MALL-spill, L2-residence, KV-layout) and upstream's sparse-FA cases are independent entries; the merged test_flash_attn_ext ctor supports both signatures. - src/models/qwen4exp.cpp: split by hunk. Kept strix's rms_norm+mul fusion (graph-adjacent gamma view, from 6130b72 "hip: optimize RDNA3.5 MoE inference paths"). Took halo-box's side for the PLE-disk-entangled hunks, following halo-box #29, which deliberately replaced --ngram-on-disk with an alias to upstream --lazy-mode. - ggml/src/ggml-vulkan/: reverted wholesale to strix. Upstream's sparse Flash Attention (ggml-org#28105) collides with strix's DYNAMIC_KV on Flags bit 16 and on the flash-attn push-constant layout. The two are rival implementations and cannot be merged mechanically. Deferred to a separate PR that can be compiled and validated on hardware. Deliberately dropped with halo-box #29 (PLE on disk): src/llama-ple-disk.{h,cpp}, the CMake wiring, C API params, and the madvise(MADV_WILLNEED) row prefetch plus host_gather graph-split avoidance, both of which referenced ple_disk. NOT BUILT. Requires CPU, ROCm/HIP and Vulkan builds before merge.
Brings pwilkin's strix-halo merge (#63) into the 2026-09-15 halo-box sync (#64). #63 landed after the sync branch was cut and conflicted with it in four files, all around the qwen4exp PLE n-gram table. #63 builds on the llama_ple_disk reader that the sync removed (following #29): it adds --lazy-mode on-direct (LLAMA_LAZY_MODE_DIRECT), which routes PLE rows through that reader page-cached, plus prefetch()/ page_cached(), qwen4exp_ple_prefetch() called from llama_context, the host_gather path and test-ple-disk. Taking the sync's deletion would have silently dropped all of that, so the reader is kept as the backend of on-direct only: - src/llama-ple-disk.{cpp,h}: kept as in #63, re-added to src/CMakeLists.txt - src/models/qwen4exp.cpp: #63's file, with the reader gated on LAZY_MODE_DIRECT alone (defaults: 64 readers, 256 MiB row cache, no O_DIRECT, which is what #63 chose for on-direct) and upstream's [n_embd, hc] TENSOR_ALLOW_RESHAPE gamma shapes from the sync re-applied, including #63's new nextn.hc_head_norm. The graph reshapes the gamma either way, so the load shape does not change the fused RMS_NORM+MUL. - src/models/models.h: ple_disk member restored - src/llama-model.cpp: init_mappings populates unless on-direct #29's user-facing cleanup stays: --ngram-on-disk is an alias for --lazy-mode on, and the --ngram-io-threads/-cache/-direct-io switches and the ple_* llama_model_params fields remain removed. Every other file #63 touched is byte-identical to strix master. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Docs only (SYNC.md). Leaves strix 0 commits behind halo-box. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
dzannotti
added a commit
that referenced
this pull request
Sep 17, 2026
#63 landed on master after #64 was cut and built on the PLE disk reader the sync had removed. Merging it into the sync kept the reader as the backend of --lazy-mode on-direct, so section 2 no longer describes madvise prefetch and host_gather as dropped. Rewritten to say what actually changed (O_DIRECT and the --ngram-* knobs are gone, on-direct and the prefetch paths stay) and to A/B the lazy modes instead. Adds a section for the prompt-processing regression reported on #63 (HIP gfx1151 ~700 -> ~400 t/s; Qwen3.8-27B IQ4_XS PP halved), unreproduced, with a bench recipe and the shared HIP files to bisect first. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
dzannotti
added a commit
that referenced
this pull request
Sep 17, 2026
PR #64 reverted ggml/src/ggml-vulkan/ wholesale rather than merge ggml-org#28105: upstream's USE_SPARSE and strix's DYNAMIC_KV claim the same Flags bit, the two declare incompatible flash-attn push-constant layouts, and they are rival implementations of overlapping functionality. Records the collision, what integration requires (renumber a bit, merge the push-constant structs and their C++ writer, decide which mechanism survives), and the test-backend-ops runs that gate it -- including that a mismatched push-constant layout yields silently wrong tensors rather than a crash, so it has to be validated numerically. Docs only. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
This was referenced Sep 29, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Brings strix-llama.cpp up to date with
halo-box/llama.cppmaster (was 77 behind /164 ahead), and brings in strix master's #63 (pwilkin/strix-halo), which landed after
this branch was cut. This is a real merge commit — do not squash it. Squashing an upstream
sync is what left halo-box stuck at "72 commits behind" and silently dropped two
upstream commits; see SYNC.md in halo-box.
Conflict resolutions (5 files)
CONTRIBUTING.md— kept strix's policy. Agents may author code, commit messagesand PR descriptions, and open PRs, per AGENTS.md. The incoming side was upstream's
stricter policy (mandatory disclosure, ~1h review per 200-400 LOC, no AI-written
posts); explicitly rejected.
tests/test-backend-ops.cpp— union, both sides kept. strix's Strix Halo probes(MALL-spill, L2-residence, KV-cache-layout) and upstream's sparse-FA cases are
independent list entries. Verified the merged
test_flash_attn_extconstructoraccepts both call signatures: it combines strix's
permute/kv_viewwith upstream'sv_is_view_of_k/n_kv_max, andVARS_TO_STR17exists to match the 17 vars.src/models/qwen4exp.cpp— resolved per hunk, not wholesale:[n_embd, hc]andforce-expanded so RMS_NORM and MUL stay graph-adjacent for the backend fuser).
Traced to
6130b7262 hip: optimize RDNA3.5 MoE inference paths. Upstream'sincoming
41abbfd59does the same thing less carefully as a one-liner.ggml/src/ggml-vulkan/— reverted wholesale to strix (5 files). See below.Vulkan: upstream sparse FA deliberately NOT integrated
Upstream
fc82583e6 vulkan: support sparse Flash Attention (ggml-org#28105)cannotbe merged mechanically with strix's Vulkan stack:
strix
DYNAMIC_KV = (Flags & 16)vs upstreamUSE_SPARSE = (Flags & 16).(strix:
m_row_len, gqa_ratio, split_k_num, output_k_num, partial_output;upstream:
sparse_base). A mismatched layout produces silent garbage, not a crash.its own gather/compaction family (
flash_attn_gather{,_dq,_union,_union_dq}.comp)and a DeepSeek-V4 sparse split path keyed on bit 31 of
gqa_ratio.Taking either side alone breaks the build: upstream's sparse code merged cleanly
into
flash_attn.comp,flash_attn_cm1.compandflash_attn_cm2.comp(strix hadnot touched those regions), so keeping strix's
flash_attn_base.glslleaves thosethree files calling a now-undefined
fa_kv_index()/USE_SPARSE/data_sparse.Reverting the whole directory to strix restores a self-consistent, known-good state.
Orphaned upstream shader
flash_attn_sparse_compact.compremoved (the shadergenerator uses an explicit list, so nothing referenced it).
Also deferred by this:
f1e44dcc1 vulkan: workaround NV queuesubmit driver bug.Follow-up PR needed to integrate sparse FA properly: renumber one feature's flag
bit, reconcile the push-constant structs and their C++ writer, and decide whether
DYNAMIC_KVandUSE_SPARSEcoexist or one supersedes. That work needs a GPU.Update 2026-09-17: #63 merged in (
82aed6017)pwilkin's #63 landed on master after this PR was opened and conflicted in 4 files, all
around the qwen4exp PLE n-gram table. #63 builds on the PLE disk reader this sync had
deleted:
--lazy-mode on-direct(LLAMA_LAZY_MODE_DIRECT) reads PLE rows through the readerwith page-cached
pread()sllama_ple_disk::prefetch()/page_cached(), andqwen4exp_ple_prefetch()calledfrom
llama_contextfor batches of 4096+ tokenshost_gatherpath and mmapmadviseprefetch, kept and extendedtests/test-ple-disk.cppTaking this branch's deletion would have silently dropped all of it. Resolution:
src/llama-ple-disk.{cpp,h}: kept as in pwilkin/strix-halo official merge #63, re-added tosrc/CMakeLists.txt.src/models/qwen4exp.cpp: pwilkin/strix-halo official merge #63's file with two changes. The reader is gated onLAZY_MODE_DIRECTalone, using its defaults (64 readers, 256 MiB cache, no O_DIRECT,which is what pwilkin/strix-halo official merge #63 already chose for on-direct). Upstream's
[n_embd, hc]TENSOR_ALLOW_RESHAPEgamma shapes are re-applied, including pwilkin/strix-halo official merge #63's newnextn.hc_head_norm. The graph reshapes the gamma either way, so RMS_NORM+MUL fusionis unaffected.
src/models/models.h:ple_diskmember restored.src/llama-model.cpp:init_mappingspopulates unless on-direct.halo-box #29's user-facing cleanup stays.
--ngram-on-diskis an alias for--lazy-mode on, and--ngram-io-threads/-cache/-direct-ioand theple_*llama_model_paramsfields are removed. Net effect versus pre-sync strix: onlyO_DIRECT reads and the tuning knobs are gone. The prefetch and
host_gatherpathsthe original description listed as collateral loss are back.
Audit: every other file #63 touched is byte-identical to master. Where the merge differs
from master, the difference is exactly this branch's own delta from the merge base.
Also merged: halo-box #31 (SYNC.md, docs only). Now 0 behind halo-box, 0 behind strix
master.
Verification
Built and tested on a Strix Halo (Ryzen AI MAX+ 395 / Radeon 8060S, gfx1151) in
containers:
rocm/dev-ubuntu-24.04:7.2.1-completefor HIP,ubuntu:26.04for Vulkan.GGML_HIP=ON,GPU_TARGETS=gfx1151, tests on)GGML_VULKAN=ON, tests on)test-backend-ops -b ROCm0(full)test-backend-ops -b Vulkan0FLASH_ATTN_EXT / MUL_MAT_ID / RMS_NORM / GET_ROWStest-llama-archs(ROCm)test-ple-disktest-qsa-prefixtest-arg-parserPP/TG: not measured yet. An A/B against master
0636c9aeeran, but the box wasserving inference at the same time, so the numbers are discarded.
Not verified: no qwen4exp GGUF was available, so the on-direct / prefetch /
host_gatherpaths are covered only bytest-ple-diskandtest-llama-archs, not bya real model run. CPU-only
test-backend-opswas not run. The PP regression reportedon #63 (HIP
700→400, Qwen3.8-27B IQ4_XS halved) predates this PR, and I didn'treproduce it. Tracked in #65's doc.
CI: the red checks on this PR are jobs that never got a runner and timed out after
exactly 24h with no logs. They aren't build failures.
Original description: PLE on disk (partly superseded by the update above)
Deliberately dropped: PLE on disk
halo-box #29 replaced the bespoke
--ngram-on-diskwith an alias to upstream--lazy-mode, removing the custom Qwen4Exp PLE disk reader, O_DIRECT handling, rowcache, worker-thread controls, C API params and build wiring. strix still carried the
full implementation; this merge follows halo-box.
Note this was not visible as a conflict. strix had not modified those files since
the merge-base, so git applied halo-box's deletions silently. Only
qwen4exp.cppconflicted — 1 of 12 affected files. What went:
Collateral loss, both from
6130b7262and both referencingple_disk, so neithersurvives the feature's removal as written:
madvise(MADV_WILLNEED)batch prefetch for random mmap PLE rows (each row isotherwise a serial page fault, ~0.2 ms from NVMe).
host_gatherpath, which avoided a CPUget_rowsnode splitting the graphinto GPU→CPU→GPU with an extra launch and device sync per ubatch.
The
host_gatheroptimization has a non-disk half (host-resident buffers) that couldbe preserved independently. Not attempted here — it would be new hybrid code, and it
needs benchmarking to prove it is worth it.