From 0853cea20a5e85705f69884feee405a9510bd1ff Mon Sep 17 00:00:00 2001 From: Jhin Lee Date: Thu, 24 Sep 2026 08:07:06 -0400 Subject: [PATCH 1/5] ci: build Windows ARM64 with Visual Studio 2026 GitHub is migrating the windows-11-arm label to the Windows 11 Arm64 with Visual Studio 2026 image between 2026-09-21 and 2026-09-30 (actions/runner-images#14602). That image has no VS 2022 instance, so the windows-arm64-full preset's hardcoded "Visual Studio 17 2022" generator fails at configure. Native release run 35976011426 lost both arm64 lanes (blas, vulkan) this way; the same lanes passed on the VS 2022 image in run 35850331498 on the same commit. Switch the preset to "Visual Studio 18 2026" and pin the arm64 jobs to windows-11-vs2026-arm so the image cannot flip mid-rollout. ClangCL, the ARM64 and x64 MSVC tools, and CMake 4.4 are all present on that image. --- .github/workflows/native_release.yml | 4 ++-- .github/workflows/validate_wrapper.yml | 2 +- CMakePresets.json | 2 +- docs/platform_backend_strategy.md | 6 ++++++ 4 files changed, 10 insertions(+), 4 deletions(-) diff --git a/.github/workflows/native_release.yml b/.github/workflows/native_release.yml index b76e433..56f6f10 100644 --- a/.github/workflows/native_release.yml +++ b/.github/workflows/native_release.yml @@ -697,7 +697,7 @@ jobs: backend_glob: "*ggml-blas*.dll" - arch: arm64 vcpkg_triplet: arm64-windows - runs_on: windows-11-arm + runs_on: windows-11-vs2026-arm backend: vulkan include_core: true needs_cuda: false @@ -706,7 +706,7 @@ jobs: backend_glob: "*ggml-vulkan*.dll" - arch: arm64 vcpkg_triplet: arm64-windows - runs_on: windows-11-arm + runs_on: windows-11-vs2026-arm backend: blas include_core: false needs_cuda: false diff --git a/.github/workflows/validate_wrapper.yml b/.github/workflows/validate_wrapper.yml index f63edae..8d7eb1c 100644 --- a/.github/workflows/validate_wrapper.yml +++ b/.github/workflows/validate_wrapper.yml @@ -134,7 +134,7 @@ jobs: windows-arm64-kleidiai: needs: changes if: needs.changes.outputs.native == 'true' - runs-on: windows-11-arm + runs-on: windows-11-vs2026-arm timeout-minutes: 30 strategy: fail-fast: false diff --git a/CMakePresets.json b/CMakePresets.json index 2ac4822..9fb51bd 100644 --- a/CMakePresets.json +++ b/CMakePresets.json @@ -175,7 +175,7 @@ "name": "windows-arm64-full", "inherits": "windows-base", "binaryDir": "${sourceDir}/build/wa64", - "generator": "Visual Studio 17 2022", + "generator": "Visual Studio 18 2026", "architecture": { "value": "ARM64", "strategy": "set" diff --git a/docs/platform_backend_strategy.md b/docs/platform_backend_strategy.md index 4047e42..bdde67f 100644 --- a/docs/platform_backend_strategy.md +++ b/docs/platform_backend_strategy.md @@ -69,6 +69,12 @@ The Android artifact check is `tools/validate_android_cpu_isa.py --help`. - `third_party/opencl-stubs/` - auto-built OpenCL ICD loader from `third_party/OpenCL-ICD-Loader` + `third_party/OpenCL-Headers` - Linux arm64 builds on x64 runners require `aarch64-linux-gnu-gcc/g++`, `libopenblas-dev:arm64`, and `libvulkan-dev:arm64`. +- Windows ARM64 uses the `Visual Studio 18 2026` generator (CMake >= 4.2) with + the ClangCL toolset. CI pins the `windows-11-vs2026-arm` runner because the + `windows-11-arm` label moved to that image in September 2026 + (actions/runner-images#14602), and that image has no VS 2022 instance. With + only VS 2022 installed, pass `-G "Visual Studio 17 2022"` after + `--preset windows-arm64-full`. - Windows MSVC builds disable IPO/LTCG for `llama-common` and `mtmd` by default. Current MSVC `link.exe` can access-violate when linking the large `llama-common` utility DLL with `/LTCG`. CMake's automatic Windows export From 2b589b7d8d18d732fbee48236cedf50b2184d393 Mon Sep 17 00:00:00 2001 From: Jhin Lee Date: Thu, 24 Sep 2026 08:14:08 -0400 Subject: [PATCH 2/5] Qualify v0.5.0 Android CPU dispatch and source policy The Android arm64 OpenCL job of native release run 35976011426 built llama.cpp v0.5.0 (7fe450e19305) and then correctly rejected its new ggml/src fingerprint at the CPU ISA source gate. Review of the v0.4.1..v0.5.0 ggml/src delta: KleidiAI stays at v1.24.0 with an unchanged kai digest, and kleidiai/, ARM cpu-feats and the CPU/ggml build files are byte-unchanged. The only ARM change is 8034c1d1f's Q1_0 repack kernels. They follow the existing pattern: compile-time DOTPROD/I8MM guards fall back to generic C, runtime selection uses ggml_cpu_has_*, and there is no SVE/SME code. The release ARMv8.2_2 variant built from the exact SHA with NDK 28.2 passes the unchanged containment policy: 260,601 instructions, with the same 2,187 scalable instructions as v0.4.1, all in the existing 18 functions. The Q1_0 I8MM GEMM compiles to its generic fallback in that non-I8MM variant. Add only the audited pair, and exact v0.5.0 rows to the Android ISA and emulated non-SVE Kleidi dispatch/compute lanes. The production pin is unchanged. Evidence: docs/v050_android_isa_qualification.md. --- .github/workflows/validate_wrapper.yml | 9 +-- docs/v050_android_isa_qualification.md | 81 ++++++++++++++++++++++++++ tests/test_android_cpu_isa.py | 1 + tools/validate_android_cpu_isa.py | 8 ++- 4 files changed, 93 insertions(+), 6 deletions(-) create mode 100644 docs/v050_android_isa_qualification.md diff --git a/.github/workflows/validate_wrapper.yml b/.github/workflows/validate_wrapper.yml index f63edae..ebcc344 100644 --- a/.github/workflows/validate_wrapper.yml +++ b/.github/workflows/validate_wrapper.yml @@ -16,6 +16,7 @@ env: # Qualification only: never moves the production submodule or release channel. LLAMADART_QUALIFICATION_SHA: 73ab7599b553c03f6f5d2db24a18ad76f2eb36a3 LLAMADART_V041_QUALIFICATION_SHA: b29c606e28a01b1bc8c1351026a0fa6e616bf6c4 + LLAMADART_V050_QUALIFICATION_SHA: 7fe450e19305b828c199d602c23a8337aaa1f03b jobs: changes: @@ -42,7 +43,7 @@ jobs: strategy: fail-fast: false matrix: - upstream: [post-v0.4.0, v0.4.1] + upstream: [post-v0.4.0, v0.4.1, v0.5.0] steps: - uses: actions/checkout@v7 with: @@ -50,7 +51,7 @@ jobs: submodules: recursive - name: Select candidate upstream env: - LLAMADART_QUALIFICATION_SHA: ${{ matrix.upstream == 'v0.4.1' && env.LLAMADART_V041_QUALIFICATION_SHA || env.LLAMADART_QUALIFICATION_SHA }} + LLAMADART_QUALIFICATION_SHA: ${{ matrix.upstream == 'v0.5.0' && env.LLAMADART_V050_QUALIFICATION_SHA || matrix.upstream == 'v0.4.1' && env.LLAMADART_V041_QUALIFICATION_SHA || env.LLAMADART_QUALIFICATION_SHA }} run: | git -C third_party/llama.cpp fetch --depth=1 origin "$LLAMADART_QUALIFICATION_SHA" git -C third_party/llama.cpp checkout --detach "$LLAMADART_QUALIFICATION_SHA" @@ -92,7 +93,7 @@ jobs: strategy: fail-fast: false matrix: - upstream: [post-v0.4.0, v0.4.1] + upstream: [post-v0.4.0, v0.4.1, v0.5.0] steps: - uses: actions/checkout@v7 with: @@ -100,7 +101,7 @@ jobs: submodules: recursive - name: Select candidate upstream env: - LLAMADART_QUALIFICATION_SHA: ${{ matrix.upstream == 'v0.4.1' && env.LLAMADART_V041_QUALIFICATION_SHA || env.LLAMADART_QUALIFICATION_SHA }} + LLAMADART_QUALIFICATION_SHA: ${{ matrix.upstream == 'v0.5.0' && env.LLAMADART_V050_QUALIFICATION_SHA || matrix.upstream == 'v0.4.1' && env.LLAMADART_V041_QUALIFICATION_SHA || env.LLAMADART_QUALIFICATION_SHA }} run: | git -C third_party/llama.cpp fetch --depth=1 origin "$LLAMADART_QUALIFICATION_SHA" git -C third_party/llama.cpp checkout --detach "$LLAMADART_QUALIFICATION_SHA" diff --git a/docs/v050_android_isa_qualification.md b/docs/v050_android_isa_qualification.md new file mode 100644 index 0000000..d1705e7 --- /dev/null +++ b/docs/v050_android_isa_qualification.md @@ -0,0 +1,81 @@ +# v0.5.0 Android CPU ISA qualification + +The Android arm64 OpenCL build in native release run `35976011426` compiled +successfully, then rejected the new complete `ggml/src` fingerprint. It +selected upstream v0.5.0 at `7fe450e19305b828c199d602c23a8337aaa1f03b`. The +arm64 Vulkan job in the same run failed earlier, in shader generation, so it +never reached this gate. This policy update does not change the production +submodule pin, the release channel, or the dispatched-function allowlist. + +## Source review + +Compared with the previously audited v0.4.1 tree +(`b29c606e28a01b1bc8c1351026a0fa6e616bf6c4`), 203 files under `ggml/src` +changed: eight CPU files, five shared files, and 190 accelerator files +(Hexagon, OpenVINO, Vulkan, Metal, CUDA, SYCL, OpenCL, WebGPU, RPC). The +following are byte-unchanged: `ggml-cpu/kleidiai/`, ARM feature detection +(`arch/arm/cpu-feats.cpp`), `ggml-cpu/CMakeLists.txt`, `ggml/src/CMakeLists.txt` +and `ggml/cmake`. KleidiAI is still v1.24.0 and its `kai` tree digest is +unchanged. The `ggml/CMakeLists.txt` delta is the ggml version bump +(0.24.0 to 0.25.1). + +Only one change touches ARM code: 8034c1d1f adds Q1_0 repack kernels +(`ggml_gemv/gemm_q1_0_4x{4,8}_q8_0`). Each NEON body is compile-time guarded +(`__ARM_FEATURE_DOTPROD` for the 4x4 kernels and the 4x8 GEMV, +`__ARM_FEATURE_MATMUL_INT8` for the 4x8 GEMM) and falls through to the +existing generic C implementation. `arch-fallback.h` maps the new names on +non-ARM targets. Runtime selection in `repack.cpp` follows the existing +repack pattern. It requires `ggml_cpu_has_neon()` plus +`ggml_cpu_has_matmul_int8()` or `ggml_cpu_has_dotprod()`, and `ne[1] % 4 == 0`. +A variant compiled without a feature therefore runs the generic body, even +when runtime detection selects the kernel. The kernels contain no SVE or SME +code. + +The other CPU changes are ISA-neutral. They add F16 `src1` to the +Hadamard/FWHT path of `ggml-cpu.c`/`.cpp`/`ops.cpp` (using the existing +`ggml_cpu_fp16_to_fp32`), add hc ops, and fix a SpacemiT RISC-V transpose. +Shared changes are: a scheduler graph-reserve failure check (#26070), meta +backend buffer-view resolution (#29266), the hc op declarations in `ggml.h`, +a `ggml_permute` stride-truncation fix (#29227), IQ1_M reference quantization +building prefix sums once per block (#28706), and GGUF data-section alignment +relative to the GGUF start (#28993). The added lines of the shared and +generic CPU files contain no intrinsics, inline assembly or new `#if` guards. +No change raises the baseline ARM ISA or alters KleidiAI kernel selection, +packing or its callers. + +The accepted pair binds the full trees, including the accelerator changes: + +- ggml/src: `68d5bc369749e78545f50dd5107368ec7ee1874619794795cd142b0043c747f6` +- kai: `64189fc613c1c4c3aaeeb6bb12b38d85dd6728cafd2261a5a88f1b77b10fe59c` + +Unknown source combinations, and scalable instructions outside the existing +18 exact ELF function ranges, remain rejected. + +## Artifact and execution evidence + +The exact upstream SHA was built locally with the release helper's +`android_armv8.2_2` CPU variant and NDK `28.2.13676358`. Before the pair was +added, the production validator failed with the identical fingerprint pair +reported by the release run. With the pair added, the unmodified instruction +containment policy inspected 260,601 instructions. All 2,187 scalable +instructions, the same count as v0.4.1, were contained in the existing 18 +functions. In this non-I8MM variant, `ggml_gemm_q1_0_4x8_q8_0` compiles to a +single tail branch into the generic kernel and contains no `smmla`. The other +three Q1_0 kernels use `sdot`, which the variant's `GGML_USE_DOTPROD` requires. +The local `libggml-cpu.so` SHA-256 was +`20543a590f8f4e92ff0087cf364e43f785331006ddabe27b1f9eb7f16e694d70` (macOS +NDK host). This CPU artifact check does not claim full Vulkan/OpenCL packaging +or hardware execution coverage. + +`Validate Wrapper` keeps the prior candidate lanes and adds exact v0.5.0 rows: + +- Android release ARMv8.2 CPU artifact build and source/instruction validation. +- Compiled selector and Q4/Q8 scalar-reference compute under QEMU `cortex-a53` + and `max,sve=off,sme=off`, both normally and with guest + `GGML_KLEIDIAI_SME=1` to exercise the unsupported-feature fallback. + +Hosted results and an independent exact-head review must pass before merge and +be linked from the PR. Physical Android and SME hardware execution is not +performed here (N/A), and emulator evidence is separate from device +qualification. Release publication and consumer adoption need their own +approval. diff --git a/tests/test_android_cpu_isa.py b/tests/test_android_cpu_isa.py index 98b19ed..76a8453 100644 --- a/tests/test_android_cpu_isa.py +++ b/tests/test_android_cpu_isa.py @@ -105,6 +105,7 @@ def test_exact_reviewed_source_pairs_are_required(self): "c4dc92a7d95ebfad7f5f55e75be2ae773b7d95faf72a9581c9479c42bc41bca0", "dcb0f04ebb9654b1fe5ac7cc45737c79e62b116a2063ceda81a7ec1ddb1b20e2", "bf7ae6d2ea861ce6cd4b56afce154a45461df79b46c02e340d1adfe0075467b5", + "68d5bc369749e78545f50dd5107368ec7ee1874619794795cd142b0043c747f6", )) self.assertEqual(audit.SOURCE_PAIRS, expected) for ggml, kai in expected: diff --git a/tools/validate_android_cpu_isa.py b/tools/validate_android_cpu_isa.py index 69835c5..5f35f9e 100644 --- a/tools/validate_android_cpu_isa.py +++ b/tools/validate_android_cpu_isa.py @@ -19,8 +19,8 @@ # Bind complete source subtree pairs, not independent hash sets: a new caller # or a different KleidiAI combination invalidates the dispatch audit. -# Candidate audit evidence: docs/post_v040_qualification.md and -# docs/v041_android_isa_qualification.md. +# Candidate audit evidence: docs/post_v040_qualification.md, +# docs/v041_android_isa_qualification.md and docs/v050_android_isa_qualification.md. SOURCE_PAIRS = frozenset({ ( # llama.cpp v0.4.0 / KleidiAI v1.24.0 "c4dc92a7d95ebfad7f5f55e75be2ae773b7d95faf72a9581c9479c42bc41bca0", @@ -34,6 +34,10 @@ "bf7ae6d2ea861ce6cd4b56afce154a45461df79b46c02e340d1adfe0075467b5", "64189fc613c1c4c3aaeeb6bb12b38d85dd6728cafd2261a5a88f1b77b10fe59c", ), + ( # llama.cpp v0.5.0 7fe450e19305b828c199d602c23a8337aaa1f03b + "68d5bc369749e78545f50dd5107368ec7ee1874619794795cd142b0043c747f6", + "64189fc613c1c4c3aaeeb6bb12b38d85dd6728cafd2261a5a88f1b77b10fe59c", + ), }) # Exact ELF STT_FUNC ranges; never allow by kai_* prefix or disassembly label. From 4a88e863340ecedc122513987fdc891a01911d29 Mon Sep 17 00:00:00 2001 From: Jhin Lee Date: Thu, 24 Sep 2026 08:19:45 -0400 Subject: [PATCH 3/5] docs: scope the VS 2022 fallback to direct cmake configures --- docs/platform_backend_strategy.md | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/docs/platform_backend_strategy.md b/docs/platform_backend_strategy.md index bdde67f..d1a0158 100644 --- a/docs/platform_backend_strategy.md +++ b/docs/platform_backend_strategy.md @@ -72,9 +72,11 @@ The Android artifact check is `tools/validate_android_cpu_isa.py --help`. - Windows ARM64 uses the `Visual Studio 18 2026` generator (CMake >= 4.2) with the ClangCL toolset. CI pins the `windows-11-vs2026-arm` runner because the `windows-11-arm` label moved to that image in September 2026 - (actions/runner-images#14602), and that image has no VS 2022 instance. With - only VS 2022 installed, pass `-G "Visual Studio 17 2022"` after - `--preset windows-arm64-full`. + (actions/runner-images#14602), and that image has no VS 2022 instance. + `tools/build.py windows --arch arm64` therefore needs VS 2026. With only + VS 2022 installed, configure directly with + `cmake --preset windows-arm64-full -G "Visual Studio 17 2022"` in a fresh + build directory. - Windows MSVC builds disable IPO/LTCG for `llama-common` and `mtmd` by default. Current MSVC `link.exe` can access-violate when linking the large `llama-common` utility DLL with `/LTCG`. CMake's automatic Windows export From 475aefa1909564bdc1e48cd60c09e0661e828608 Mon Sep 17 00:00:00 2001 From: Jhin Lee Date: Thu, 24 Sep 2026 08:21:40 -0400 Subject: [PATCH 4/5] ci: qualify exact v0.5.0 on the Windows ARM64 Kleidi lane AGENTS.md asks for the candidate Windows gate on ARM64 upstream upgrades. The v0.5.0 release run never got past configure on Windows ARM64, and the existing rows build only v0.4.1 and the post-v0.4.0 candidate. The lane's selector also named a v0.4.1 row the matrix no longer has. --- .github/workflows/validate_wrapper.yml | 4 ++-- docs/v050_android_isa_qualification.md | 2 ++ tests/test_ci_scope.py | 2 +- 3 files changed, 5 insertions(+), 3 deletions(-) diff --git a/.github/workflows/validate_wrapper.yml b/.github/workflows/validate_wrapper.yml index 020bd06..cfafd01 100644 --- a/.github/workflows/validate_wrapper.yml +++ b/.github/workflows/validate_wrapper.yml @@ -140,7 +140,7 @@ jobs: strategy: fail-fast: false matrix: - upstream: [pinned, post-v0.4.0] + upstream: [pinned, post-v0.4.0, v0.5.0] steps: - uses: actions/checkout@v7 with: @@ -152,7 +152,7 @@ jobs: arch: arm64 - name: Select candidate upstream env: - LLAMADART_QUALIFICATION_SHA: ${{ matrix.upstream == 'v0.4.1' && env.LLAMADART_V041_QUALIFICATION_SHA || env.LLAMADART_QUALIFICATION_SHA }} + LLAMADART_QUALIFICATION_SHA: ${{ matrix.upstream == 'v0.5.0' && env.LLAMADART_V050_QUALIFICATION_SHA || env.LLAMADART_QUALIFICATION_SHA }} if: matrix.upstream != 'pinned' shell: bash run: | diff --git a/docs/v050_android_isa_qualification.md b/docs/v050_android_isa_qualification.md index d1705e7..ecb674f 100644 --- a/docs/v050_android_isa_qualification.md +++ b/docs/v050_android_isa_qualification.md @@ -73,6 +73,8 @@ or hardware execution coverage. - Compiled selector and Q4/Q8 scalar-reference compute under QEMU `cortex-a53` and `max,sve=off,sme=off`, both normally and with guest `GGML_KLEIDIAI_SME=1` to exercise the unsupported-feature fallback. +- Windows ARM64 release preset (ClangCL, VS 2026 image) with the optimized + Kleidi CPU and the native wrapper contracts, built from the exact v0.5.0 SHA. Hosted results and an independent exact-head review must pass before merge and be linked from the PR. Physical Android and SME hardware execution is not diff --git a/tests/test_ci_scope.py b/tests/test_ci_scope.py index 15521bc..40d81b4 100644 --- a/tests/test_ci_scope.py +++ b/tests/test_ci_scope.py @@ -90,7 +90,7 @@ def test_actual_workflow_routes_every_compiler_and_aggregate(self): self.assertIn('github.event.pull_request.base.sha || github.event.before',changes) self.assertIn('tools/ci_scope.py',changes) windows=workflow_job(workflow,'windows-arm64-kleidiai') - self.assertIn('upstream: [pinned, post-v0.4.0]',windows) + self.assertIn('upstream: [pinned, post-v0.4.0, v0.5.0]',windows) for filename in ('validate_wrapper.yml','validate_release_provenance.yml'): content=(ROOT/'.github/workflows'/filename).read_text() triggers=content.split('permissions:')[0] From c296d4f516ebfa087abc3cd640b17d857477aea7 Mon Sep 17 00:00:00 2001 From: Jhin Lee Date: Thu, 24 Sep 2026 08:24:05 -0400 Subject: [PATCH 5/5] docs: correct the v0.5.0 Q1_0 selection argument and state its coverage gap --- docs/v050_android_isa_qualification.md | 29 +++++++++++++++++++------- 1 file changed, 21 insertions(+), 8 deletions(-) diff --git a/docs/v050_android_isa_qualification.md b/docs/v050_android_isa_qualification.md index ecb674f..888d2c0 100644 --- a/docs/v050_android_isa_qualification.md +++ b/docs/v050_android_isa_qualification.md @@ -24,18 +24,20 @@ Only one change touches ARM code: 8034c1d1f adds Q1_0 repack kernels (`__ARM_FEATURE_DOTPROD` for the 4x4 kernels and the 4x8 GEMV, `__ARM_FEATURE_MATMUL_INT8` for the 4x8 GEMM) and falls through to the existing generic C implementation. `arch-fallback.h` maps the new names on -non-ARM targets. Runtime selection in `repack.cpp` follows the existing +non-ARM targets. Kernel selection in `repack.cpp` follows the existing repack pattern. It requires `ggml_cpu_has_neon()` plus `ggml_cpu_has_matmul_int8()` or `ggml_cpu_has_dotprod()`, and `ne[1] % 4 == 0`. -A variant compiled without a feature therefore runs the generic body, even -when runtime detection selects the kernel. The kernels contain no SVE or SME -code. +On ARM those predicates are compile-time constants of the variant being built +(`#if __ARM_FEATURE_*` in `ggml-cpu.c`), not HWCAP queries, so each isolated +variant selects only kernels its own flags allow. Which variant loads remains +the job of the unchanged `cpu-feats.cpp` scoring. The kernels contain no SVE or +SME code. The other CPU changes are ISA-neutral. They add F16 `src1` to the Hadamard/FWHT path of `ggml-cpu.c`/`.cpp`/`ops.cpp` (using the existing `ggml_cpu_fp16_to_fp32`), add hc ops, and fix a SpacemiT RISC-V transpose. Shared changes are: a scheduler graph-reserve failure check (#26070), meta -backend buffer-view resolution (#29266), the hc op declarations in `ggml.h`, +backend buffer-view resolution (#29266), the hc ops (declared in `ggml/include/ggml.h`, outside the fingerprinted tree), a `ggml_permute` stride-truncation fix (#29227), IQ1_M reference quantization building prefix sums once per block (#28706), and GGUF data-section alignment relative to the GGUF start (#28993). The added lines of the shared and @@ -62,9 +64,12 @@ instructions, the same count as v0.4.1, were contained in the existing 18 functions. In this non-I8MM variant, `ggml_gemm_q1_0_4x8_q8_0` compiles to a single tail branch into the generic kernel and contains no `smmla`. The other three Q1_0 kernels use `sdot`, which the variant's `GGML_USE_DOTPROD` requires. -The local `libggml-cpu.so` SHA-256 was -`20543a590f8f4e92ff0087cf364e43f785331006ddabe27b1f9eb7f16e694d70` (macOS -NDK host). This CPU artifact check does not claim full Vulkan/OpenCL packaging +The instruction counts are the reproducible evidence: an independent rebuild +and the hosted `android-arm64-isa (v0.5.0)` job printed the same PASS line. The +`libggml-cpu.so` bytes depend on the build path and are not recorded. An +independent rebuild of `android_armv8.0_1` found all four Q1_0 kernels reduced +to a branch into the generic code, and no `sdot`, `smmla` or FP16 vector +arithmetic outside KleidiAI. This CPU artifact check does not claim full Vulkan/OpenCL packaging or hardware execution coverage. `Validate Wrapper` keeps the prior candidate lanes and adds exact v0.5.0 rows: @@ -76,6 +81,14 @@ or hardware execution coverage. - Windows ARM64 release preset (ClangCL, VS 2026 image) with the optimized Kleidi CPU and the native wrapper contracts, built from the exact v0.5.0 SHA. +Coverage gap: this is ISA-safety evidence, not numerical evidence for the new +Q1_0 fast paths. The emulated dispatch test builds `armv8-a` (the Q1_0 kernels +compile to their generic fallback) and exercises KleidiAI Q4/Q8 only. No gate +here compares the DOTPROD Q1_0 kernels with the generic reference, and they +only matter for Q1_0 models. The containment validator also checks SVE/SME +only, so I8MM or DOTPROD in a lower variant was checked by hand for this +release. + Hosted results and an independent exact-head review must pass before merge and be linked from the PR. Physical Android and SME hardware execution is not performed here (N/A), and emulator evidence is separate from device