Skip to content

ggml 0.25.1 - #313141

Merged
BrewTestBot merged 2 commits into
mainfrom
bump-ggml-0.25.1
Sep 23, 2026
Merged

BrewTestBot merged 2 commits into
mainfrom
bump-ggml-0.25.1

Conversation

@BrewTestBot

Copy link
Copy Markdown
Contributor

Created by brew bump


Created with brew bump-formula-pr.

  • TODO and FIXME comments have been checked.
release notes
## Overview

A hotfix release focused on the CUDA backend: sparse flash attention, which had been disabled due to a batch-dependent gate, is re-enabled for long-context prefill (up to ~1.4x prefill speedup at 131k context), with the supporting sparse mask scan kernel significantly sped up. Additionally, Metal gains the missing f32 × bf16 matmul kernels (fixing depthwise convolutions with bf16 weights) and its flash-attention tuning tables are now keyed by GPU family instead of individual SKU, while Vulkan adds MMQ/MMV matmul kernels for the IQ4_XS quantization type.

Backend changes

CUDA

  • Re-enable sparse flash attention for long-context prefill (switched off by llama.cpp#28770 due to a
    batch-dependent gate), with up to ~1.4x prefill speedup at 131k context (llama.cpp#29298)
  • Speed up the sparse mask scan kernel via compile-time loop unrolling and host-side bound selection (586 → 244 us
    for the batched sparse op at 49k context) (llama.cpp#29298)

Metal

  • Add the missing f32 × bf16 mul_mv kernel variants, fixing depthwise convolutions with bf16 weights (llama.cpp#28741)
  • Key the flash-attention vector tuned tables by GPU family instead of individual SKU, falling back to the
    baseline for untuned families and substantially shrinking the tuning tables (llama.cpp#29075)

Vulkan

  • Add MMQ/MMV matmul kernels for the IQ4_XS quantization type (llama.cpp#28415)

More info

Changelog since v0.25.0

e565a8f4 ggml : bump version to 0.25.1 (#1637)
ef97dbf9 sync : llama.cpp
2b3b7c73 CUDA: add a reserve to avoid spurious warning on older GCC builds (llama/29317)
60e21b46 metal: add the missing f32 x bf16 mul_mv variants (llama/28741)
4506ab1d CUDA: enable sparse-fa for dsv4 prefill (again) (llama/29298)
a70078ad metal : key the fa-vec tuned table by family instead of SKU (llama/29075)
8b93b8f0 vulkan: add IQ4_XS MMQ/MMV matmul kernels (llama/28415)

View the full release notes at https://github.com/ggml-org/ggml/releases/tag/v0.25.1.


@github-actions github-actions Bot added the bump-formula-pr PR was created using `brew bump-formula-pr` label Sep 23, 2026
@github-actions

Copy link
Copy Markdown
Contributor

🤖 An automated task has requested bottles to be published to this PR.

Caution

Please do not push to this PR branch before the bottle commits have been pushed, as this results in a state that is difficult to recover from. If you need to resolve a merge conflict, please use a merge commit. Do not force-push to this PR branch.

@github-actions github-actions Bot added the CI-published-bottle-commits The commits for the built bottles have been pushed to the PR branch. label Sep 23, 2026
@BrewTestBot
BrewTestBot added this pull request to the merge queue Sep 23, 2026
Merged via the queue into main with commit 9ebd6f3 Sep 23, 2026
22 checks passed
@BrewTestBot
BrewTestBot deleted the bump-ggml-0.25.1 branch September 23, 2026 23:11
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

bump-formula-pr PR was created using `brew bump-formula-pr` CI-published-bottle-commits The commits for the built bottles have been pushed to the PR branch.

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants