ggml 0.25.1 - #313141
Merged
Merged
ggml 0.25.1#313141
Conversation
botantony
approved these changes
Sep 23, 2026
Contributor
|
🤖 An automated task has requested bottles to be published to this PR. Caution Please do not push to this PR branch before the bottle commits have been pushed, as this results in a state that is difficult to recover from. If you need to resolve a merge conflict, please use a merge commit. Do not force-push to this PR branch. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Created by
brew bumpCreated with
brew bump-formula-pr.release notes
batch-dependent gate), with up to ~1.4x prefill speedup at 131k context (llama.cpp#29298)
for the batched sparse op at 49k context) (llama.cpp#29298)
Metal
mul_mvkernel variants, fixing depthwise convolutions with bf16 weights (llama.cpp#28741)baseline for untuned families and substantially shrinking the tuning tables (llama.cpp#29075)
Vulkan
More info
ggml-orgprojectsChangelog since v0.25.0
e565a8f4 ggml : bump version to 0.25.1 (#1637)
ef97dbf9 sync : llama.cpp
2b3b7c73 CUDA: add a reserve to avoid spurious warning on older GCC builds (llama/29317)
60e21b46 metal: add the missing f32 x bf16 mul_mv variants (llama/28741)
4506ab1d CUDA: enable sparse-fa for dsv4 prefill (again) (llama/29298)
a70078ad metal : key the fa-vec tuned table by family instead of SKU (llama/29075)
8b93b8f0 vulkan: add IQ4_XS MMQ/MMV matmul kernels (llama/28415)
View the full release notes at https://github.com/ggml-org/ggml/releases/tag/v0.25.1.