Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
104 commits
Select commit Hold shift + click to select a range
be12d39
metal : null-check buffer alloc to fix OOM crash (llama/25371)
ykhrustalev Aug 25, 2026
e820c28
rpc: support apple RDMA as an RPC transport (llama/26421)
ryan5rdx Aug 25, 2026
482956e
kleidiai: Rework KleidiAI Build System/Integration (llama/26077)
JonathanC-ARM Aug 25, 2026
8df657a
ggml-meta: propagate buffer usage and call init on the new tensors (l…
max-krasnyansky Aug 26, 2026
8c0adb0
ggml-metal: add chunked SSD MMA for Mamba-2 prefill optimization (lla…
dpantaleoni Aug 26, 2026
9d8e6b9
cuda: unblock mmq for MoE on sm_60 (llama/26264)
dfriehs Aug 30, 2026
82f5f85
rpc : implement event and async backend APIs (llama/18626)
rgerganov Aug 26, 2026
0a02697
Implemented vulkan cross_entropy_loss and cross_entropy_loss_back (ll…
PranavUttarkar Aug 26, 2026
5271734
vulkan: warptiles currently assume warp sizes <= 64, clamp to work ar…
0cc4m Aug 26, 2026
3fea10d
hexagon: support for multi-NPU devices (IQ9, IQ10) and fully asynchro…
max-krasnyansky Aug 27, 2026
a5db1d6
metal : fix memory leaks due to missing autoreleasepools (llama/27758)
nikwen Aug 27, 2026
55ab1e5
Feature: Added LIGHTNING_INDEXER support for Deepseek V4 ops on Vulka…
shenron0101 Aug 27, 2026
8529971
opencl: add bin kernels `kernel_gemm_moe_q4_0_q8_1_dp4a_bin`, `kernel…
shawngu-quic Aug 27, 2026
a0614d9
hex-unary: fix RMS_NORM_MUL weight-offset bugs for grouped/broadcast …
aparmp-quic Aug 27, 2026
b6571e4
ggml-hexagon: add HTP unary ops for ABS and LOG (llama/27786)
cqderek Aug 27, 2026
ff38b98
metal : add fa-vec tunings for M4 Pro (llama/27824)
infinitewarp Aug 28, 2026
530e3f4
metal : add fa-vec tunings for M3 Max, M5 and M5 Pro (llama/27863)
ggerganov Aug 28, 2026
97d0da2
sycl: bind the f16 KV cache in place for the oneDNN SDPA path (llama/…
Titaniumtown Aug 28, 2026
fa4d244
sycl: use TILE for quantized KV decode on BMG (llama/26689)
johnkarlhill Aug 28, 2026
7f78e1b
OpenVINO: Update OV to 2026.3.1, whisper.cpp support, Qwen3.5 on NPU,…
wine99 Aug 28, 2026
0a15087
metal : add fa-vec tunings for M4 (llama/27875)
Strongtut Aug 28, 2026
ba99c09
Vulkan: add hoisting support for row IDs and expert count in shaders …
ravel7524 Aug 28, 2026
caea96f
ggml : fix conv_transpose_2d for multiple batches (llama/26132)
tekinertekin Aug 28, 2026
4c38040
vulkan: fix missing view-alias dependencies in ggml_vk_graph_optimize…
Eric-A-Stalee Aug 28, 2026
d501a0a
vulkan: Change mul_mat_id to pad K rather than N (llama/27925)
jeffbolznv Aug 29, 2026
590fe18
metal : add fa-vec tunings for M1 Max (llama/27932)
jhen0409 Aug 29, 2026
e644752
vulkan: combine duplicated fastdiv functions, rename the one optimizi…
jeffbolznv Aug 29, 2026
325c8d1
sycl: make --fit respect --fit-target better (llama/27629)
nicois Aug 29, 2026
308fa4f
metal : add remaining fa-vec tunings for M4 Pro (llama/27915)
nikwen Aug 29, 2026
285f1ff
metal : assert shared memory padding (llama/27951)
ggerganov Aug 29, 2026
2a11026
opencl: use a better matmul path on two Adreno GPU generations (llama…
wanghqc Aug 29, 2026
b33bbc5
metal : add fa-vec tunings for M2 (llama/27940)
ring2003 Aug 29, 2026
c969c68
ggml: allow passing alloc dependencies in graph_optimize (llama/27301)
am17an Aug 30, 2026
c68f205
metal : fix null-pipeline crash for F16 src1 mul_mat/mul_mat_id (llam…
QuintinShaw Aug 30, 2026
3d4e0e9
sycl: split long rows in TOP_K instead of one work-group per row (lla…
Titaniumtown Aug 30, 2026
3ad8b9b
hexagon: support for device discovery and create sessions on demand (…
max-krasnyansky Aug 30, 2026
5e49459
rpc : fix pre-rdma macOS versions (llama/27815)
ryan5rdx Aug 30, 2026
b66593e
metal : Add fa-vec tuning for M3 Pro (llama/27963)
addianto Aug 30, 2026
1e0f382
metal: add fa-vec tunings for M3 Ultra (llama/27999)
ngladitz Aug 30, 2026
43acf3d
rpc: fix apple rdma error spew on teardown (llama/27908)
ryan5rdx Aug 30, 2026
4b2243a
ggml : add ggml_backend_op_alloc_size_may_expand, use it in RPC (llam…
ggerganov Aug 30, 2026
e5c9e3e
hip : optimize Q2_0 dot-product path for gfx1201 (llama/26753)
LunalFresh Aug 30, 2026
35d9e22
hip: tune rdna 3 mmq config (llama/26284)
itterative Aug 30, 2026
e900a73
CUDA: use the fast mm_ids_helper path for any n_expert_used (llama/27…
ServeurpersoCom Aug 30, 2026
e9583f0
ggml: add SWIGLU_CLAMP (llama/27930)
am17an Aug 30, 2026
749683d
ggml : fix ggml_backend_buft_get_alloc_size() guard (llama/28038)
ggerganov Aug 30, 2026
4089fa6
rpc: avoid serializing buffers from other servers (llama/26500)
hmirin Aug 30, 2026
e5c96ca
metal : add remaining Q4_1/Q5_0/Q5_1 fa-vec tunings for M2 (llama/28017)
ring2003 Aug 30, 2026
01ebd22
hexagon: fix CPY fence bug (llama/28033)
yshsharke Aug 30, 2026
db00b01
vulkan: top_k radix select for k >= 1024 for Qwen 3.8 Flash Next (lla…
0cc4m Aug 31, 2026
96dddd8
ggml : add MUL_MAT to the list of ops that may need additional memory…
fairydreaming Aug 31, 2026
6ce7b89
vulkan: tune mat-vec rows for batched inference on Strix Halo (llama/…
SimonTeixidor Aug 31, 2026
76a51e8
sycl : Enhance to get the free memory of Intel GPU (llama/27968)
arthw Aug 31, 2026
b0f4bc0
CUDA: extend MOE fusion to specdec, earlier MOE glu fusion and topk-r…
ynankani Aug 31, 2026
7614a4c
metal : add fa-vec tunings for M1 (llama/28078)
nikwen Aug 31, 2026
c1be45b
ROCm: add radix TOP_K for long rows (llama/27466)
jadenmach2 Aug 31, 2026
c648b9a
webgpu : avoid crash when offset is not multiple of 4 in WebGPU ggml_…
fairydreaming Aug 31, 2026
088c603
opencl: tune the quant paths for Intel Xe-LP GPUs to improve its TG a…
wanghqc Aug 31, 2026
c6934d0
metal : add top-k radix implementation (llama/28073)
ggerganov Aug 31, 2026
2f608ab
AVX2: Speed up large batch size prompt processing of IQ models (llama…
bartowski1182 Aug 31, 2026
dbc40ef
metal : add concat support for quantized types (llama/28116)
ggerganov Aug 31, 2026
f22bb2e
CUDA: XOR swizzle flash attn K,V smem fp16 tiles (llama/25635)
ynankani Aug 31, 2026
8e54c65
metal : add fa-vec tunings for M1 Ultra (llama/28088)
ozgursoy Aug 31, 2026
5032008
metal: enable Metal 4.0 tensor API on M5+/A19+ (llama/27461)
JamesFranc Sep 1, 2026
4f3a2a4
sycl : support limit max alloc memory within 2GB for host-pinned memo…
arthw Sep 1, 2026
870db2a
metal : add fa-vec tuning for M2 Max (llama/28015)
ggerganov Sep 1, 2026
a245a8f
metal : fix more leaks due to missing autoreleasepools (llama/27883)
nikwen Sep 1, 2026
8cca1a3
metal : add fa-vec tunings for A18 Pro (MacBook Neo) (llama/28152)
jhen0409 Sep 1, 2026
5f07f85
metal : add fa-vec tuning for M2 Pro (llama/28122)
lstolcman Sep 1, 2026
408faaa
sycl : add Kronecker product FWHT support for sizes 384, 640, 768, 12…
philip-jingxin Sep 1, 2026
f162a19
Revert "sycl : add Kronecker product FWHT support for sizes 384, 640,…
Titaniumtown Sep 1, 2026
2c48678
cuda: fuse MoE weighted expert reduction (llama/25952)
anujj Sep 1, 2026
fcc2fee
metal : add metallib build support for xcframework (llama/28163)
jhen0409 Sep 1, 2026
c94921f
hexagon: add missing FARF logs for cpy/get_rows/set_rows/gdn ops (lla…
trivikram-reddy1 Sep 2, 2026
35133c9
opencl: fix out‐of‐bound reads in the Adreno image kernels (#27632)
wanghqc Sep 2, 2026
d57ae98
ggml-cpu : conditionally add SpacemiT IME kernel sources (llama/27961)
alanhc Sep 2, 2026
a9e5861
vulkan : only request VK_KHR_shader_bfloat16 extension if supported (…
madsmtm Sep 2, 2026
1c7d35e
vulkan: handle larger batch sizes (>4) efficiently for IQ3_S mat-vec …
LaurentZuijdwijk Sep 2, 2026
dc70853
hexagon: MUL_MAT and MUL_MAT_ID fusion and fixes (llama/28202)
max-krasnyansky Sep 2, 2026
c2b4007
ggml: avoid KleidiAI buffer type init on dispatch (llama/27891)
ac-mmi Sep 2, 2026
519df61
CUDA + ggml: add sparse-fa for DSV4/GLM (llama/27970)
am17an Sep 2, 2026
4d343d7
ggml-cuda : remove unused vars (llama/28235)
angt Sep 2, 2026
3a1c7d6
metal : fix memory query under low-memory conditions (llama/27701)
madsmtm Sep 2, 2026
1bdda1e
metal : add fa-vec tunings for M3 (llama/28236)
init-22 Sep 2, 2026
37f0f44
ggml-hexagon: add F16 support for unary ops (llama/28228)
cqderek Sep 2, 2026
e560569
finetune: fix no KV cache (llama/27199)
ngxson Sep 2, 2026
a704770
sycl: reduce redundant work in Q4_K multi-column MMVQ (llama/27062)
Eurekatic Sep 3, 2026
f24a386
sycl : enhance the api to support peer-to-peer copy (llama/27550)
arthw Sep 3, 2026
47d348a
vulkan: fix FA dequant path engagement (llama/28190)
Nathanw1014 Sep 3, 2026
25350b5
CUDA: Allow concurrent streams per split for multi-GPU (llama/28198)
tannerbruhn Sep 3, 2026
d55d345
metal : fix glu dispatch with ne00 = 1 (llama/28306)
ggerganov Sep 3, 2026
4dd48dd
metal : add sparse FA (llama/28098)
ggerganov Sep 3, 2026
0a4a95c
tune MMVQ to MMQ crossover for SM87 (llama/28285)
kbenkhaled Sep 3, 2026
d784add
opencl: quant lm_head / decode GEMV and medium-batch GEMM optimizatio…
wanghqc Sep 3, 2026
36f170e
SYCL: Refactor GGML_SYCL_ENABLE_MKL_FA to global var (llama/26863)
johnkarlhill Sep 4, 2026
d1e0e64
sycl: fuse rms_norm+mul+add and add+add residual chains (llama/27610)
newjordan Sep 4, 2026
e1bbe40
ggml-cpu(s390x) : fix q5_1 uninitialized v_acc (llama/28332)
taronaeo Sep 4, 2026
f32e6fa
ggml : remove GGML_CUDA_PEER_MAX_BATCH_SIZE (llama/28177)
angt Sep 4, 2026
1b37bea
ggml : don't crash when backend search path can't be read (llama/28271)
angt Sep 4, 2026
e2389eb
ggml : rename and make private ggml_op_alloc_size_may_expand() (ggml/0)
ggerganov Sep 4, 2026
140e57a
ggml : replace compile definitions with version.h.in (llama/28364)
danbev Sep 4, 2026
11d4eec
metal : add remaining fa-vec tunings for M3 Max (llama/28373)
nikwen Sep 4, 2026
a937f4e
ggml : bump version to 0.23.0 (ggml/1618)
ggerganov Sep 4, 2026
52a939a
sync : ggml
ggerganov Sep 4, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
10 changes: 3 additions & 7 deletions ggml/CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -4,7 +4,7 @@ project("ggml" C CXX ASM)

### GGML Version
set(GGML_VERSION_MAJOR 0)
set(GGML_VERSION_MINOR 22)
set(GGML_VERSION_MINOR 23)
set(GGML_VERSION_PATCH 0)
set(GGML_VERSION_BASE "${GGML_VERSION_MAJOR}.${GGML_VERSION_MINOR}.${GGML_VERSION_PATCH}")

Expand Down Expand Up @@ -200,8 +200,6 @@ option(GGML_CUDA "ggml: use CUDA"
option(GGML_MUSA "ggml: use MUSA" OFF)
option(GGML_CUDA_FORCE_MMQ "ggml: use mmq kernels instead of cuBLAS" OFF)
option(GGML_CUDA_FORCE_CUBLAS "ggml: always use cuBLAS instead of mmq kernels" OFF)
set (GGML_CUDA_PEER_MAX_BATCH_SIZE "128" CACHE STRING
"ggml: max. batch size for using peer access")
option(GGML_CUDA_NO_PEER_COPY "ggml: do not use peer to peer copies" OFF)
option(GGML_CUDA_NO_VMM "ggml: do not try to use CUDA VMM" OFF)
option(GGML_CUDA_FA "ggml: compile ggml FlashAttention CUDA kernels" ON)
Expand Down Expand Up @@ -242,6 +240,8 @@ option(GGML_METAL_EMBED_LIBRARY "ggml: embed Metal library"
set (GGML_METAL_MACOSX_VERSION_MIN "" CACHE STRING
"ggml: metal minimum macOS version")
set (GGML_METAL_STD "" CACHE STRING "ggml: metal standard version (-std flag)")
set (GGML_METAL_TARGET_OS "macos" CACHE STRING
"ggml: metal -mtargetos OS name (macos, ios, xros, tvos)")
option(GGML_OPENMP "ggml: use OpenMP" ON)
option(GGML_OPENMP_FETCH "ggml: fetch LLVM OpenMP" OFF)
option(GGML_RPC "ggml: use RPC" OFF)
Expand Down Expand Up @@ -404,10 +404,6 @@ write_basic_package_version_file(
VERSION ${GGML_INSTALL_VERSION}
COMPATIBILITY SameMajorVersion)

target_compile_definitions(ggml-base PRIVATE
GGML_VERSION="${GGML_INSTALL_VERSION}"
GGML_COMMIT="${GGML_BUILD_COMMIT}"
)
message(STATUS "ggml version: ${GGML_INSTALL_VERSION}")
message(STATUS "ggml commit: ${GGML_BUILD_COMMIT}")

Expand Down
10 changes: 10 additions & 0 deletions ggml/cmake/ggml-config.cmake.in
Original file line number Diff line number Diff line change
Expand Up @@ -110,6 +110,16 @@ set_and_check(GGML_INCLUDE_DIR "@PACKAGE_GGML_INCLUDE_INSTALL_DIR@")
set_and_check(GGML_LIB_DIR "@PACKAGE_GGML_LIB_INSTALL_DIR@")
#set_and_check(GGML_BIN_DIR "@PACKAGE_GGML_BIN_INSTALL_DIR@")

if (NOT GGML_SHARED_LIB AND GGML_CPU_KLEIDIAI)
unset(KLEIDIAI_LIBRARY CACHE)
unset(KLEIDIAI_LIBRARY)
find_library(KLEIDIAI_LIBRARY kleidiai
REQUIRED
HINTS ${GGML_LIB_DIR}
NO_CMAKE_FIND_ROOT_PATH)
list(APPEND GGML_CPU_INTERFACE_LINK_LIBRARIES ${KLEIDIAI_LIBRARY})
endif()

if(NOT TARGET ggml::ggml)
find_package(Threads REQUIRED)

Expand Down
4 changes: 2 additions & 2 deletions ggml/include/ggml-rpc.h
Original file line number Diff line number Diff line change
Expand Up @@ -6,8 +6,8 @@
extern "C" {
#endif

#define RPC_PROTO_MAJOR_VERSION 5
#define RPC_PROTO_MINOR_VERSION 1
#define RPC_PROTO_MAJOR_VERSION 6
#define RPC_PROTO_MINOR_VERSION 0
#define RPC_PROTO_PATCH_VERSION 0

#ifdef __cplusplus
Expand Down
13 changes: 13 additions & 0 deletions ggml/include/ggml.h
Original file line number Diff line number Diff line change
Expand Up @@ -627,6 +627,7 @@ extern "C" {
GGML_GLU_OP_SWIGLU_OAI,
GGML_GLU_OP_GEGLU_ERF,
GGML_GLU_OP_GEGLU_QUICK,
GGML_GLU_OP_SWIGLU_CLAMP,

GGML_GLU_OP_COUNT,
};
Expand Down Expand Up @@ -1367,6 +1368,12 @@ extern "C" {
float alpha,
float limit);

GGML_API struct ggml_tensor * ggml_swiglu_clamp(
struct ggml_context * ctx,
struct ggml_tensor * a,
struct ggml_tensor * b,
float limit);

// normalize along rows
GGML_API struct ggml_tensor * ggml_norm(
struct ggml_context * ctx,
Expand Down Expand Up @@ -2446,6 +2453,12 @@ extern "C" {
GGML_API enum ggml_prec ggml_flash_attn_ext_get_prec(
const struct ggml_tensor * a);

// Use finite mask entries as a sparse K/V set. Set 0 to disable.
// n_kv_max must bound the number of finite entries in every mask row.
GGML_API void ggml_flash_attn_ext_set_n_kv_max(
struct ggml_tensor * a,
int32_t n_kv_max);

GGML_API void ggml_flash_attn_ext_add_sinks(
struct ggml_tensor * a,
struct ggml_tensor * sinks);
Expand Down
4 changes: 3 additions & 1 deletion ggml/src/CMakeLists.txt
Original file line number Diff line number Diff line change
Expand Up @@ -213,7 +213,9 @@ set_target_properties(ggml-base PROPERTIES
SOVERSION ${GGML_VERSION_MAJOR}
)

target_include_directories(ggml-base PRIVATE .)
configure_file(ggml-version.h.in ${CMAKE_CURRENT_BINARY_DIR}/ggml-version.h @ONLY)

target_include_directories(ggml-base PRIVATE . ${CMAKE_CURRENT_BINARY_DIR})
if (GGML_BACKEND_DL)
target_compile_definitions(ggml-base PUBLIC GGML_BACKEND_DL)
endif()
Expand Down
18 changes: 17 additions & 1 deletion ggml/src/ggml-backend-impl.h
Original file line number Diff line number Diff line change
Expand Up @@ -34,6 +34,11 @@ extern "C" {
void * context;
};

// [TAG_ALLOC_SIZE_EXPAND]
// returns true for ops that may require additional memory for fleeting data on some backends,
// i.e. the backend buffer type's get_alloc_size may return more than ggml_nbytes for the output tensor
GGML_API bool ggml_op_alloc_size_may_expand(enum ggml_op op);

//
// Backend buffer
//
Expand Down Expand Up @@ -83,6 +88,7 @@ extern "C" {
GGML_API ggml_backend_buffer_t ggml_backend_multi_buffer_alloc_buffer(ggml_backend_buffer_t * buffers, size_t n_buffers);
GGML_API bool ggml_backend_buffer_is_multi_buffer(ggml_backend_buffer_t buffer);
GGML_API void ggml_backend_multi_buffer_set_usage(ggml_backend_buffer_t buffer, enum ggml_backend_buffer_usage usage);
GGML_API void ggml_backend_meta_buffer_set_usage (ggml_backend_buffer_t buffer, enum ggml_backend_buffer_usage usage);

//
// Backend (meta)
Expand All @@ -102,6 +108,16 @@ extern "C" {
// Backend (stream)
//

// passed to graph_optimize so the backend can add allocation dependencies:
// if the backend executes parts of the graph out of order (e.g. on concurrent streams),
// it must keep the affected tensors allocated until a node where execution is known to have joined
struct ggml_backend_graph_optimize_params {
// keep `tensor` allocated at least until `until` (a node of the same graph) has been computed
// can be called multiple times for the same tensor: the longest lifetime applies
void (*add_alloc_dep)(void * user_data, struct ggml_tensor * tensor, struct ggml_tensor * until);
void * user_data;
};

struct ggml_backend_i {
const char * (*get_name)(ggml_backend_t backend);

Expand Down Expand Up @@ -136,7 +152,7 @@ extern "C" {
void (*event_wait) (ggml_backend_t backend, ggml_backend_event_t event);

// (optional) sort/optimize the nodes in the graph
void (*graph_optimize) (ggml_backend_t backend, struct ggml_cgraph * cgraph);
void (*graph_optimize) (ggml_backend_t backend, struct ggml_cgraph * cgraph, struct ggml_backend_graph_optimize_params * params);
};

struct ggml_backend {
Expand Down
20 changes: 18 additions & 2 deletions ggml/src/ggml-backend-meta.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -1168,7 +1168,6 @@ static struct ggml_backend_meta_split_state ggml_backend_meta_get_split_state(
}

static struct ggml_backend_meta_split_state ggml_backend_meta_get_split_state(const struct ggml_tensor * tensor, bool assume_sync) {
GGML_ASSERT(ggml_backend_buffer_is_meta(tensor->buffer));
ggml_backend_meta_buffer_context * buf_ctx = (ggml_backend_meta_buffer_context *) tensor->buffer->context;
return ggml_backend_meta_get_split_state(buf_ctx->get_simple_tensor_container(tensor), tensor, assume_sync);
}
Expand Down Expand Up @@ -1259,7 +1258,14 @@ static enum ggml_status ggml_backend_meta_buffer_init_tensor_impl(ggml_backend_m
t_ij->data = (char *) ggml_backend_buffer_get_base(simple_buf)
+ size_t(tensor->data) - size_t(ggml_backend_buffer_get_base(tensor->buffer));
}
t_ij->extra = tensor->extra;

if (simple_buf) {
// the backend that owns the buffer will set .extra
ggml_backend_buffer_init_tensor(simple_buf, t_ij);
} else {
t_ij->extra = tensor->extra;
}

for (int i = 0; i < GGML_MAX_SRC; i++) {
t_ij->src[i] = tensor->src[i];
if (tensor->src[i] == tensor) {
Expand Down Expand Up @@ -1668,6 +1674,16 @@ bool ggml_backend_buffer_is_meta(ggml_backend_buffer_t buf) {
return buf != nullptr && buf->iface.free_buffer == ggml_backend_meta_buffer_iface.free_buffer;
}

void ggml_backend_meta_buffer_set_usage(ggml_backend_buffer_t buffer, enum ggml_backend_buffer_usage usage) {
GGML_ASSERT(ggml_backend_buffer_is_meta(buffer));
ggml_backend_meta_buffer_context * buf_ctx = (ggml_backend_meta_buffer_context *) buffer->context;
for (size_t i = 0; i < buf_ctx->bufs.size(); i++) {
if (buf_ctx->bufs[i]) {
ggml_backend_buffer_set_usage(buf_ctx->bufs[i].get(), usage);
}
}
}

static ggml_backend_buffer_t ggml_backend_meta_buffer_type_alloc_buffer(ggml_backend_buffer_type_t buft, size_t size) {
const size_t n_simple_bufts = ggml_backend_meta_buft_n_bufts(buft);

Expand Down
18 changes: 15 additions & 3 deletions ggml/src/ggml-backend-reg.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -490,7 +490,13 @@ static ggml_backend_reg_t ggml_backend_load_best(const char * name, bool silent,
#endif
// default search paths: executable directory, current directory
search_paths.push_back(get_executable_path());
search_paths.push_back(fs::current_path());
std::error_code cwd_ec;
const fs::path cwd = fs::current_path(cwd_ec);
if (cwd_ec) {
GGML_LOG_DEBUG("%s: current_path() failure, error-message: %s\n", __func__, cwd_ec.message().c_str());
} else {
search_paths.push_back(cwd);
}
} else {
search_paths.push_back(fs::u8path(user_search_path));
}
Expand All @@ -508,8 +514,14 @@ static ggml_backend_reg_t ggml_backend_load_best(const char * name, bool silent,
}
continue;
}
fs::directory_iterator dir_it(search_path, fs::directory_options::skip_permission_denied);
for (const auto & entry : dir_it) {
std::error_code dir_ec;
fs::directory_iterator dir_it(search_path, fs::directory_options::skip_permission_denied, dir_ec);
if (dir_ec) {
GGML_LOG_DEBUG("%s: failed to enumerate %s: %s\n", __func__, path_str(search_path).c_str(), dir_ec.message().c_str());
continue;
}
for (const fs::directory_iterator end; dir_it != end; dir_it.increment(dir_ec)) {
const auto & entry = *dir_it;
if (entry.is_regular_file(ec)) {
auto filename = entry.path().filename();
auto ext = entry.path().extension();
Expand Down
90 changes: 82 additions & 8 deletions ggml/src/ggml-backend.cpp
Original file line number Diff line number Diff line change
Expand Up @@ -20,6 +20,7 @@
#include <stdlib.h>
#include <string.h>
#include <algorithm>
#include <unordered_map>
#include <vector>

#ifdef __APPLE__
Expand Down Expand Up @@ -64,6 +65,14 @@ size_t ggml_backend_buft_get_alloc_size(ggml_backend_buffer_type_t buft, const s
if (buft->iface.get_alloc_size) {
size_t size = buft->iface.get_alloc_size(buft, tensor);
assert(size >= ggml_nbytes(tensor));

// [TAG_ALLOC_SIZE_EXPAND]
// if you hit this assert, update ggml_backend_op_alloc_size_may_expand() accordingly
GGML_ASSERT(size <= ggml_nbytes(tensor) ||
ggml_op_is_empty(tensor->op) ||
ggml_is_quantized(tensor->type) || // [TAG_ALLOC_SIZE_EXPAND]
ggml_op_alloc_size_may_expand(tensor->op));

return size;
}
return ggml_nbytes(tensor);
Expand Down Expand Up @@ -182,6 +191,8 @@ void ggml_backend_buffer_set_usage(ggml_backend_buffer_t buffer, enum ggml_backe
// FIXME: add a generic callback to the buffer interface
if (ggml_backend_buffer_is_multi_buffer(buffer)) {
ggml_backend_multi_buffer_set_usage(buffer, usage);
} else if (ggml_backend_buffer_is_meta(buffer)) {
ggml_backend_meta_buffer_set_usage(buffer, usage);
}
}

Expand Down Expand Up @@ -556,10 +567,10 @@ void ggml_backend_event_wait(ggml_backend_t backend, ggml_backend_event_t event)
backend->iface.event_wait(backend, event);
}

static void ggml_backend_graph_optimize(ggml_backend_t backend, struct ggml_cgraph * cgraph) {
static void ggml_backend_graph_optimize(ggml_backend_t backend, struct ggml_cgraph * cgraph, struct ggml_backend_graph_optimize_params * params) {
GGML_ASSERT(backend);
if (backend->iface.graph_optimize != NULL) {
backend->iface.graph_optimize(backend, cgraph);
backend->iface.graph_optimize(backend, cgraph, params);
}
}

Expand Down Expand Up @@ -1439,11 +1450,40 @@ void ggml_backend_sched_split_graph(ggml_backend_sched_t sched, struct ggml_cgra
sched->prev_leaf_backend_ids = tmp;
}

// optimize the split graphs and collect the allocation dependencies added by the backends
// this needs to happen before we make graph_copy, so they are in sync
// TODO: this may create many small allocations in the scheduler, restructure to use a flat array
std::unordered_map<ggml_tensor *, std::vector<ggml_tensor *>> alloc_deps;

struct ggml_backend_graph_optimize_params opt_params = {
/* .add_alloc_dep = */ [](void * user_data, ggml_tensor * tensor, ggml_tensor * until) {
auto & deps = *(std::unordered_map<ggml_tensor *, std::vector<ggml_tensor *>> *) user_data;
std::vector<ggml_tensor *> & keep = deps[until];
if (std::find(keep.begin(), keep.end(), tensor) == keep.end()) {
keep.push_back(tensor);
}
},
/* .user_data = */ &alloc_deps,
};

for (int i = 0; i < sched->n_splits; i++) {
struct ggml_backend_sched_split * split = &sched->splits[i];
split->graph = ggml_graph_view(graph, split->i_start, split->i_end);

ggml_backend_graph_optimize(sched->backends[split->backend_id], &split->graph, &opt_params);
}

// each dep is added to graph_copy as a GGML_OP_NONE node with the kept tensors as srcs
int n_dep_nodes = 0;
for (const auto & it : alloc_deps) {
n_dep_nodes += (it.second.size() + GGML_MAX_SRC - 1) / GGML_MAX_SRC;
}

int total_inputs = sched->n_graph_inputs;
for (int i = 0; i < sched->n_splits; i++) {
total_inputs += sched->splits[i].n_inputs;
}
int graph_size = std::max(graph->n_nodes, graph->n_leafs) + total_inputs * 2 * sched->n_copies;
int graph_size = std::max(graph->n_nodes, graph->n_leafs) + total_inputs * 2 * sched->n_copies + n_dep_nodes;

// remember the actual graph_size for performing reallocation checks later [GGML_SCHED_DEBUG_REALLOC]
sched->debug_prev_graph_size = sched->debug_graph_size;
Expand All @@ -1461,13 +1501,10 @@ void ggml_backend_sched_split_graph(ggml_backend_sched_t sched, struct ggml_cgra

struct ggml_cgraph * graph_copy = &sched->graph;

int n_dep_nodes_added = 0;

for (int i = 0; i < sched->n_splits; i++) {
struct ggml_backend_sched_split * split = &sched->splits[i];
split->graph = ggml_graph_view(graph, split->i_start, split->i_end);

// Optimize this split of the graph. This needs to happen before we make graph_copy,
// so they are in sync.
ggml_backend_graph_optimize(sched->backends[split->backend_id], &split->graph);

// add inputs to the graph copy so that they are allocated by ggml-alloc at the start of the split
for (int j = 0; j < split->n_inputs; j++) {
Expand All @@ -1492,9 +1529,32 @@ void ggml_backend_sched_split_graph(ggml_backend_sched_t sched, struct ggml_cgra
assert(graph_copy->size > graph_copy->n_nodes);
sched->node_backend_ids[graph_copy->n_nodes] = tensor_backend_id(graph->nodes[j]);
graph_copy->nodes[graph_copy->n_nodes++] = graph->nodes[j];

if (alloc_deps.empty()) {
continue;
}

// add a dependency node so that the kept tensors are not freed before this node is computed
auto it = alloc_deps.find(graph->nodes[j]);
if (it != alloc_deps.end()) {
const std::vector<ggml_tensor *> & keep = it->second;
for (size_t k = 0; k < keep.size(); k += GGML_MAX_SRC) {
struct ggml_tensor * dep = ggml_view_tensor(sched->ctx, keep[k]);
for (size_t s = 0; s < GGML_MAX_SRC && k + s < keep.size(); s++) {
dep->src[s] = keep[k + s];
}
assert(graph_copy->size > graph_copy->n_nodes);
sched->node_backend_ids[graph_copy->n_nodes] = split->backend_id;
graph_copy->nodes[graph_copy->n_nodes++] = dep;
n_dep_nodes_added++;
}
}
}
}

// a mismatch means a backend added a dep with an `until` tensor that is not a node of the optimized graph
GGML_ASSERT(n_dep_nodes_added == n_dep_nodes);

if (sched->n_copies > 1) {
// add input copies as leafs so that they are allocated first
for (int i = 0; i < sched->n_graph_inputs; i++) {
Expand Down Expand Up @@ -2049,6 +2109,20 @@ ggml_backend_t ggml_backend_sched_get_tensor_backend(ggml_backend_sched_t sched,

// utils

bool ggml_op_alloc_size_may_expand(enum ggml_op op) {
switch (op) {
case GGML_OP_FLASH_ATTN_EXT:
case GGML_OP_MUL_MAT:
case GGML_OP_MUL_MAT_ID:
case GGML_OP_CUMSUM:
case GGML_OP_ARGSORT:
case GGML_OP_TOP_K:
return true;
default:
return false;
}
}

enum ggml_status ggml_backend_view_init(struct ggml_tensor * tensor) {
GGML_ASSERT(tensor);
GGML_ASSERT(tensor->buffer == NULL);
Expand Down
Loading
Loading