Skip to content

llama: GPU-resident LRU cache for host-offloaded MoE expert weights - #27861

Draft
csantiago78 wants to merge 1 commit into
ggml-org:masterfrom
csantiago78:moe-expert-cache
Draft

csantiago78 wants to merge 1 commit into
ggml-org:masterfrom
csantiago78:moe-expert-cache

Commits

Commits on Aug 28, 2026