Prerequisites
Feature Description
Would llama.cpp consider implementing expert-aware streaming for MoE models, where only the experts selected by the router are fetched from storage for the current token/layer?
The idea would be to keep large MoE models file-backed, fetch only selected expert weight slices just in time, optionally cache frequently used experts in RAM, and overlap storage reads with CPU computation.
A project called Big MoE on Edge is already experimenting with this approach, including Android/CPU-only use. This could make MoE models substantially larger than available RAM practical on devices with fast NVMe/UFS storage.
It would be especially interesting as an upstream llama.cpp feature with options for expert caching and asynchronous I/O/prefetching.
Motivation
llama.cpp currently supports memory-mapped model loading, but sparse MoE models present an opportunity to reduce RAM requirements further by making storage access expert-aware.
For each token and MoE layer, only a subset of experts is selected by the router. Instead of relying entirely on normal OS page caching/mmap behavior, llama.cpp could optionally fetch only the weight slices belonging to the selected experts from storage just in time, while caching frequently used experts in RAM.
This could make MoE models larger than available physical RAM more practical on memory-constrained systems, particularly Android devices with fast UFS storage and PCs with fast NVMe SSDs.
An additional optimization could overlap reads for upcoming expert weights with computation on already-loaded weights, reducing the amount of time inference stalls waiting for storage.
A proof-of-concept project, “Big MoE on Edge,” appears to implement this approach using llama.cpp's evaluation callback, with an optional small hook for per-expert synchronization/overlap. It would be useful to have similar functionality supported upstream in llama.cpp.
Possible Implementation
Add an optional expert-streaming backend for MoE tensors:
Keep expert weights file-backed rather than requiring all experts to remain resident in RAM.
After routing, determine the experts selected for the current layer/token.
Fetch only the required expert weight slices.
Maintain a configurable RAM cache/LRU for frequently selected experts.
Optionally prefetch/overlap storage I/O with computation.
Fall back to the existing mmap behavior when expert streaming is disabled.
Ideally this would remain optional so existing llama.cpp behavior and performance are unchanged by default.
Reference implementation / prior work: https://github.com/Helldez/BigMoeOnEdge
Prerequisites
Feature Description
Would llama.cpp consider implementing expert-aware streaming for MoE models, where only the experts selected by the router are fetched from storage for the current token/layer?
The idea would be to keep large MoE models file-backed, fetch only selected expert weight slices just in time, optionally cache frequently used experts in RAM, and overlap storage reads with CPU computation.
A project called Big MoE on Edge is already experimenting with this approach, including Android/CPU-only use. This could make MoE models substantially larger than available RAM practical on devices with fast NVMe/UFS storage.
It would be especially interesting as an upstream llama.cpp feature with options for expert caching and asynchronous I/O/prefetching.
Motivation
llama.cpp currently supports memory-mapped model loading, but sparse MoE models present an opportunity to reduce RAM requirements further by making storage access expert-aware.
For each token and MoE layer, only a subset of experts is selected by the router. Instead of relying entirely on normal OS page caching/mmap behavior, llama.cpp could optionally fetch only the weight slices belonging to the selected experts from storage just in time, while caching frequently used experts in RAM.
This could make MoE models larger than available physical RAM more practical on memory-constrained systems, particularly Android devices with fast UFS storage and PCs with fast NVMe SSDs.
An additional optimization could overlap reads for upcoming expert weights with computation on already-loaded weights, reducing the amount of time inference stalls waiting for storage.
A proof-of-concept project, “Big MoE on Edge,” appears to implement this approach using llama.cpp's evaluation callback, with an optional small hook for per-expert synchronization/overlap. It would be useful to have similar functionality supported upstream in llama.cpp.
Possible Implementation
Add an optional expert-streaming backend for MoE tensors:
Keep expert weights file-backed rather than requiring all experts to remain resident in RAM.
After routing, determine the experts selected for the current layer/token.
Fetch only the required expert weight slices.
Maintain a configurable RAM cache/LRU for frequently selected experts.
Optionally prefetch/overlap storage I/O with computation.
Fall back to the existing mmap behavior when expert streaming is disabled.
Ideally this would remain optional so existing llama.cpp behavior and performance are unchanged by default.
Reference implementation / prior work: https://github.com/Helldez/BigMoeOnEdge