Prerequisites
Feature Description
When a MoE model is larger than RAM, prompt processing with op offload (copying experts to the GPU) is much slower when the experts are read through mmap than when they are read with O_DIRECT.
Motivation
With 12GB VRAM, 32GB RAM and a 106.6GB model, the speed of master vs the patch is:
GPU: RTX 5070 Ti Laptop 12GB, RAM: 32GB, CPU: i7-14650HX
Disk: ZHITAI TiPlus7100 2TB (PCIe 4.0 NVMe), xfs, Linux 7.0
Model: 106.6GB MoE GGUF (125B-A6B, 512 experts, 10 active), mmap
Base: 4fea119de
Args: -ngl 99 -ncmoe 45 -c 65536 -b 2048 -ub 2048 -t 16 -fa on
Page cache dropped before each run.
prompt tokens | master (t/s) | patched (t/s)
80 | 4.6 | 16.3
412 | 13.0 | 48.4
1332 | 35.3 | 117.5
3502 | 35.2 | 140.3
disk read during prompt processing: ~1.1 GB/s (master) vs ~4.1-4.9 GB/s (patched)
tg after long prompts: 11.7 t/s (master) vs 14.4 t/s (patched)
correctness: 1332-token prompt, top-3 token probs bit-identical to master over 14 runs
In addition, long prompts flush the page cache, which makes token generation slower afterwards.
Possible Implementation
This is a patch made with Claude Opus 5 xhigh: https://gist.github.com/A233S/64d391631a1ca1341b306bd53aa042c4
When copying experts, the patch does not read them through mmap. It reads them directly with O_DIRECT, with multiple threads, into pinned memory, and uses double buffering to transfer them to the GPU. Experts that are already in the page cache are still read from the cache.
(See above for results.)
Drawbacks:
- Linux only
- The patch currently finds the file through /proc/self/maps. A better way would be for the model loader to pass it in.
- The patch is currently controlled by environment variables, not by parameters.
Related:
The patch was generated by Claude Opus 5 xhigh. The text was written by me in Chinese and translated to English with Claude.
Prerequisites
Feature Description
When a MoE model is larger than RAM, prompt processing with op offload (copying experts to the GPU) is much slower when the experts are read through mmap than when they are read with O_DIRECT.
Motivation
With 12GB VRAM, 32GB RAM and a 106.6GB model, the speed of master vs the patch is:
In addition, long prompts flush the page cache, which makes token generation slower afterwards.
Possible Implementation
This is a patch made with Claude Opus 5 xhigh: https://gist.github.com/A233S/64d391631a1ca1341b306bd53aa042c4
When copying experts, the patch does not read them through mmap. It reads them directly with O_DIRECT, with multiple threads, into pinned memory, and uses double buffering to transfer them to the GPU. Experts that are already in the page cache are still read from the cache.
(See above for results.)
Drawbacks:
Related:
The patch was generated by Claude Opus 5 xhigh. The text was written by me in Chinese and translated to English with Claude.