Skip to content

Feature Request: Slow prompt processing when MoE experts are read from disk via mmap - an optimization using O_DIRECT #29130

Description

@A233S

Prerequisites

  • I am running the latest code. Mention the version if possible as well.
  • I carefully followed the README.md.
  • I searched using keywords relevant to my issue to make sure that I am creating a new issue that is not already open (or closed).
  • I reviewed the Discussions, and have a new and useful enhancement to share.

Feature Description

When a MoE model is larger than RAM, prompt processing with op offload (copying experts to the GPU) is much slower when the experts are read through mmap than when they are read with O_DIRECT.

Motivation

With 12GB VRAM, 32GB RAM and a 106.6GB model, the speed of master vs the patch is:

GPU: RTX 5070 Ti Laptop 12GB, RAM: 32GB, CPU: i7-14650HX
Disk: ZHITAI TiPlus7100 2TB (PCIe 4.0 NVMe), xfs, Linux 7.0
Model: 106.6GB MoE GGUF (125B-A6B, 512 experts, 10 active), mmap
Base: 4fea119de
Args: -ngl 99 -ncmoe 45 -c 65536 -b 2048 -ub 2048 -t 16 -fa on
Page cache dropped before each run.

prompt tokens | master (t/s) | patched (t/s)
           80 |          4.6 |          16.3
          412 |         13.0 |          48.4
         1332 |         35.3 |         117.5
         3502 |         35.2 |         140.3

disk read during prompt processing: ~1.1 GB/s (master) vs ~4.1-4.9 GB/s (patched)
tg after long prompts: 11.7 t/s (master) vs 14.4 t/s (patched)
correctness: 1332-token prompt, top-3 token probs bit-identical to master over 14 runs

In addition, long prompts flush the page cache, which makes token generation slower afterwards.

Possible Implementation

This is a patch made with Claude Opus 5 xhigh: https://gist.github.com/A233S/64d391631a1ca1341b306bd53aa042c4

When copying experts, the patch does not read them through mmap. It reads them directly with O_DIRECT, with multiple threads, into pinned memory, and uses double buffering to transfer them to the GPU. Experts that are already in the page cache are still read from the cache.
(See above for results.)

Drawbacks:

  1. Linux only
  2. The patch currently finds the file through /proc/self/maps. A better way would be for the model loader to pass it in.
  3. The patch is currently controlled by environment variables, not by parameters.

Related:

The patch was generated by Claude Opus 5 xhigh. The text was written by me in Chinese and translated to English with Claude.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    enhancementNew feature or request

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions