Router-weighted Expert Activation Pruning for MoE models on CUDA / PyTorch.
Quick Start · Workflow · Supported Models · Memory · Pruning · Backends · CLI · Data · Docs · Development · License
Fork of
egesabanci/reap-cuda. This fork carries one fix forqwen3_5_moe-class MoEs (KAT-Coder-V2.5, Ornith-1.5): the router-renormalization state is detected from the loaded adapter rather than the config, and the pruning-metrics one-hot is cast to float before the matmul (CUDA has no integer GEMM). Submitted upstream asegesabanci/reap-cuda#92. Everything else on this page is upstream's documentation.
REAP CUDA compresses HuggingFace Mixture-of-Experts LLMs by pruning or merging routed experts. A short calibration pass records router-weighted activation statistics; low-saliency experts are removed (or clustered and fused) and the mutated model is saved as a standard transformers checkpoint.
Use it for one-shot MoE compression on NVIDIA GPUs — including single-GPU layerwise calibration for 30B-class models on ~46 GB cards, and GPU-resident weights on small-RAM hosts (e.g. g6.xlarge 16 GiB RAM + L4).
- Typer CLI —
reap prune|merge×full|layerwise, plusversionandkernels - Adapter-based MoE support — Qwen3 / Qwen3.5–3.6 / Llama4 / Mixtral / LFM2.5
- Weight residency —
--residency auto|gpu_full|layerwise|cpu_fullavoids full-CPU pins when VRAM fits but host RAM is tight; stream-save from GPU - FREA profitability —
--frea-backend autoprobes Triton vs cuBLAS per host/shape (L4-safe throughput); forcetritonorpytorchwhen needed - GPU-first observation — saliency stays on the compute device; routed-only
backends avoid
(E, T, H)activation materialization - Layerwise mode — one decoder block on GPU at a time for large MoEs
- Prune + merge — REAP/EAN/frequency saliency; agglomerative / TIES / …
- Layout-normalized kernels — F4 weight cache, F5 / native router, grouped bmm / FREA / F2 (Triton when profitable; always safe fallbacks)
- Hermetic tests — tiny in-memory models, mocked CLI dispatch (no Hub)
Prerequisites: Python 3.12+, uv,
NVIDIA GPU for real runs (CPU/MPS for control-flow and unit tests).
git clone https://github.com/egesabanci/reap-cuda.git
cd reap-cuda
uv venv .venv --seed --python 3.12
uv pip install --editable .
uv pip install pytest
# Optional on CUDA hosts
uv pip install -e '.[cuda]' # triton
uv pip install -e '.[eval]' # lm-evalSanity-check the CLI (no model download):
uv run reap --help
uv run reap version
uv run reap kernels # CUDA / Triton / auto-backend readiness
uv run pytest tests/ -quv run reap prune layerwise \
--model Qwen/Qwen3-30B-A3B \
--dataset theblackcat102/evol-codealpaca-v1 \
--prune-method reap \
--compression-ratio 0.5 \
--observe-backend bmm \
--residency auto \
--batches-per-category 64 \
--batch-size 1uv run reap prune full \
--model Qwen/Qwen3-30B-A3B \
--dataset theblackcat102/evol-codealpaca-v1 \
--prune-method reap \
--compression-ratio 0.5 \
--residency autoDefaults match the Typer CLI: model Qwen/Qwen3-30B-A3B, dataset
theblackcat102/evol-codealpaca-v1, --prune-method reap,
--compression-ratio 0.5, --observe-backend auto, --frea-backend auto,
--residency auto, --batches-per-category 1024. Full path default
--batch-size is 8; layerwise default is 4 (examples above often use 1
for low VRAM).
When the model fits VRAM but is large vs host RAM, prefer GPU-resident weights — do not pin the full model on CPU:
uv run reap prune full \
--model LiquidAI/LFM2.5-8B-A1B \
--dataset theblackcat102/evol-codealpaca-v1 \
--prune-method reap \
--compression-ratio 0.5 \
--residency gpu_full \
--observe-backend bmm \
--batches-per-category 8 \
--batch-size 1--residency auto usually picks gpu_full in this situation. Offline calib:
uv run reap prune full \
-m LiquidAI/LFM2.5-8B-A1B \
-d theblackcat102/evol-codealpaca-v1 \
--dataset-path /data/datasets/evol-codealpaca-calib-200 \
--artifacts-dir /data/reap-artifacts \
--residency gpu_full \
--observe-backend auto \
--frea-backend auto \
--batches-per-category 8 \
--batch-size 1uv run reap merge layerwise \
--model Qwen/Qwen3-30B-A3B \
--dataset theblackcat102/evol-codealpaca-v1 \
--expert-sim characteristic_activation \
--cluster-method agglomerative \
--compression-ratio 0.5Merge defaults: --expert-sim characteristic_activation,
--cluster-method agglomerative, --merge-method frequency_weighted_average,
--distance angular, --linkage average.
Artifacts land under artifacts/<model>/<dataset>/ (or --artifacts-dir /
REAP_ARTIFACTS_DIR): pruned or merged checkpoints, observer .pt files,
optional eval outputs. On reap prune full, smoke generate is on by
default (--no-smoke-test to skip). Layerwise prune does not run smoke.
| Step | What happens | Evidence |
|---|---|---|
| 0. Residency | Resolve --residency → load/save plan (GPU / offload / CPU) |
log: Residency resolved: … |
| 1. Load | HF model + tokenizer per plan (device_map auto / offload / cpu) |
model class / config |
| 2. Calibrate | Tokenize calibration batches (single, composite, or offline path) | batch count |
| 3. Observe | Routed expert stats via hooks or block replay | observer .pt |
| 4. Decide | Rank experts (prune) or cluster (merge) | saliency / labels |
| 5. Mutate | slice_experts or in-place merge |
live module shapes |
| 6. Save | Stream save_pretrained (hooks stripped; no full CPU dump) |
safetensors |
| 7. Validate | Optional smoke generate / lm-eval | logs / eval/ |
residency → load → calibrate → observe → prune|merge → stream-save → smoke|eval| Adapter | Family | Experts layout | Notes |
|---|---|---|---|
Qwen3MoeModelAdapter |
Qwen3-MoE | Fused gate_up / down (TF ≥5) |
Default fused path |
Qwen3_5MoeModelAdapter |
Qwen3.5 / 3.6 MoE | Fused + shared expert | Shared expert kept |
Llama4MoeModelAdapter |
Llama4 Text MoE | Fused bmm layout | Router attr .router |
MixtralMoeModelAdapter |
Mixtral / PhiMoE | Non-fused ModuleList | num_local_experts |
Lfm2MoeModelAdapter |
LFM2.5 MoE | Fused linear | Slices expert_bias |
Requires transformers>=5.5.0 for current fused Qwen stacks. Layout
detection is runtime-based (infer_model_adapter). See
docs/model-adapters.md.
Example hub ids that have been exercised end-to-end: Qwen/Qwen3-30B-A3B,
LiquidAI/LFM2.5-8B-A1B (local path or hub).
Two orthogonal controls:
| Command | Peak VRAM during observe (order of magnitude) | Use when |
|---|---|---|
reap prune full / reap merge full |
Whole model (~60 GB bf16 for 30B-class) | Multi-GPU / A100-80 / H100 |
reap prune layerwise / reap merge layerwise |
One block (~1–2 GB + routed transients) | Single L40S 46 GB, 30B+ |
Layerwise still reloads the full model for the final prune mutate/save step
(via gpu_full plan). Plan that VRAM separately from calibration.
Where parameters live during load / save — critical on low host RAM:
| Mode | Load | Save | Typical host |
|---|---|---|---|
auto |
Heuristic from host/GPU + model size | per resolved mode | Default |
gpu_full |
device_map="auto" on GPU |
Stream from device (no full CPU pin) | g6.xlarge-class 16 GiB RAM + GPU |
layerwise |
auto + disk offload (not full CPU pin) |
Reload gpu_full then stream |
Large MoE, mid GPU |
cpu_full |
device_map="cpu" |
Normal | Ample host RAM / debug |
# Prefer GPU weights when model fits VRAM but is large vs RAM
uv run reap prune full --residency gpu_full ...
# Or let auto pick (g6-like hosts often → gpu_full)
uv run reap prune full --residency auto ...Full policy, heuristics, delegation (full↔layerwise), and env knobs: docs/residency.md.
--prune-method ranks experts; higher scores are kept.
| Method | Meaning |
|---|---|
reap |
Router-weighted activation-norm mean (default) |
frequency |
Top-k assignment counts |
ean_sum / ean_mean |
Sum / mean of routed L2 norms |
weighted_ean_sum |
Sum of norm × router_weight |
weighted_frequency_sum |
Sum of router weights |
max_activations |
Max activation element over routed outputs |
ean_ca |
Norm of routed characteristic activation (needs full metrics) |
--compression-ratio in [0, 1) removes int(E × ratio) experts per layer
(always keeps ≥1). Or set --n-experts-to-prune (overrides the ratio).
--observe-backend |
Role |
|---|---|
auto |
f2 if CUDA + Triton runtime available, else bmm (default) |
bmm |
Grouped routed-only matmuls (recommended first EC2 / bring-up path) |
frea |
FREA expert MLP only |
f2 |
FREA expert MLP + F2 scatter reduce |
loop |
Legacy / parity oracle |
FREA sub-policy (orthogonal; applies when the resolved path uses FREA —
i.e. auto→f2, or explicit frea / f2):
--frea-backend |
Role |
|---|---|
auto |
Probe Triton vs cuBLAS once per shape; keep winner (default; L4 often → pytorch) |
triton |
Force Triton when tiles fit (L4 max often 128×64, not 128×128) |
pytorch |
Force cuBLAS grouped path (usually best throughput on L4/T4) |
uv run reap prune layerwise --observe-backend bmm ...
uv run reap prune full --observe-backend auto --frea-backend auto ...
uv run reap prune full --frea-backend pytorch # prefer throughput on small-SM GPUs
uv run reap kernels # print CUDA / Triton / auto resolutionSaliency tensors stay on GPU until save. Design / ops: docs/gpu-and-backends.md, docs/frea-throughput.md, docs/kernels/.
uv run reap --help
uv run reap prune --help
uv run reap prune full --help
uv run reap merge full --help
uv run reap kernels
uv run reap versionCommand tree:
reap
├── prune
│ ├── full # whole-model GPU observe → prune → save
│ └── layerwise # one block on GPU at a time (30B+ on single L40S-class)
├── merge
│ ├── full
│ └── layerwise
├── kernels # CUDA / Triton / auto-backend status (no model load)
└── version| Command | Memory during observe | Purpose |
|---|---|---|
reap prune full |
Full GPU | Observe → prune → stream-save |
reap prune layerwise |
One block | Same, layerwise calibration |
reap merge full |
Full GPU | Observe → cluster → merge → save |
reap merge layerwise |
One block | Same, layerwise calibration |
reap kernels |
— | Triton / auto-backend readiness |
reap version |
— | Package version (0.1.0) |
Common flags (all prune / merge subcommands unless noted):
| Flag | Default | Notes |
|---|---|---|
-m / --model |
Qwen/Qwen3-30B-A3B |
Hub id or local path |
-d / --dataset |
theblackcat102/evol-codealpaca-v1 |
Hub id, composite, or combined |
--dataset-path |
unset | Offline arrow/json/dir; processor still from -d |
--compression-ratio |
0.5 |
Fraction of experts to drop (or merge down by) |
--prune-method |
reap |
Prune only |
--observe-backend |
auto |
auto | loop | bmm | frea | f2 |
--frea-backend |
auto |
auto | triton | pytorch |
--residency |
auto |
auto | gpu_full | layerwise | cpu_full |
--batches-per-category |
1024 |
Composite :N overrides per component |
--batch-size |
8 (full) / 4 (layerwise) |
Sequences per batch |
--artifacts-dir |
./artifacts or REAP_ARTIFACTS_DIR |
Output root |
--observe-only |
off | Calibrate only; skip mutate/save |
--smoke-test / --no-smoke-test |
smoke on (prune full only) |
Layerwise has no smoke flag |
--eval / --no-eval |
no eval | lm-eval (needs [eval] extra) |
--seed |
42 |
|
-v / --verbose |
off | DEBUG logging (global) |
Full flag tables (merge cluster options, layerwise knobs, preserve flags): docs/cli.md.
Legacy console scripts (reap-prune, reap-layerwise, reap-merge,
reap-layerwise-merge) remain for HfArgumentParser workflows; prefer reap ….
| Mode | Example |
|---|---|
| Single (hub) | --dataset theblackcat102/evol-codealpaca-v1 |
| Offline local | --dataset theblackcat102/evol-codealpaca-v1 --dataset-path /data/… |
| Composite | --dataset "ds_a:64,ds_b[code]:64" (:N = batch count, not samples) |
| Composite offline | name:N@/local/path and/or shared --dataset-path root |
| Cached observations | --dataset combined (requires prior .pt) |
--dataset always selects the field-mapping processor (columns must match);
--dataset-path only chooses the files. Offline env vars
HF_HUB_OFFLINE / HF_DATASETS_OFFLINE need a local path or they fail with a
hint. Full rules: docs/calibration.md.
Maintainer documentation (one concern per file):
| Doc | Topic |
|---|---|
| Setup | Install, CUDA/Triton, first run |
| Index | Full documentation map |
| Architecture | Modules, data flow, invariants |
| Pipeline | Phase-by-phase prune/merge |
| CLI | Full command and flag reference |
| Calibration | Datasets, offline path, composite @path |
| Model adapters | Families, slice contract |
| Observation & metrics | Saliency state |
| GPU & backends | Device policy, F4/F5/FREA/Triton |
| FREA throughput | --frea-backend, probe, tiles, L4 tradeoff |
| Weight residency | --residency, auto heuristics, stream save |
| Pruning | Ranking and save |
| Merging | Cluster + fuse |
| Layerwise | Block-replay memory mode |
| Evaluation | Smoke + lm-eval |
| Development | Tests and extension |
| Kernels design | Kernel phase design (SoC) |
reap-cuda/
README.md
LICENSE
pyproject.toml
docs/ # user + maintainer docs
kernels/ # kernel design (SoC phases)
src/reap/
cli/ # Typer app (prune / merge / kernels / version)
kernels/ # observe backends (bmm, FREA, F2, F4, F5)
residency.py # weight load/save policy
data.py # calibration loaders + processors
model_adapters.py
observer.py / layerwise_*.py
prune.py / merge*.py / pipeline.py
...
tests/ # hermetic suite (no Hub)
scripts/ # instrumented EC2 helpers
data/ # small fixtures (e.g. smoke jsonl)uv pip install --editable . pytest
uv run pytest tests/ -q
uv run reap --help
uv run reap kernels
git diff --checkFocused suites: tests/test_residency.py, tests/test_run_findings_fixes.py,
tests/test_dataset_loading.py, tests/test_cli.py,
tests/test_triton_kernels.py.
Conventional Commits are used (feat:, fix:, test:, docs:, …).
This repository has been used to produce a published REAP + AWQ compression study on LiquidAI's LFM2.5-8B-A1B MoE model. The full analysis lives in reports/:
| Report | Contents |
|---|---|
reports/architecture-analysis.md |
Per-layer expert pruning map, survival rates, empirical bit-for-bit subset validation (1,202/1,202 tensors identical) |
reports/benchmark-results.md |
MATH500 + BFCLv3 scores for pruned & quantized models |
reports/compression-metrics.md |
Size / VRAM / throughput numbers (vLLM 4,809 tok/s peak) |
reports/quantization-details.md |
AWQ config, custom LFM2.5 mappings, INT4 loading fix |
reports/published-artifacts.md |
Published HuggingFace model artifacts + loading instructions |
Published models (HuggingFace, konic-labs):
konic-labs/LFM2.5-8B-A1B-REAP-50— 50% expert pruning, 8.57 GBkonic-labs/LFM2.5-8B-A1B-REAP-50-AWQ-INT4— pruning + AWQ INT4, 2.79 GB (83.5% total compression)
Results: 81–88% quality retention (MATH500 72.0%, BFCLv3 single-turn 57.36%) at 83.5% compression vs the 16.94 GB base.
Upstream contribution: Two transformers bugs blocking asymmetric MoE INT4 loading were fixed in huggingface/transformers#47430.
- reap-mlx — Apple Silicon / MLX port
- Paper: REAP the Experts (arXiv 2510.13999)
- Upstream inspiration: CerebrasResearch/reap
Apache License 2.0. See LICENSE.