forked from halo-box/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 14
Question about this repo? #33
Copy link
Copy link
Open
Description
aahmozart
opened on Sep 8, 2026
Issue body actions
- Is there a prebuilt MTP sidecar anywhere (HF, release asset, Discord pin)? Their README documents --spec-draft-adaptive as speeding up MTP, so someone must be running it — I just can't tell what file they feed -md.
- What is the exact 37-tensor set the loader expects? If it's the 33 non-indexer tensors plus four more, I may be able to build a compatible sidecar from the target GGUF without the 250 GB safetensors download.
- Would they accept upstream's exports — i.e. support the qwen4exp.nextn_shared_target_tensors flag from PR models: Qwen3.8-Flash-Next MTP ggml-org/llama.cpp#28243, which would make the fork drop-in for anyone already on unsloth's GGUFs.
Reactions are currently unavailable
Activity
Metadata
Metadata
Assignees
Labels
No labels