Skip to content

Question about this repo? #33

Description

@aahmozart
  1. Is there a prebuilt MTP sidecar anywhere (HF, release asset, Discord pin)? Their README documents --spec-draft-adaptive as speeding up MTP, so someone must be running it — I just can't tell what file they feed -md.
  2. What is the exact 37-tensor set the loader expects? If it's the 33 non-indexer tensors plus four more, I may be able to build a compatible sidecar from the target GGUF without the 250 GB safetensors download.
  3. Would they accept upstream's exports — i.e. support the qwen4exp.nextn_shared_target_tensors flag from PR models: Qwen3.8-Flash-Next MTP ggml-org/llama.cpp#28243, which would make the fork drop-in for anyone already on unsloth's GGUFs.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions