Where is the MTP?

#3
by suitup91 - opened

Should I use the existing qwen 3.8 27b MTP?

Seems to work fine using https://huggingface.co/unsloth/Qwen3.8-27B-GGUF/blob/main/MTP/mtp-Qwen3.8-27B-Q4_0.gguf, I see double token generation speed with --spec-type draft-mtp --spec-draft-n-max 3

Problem with external MTP head is size.
External: ~1.37 Gb of precious VRAM.
Embedded: ~250-300 Mb.
So MTP should go as a part of a model rather than external.

Problem with external MTP head is size.
External: ~1.37 Gb of precious VRAM.
Embedded: ~250-300 Mb.

The MTP head is a fixed size whether it is bundled or a separate GGUF sidecar. You don't save 1GB just by using the included MTP. For example, I made an external DFlash2 drafter that is only 561MB, which is smaller than most MTP heads.

Problem with external MTP head is size.
External: ~1.37 Gb of precious VRAM.
Embedded: ~250-300 Mb.

The MTP head is a fixed size whether it is bundled or a separate GGUF sidecar. You don't save 1GB just by using the included MTP. For example, I made an external DFlash2 drafter that is only 561MB, which is smaller than most MTP heads.

Well, MTP head in my GGUF quants for Qwen3.8-27B is ~250-300 Mb. IQ4_XS quants. Somehow Unsloth gives Q4_0 (nearly same as IQ4_XS) standalone MTP that is 1.37 Gb. How is that possible?

UPD: Ah, now I see. Standalone MTP includes heavy embedding and output tensors. So not the same as using MTP as a part of a model, where it shares these tensors.

Sign up or log in to comment