MTP / dflash2 for LLM completions API

#2
by BebopVox - opened

Thank you for your work!

Is it possible to use MTP or dflash draft heads with this model? If yes which ones to download?

Thanks!

Yes! MTP support is now included in the latest Winnow inference release. I’ve uploaded the matching Gemma 4 12B assistant GGUF here too, and the downloader can fetch it when you enable MTP.
Use --mtp on with the download and serve commands. It accelerates chat/reasoning generation; direct decisions read answer logits without generating tokens, so MTP doesn’t speed up that part.
DFlash/DFlash2 isn’t included in this release. The supported path is the matching MTP assistant.
https://github.com/EldanRing/winnow-inference/blob/main/docs/QUICKSTART.md

Amazing, will give it a shot later!
Thank you!

EldanRing changed discussion status to closed

Sign up or log in to comment