MTP support?
The model card says to run it with MTP, but I don't think Ling-3.0-tiny includes MTP, only Ling-3.0-flash supports that.
Is it possible for inclusionAI to release a version with MTP?
or a dspark/dflash2
Bumping this. Might not be so useful due to MOE arch but still would give some speed boost.
There is a free MTP draft for tiny already on the Hub, it just is not in this checkpoint. All three base checkpoints (Ling-3.0-tiny-base, -base-midtrain, -base-30T) ship the pretraining MTP layer (num_nextn_predict_layers = 1, model.layers.24.*, 403 tensors, 0.63 GB bf16); the post-trained release dropped it. That layer works on the post-trained trunk with no training at all.
Recipe (CPU, a minute): copy every model.layers.24.* tensor out of Ling-3.0-tiny-base into a sidecar model-mtp.safetensors next to a copy of this repo, add those entries to model.safetensors.index.json, and set "num_nextn_predict_layers": 1 in config.json. All other files stay tiny's (the two checkpoints share every non-MTP tensor name). Then serve with SGLang main:
python -m sglang.launch_server --model-path <grafted dir> --trust-remote-code \
--speculative-algorithm NEXTN --speculative-num-steps 5 \
--speculative-eagle-topk 1 --speculative-num-draft-tokens 6
Acceptance length on tiny's own outputs (greedy, 100 prompts per set, one RTX 4090; vanilla decoding = 1.0):
| draft | mtbench | humaneval | gsm8k |
|---|---|---|---|
| NGRAM (no model) | 1.42 | 1.68 | 1.64 |
| grafted MTP layer, NEXTN 3 steps | 2.69 | 3.12 | 2.97 |
| grafted MTP layer, NEXTN 5 steps | 2.87 | 3.37 | 3.20 |
| same, after a 400-step head-only fine-tune on tiny's outputs | 3.00 | 3.48 | 3.35 |
| a DSpark draft trained on tiny (11k conversations x 20 epochs, block 8) | 2.97 | 4.23 | 4.76 |
So zero-training it matches a trained DSpark on chat and trails it on code/math. One property the trained drafts do not have: it is neutral between thinking and non-thinking serving (5 steps, thinking on: 2.83 / 3.02 / 3.05, i.e. -1 / -10 / -5 %, versus -8 / -20 / -23 % for a non-thinking-trained DSpark).
Two notes for anyone reproducing: (1) the base checkpoints declare a 262k context, so serving them directly needs SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1; the graft keeps tiny's config and does not. (2) If you fine-tune the head through transformers, pass an explicit causal mask or load with attn_implementation="eager"; the remote code is otherwise non-causal for unpadded input (https://github.com/inclusionAI/Ling/issues/27), and a head tuned without the mask serves worse than the untuned one.
@inclusionAI: shipping the MTP layer with the post-trained tiny (or as a sidecar) would give everyone this for 0.63 GB.