MTP support?

#11
by coder543 - opened

The model card says to run it with MTP, but I don't think Ling-3.0-tiny includes MTP, only Ling-3.0-flash supports that.

Is it possible for inclusionAI to release a version with MTP?

or a dspark/dflash2

Bumping this. Might not be so useful due to MOE arch but still would give some speed boost.

There is a free MTP draft for tiny already on the Hub, it just is not in this checkpoint. All three base checkpoints (Ling-3.0-tiny-base, -base-midtrain, -base-30T) ship the pretraining MTP layer (num_nextn_predict_layers = 1, model.layers.24.*, 403 tensors, 0.63 GB bf16); the post-trained release dropped it. That layer works on the post-trained trunk with no training at all.

Recipe (CPU, a minute): copy every model.layers.24.* tensor out of Ling-3.0-tiny-base into a sidecar model-mtp.safetensors next to a copy of this repo, add those entries to model.safetensors.index.json, and set "num_nextn_predict_layers": 1 in config.json. All other files stay tiny's (the two checkpoints share every non-MTP tensor name). Then serve with SGLang main:

python -m sglang.launch_server --model-path <grafted dir> --trust-remote-code \
  --speculative-algorithm NEXTN --speculative-num-steps 5 \
  --speculative-eagle-topk 1 --speculative-num-draft-tokens 6

Acceptance length on tiny's own outputs (greedy, 100 prompts per set, one RTX 4090; vanilla decoding = 1.0):

draft mtbench humaneval gsm8k
NGRAM (no model) 1.42 1.68 1.64
grafted MTP layer, NEXTN 3 steps 2.69 3.12 2.97
grafted MTP layer, NEXTN 5 steps 2.87 3.37 3.20
same, after a 400-step head-only fine-tune on tiny's outputs 3.00 3.48 3.35
a DSpark draft trained on tiny (11k conversations x 20 epochs, block 8) 2.97 4.23 4.76

So zero-training it matches a trained DSpark on chat and trails it on code/math. One property the trained drafts do not have: it is neutral between thinking and non-thinking serving (5 steps, thinking on: 2.83 / 3.02 / 3.05, i.e. -1 / -10 / -5 %, versus -8 / -20 / -23 % for a non-thinking-trained DSpark).

Two notes for anyone reproducing: (1) the base checkpoints declare a 262k context, so serving them directly needs SGLANG_ALLOW_OVERWRITE_LONGER_CONTEXT_LEN=1; the graft keeps tiny's config and does not. (2) If you fine-tune the head through transformers, pass an explicit causal mask or load with attn_implementation="eager"; the remote code is otherwise non-causal for unpadded input (https://github.com/inclusionAI/Ling/issues/27), and a head tuned without the mask serves worse than the untuned one.

@inclusionAI: shipping the MTP layer with the post-trained tiny (or as a sidecar) would give everyone this for 0.63 GB.

Sign up or log in to comment