Stage-1 KD draft (xLAM contexts)
A draft model for speculative decoding against the frozen target
Qwen/Qwen2.5-Coder-14B-Instruct (vLLM speculative decoding with
method=draft_model), produced by Stage-1 supervised knowledge distillation. Full project, training
scripts, and the complete acceptance / wall-clock / exactness evaluation:
Speculative_Decoding.
Qwen2.5-Coder-0.5B-Instruct distilled via supervised KD (top-20 logprobs, KL + 0.1·CE) on the Qwen2.5-Coder-14B-Instruct target s own greedy generations over 5k xLAM function-calling contexts. Stage-1 draft from the Speculative_Decoding project.
Measured (k=5, greedy, 100 frozen prompts, instrumented HF loop)
| eval | tau | alpha | exactness vs AR |
|---|---|---|---|
| xLAM-500 | 4.01 | 0.942 | 49/50 token-identical |
| TB-500 (held-out tools) | 2.65 | 0.812 | 49/50 token-identical |
Outputs are token-identical to plain greedy decoding of the target in
the HF instrumented loop (exactness gates; the vLLM bf16 near-tie caveat
is documented in the repo README). The vocabulary is zero-padded to
152,064 rows to satisfy vLLM s SpeculativeConfig same-vocab-size
requirement; no real token id lives in the padding band, and the padding
is gated on greedy parity (see src/serving/prepare_draft.py).
Benchmark context: vLLM 0.30.0, target bf16, k=5, batch 1. This model is a research artifact from a 1-day H100 study; no safety fine-tuning was performed (the base model s policies apply).
- Downloads last month
- 5
Model tree for vaishaalli/stage1-kd
Base model
Qwen/Qwen2.5-0.5B