Stage-1 KD draft (xLAM contexts)

A draft model for speculative decoding against the frozen target Qwen/Qwen2.5-Coder-14B-Instruct (vLLM speculative decoding with method=draft_model), produced by Stage-1 supervised knowledge distillation. Full project, training scripts, and the complete acceptance / wall-clock / exactness evaluation: Speculative_Decoding.

Qwen2.5-Coder-0.5B-Instruct distilled via supervised KD (top-20 logprobs, KL + 0.1·CE) on the Qwen2.5-Coder-14B-Instruct target s own greedy generations over 5k xLAM function-calling contexts. Stage-1 draft from the Speculative_Decoding project.

Measured (k=5, greedy, 100 frozen prompts, instrumented HF loop)

eval tau alpha exactness vs AR
xLAM-500 4.01 0.942 49/50 token-identical
TB-500 (held-out tools) 2.65 0.812 49/50 token-identical

Outputs are token-identical to plain greedy decoding of the target in the HF instrumented loop (exactness gates; the vLLM bf16 near-tie caveat is documented in the repo README). The vocabulary is zero-padded to 152,064 rows to satisfy vLLM s SpeculativeConfig same-vocab-size requirement; no real token id lives in the padding band, and the padding is gated on greedy parity (see src/serving/prepare_draft.py).

Benchmark context: vLLM 0.30.0, target bf16, k=5, batch 1. This model is a research artifact from a 1-day H100 study; no safety fine-tuning was performed (the base model s policies apply).

Downloads last month
5
Safetensors
Model size
0.5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for vaishaalli/stage1-kd

Finetuned
(109)
this model