Up to 3.6x speedup with a DSpark speculative decoding draft model for Phi-4-mini-instruct (134 --> 486 tokens/sec; 4.95 accept length on math)

#48
by rasyosef - opened

Phi-4-mini-instruct-DSpark

Model link: https://huggingface.co/rasyosef/Phi-4-mini-instruct-DSpark

A DSpark draft model for speculative decoding with microsoft/Phi-4-mini-instruct as the verifier, trained with speculators. The drafter proposes 8 tokens at a time and the verifier checks them in one forward pass, so output is identical to running the verifier alone β€” a lossless speedup. Mean acceptance length is 3.23 tokens committed per verification step, up to 4.95 on math_reasoning. That gives a throughput speedup over the verifier alone of 3.61Γ— on math_reasoning and 3.44Γ— on HumanEval, and 2.20Γ— averaged across nine task types.

Training code: rasyosef/train-dspark-draft-models.

Trained on 116,000 samples.

Usage

vLLM loads the verifier automatically from the config β€” don't pass it separately.

vllm serve rasyosef/Phi-4-mini-instruct-DSpark \
  --port 8000 \
  --gpu-memory-utilization 0.8 \
  --override-generation-config '{"temperature": 0}'

Then query the OpenAI-compatible endpoint at http://localhost:8000/v1. The override sets the server's default temperature to 0, matching the evaluation below; a request that passes its own temperature still takes precedence.

Details

4 Qwen3 layers (hidden size 3072, intermediate size 8192, 24 attention heads over 8 KV heads, head dim 128). The first three layers use sliding-window attention with a 2048-token window and the last layer uses full attention. Block size 8, aux hidden-state layers 2/10/22/30. The drafter predicts over a 50,000-token draft vocabulary, a subset of the verifier's 200,064 tokens.

Trained for 3 epochs at lr 3e-4 (cosine schedule with 4% warmup) on 116,000 Open PerfectBlend prompts regenerated by the verifier itself, split 96/4 into train and validation, with a {"ce": 0.1, "tv": 0.9} loss. Prompts prepared at 2048 tokens; training sequence length 4096, up to 512 anchors per sample. Verifier hidden states were pulled on demand from a running vLLM server during training and deleted after use rather than staged to disk up front.

Evaluation

evaluate.py throughput across the nine RedHatAI/speculator_benchmarks subsets, at temperature 0. acceptance_length is mean tokens committed per verification step, including the bonus token β€” floor 1.0, ceiling 9.0 at block size 8.

subset acceptance_length pos_0 pos_1 pos_2 pos_3 pos_4 pos_5 pos_6 pos_7
math_reasoning 4.954 86.8% 74.1% 62.3% 51.1% 41.3% 33.1% 26.5% 20.2%
HumanEval 4.736 85.1% 70.8% 58.6% 47.3% 38.6% 30.2% 24.0% 18.9%
rag 3.157 70.2% 49.2% 34.3% 25.6% 19.4% 8.8% 4.8% 3.2%
writing 3.073 68.7% 45.9% 31.3% 22.0% 15.8% 11.6% 7.5% 4.6%
tool_call 2.829 69.0% 44.8% 28.8% 18.0% 10.5% 6.1% 3.5% 2.1%
qa 2.646 63.2% 39.6% 25.0% 17.8% 14.2% 2.7% 1.5% 0.7%
question 2.618 65.1% 39.2% 23.2% 14.0% 8.6% 5.5% 3.7% 2.4%
summarization 2.327 62.8% 35.5% 19.0% 9.2% 4.1% 1.6% 0.4% 0.1%
translation 2.151 60.4% 32.9% 13.0% 5.3% 1.9% 0.8% 0.5% 0.3%

Weighted across all subsets: 3.234 over 69,478 verification steps.

Acceptance is highest where the verifier's next token is most predictable β€” math and code. math_reasoning leads HumanEval by about 0.2 tokens and holds a margin at every position in the block, and the two sit well clear of everything else: the next subset, rag, is more than 1.5 tokens behind HumanEval. math_reasoning's pos_4 (41.3%) is above pos_1 on qa, question, summarization, and translation (39.6% or lower). rag and writing follow at 3.16 and 3.07, tool_call at 2.83, and qa and question at 2.65 and 2.62. summarization (2.33) and translation (2.15) are lowest; on translation fewer than one draft in seven survives to pos_2. qa and rag also drop sharply between pos_4 and pos_5 (14.2% β†’ 2.7% and 19.4% β†’ 8.8%) rather than decaying smoothly.

Throughput (tokens/s)

Mean output throughput (tokens/s) across the same nine subsets, measured on a single A100 at concurrency 1, with DSpark at temperature 0. Baseline is Phi-4-mini-instruct with no speculative decoding, measured on every subset; without a drafter, throughput varies only modestly with the task (123.4–135.4 tokens/s). Speedup is the mean throughput relative to that subset's baseline.

subset baseline (no drafter) DSpark (this model) speedup
math_reasoning 134.6 486.0 3.61Γ—
HumanEval 133.5 459.3 3.44Γ—
rag 123.4 219.1 1.78Γ—
writing 131.8 240.5 1.82Γ—
tool_call 123.8 287.3 2.32Γ—
qa 135.4 225.9 1.67Γ—
question 132.8 240.5 1.81Γ—
summarization 124.6 228.2 1.83Γ—
translation 127.9 190.6 1.49Γ—
average speedup 1.00Γ— β€” 2.20Γ—

DSpark averages a 2.20Γ— speedup (unweighted mean of per-subset speedups). The biggest gains are on math_reasoning at 3.61Γ— (134.6 β†’ 486.0 tokens/s) and HumanEval at 3.44Γ—. tool_call is next at 2.32Γ—; summarization, writing, question, and rag sit in a tight band between 1.78Γ— and 1.83Γ—, with qa at 1.67Γ—. translation is lowest at 1.49Γ—. Outside math and code, speedup does not follow acceptance length closely: tool_call is third on throughput but fifth on acceptance, and rag has the third-highest acceptance but one of the smaller speedups.

Means and medians agree to within about 6% on every DSpark subset. The mean runs about 5% above the median on tool_call, writing, and question, and 5–6% below it on math_reasoning and translation. Measured on medians instead, the speedup is 3.72Γ— on math_reasoning, 3.51Γ— on HumanEval, 2.06Γ— on tool_call, and 2.15Γ— on average; translation falls to 1.42Γ—, because its baseline median (143.0 tokens/s) is well above its baseline mean (127.9).

Limitations

Works only with Phi-4-mini-instruct and is not usable as a standalone model. Acceptance falls off steeply past the first few positions on prose-like traffic, and on summarization and translation in particular a block size of 8 is mostly wasted β€” the gains concentrate in math and code. Translation is the weakest subset on both acceptance and throughput; the 50,000-token draft vocabulary, which covers a quarter of the verifier's multilingual vocabulary, may contribute to that. Throughput was measured at concurrency 1 on one A100 at temperature 0; real-world speedup depends on your traffic mix, hardware, sampling settings, and load, and because verification is lossless, the verifier's own behavior and biases carry through unchanged.

License

MIT, matching the verifier. The speculators training code is Apache-2.0.

Sign up or log in to comment