Baseline drafter for the RSI Bench task speculative-decoder-discovery
A DSpark speculative-decoding drafter for Qwen/Qwen3-8B, with the architecture of DeepSeek's
dspark_qwen3_8b_block7 (5 layers, block size 7, target layers 1/9/17/25/33, Markov head rank 256, confidence head,
full vocabulary), trained from scratch with speculators 0.8.0 on
yemara/specdec-discovery-starter-data for 12 H100-hours (hidden-state server on one H100, trainer on another,
6 hours, 44,346 steps of 8,192 packed tokens, seed 0). train.sh is the exact recipe.
No pretrained drafter weights were used. It is the equal-compute baseline of the task.
Serve with vLLM 0.30.0:
vllm serve Qwen/Qwen3-8B --speculative-config '{"method": "dspark", "model": "<this repo>", "num_speculative_tokens": 7}'
- Downloads last month
- 32
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support