AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters
Paper • 2607.19223 • Published
How to use AdaFlash/Qwen3-8B-AdaFlash with Transformers:
# Use a pipeline as a high-level helper
from transformers import pipeline
pipe = pipeline("feature-extraction", model="AdaFlash/Qwen3-8B-AdaFlash", trust_remote_code=True) # Load model directly
from transformers import AutoTokenizer, AutoModel
tokenizer = AutoTokenizer.from_pretrained("AdaFlash/Qwen3-8B-AdaFlash", trust_remote_code=True)
model = AutoModel.from_pretrained("AdaFlash/Qwen3-8B-AdaFlash", trust_remote_code=True, device_map="auto")
This is the AdaFlash drafter for Qwen3-8B — an on-policy distilled diffusion drafter with an adaptive length head, trained using PerfectBlend data and reverse-KL loss. It is initialized from Qwen3-8B-DFlash-b16.
AdaFlash accelerates LLM inference through speculative decoding. The adaptive length head dynamically adjusts the number of draft candidates per step, balancing speed and acceptance rate.
AdaFlash requires the patched SGLang PR #34171 for serving.
# Install
conda create -n adaflash python=3.12 -y && conda activate adaflash
pip install "sglang[all] @ git+https://github.com/sgl-project/sglang.git@refs/pull/34171/head#subdirectory=python"
# Serve
export MODEL_PATH=Qwen/Qwen3-8B
export DRAFT_MODEL_PATH=AdaFlash/Qwen3-8B-AdaFlash
# Launch the server & benchmark
bash scripts/tools/benchmark.sh
For full setup, training, and benchmark details, see the GitHub repo.
@article{arxiv'26:adaflash,
title = {AdaFlash: Adaptive Speculative Decoding via On-Policy Distilled Diffusion Drafters},
author = {Qian, Yu-Yang and Wu, Hao-Cong and Chen, Chen and Sun, Jiacheng and Dong, Zhenhua and Zhao, Peng and Zhou, Zhi-Hua},
journal = {ArXiv preprint},
volume = {arXiv:2607.19223},
year = {2026}
}
Base model
z-lab/Qwen3-8B-DFlash-b16