spec-b300 / source /docs /user_guide /algorithms /decision_guide.md
khazic's picture
Archive three-epoch run: logs and provenance part 1
a181ec9 verified
|
Raw History Blame Contribute Delete
3.64 kB

Algorithm Decision Guide

Speculators currently supports six speculative decoding algorithms: Eagle-3, P-EAGLE, DFlash, DFlash2, DSpark, and MTP. All are lossless -- they produce output from the same distribution as the target model.

How They Differ

Eagle-3 predicts draft tokens autoregressively, one at a time.

P-EAGLE extends Eagle-3 with parallel multi-token prediction across multiple depths, using COD sampling for memory-efficient training.

DFlash predicts all draft tokens in a single forward pass using block-based prediction with anchor points.

DFlash2 extends DFlash with local dynamic convolutions and a low-rank candidate selector. Its training objective is experimental and follows the upstream DFlash2 candidate-set contract.

DSpark extends DFlash with a Markov head, so positions inside a block depend on earlier ones, plus a confidence head that predicts per-position acceptance.

MTP finetunes the model's native multi-token prediction head on domain-specific data. Unlike the other algorithms, MTP does not train from scratch -- it starts from pre-existing MTP layers and is only available for models with native MTP support.

Eagle-3, P-EAGLE, DFlash, and DSpark can be paired with any supported verifier model (including quantized variants). DFlash2 requires the verifier's full vocabulary because its selector operates on the complete vocabulary. MTP requires a model with native MTP layers (e.g. Qwen3-Next, Qwen3.5).

Current Support

Eagle-3 P-EAGLE DFlash DFlash2 DSpark MTP
Draft layers Llama-style Llama-style Qwen3-style Qwen3-style Qwen3-style Native MTP layers
Verifier models Any supported Any supported Any supported Full vocabulary Any supported Models with native MTP only
Training mode From scratch From scratch From scratch Experimental From scratch Finetune existing MTP head
Speculators Mature Newer, growing fast Newer, growing fast Experimental Newer, growing fast Newer, growing fast
vLLM Mature Newer, growing fast Newer, growing fast Native/experimental Newer, growing fast Newer, growing fast

Eagle-3 has been available longer and has broader support in both Speculators and vLLM. The others were added more recently and support is improving rapidly; DFlash2 remains experimental.

Which Should I Use?

If you're unsure, start with Eagle-3 -- it has the most mature tooling and documentation. If you want parallel multi-token prediction with an Eagle-3-based architecture, try P-EAGLE. If you want to experiment with DFlash's single-forward-pass block prediction approach, the training workflow is the same. DFlash2 adds local convolution and candidate reranking, while DSpark adds intra-block dependencies and confidence scheduling. If your model already has native MTP layers (e.g. Qwen3-Next, Qwen3.5), MTP finetuning lets you improve the existing MTP head on domain-specific data without training a separate draft model.

For more details on each algorithm, see: