Quantized-Drafters
Collection
20 items โข Updated
Block8 is a quantized DFlash drafter for the Qwen3-8B target, derived from RedHatAI/Qwen3-8B-speculator.dflash at pinned revision 1a11b170eb65c8a62c80ecd01dfe22a5907298e6. This repository contains the drafter component; it is not a standalone chat model.
model_free_ptq(scheme="FP8_BLOCK") with static 128x128 blocks and dynamic 128 groups).1a11b170eb65c8a62c80ecd01dfe22a5907298e6 (block size 16, draft vocabulary 151,936, sliding-window attention).Pair this drafter with the Qwen3-8B target and a DFlash-capable vLLM build:
vllm serve Qwen/Qwen3-8B \
--spec-model inference-optimization/Qwen3-8B-DFlash-FP8-BLOCK \
--spec-tokens 7 \
--spec-method dflash
config.py provides the custom drafter configuration. Serving command and runtime patch are preserved in provenance/evaluation/.
Evaluated with 7 speculative tokens per draft event against the Qwen/Qwen3-8B target:
| Benchmark | Mean Acceptance Length ($L = 1 + A/D$) | Mean Accepted Ratio ($A/P$) |
|---|---|---|
| SPEED-Bench Qualitative (11 categories, 880 requests) | 3.1535 tokens | 30.76% |
| RedHatAI / speculator_benchmarks (9 subsets, 924 requests) | 3.1410 tokens | 30.59% |
Full provenance and reproducibility artifacts are preserved:
provenance/ contains:train_command.txt: Source model training command.train_command_scope.txt: Clarification of source drafter snapshot lineage.quantization_command.txt: Exact command and environment used to run quantization.quantization_source/: Source code snapshot of quantizer utilities.evaluation/: Contains vLLM serving commands (vllm_command.txt), vLLM runtime patch (vllm.patch), speculators patch (speculators.patch), target/drafter SHA-256 hashes, and per-subset evaluation commands.publication_sha256.txt: SHA-256 hashes of all published files.The source drafter lists Apache-2.0 licensing on its Hugging Face model card.