Drift8-FP8-DYNAMIC โ€” Data-free FP8_DYNAMIC quantization

Drift8 is a quantized DFlash drafter for the Qwen3-8B target, derived from RedHatAI/Qwen3-8B-speculator.dflash. This repository contains the drafter component; it is not a standalone chat model.

Variant

  • Quantization: FP8_DYNAMIC (per-channel weights with token-dynamic activations).
  • Calibration: Data-free quantization: no calibration data or calibration samples were used; manifest seed 0. This is the single FP8_DYNAMIC checkpoint used in the evaluation.
  • Calibration seed: 0.
  • Quantization settings and source revisions: quant_run_manifest.json.

Use with vLLM

Pair this drafter with the Qwen3-8B target and a DFlash-capable vLLM build:

vllm serve Qwen/Qwen3-8B \
  --spec-model inference-optimization/Qwen3-8B-DFlash-Drift8-FP8-DYNAMIC \
  --spec-tokens 7 \
  --spec-method dflash

config.py provides the custom drafter configuration. The experiment's serving command and runtime patch are in provenance/evaluation/.

Reproducibility

The manifests are included at the repository root. provenance/ contains the source drafter's captured train_command.txt, a quantization command explicitly marked as reconstructed, the quantizer and calibration source snapshot, the vLLM command and patch, both target and drafter checkpoint hashes, and the nine per-subset evaluation commands. The selected seed-0 checkpoint is the same checkpoint used in the 2026-09-24 all-subset evaluation.

The PerfectBlend preparation, prompts, and hidden-state cache remain local because the prepared prompts are not redistributable. The cache sample counts and content hashes are recorded in calibration_manifest.json; no prompts or hidden-state tensors are uploaded.

The source drafter lists Apache-2.0 licensing on its Hugging Face model card.

Downloads last month
42
Safetensors
Model size
2B params
Tensor type
BF16
ยท
F8_E4M3
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for inference-optimization/Qwen3-8B-DFlash-FP8-DYNAMIC

Finetuned
Qwen/Qwen3-8B
Quantized
(455)
this model

Collection including inference-optimization/Qwen3-8B-DFlash-FP8-DYNAMIC