Block8-FP8-BLOCK โ€” Data-free FP8 Block quantization

Block8 is a quantized DFlash drafter for the Qwen3-8B target, derived from RedHatAI/Qwen3-8B-speculator.dflash at pinned revision 1a11b170eb65c8a62c80ecd01dfe22a5907298e6. This repository contains the drafter component; it is not a standalone chat model.

Variant

  • Quantization: FP8 Block (model_free_ptq(scheme="FP8_BLOCK") with static 128x128 blocks and dynamic 128 groups).
  • Calibration: Data-free quantization; no calibration data or samples were used.
  • Source snapshot: Pinned newer BF16 drafter snapshot 1a11b170eb65c8a62c80ecd01dfe22a5907298e6 (block size 16, draft vocabulary 151,936, sliding-window attention).

Use with vLLM

Pair this drafter with the Qwen3-8B target and a DFlash-capable vLLM build:

vllm serve Qwen/Qwen3-8B \
  --spec-model inference-optimization/Qwen3-8B-DFlash-FP8-BLOCK \
  --spec-tokens 7 \
  --spec-method dflash

config.py provides the custom drafter configuration. Serving command and runtime patch are preserved in provenance/evaluation/.

Evaluation Performance

Evaluated with 7 speculative tokens per draft event against the Qwen/Qwen3-8B target:

Benchmark Mean Acceptance Length ($L = 1 + A/D$) Mean Accepted Ratio ($A/P$)
SPEED-Bench Qualitative (11 categories, 880 requests) 3.1535 tokens 30.76%
RedHatAI / speculator_benchmarks (9 subsets, 924 requests) 3.1410 tokens 30.59%

Reproducibility & Provenance

Full provenance and reproducibility artifacts are preserved:

  • Root directory contains model weights, configuration, tokenizer, quantization manifests, and recipe.
  • provenance/ contains:
    • train_command.txt: Source model training command.
    • train_command_scope.txt: Clarification of source drafter snapshot lineage.
    • quantization_command.txt: Exact command and environment used to run quantization.
    • quantization_source/: Source code snapshot of quantizer utilities.
    • evaluation/: Contains vLLM serving commands (vllm_command.txt), vLLM runtime patch (vllm.patch), speculators patch (speculators.patch), target/drafter SHA-256 hashes, and per-subset evaluation commands.
    • publication_sha256.txt: SHA-256 hashes of all published files.

The source drafter lists Apache-2.0 licensing on its Hugging Face model card.

Downloads last month
42
Safetensors
Model size
2B params
Tensor type
BF16
ยท
F8_E4M3
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for inference-optimization/Qwen3-8B-DFlash-FP8-BLOCK

Finetuned
Qwen/Qwen3-8B
Quantized
(457)
this model

Collection including inference-optimization/Qwen3-8B-DFlash-FP8-BLOCK