Qwen3.6-27B DFlash drafter โ€” safetensors (vLLM)

Mirror of z-lab/Qwen3.6-27B-DFlash (MIT), the block-diffusion DFlash drafter for Qwen3.6-27B targets. Kept under this account so the full vLLM speculative-decoding stack pairs with the uncensored finetune repos below.

Serve with vLLM (FP8 target + DFlash, ~2x decode)

vllm serve zeeksa/Qwen3.6-27B-Uncensored-HauhauCS-Aggressive-FP8-Dynamic \
  --speculative-config '{"method": "dflash", "model": "zeeksa/Qwen3.6-27B-DFlash", "num_speculative_tokens": 15}' \
  --max-num-batched-tokens 32768 --max-num-seqs 64

DFlash is lossless โ€” output is identical to serving the target alone; only speed changes.

Related

Downloads last month
85
Safetensors
Model size
2B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for zeeksa/Qwen3.6-27B-DFlash

Finetuned
(6)
this model