GLM-5.2 FP8 DFLASH v2

Overview

This is a DFLASH speculative draft model for GLM-5.2 FP8 serving. The checkpoint uses DFLASH block size 12 and is intended to be loaded as the draft model in SGLang speculative decoding.

This model is fine-tuned on top of SubconsciousDev/glm-5.2-fp8-dflash-v1 using SubconsciousDev/Subconscious-Dflash-Training-Dataset-mix-glm52-25k.

SGLang Usage

Add these arguments to the SGLang launch command:

--speculative-algorithm DFLASH \
--speculative-draft-model-path SubconsciousDev/glm-5.2-fp8-dflash-v2 \
--speculative-num-draft-tokens 12 \
--speculative-draft-kv-cache-dtype bfloat16
Downloads last month
15
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support