Hugging Face's logo Hugging Face
  • Models
  • Datasets
  • Spaces
  • Buckets new
  • Docs
  • Enterprise
  • Pricing
    • Website
      • Tasks
      • HuggingChat
      • Collections
      • Languages
      • Organizations
    • Community
      • Blog
      • Posts
      • Daily Papers
      • Hardware
      • Learn
      • Discord
      • Forum
      • GitHub
    • Solutions
      • Team & Enterprise
      • Hugging Face PRO
      • Enterprise Support
      • Inference Providers
      • Inference Endpoints
      • Storage Buckets

  • Log In
  • Sign Up

canada-quant
/
GLM-5.3-Flash-DFlash2-F

Text Generation
Safetensors
vllm
qwen3
speculative-decoding
dflash2
draft-model
draft
glm-5.3-flash
glm5_next
w4a16
dgx-spark
h200
b300
Model card Files Files and versions
xet
Community
GLM-5.3-Flash-DFlash2-F
6.22 GB
Ctrl+K
Ctrl+K
  • 1 contributor
History: 6 commits
pastapaul's picture
pastapaul
Card + launcher: drafter memory cost, MTP guidance, KV default 8 GiB (2026-10-04)
f8e7677 verified 5 days ago
  • .gitattributes
    1.52 kB
    initial commit 15 days ago
  • PROVENANCE.txt
    881 Bytes
    dflash2F: DFlash2 drafter for GLM-5.3-Flash W4A16 (3.626 @K7, parity with the incoai reference on the same B300) 15 days ago
  • README.md
    10.4 kB
    Card + launcher: drafter memory cost, MTP guidance, KV default 8 GiB (2026-10-04) 5 days ago
  • config.json
    1.97 kB
    dflash2F: DFlash2 drafter for GLM-5.3-Flash W4A16 (3.626 @K7, parity with the incoai reference on the same B300) 15 days ago
  • launch_dflash2_tp2.sh
    6.47 kB
    Card + launcher: drafter memory cost, MTP guidance, KV default 8 GiB (2026-10-04) 5 days ago
  • mask_embedding.pt
    9.88 kB
    xet
    dflash2F: DFlash2 drafter for GLM-5.3-Flash W4A16 (3.626 @K7, parity with the incoai reference on the same B300) 15 days ago
  • model.safetensors
    6.22 GB
    xet
    dflash2F: DFlash2 drafter for GLM-5.3-Flash W4A16 (3.626 @K7, parity with the incoai reference on the same B300) 15 days ago