All QUASAR Models Collection All QUASAR checkpoints in one place: 4-bit QAT for Qwen, Gemma and Muse — NVFP4 W4A16/W4A4 for vLLM, Q4_0 GGUF for llama.cpp / Ollama. • 12 items • Updated 19 days ago • 2
Muse Glimmer 30B — QUASAR 4-bit QAT Collection Muse Glimmer 30B at 4 bits with QUASAR: W4A16 72% lower KL than Red Hat / 39% than NVIDIA · W4A4 20% lower than Red Hat · Q4_0 GGUF beats Meta Q4_K_M. • 4 items • Updated 19 days ago • 1
Gemma 4 E4B & 12B — QUASAR 4-bit QAT Collection Native 4-bit Gemma 4 E4B & 12B with QUASAR: lower KL to BF16 than Google's official QAT at both sizes, in Q4_0 GGUF (llama.cpp) and W4A16 (vLLM). • 5 items • Updated 19 days ago • 1
Qwen3.5-4B & Qwen3.8-27B — QUASAR 4-bit QAT Collection Full-decoder 4-bit Qwen with QUASAR: 27B leads the compared NVFP4 builds on reasoning; 4B ships W4A16, Blackwell W4A4, and native Q4_0 GGUF. • 5 items • Updated 19 days ago • 1
QUASAR: Lowering the Loss Floor of Quantization-Aware Training with Loss-Aware Reconstruction Paper • 2608.13966 • Published Aug 14 • 6