GLM-4.7-Flash-Coder โ€” Dynamic 1-bit / 2-bit GGUF quants

Ultra-low-bit GGUF quantizations of whitecircle/GLM-4.7-Flash-Coder (MIT, a coding-agent finetune of zai-org/GLM-4.7-Flash, 29.9B params, MoE: 47 layers, 64 routed experts top-4 + 1 shared, MLA attention, layer 0 dense), following the logic of Unsloth's Dynamic 3.0 GGUFs: per-tensor mixed precision from an imatrix, not a uniform bit width.

Method

  1. Per-tensor type plan: read from the public headers of unsloth/GLM-4.7-Flash-GGUF (their quant of the base model, plan/unsloth_*.json) and re-applied to this finetune via llama-quantize --tensor-type-file. Pattern: norms/router/biases stay F32, MLA attention projections stay Q8_0/Q5_K/Q6_K, the shared expert and embeddings/output stay 4โ€“6 bit, and routed-expert gate/up/down tensors carry the bulk of the bit-budget cut (down to IQ1_S/IQ1_M in the most compressible layers, IQ2_Sโ€“IQ4_XS elsewhere).
  2. Our own imatrix, computed on our own calibration mix (calib/calib.txt, ~410K tokens): 55% agentic coding traces (whitecircle/swe-rebench-v2-glm-5.1-pi-agent-successful-traces, rendered through the model's own chat template with tool definitions), 20% general chat (SmolTalk), 15% general text (Cosmopedia-v2), 10% multilingual (FineWeb-2, 6 languages) โ€” no Unsloth weights or imatrix data reused.
  3. Quality check: KL-divergence and top-1 agreement of each quant vs. the Q8_0 reference, on a held-out slice (calib/heldout.txt, ~24.6K tokens, different documents/traces than calibration).

Results (12 chunks ร— 2048 tokens, held-out; Q8_0 baseline PPL 13.98)

variant size bits mean PPL ฮ”PPL vs Q8_0 mean KL divergence same top-1 token
UD-IQ1_S (1-bit class) 8.6 GiB 2.47 BPW 15.51 +17.6% 1.241 68.9%
UD-IQ2_XXS (2-bit class) 9.8 GiB ~2.8 BPW 14.00 +6.1% 1.068 74.9%
UD-Q2_K_XL (2-bit class) 11.1 GiB ~3.1 BPW 15.22 +15.4% 1.085 75.8%

Caveat on these numbers: the held-out set is small (~24.6K tokens) and mixed-domain, so the ยฑ0.5 PPL error bars are non-trivial relative to the differences between quants โ€” treat this as a directional signal, not a precise ranking. Notably, IQ2_XXS scored best on PPL/ฮ”PPL despite being smaller than Q2_K_XL, which is not what Unsloth's own (much larger-scale) evaluation found for the base model; the tensor-type plan was derived from the base model and may not transfer perfectly to this coding-agent finetune's expert routing.

Generation sanity check

Both UD-IQ1_S and UD-IQ2_XXS produce coherent, on-topic reasoning for a coding prompt ("write a function that reverses a singly linked list") โ€” thinking-mode traces stay on task and don't devolve into repetition or gibberish in a short sample. Not a rigorous eval; see Limitations.

variant prompt eval generation threads
UD-IQ1_S 43.5 tok/s 8.3 tok/s 56 (CPU)
UD-IQ2_XXS 32.0 tok/s 6.3 tok/s 56 (CPU)

Limitations

  • 1-bit quantization measurably degrades this model (+17.6% perplexity, only 69% of tokens agree with the fp16-equivalent reference at top-1). Per Unsloth's own published guidance, expect noticeably worse tool-calling and multi-step reasoning at this level โ€” this is a research/fits-in-less-VRAM option, not a quality-preserving one.
  • 2-bit (IQ2_XXS/Q2_K_XL) is the practical floor for this size/architecture if you want the model to still behave close to its original self.
  • No agentic tool-calling benchmark (SWE-bench-style) was run โ€” only a KL-divergence/perplexity proxy and a short generation spot-check. If you depend on tool-calling accuracy, evaluate that directly before relying on these quants.
  • The per-tensor type plan is transplanted from the base model's Unsloth quant, not derived from this finetune's own imatrix-guided layer sensitivity analysis โ€” a plan computed natively for this checkpoint might do better, especially for the routed experts most reshaped by the coding-agent finetuning.

Files

  • GLM-4.7-Flash-Coder-UD-IQ1_S.gguf โ€” 8.6 GiB, 1-bit class
  • GLM-4.7-Flash-Coder-UD-IQ2_XXS.gguf โ€” 9.8 GiB, 2-bit class, best measured quality/size here
  • GLM-4.7-Flash-Coder-UD-Q2_K_XL.gguf โ€” 11.1 GiB, 2-bit class, Unsloth's usual flagship 2-bit recipe

License

MIT (inherited from the base model and this finetune). Calibration data license terms are their own (SmolTalk: Apache-2.0; Cosmopedia-v2: Apache-2.0/ODC-By; FineWeb-2: ODC-By; the agentic traces dataset: see its own card) โ€” none of it is redistributed here, only used to compute the imatrix.

Downloads last month
813
GGUF
Model size
30B params
Architecture
deepseek2
Hardware compatibility
Log In to add your hardware

1-bit

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for LakoMoor/GLM-4.7-Flash-Coder-GGUF

Quantized
(3)
this model