grug-35b GGUF

grug brain inside rock. no custom Grug instruction required.

July 15, 2026 GGUF replacement: same public repo, all popular quants rebuilt from the corrected parent. Old files preserved on pre-intrinsic-style-fix-2026-07-14.

These files were rebuilt from the intrinsic-style-gated ProCreations/grug-35b checkpoint. The parent and embedded template contain no hidden Grug/style prompt; ordinary chat and tool-enabled agent scaffolds rely on the tuned weights.

27b and 35b hunt same prey

both parent grug hunt HumanEval and sanitized MBPP. number below come from big parent brain, NOT squeezed GGUF rock. grug not claim rock test it never get. number show pass@1 percent. bold grug win that hunt.

hunt grug-27b v2.1 grug-35b rebuilt
HumanEval (164) 87.2 80.5
MBPP sanitized (100) 85.0 88.0
file size
grug-35b-Q4_K_M.gguf 19.71 GiB
grug-35b-Q5_K_M.gguf 23.03 GiB
grug-35b-Q6_K.gguf 26.56 GiB
grug-35b-Q8_0.gguf 34.37 GiB
  • Q4_K_M: recommended local default
  • Q5_K_M: more accuracy meat
  • Q6_K: high-quality rock
  • Q8_0: closest popular quantized weight

All four loaded and generated successfully with llama.cpp f955e394bf94e01e5e36186d13c985727e5ef5b5 before upload. Hashes live in SHA256SUMS; smoke evidence lives in smoke-results.json.

Full BF16 parent: HumanEval 80.5%, MBPP 88.0%, broad valid/strict/right 100.0/100.0/95.0%, unprompted dialect-clean 100.0%.

Quant-specific full benchmark not claimed. Runtime must honor the embedded Qwen3.5 template. Reasoning stays in <think>...</think>; native XML tool calls stay intact.

Downloads last month
837
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ProCreations/grug-35b-gguf

Quantized
(2)
this model