Smaug-Mini-GGUF

Smaug-Mini is Abacus.AI's agentic finetune of Qwen3.8-27B, trained via on-policy reinforcement learning (GRPO, LoRA-merged into the language trunk only, vision tower left bitwise-identical to the base) over multi-turn, tool-using automation episodes with verified, outcome-based rewards, targeting more reliable end-to-end tool use and automation performance while holding general capabilities at parity with the base model. It retains Qwen3.8-27B's architecture, layout, 262,144-token context, and xhigh/medium/low reasoning-effort interface exactly, functioning as a drop-in replacement, while delivering notable agentic gains — +4.5 on AutomationBench (41.8 vs. 37.3), +17.1 on JobBench (50.5 vs. 33.4), +13.5 on NL2Repo-Bench (55.8 vs. 42.3), and +2.0 overall on LiveBench (76.9 vs. 75.3) — alongside modest improvements on reasoning benchmarks like HLE and IFBench, with GPQA-diamond and MMMU-Pro held essentially at parity. Its key behavioral shift is redistributing deliberation rather than adding it: episodes finish about three steps sooner at roughly unchanged total reasoning volume, and episodes that exhaust their step budget without completing drop from 3.4% to 1.0%. It's served via vLLM with a Qwen3 reasoning parser and tool-call parser (recommended sampling: temperature 1.0, top_p 0.95, reasoning effort xhigh), with the inherited MTP head left untrained against the updated trunk — so speculative decoding via MTP should stay disabled — and is released under Apache 2.0, inherited from Qwen3.8-27B.

Model Files

File Name Quant Type File Size File Link Description
Smaug-Mini.BF16.gguf BF16 53.8 GB Link Full BF16 weights. Highest quality, largest file size.
Smaug-Mini.Q4_K_M.gguf Q4_K_M 16.5 GB Link Good quality, default size for most use cases, recommended.
Smaug-Mini.Q5_K_M.gguf Q5_K_M 19.2 GB Link High quality, recommended.
Smaug-Mini.mmproj-bf16.gguf mmproj-bf16 931 MB Link Multimodal projection file in BF16 format. Used for vision/language models.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
431
GGUF
Model size
27B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/Smaug-Mini-GGUF

Base model

Qwen/Qwen3.8-27B
Quantized
(1)
this model

Collection including prithivMLmods/Smaug-Mini-GGUF