--- license: apache-2.0 library_name: transformers base_model: Qwen/Qwen3.8-27B base_model_relation: finetune pipeline_tag: image-text-to-text tags: - agentic - smaug - abacusai ---
| Architecture | Dense hybrid-attention transformer + vision tower |
| Total Parameters | 27B |
| Number of Layers | 64 (48 linear-attention + 16 full-attention, 3:1 interleave) |
| Attention Mechanism | Gated linear attention & full attention (GQA) |
| Hidden Dimension | 5120 |
| Number of Attention Heads | 24 (4 KV heads) |
| Vision Encoder | 27-layer ViT, patch 16 |
| Vocabulary Size | ~248K |
| Context Length | 262,144 |
| Multi-Token Prediction | 1-layer MTP head (inherited; leave speculative decoding off) |
| Precision | bfloat16 |
| Modality | Text, Image |
| Base Model | Qwen/Qwen3.8-27B |
| Adaptation | On-policy RL (GRPO), LoRA merged as full delta (language trunk only) |
| Smaug-Mini | Qwen3.8-27B (base) |
Qwen3.6-27B | Qwen3.7-Plus | Opus4.6 Max | |
|---|---|---|---|---|---|
| Agentic | |||||
| AutomationBench | 41.8 | 37.3 | — | 20.4 | 25.5 |
| JobBench | 50.5 | 33.4 | 21.8 | 27.6 | 36.9 |
| LiveBench agentic coding | 60.8 | 61.4 | 39.3 | — | 49.0 |
| NL2Repo-Bench | 55.8 | 42.3 | 36.2 | 41.1 | 47.6 |
| Reasoning, knowledge & instruction following | |||||
| GPQA-diamond | 89.4 | 89.2 | 87.8 | 90.3 | 91.3 |
| HLE | 34.2 | 30.8 | 24.0 | 34.7 | 40.0 |
| IFBench | 82.0 | 79.5 | 69.1 | 79.1 | 62.5 |
| LiveBench overall | 76.9 | 75.3 | 64.0 | — | 74.5 |
| Vision | |||||
| MMMU-Pro | 75.6 | 76.3 | 75 | 80 | 75 |