grug-35b-v2 — ROCmFP4-STRIX_LEAN (Strix Halo / gfx1151)

Version 1.0 — 2026-08-11

TL;DR

grug-35b-v2 (35B params, 3B active per token, Qwen3.5-VL-MoE family) quantized to Q4_0_ROCMFP4_STRIX_LEAN (type 106 preset, ~4.29 BPW). Tuned for AMD Strix Halo (gfx1151 / RDNA 3.5) with the charlie12345/ROCmFPX fork of llama.cpp. Runs the full vision + text multimodal model in ~17.3 GiB.

⚠️ Critical warnings — read before downloading

  • Requires charlie12345/ROCmFPX fork of llama.cpp (built via the kyuz0/amd-strix-halo-toolboxes container). The type 106 (Q4_0_ROCMFP4_STRIX_LEAN) tensor format is INVALID in stock llama.cpp — it will refuse to load. See Usage below.
  • Profiled for gfx1151 only (Strix Halo / Ryzen AI Max+ 395, RDNA 3.5). Not tested on other GPUs.
  • FP4 here is software on RDNA 3.5 (no FP4 silicon units): the win is bandwidth / memory, not raw compute throughput. The Strix Halo ceiling on this MoE is bandwidth-bound, which is exactly where FP4 helps.

Benchmarks

Tested on Strix Halo (AMD Ryzen AI Max+ 395, 128 GB LPDDR5X). Methodology: llama-bench -ngl 999 -fa on -p 512 -n 128 -mmap 0.

Model Quant Size tg128 (tok/s) pp512 (tok/s)
grug-35b-v2 ROCmFP4-STRIX_LEAN 17.31 GiB 70.92 1418
grug-35b-v2 Q4_K_M (baseline) 19.70 GiB 61.18
Qwen3.6-35B-A3B (production ref) ROCmFP4-STRIX_LEAN 17.31 GiB 63

Speed-up: +16% vs Q4_K_M (70.92 vs 61.18 tok/s tg128) at −12% size (17.31 vs 19.70 GiB). +12% vs the production Qwen3.6-35B-A3B reference.

System configuration at bench time

Declared for reproducibility:

  • Bare metal host: Bosgame BeyondMax Series (bosgame-m5), Ubuntu 24.04.4 LTS, kernel 7.0.0-28-generic
  • CPU power profile: balanced (powerprofilesctl get) — default, NOT forced to performance. Representative of an out-of-the-box setup.
  • CPU scaling driver: amd-pstate-epp, scaling_governor performance (amd-pstate-epp default), EPP performance
  • IOMMU / iGPU power: auto (no manual tuning)

Note: tok/s above were measured on a non-tuned system (power profile balanced). Users who set powerprofilesctl set performance may see slightly higher numbers.

Quantization details

Preset Q4_0_ROCMFP4_STRIX_LEAN (GGUF file_type 106, ~4.29 bits/weight):

  • Attention K/V (blk.*.attn_qkv.weight, blk.*.attn_v.weight) → q4_0_rocmfp4 (high-precision path for attention state)
  • Token embeddings (token_embd.weight) → Q5_K (preserve vocab fidelity)
  • Expert FFN (blk.*.ffn_*_exps.weight) → q4_0_rocmfp4_fast (max speed path; the bulk of MoE weights)
  • Other tensors → F32 / Q4_0_ROCMFP4_FAST as appropriate

Reference fork: charlie12345/ROCmFPX commit 00d5452.

imatrix methodology

Generated with llama-imatrix (256 chunks, 16 threads, CPU-only). Calibration text from ProCreations/grug-think-v3-10kpublic Apache-2.0 dataset, not gated: anyone can download it to replicate. Many thanks to the grug team for publishing both the model and a clean calibration set.

  • 510 entries over 733 tensors
  • Warning partial data 99.61% during quantization = 1/256 expert not activated in calibration (normal for MoE — see tools/imatrix/imatrix.cpp in llama.cpp). Negligible impact.

Files

File Size Description
grug-35b-v2-ROCmFP4-STRIX_LEAN.gguf ~17.32 GiB Main model (type 106)
mmproj-grug-35b-v2-f16.gguf ~857 MB Vision projector (F16)
imatrix-grug-35b-v2.gguf ~183 MB Importance matrix (for re-quantization)

Usage

# Requires the kyuz0 Strix Halo toolbox (which builds charlie12345/ROCmFPX)
docker run --rm -p 1234:1234 --device /dev/kfd --device /dev/dri \
  -v /path/to/models:/models rocmfpx-llm-service \
  llama-server \
    -m /models/grug-35b-v2-ROCmFP4-STRIX_LEAN.gguf \
    --mmproj /models/mmproj-grug-35b-v2-f16.gguf \
    -ngl 999 -fa on --jinja -c 32768 --host 0.0.0.0 --port 1234

Notes:

  • MTP not enabled for grug. The mtp_num_hidden_layers field is 0 in this model (MTP was removed during fine-tuning), so it cannot be activated.
  • The --mmproj flag is required for the vision tower (multimodal). Without it, text-only still works.

How to replicate

Textual pipeline only (no published scripts):

  1. Build the docker-llm-service-convert image from kyuz0/amd-strix-halo-toolboxes + charlie12345/ROCmFPX (commit 00d5452 or later main HEAD — must contain MODEL_ARCH.QWEN35MOE).
  2. Download the BF16 safetensors from ProCreations/grug-35b-v2.
  3. Convert to GGUF with convert_hf_to_gguf.py (inside the container).
  4. Generate the imatrix with llama-imatrix using ProCreations/grug-think-v3-10k (256 chunks).
  5. Quantize: llama-quantize <bf16>.gguf <out>.gguf Q4_0_ROCMFP4_STRIX_LEAN 16.

Attribution & model tree

Qwen3.5-VL-MoE (base architecture)
    └── ornith-ai/Ornith-1.0-35B (MIT)
            └── ProCreations/grug-35b-v2 (Apache-2.0)
                    └── this GGUF (ROCmFP4-STRIX_LEAN)

License

Apache-2.0 (inherited from ProCreations/grug-35b-v2). Derivative work: original model and its license are preserved. See LICENSE and NOTICE.

Acknowledgements

Built on the shoulders of giants:

Limitations & community feedback

  • Speed benchmark only. No perplexity / MMLU / quality eval is included in this release. The MoE structure is preserved bit-for-bit from the BF16 source except for the quantized tensor formats above; quality is expected to track standard Q4_K_M-class with the ROCmFP4 attention/K-V choices, but this is not measured here.
  • Profiled for gfx1151 only. Not tested on other GPUs (no Navi 3 / Navi 4 / data-center MI series numbers — feel free to share yours).
  • MTP not activated (plain inference).

We invite the community — especially fellow Strix Halo owners — to test and share quality results. Open a Discussion on this repo.

Citation

@misc{grug35b2026,
  title  = {grug-35b-v2},
  author = {ProCreations},
  year   = {2026},
  url    = {https://huggingface.co/ProCreations/grug-35b-v2}
}

Disclaimer

No affiliation with AMD, Qwen, ProCreations, DeepReinforce, unsloth, kyuz0, or charlie12345. Provided as-is, without warranty. Users must comply with the base model license (Apache-2.0).

Downloads last month
-
GGUF
Model size
35B params
Architecture
qwen35moe
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for pugant/grug-35b-v2-ROCmFP4-STRIX_LEAN

Quantized
(8)
this model