Wren logo

Wren - GGUF

Wren is a compact derivative of Swift 1.5 Qwen3.8-Flash-Next by UkisAI, which is itself a derivative of Qwen3.8-Flash-Next. Like the wren, one of the smallest birds, it is built to be small and agile: half of the experts, measured mixed precision, made for a consumer GPU. "Swift" and "UkisAI" are trademarks of UkisAI and are used here only to describe the model's origin.

GGUF quantisations of ohmychemo/Wren.

  • The base is Swift1.5-Qwen3.8-Flash-Next with 50% of the routed experts removed: 256 of 512 per layer, chosen by maximin over REAP saliency on 13 calibration areas, including agentic ones.
  • Routers and norms were healed by distillation from Swift's own logits.
  • All files use an imatrix computed from Swift's own outputs.
  • They are built to run on a consumer GPU, with the experts and the n-gram table in RAM.
  • In order to use Wren models you need to install this llama.cpp fork.

The research behind these files

📄 Paper: paper/paper.md, "Wren_T3: measured-sensitivity, per-expert mixed precision for a pruned MoE on a 12 GB consumer GPU" (about 14k words).

  • It documents every phase, with motivations, all measurements and the negative results.
  • It was written by T3 itself, running locally on the RX 6700 XT through Claude Code, from the project record in paper/context/. That record includes the full research summary, the local measurements and the raw data the paper cites.
  • It was then fact-checked against that record and corrected.

Goal: "intelligence density". Derive from a strong open MoE a model that runs on one consumer machine (12 GB VRAM, 64 GB RAM), loses as little measured intelligence as possible and keeps Swift's short reasoning. Every choice is driven by measurements on a multidisciplinary quiz, not by heuristics.

Phase What we did Key result
1. Measurement Swift run in bf16 one layer at a time (exact to 5e-8). Expert frequency, router weight and REAP saliency on 13 areas: 8 topical + 5 agentic (terminal, code agent, tool use, search, office). Pruning simulated by masking experts in the router Maximin (protect the worst-served area) at 50%: quiz 84.0, the best criterion; it keeps 77–82% of REAP mass in every area
2. Pruning + healing Experts removed physically in stages (100 → 75 → 60 → 50%). Routers and norms healed on Swift's own top-20 logits (3.58M tokens) Quiz 84.0 → 83.3, KL 0.338 → 0.191 (intact Swift 88.0). 35% failed: dropping 74.3 after healing; merging experts 67.0, worse than dropping. Training more parameters made KL worse
3. Quantisation Fake-RTN sweep, then GGUF variants with imatrix Breaking point between 3 and 2 bits. Down projections are the most sensitive tensors and, with 640 columns, admit only block-32 types
4. Measured mixed precision 48-layer sensitivity scan (quiz-KL, in-place byte patching). Quiz-weighted per-expert error. A llama.cpp fork with two expert groups per layer: the routing is unchanged, and the CUDA bug it exposed is fixed. Knapsack/Lagrangian allocations (T1, T3). Fast proxy pipeline and a 60-configuration random design with ridge regression (T4) The middle layers (11–29) are the most sensitive, the deep ones (30–46) the most robust, contrary to the "protect the first/last layers" heuristic. Within a layer, the measured expert error ranks experts twice as well as REAP. Damage is not additive. Regression R² 0.88 on KLD
5. Local deployment T3 on an RX 6700 XT: ROCm (with a flash-attention fix for RDNA2) and Vulkan Same quiz across backends within noise. About 11 tokens/s generation, 234 tokens/s prefill, 262k context in 10.1 GB VRAM

Main findings.

  1. Measured sensitivity beats position heuristics.
  2. Optimising quiz answers and optimising the next-token distribution give different models. T3 keeps math at 72; the KLD-optimised T4 drops it to 62 at the same size.
  3. IQ2_XXS is too destructive to be a useful tier. Uniform iq3 gate/up beats every iq2 mix in KLD, for about 0.7 GB more.
  4. Fast proxies rank configurations correctly but compress the gaps by about 2.5×.

The negative results are published together with the files that worked.

Files

File Size KLD vs Q8_0 Same top-1 Quiz (300 q) Runs on
Wren_Q4_K_M.gguf 77.6 GB 0.0185 95.6% 83.3 upstream llama.cpp
Wren_T1.gguf 62.7 GB 0.0499 92.8% 81.0 upstream llama.cpp
Wren_T3.gguf 60.4 GB 0.0788 90.9% 82.3 tiered-experts fork
Wren_T3-v2.gguf 60.4 GB — ¹ — ¹ 83.0 ² (T3: 80.7 ²) tiered-experts fork
Wren_T4-2.92.gguf 60.4 GB 0.0745 91.5% 80.7 tiered-experts fork
Wren_T4-2.6.gguf 58.7 GB 0.0973 90.2% 80.0 tiered-experts fork

¹ T3-v2 is re-healed toward intact Swift, so it is meant to differ from the pruned model's Q8_0: KLD against that Q8_0 does not measure it. ² Measured on ROCm (RX 6700 XT); the other quiz values in this table were measured with CUDA on an RTX 4090. On ROCm the same T3 scores 80.7, so compare 83.0 with 80.7.

Reference values (not published here):

  • the Q8_0 of the same pruned model: about 124 GB, quiz 82.0;
  • intact Swift: quiz 88.0;
  • the healed 50% model in bf16: quiz 83.3.

What the variants are. All files have the attention, linear attention, indexer, shared experts, hyper-connections, token embedding and output at q8_0. They differ only in the routed experts and the n-gram table:

File Expert gate/up Expert down N-gram table
Q4_K_M standard Q4_K_M standard Q4_K_M q4_1
T1 per layer, from a 48-layer sensitivity scan measured on the quiz: q4_K on the 14 most sensitive layers, iq2_xxs on layers 44–46, iq3_xxs elsewhere iq4_nl q4_0
T3 per expert: two groups per layer, sized by a Lagrangian allocation over the measured layer sensitivity and a quiz-weighted per-expert error. 6.5% of experts at q4_K, 69.8% at iq3_xxs, 23.7% at iq2_xxs iq4_nl q4_0
T4-2.92 / T4-2.6 per expert, fitted by ridge regression on 64 quickly evaluated configurations (2.92 and 2.6 bpw on gate/up) iq4_nl q4_0

Which one to use:

  • Most faithful: Q4_K_M. It is effectively lossless vs Q8_0.
  • About 60 GB, best task answers: T3-v2 (same size and speed as T3; see below).
  • T3 is the same file before the re-healing.
  • About 60 GB, closest next-token distribution: T1 (62.7 GB, no fork needed) or T4-2.92.
  • T4-2.6 is the smallest; it is mainly a data point.

Wren_T3-v2: re-healed on the compressed model

Same routed experts, n-gram table and quant types as T3; only the routers, the norms, the shared experts and the attention / linear-attention output projections (Q8_0) changed:

  • What was trained: routers (F32), all norms, and rank-16 LoRA on the shared experts and the attention / linear-attention output projections. The LoRA is merged into the Q8_0 tensors.
  • On what: the compressed T3 itself, in PyTorch. The experts stay as GGUF bytes and are decoded by Triton kernels. Activations and the KV cache are quantised as llama.cpp does at inference, so the training sees the model that actually runs.
  • Teacher: intact Swift's top-20 log-probs. The first stage (LoRA only, 6M tokens) used the original distillation set. The second stage (routers + norms + LoRA, 12.1M tokens, 4×A100 data-parallel) used new Swift xhigh generations for the domains that set lacked:
Data of the second stage Docs
Math: Swift xhigh solutions + scored problem/solution texts 1,171 + 5,134
Logic: Swift xhigh answers 800
Philosophy: Swift xhigh analysis of passages 400
Science: ARC-Challenge train split, Swift xhigh 400
Commonsense: CommonsenseQA train split, Swift xhigh 400
General share (30% of tokens) original set

None of these sources is the quiz's split (the quiz uses ARC-Challenge test, HellaSwag and MMLU).

Held-out KL to intact Swift (second stage, before → after):

Commonsense Logic Science Philosophy Math General All
0.199 → 0.136 0.154 → 0.114 0.195 → 0.150 0.262 → 0.220 0.130 → 0.114 0.192 → 0.172 0.158 → 0.134

llama.cpp, same machine and backend (ROCm):

Perplexity (8 × 512) Quiz Code Commonsense Logic Math Philosophy Science
T3 1.823 80.7 78 92 90 70 74 80
T3 + LoRA (stage 1, not published) 1.764 82.3 80 92 92 68 80 82
T3-v2 1.698 83.0 82 92 90 68 82 84

Against T3, 12 answers turn right and 5 turn wrong. At 300 questions the 2.3-point gain is about one noise width, so the perplexity and KL drops are the stronger evidence. The quiz answers without thinking, while the new math data are reasoned xhigh solutions: on a 100-question held-out math set (no thinking) the score went 59 → 65, and the effect on xhigh reasoning is not measured yet.

Full details of T3: MODEL_CARD_Wren_T3.md. The whole study: paper/paper.md.

Running

Keep the experts and the n-gram table in RAM and everything else on the GPU. Pass the overrides as one comma-separated -ot: with repeated -ot flags only the last one is applied.

llama-server -m <file>.gguf -ngl 99 -ot "exps=CPU,per_layer_token_embd=CPU" -fa on \
  -c 262144 -ctk q8_0 -ctv q8_0 --jinja --reasoning-effort xhigh \
  --temp 1.0 --top-p 0.95 --top-k 20 --cache-ram 2048
  • Q4_K_M and T1: any llama.cpp build with qwen4exp support from October 2026 or later (the MTP head was removed from these files).
  • T3 and T4 need the tiered-experts fork:
    • commit 0fb30d938 or later for CUDA/HIP;
    • daaf72c for flash attention on RDNA2.
    • Vulkan needs --no-op-offload, otherwise expert matmuls on the GPU produce NaN.
  • Reasoning effort: use xhigh (the template default). It is the mode in which Swift was trained to reason concisely.
  • Memory: keep --cache-ram small on 64 GB machines. A large prompt cache on top of a 60 GB model pushes the system into swap.

Measured on an RX 6700 XT 12 GB + i9-14900KS + 61 GB RAM (T3, ROCm):

  • prefill 234 tokens/s;
  • generation about 11 tokens/s at steady state;
  • the native 262,144-token context fits in 10.1 GB of VRAM with q8_0 KV cache.

Evaluation

Metric How it is measured Noise
Quiz 300 multiple-choice questions, 6 domains × 50 (code, commonsense, logic, math, philosophy, science); chat template, thinking off; answer = the letter with the highest next-token log-probability ±2.6 points
KLD llama-perplexity --kl-divergence over 25 × 2048 tokens of calibration text, against the Q8_0 of the same pruned model; it measures quantisation only, not pruning —

Per-domain results are in quiz_gguf.jsonl, the KLD logs in kld_*.log, the allocations in alloc.json, alloc3.log, tiers3.json, types3.json, cfg_T4*.json, tiers_T4*.json, types_T4*.json, the scans in sens.jsonl and expert_err.json, and the regression data in t4.jsonl and t4_cfgs.jsonl.

These KLD values are relative to this pruned model's own Q8_0. They are not comparable with KLD figures measured against the unpruned Swift BF16.

License

The licenses of the base models apply:

  • Swift Open License v1.0 for the UkisAI contribution, see LICENSE;
  • Qwen Community License 1.0 for Qwen3.8-Flash-Next, see LICENSE-QWEN.

Attribution and the list of changes are in NOTICE. Under the Swift Open License, commercial use by organizations with more than US$1,000,000 in gross annual revenue needs a separate license from UkisAI; the Qwen license has its own thresholds. See the license texts.

Downloads last month
2,165
GGUF
Model size
117B params
Architecture
qwen4exp
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ohmysimo/Wren-GGUF

Finetuned
ohmysimo/Wren
Quantized
(1)
this model