Instructions to use ohmysimo/Wren-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use ohmysimo/Wren-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf ohmysimo/Wren-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ohmysimo/Wren-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf ohmysimo/Wren-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf ohmysimo/Wren-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf ohmysimo/Wren-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf ohmysimo/Wren-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf ohmysimo/Wren-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf ohmysimo/Wren-GGUF:Q4_K_M
Use Docker
docker model run hf.co/ohmysimo/Wren-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use ohmysimo/Wren-GGUF with Ollama:
ollama run hf.co/ohmysimo/Wren-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use ohmysimo/Wren-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ohmysimo/Wren-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "ohmysimo/Wren-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use ohmysimo/Wren-GGUF with Docker Model Runner:
docker model run hf.co/ohmysimo/Wren-GGUF:Q4_K_M
- Lemonade
How to use ohmysimo/Wren-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull ohmysimo/Wren-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.Wren-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use ohmysimo/Wren-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ohmysimo/Wren-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default ohmysimo/Wren-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use ohmysimo/Wren-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf ohmysimo/Wren-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "ohmysimo/Wren-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Wren - GGUF
Wren is a compact derivative of Swift 1.5 Qwen3.8-Flash-Next by UkisAI, which is itself a derivative of Qwen3.8-Flash-Next. Like the wren, one of the smallest birds, it is built to be small and agile: half of the experts, measured mixed precision, made for a consumer GPU. "Swift" and "UkisAI" are trademarks of UkisAI and are used here only to describe the model's origin.
GGUF quantisations of ohmychemo/Wren.
- The base is Swift1.5-Qwen3.8-Flash-Next with 50% of the routed experts removed: 256 of 512 per layer, chosen by maximin over REAP saliency on 13 calibration areas, including agentic ones.
- Routers and norms were healed by distillation from Swift's own logits.
- All files use an imatrix computed from Swift's own outputs.
- They are built to run on a consumer GPU, with the experts and the n-gram table in RAM.
- In order to use Wren models you need to install this llama.cpp fork.
The research behind these files
📄 Paper: paper/paper.md, "Wren_T3: measured-sensitivity, per-expert mixed precision for a pruned MoE on a 12 GB consumer GPU" (about 14k words).
- It documents every phase, with motivations, all measurements and the negative results.
- It was written by T3 itself, running locally on the RX 6700 XT through Claude Code, from the project record in paper/context/. That record includes the full research summary, the local measurements and the raw data the paper cites.
- It was then fact-checked against that record and corrected.
Goal: "intelligence density". Derive from a strong open MoE a model that runs on one consumer machine (12 GB VRAM, 64 GB RAM), loses as little measured intelligence as possible and keeps Swift's short reasoning. Every choice is driven by measurements on a multidisciplinary quiz, not by heuristics.
| Phase | What we did | Key result |
|---|---|---|
| 1. Measurement | Swift run in bf16 one layer at a time (exact to 5e-8). Expert frequency, router weight and REAP saliency on 13 areas: 8 topical + 5 agentic (terminal, code agent, tool use, search, office). Pruning simulated by masking experts in the router | Maximin (protect the worst-served area) at 50%: quiz 84.0, the best criterion; it keeps 77–82% of REAP mass in every area |
| 2. Pruning + healing | Experts removed physically in stages (100 → 75 → 60 → 50%). Routers and norms healed on Swift's own top-20 logits (3.58M tokens) | Quiz 84.0 → 83.3, KL 0.338 → 0.191 (intact Swift 88.0). 35% failed: dropping 74.3 after healing; merging experts 67.0, worse than dropping. Training more parameters made KL worse |
| 3. Quantisation | Fake-RTN sweep, then GGUF variants with imatrix | Breaking point between 3 and 2 bits. Down projections are the most sensitive tensors and, with 640 columns, admit only block-32 types |
| 4. Measured mixed precision | 48-layer sensitivity scan (quiz-KL, in-place byte patching). Quiz-weighted per-expert error. A llama.cpp fork with two expert groups per layer: the routing is unchanged, and the CUDA bug it exposed is fixed. Knapsack/Lagrangian allocations (T1, T3). Fast proxy pipeline and a 60-configuration random design with ridge regression (T4) | The middle layers (11–29) are the most sensitive, the deep ones (30–46) the most robust, contrary to the "protect the first/last layers" heuristic. Within a layer, the measured expert error ranks experts twice as well as REAP. Damage is not additive. Regression R² 0.88 on KLD |
| 5. Local deployment | T3 on an RX 6700 XT: ROCm (with a flash-attention fix for RDNA2) and Vulkan | Same quiz across backends within noise. About 11 tokens/s generation, 234 tokens/s prefill, 262k context in 10.1 GB VRAM |
Main findings.
- Measured sensitivity beats position heuristics.
- Optimising quiz answers and optimising the next-token distribution give different models. T3 keeps math at 72; the KLD-optimised T4 drops it to 62 at the same size.
- IQ2_XXS is too destructive to be a useful tier. Uniform iq3 gate/up beats every iq2 mix in KLD, for about 0.7 GB more.
- Fast proxies rank configurations correctly but compress the gaps by about 2.5×.
The negative results are published together with the files that worked.
Files
| File | Size | KLD vs Q8_0 | Same top-1 | Quiz (300 q) | Runs on |
|---|---|---|---|---|---|
Wren_Q4_K_M.gguf |
77.6 GB | 0.0185 | 95.6% | 83.3 | upstream llama.cpp |
Wren_T1.gguf |
62.7 GB | 0.0499 | 92.8% | 81.0 | upstream llama.cpp |
Wren_T3.gguf |
60.4 GB | 0.0788 | 90.9% | 82.3 | tiered-experts fork |
Wren_T3-v2.gguf |
60.4 GB | — ¹ | — ¹ | 83.0 ² (T3: 80.7 ²) | tiered-experts fork |
Wren_T4-2.92.gguf |
60.4 GB | 0.0745 | 91.5% | 80.7 | tiered-experts fork |
Wren_T4-2.6.gguf |
58.7 GB | 0.0973 | 90.2% | 80.0 | tiered-experts fork |
¹ T3-v2 is re-healed toward intact Swift, so it is meant to differ from the pruned model's Q8_0: KLD against that Q8_0 does not measure it. ² Measured on ROCm (RX 6700 XT); the other quiz values in this table were measured with CUDA on an RTX 4090. On ROCm the same T3 scores 80.7, so compare 83.0 with 80.7.
Reference values (not published here):
- the Q8_0 of the same pruned model: about 124 GB, quiz 82.0;
- intact Swift: quiz 88.0;
- the healed 50% model in bf16: quiz 83.3.
What the variants are. All files have the attention, linear attention, indexer, shared experts, hyper-connections, token embedding and output at q8_0. They differ only in the routed experts and the n-gram table:
| File | Expert gate/up | Expert down | N-gram table |
|---|---|---|---|
| Q4_K_M | standard Q4_K_M | standard Q4_K_M | q4_1 |
| T1 | per layer, from a 48-layer sensitivity scan measured on the quiz: q4_K on the 14 most sensitive layers, iq2_xxs on layers 44–46, iq3_xxs elsewhere | iq4_nl | q4_0 |
| T3 | per expert: two groups per layer, sized by a Lagrangian allocation over the measured layer sensitivity and a quiz-weighted per-expert error. 6.5% of experts at q4_K, 69.8% at iq3_xxs, 23.7% at iq2_xxs | iq4_nl | q4_0 |
| T4-2.92 / T4-2.6 | per expert, fitted by ridge regression on 64 quickly evaluated configurations (2.92 and 2.6 bpw on gate/up) | iq4_nl | q4_0 |
Which one to use:
- Most faithful: Q4_K_M. It is effectively lossless vs Q8_0.
- About 60 GB, best task answers: T3-v2 (same size and speed as T3; see below).
- T3 is the same file before the re-healing.
- About 60 GB, closest next-token distribution: T1 (62.7 GB, no fork needed) or T4-2.92.
- T4-2.6 is the smallest; it is mainly a data point.
Wren_T3-v2: re-healed on the compressed model
Same routed experts, n-gram table and quant types as T3; only the routers, the norms, the shared experts and the attention / linear-attention output projections (Q8_0) changed:
- What was trained: routers (F32), all norms, and rank-16 LoRA on the shared experts and the attention / linear-attention output projections. The LoRA is merged into the Q8_0 tensors.
- On what: the compressed T3 itself, in PyTorch. The experts stay as GGUF bytes and are decoded by Triton kernels. Activations and the KV cache are quantised as llama.cpp does at inference, so the training sees the model that actually runs.
- Teacher: intact Swift's top-20 log-probs. The first stage (LoRA only, 6M tokens) used the original distillation set. The second stage (routers + norms + LoRA, 12.1M tokens, 4×A100 data-parallel) used new Swift xhigh generations for the domains that set lacked:
| Data of the second stage | Docs |
|---|---|
| Math: Swift xhigh solutions + scored problem/solution texts | 1,171 + 5,134 |
| Logic: Swift xhigh answers | 800 |
| Philosophy: Swift xhigh analysis of passages | 400 |
| Science: ARC-Challenge train split, Swift xhigh | 400 |
| Commonsense: CommonsenseQA train split, Swift xhigh | 400 |
| General share (30% of tokens) | original set |
None of these sources is the quiz's split (the quiz uses ARC-Challenge test, HellaSwag and MMLU).
Held-out KL to intact Swift (second stage, before → after):
| Commonsense | Logic | Science | Philosophy | Math | General | All |
|---|---|---|---|---|---|---|
| 0.199 → 0.136 | 0.154 → 0.114 | 0.195 → 0.150 | 0.262 → 0.220 | 0.130 → 0.114 | 0.192 → 0.172 | 0.158 → 0.134 |
llama.cpp, same machine and backend (ROCm):
| Perplexity (8 × 512) | Quiz | Code | Commonsense | Logic | Math | Philosophy | Science | |
|---|---|---|---|---|---|---|---|---|
| T3 | 1.823 | 80.7 | 78 | 92 | 90 | 70 | 74 | 80 |
| T3 + LoRA (stage 1, not published) | 1.764 | 82.3 | 80 | 92 | 92 | 68 | 80 | 82 |
| T3-v2 | 1.698 | 83.0 | 82 | 92 | 90 | 68 | 82 | 84 |
Against T3, 12 answers turn right and 5 turn wrong. At 300 questions the 2.3-point gain is about one noise width, so the perplexity and KL drops are the stronger evidence. The quiz answers without thinking, while the new math data are reasoned xhigh solutions: on a 100-question held-out math set (no thinking) the score went 59 → 65, and the effect on xhigh reasoning is not measured yet.
Full details of T3: MODEL_CARD_Wren_T3.md. The whole study: paper/paper.md.
Running
Keep the experts and the n-gram table in RAM and everything else on the GPU. Pass the overrides as one comma-separated -ot: with repeated -ot flags only the last one is applied.
llama-server -m <file>.gguf -ngl 99 -ot "exps=CPU,per_layer_token_embd=CPU" -fa on \
-c 262144 -ctk q8_0 -ctv q8_0 --jinja --reasoning-effort xhigh \
--temp 1.0 --top-p 0.95 --top-k 20 --cache-ram 2048
- Q4_K_M and T1: any llama.cpp build with qwen4exp support from October 2026 or later (the MTP head was removed from these files).
- T3 and T4 need the tiered-experts fork:
- commit
0fb30d938or later for CUDA/HIP; daaf72cfor flash attention on RDNA2.- Vulkan needs
--no-op-offload, otherwise expert matmuls on the GPU produce NaN.
- commit
- Reasoning effort: use xhigh (the template default). It is the mode in which Swift was trained to reason concisely.
- Memory: keep
--cache-ramsmall on 64 GB machines. A large prompt cache on top of a 60 GB model pushes the system into swap.
Measured on an RX 6700 XT 12 GB + i9-14900KS + 61 GB RAM (T3, ROCm):
- prefill 234 tokens/s;
- generation about 11 tokens/s at steady state;
- the native 262,144-token context fits in 10.1 GB of VRAM with q8_0 KV cache.
Evaluation
| Metric | How it is measured | Noise |
|---|---|---|
| Quiz | 300 multiple-choice questions, 6 domains × 50 (code, commonsense, logic, math, philosophy, science); chat template, thinking off; answer = the letter with the highest next-token log-probability | ±2.6 points |
| KLD | llama-perplexity --kl-divergence over 25 × 2048 tokens of calibration text, against the Q8_0 of the same pruned model; it measures quantisation only, not pruning |
— |
Per-domain results are in quiz_gguf.jsonl, the KLD logs in kld_*.log, the allocations in alloc.json, alloc3.log, tiers3.json, types3.json, cfg_T4*.json, tiers_T4*.json, types_T4*.json, the scans in sens.jsonl and expert_err.json, and the regression data in t4.jsonl and t4_cfgs.jsonl.
These KLD values are relative to this pruned model's own Q8_0. They are not comparable with KLD figures measured against the unpruned Swift BF16.
License
The licenses of the base models apply:
- Swift Open License v1.0 for the UkisAI contribution, see LICENSE;
- Qwen Community License 1.0 for Qwen3.8-Flash-Next, see LICENSE-QWEN.
Attribution and the list of changes are in NOTICE. Under the Swift Open License, commercial use by organizations with more than US$1,000,000 in gross annual revenue needs a separate license from UkisAI; the Qwen license has its own thresholds. See the license texts.
- Downloads last month
- 2,165
4-bit
Model tree for ohmysimo/Wren-GGUF
Base model
Qwen/Qwen3.8-Flash-Next