Wren logo

Wren

Wren is a compact derivative of Swift 1.5 Qwen3.8-Flash-Next by UkisAI, which is itself a derivative of Qwen3.8-Flash-Next. This repository holds the bf16 weights with 50% of the routed experts removed and the routers and norms healed. "Swift" and "UkisAI" are trademarks of UkisAI and are used here only to describe the model's origin.

  • 256 of 512 routed experts are kept per layer, chosen by maximin over REAP saliency measured on 13 areas, including agentic data. The MTP head was removed.
  • Routers and norms were healed by logit distillation from Swift's top-20 log-probs (reasoning effort xhigh), on 3.6M tokens.
  • Evaluation, on the same 300 questions: all-domain quiz 83.3 vs 88.0 for intact Swift; synthetic logic 70.0 vs 74.7.
  • Training logs, masks and evaluations are in training/.

Quantised versions (GGUF, including the per-expert mixed-precision Wren_T3 for a 12 GB GPU) and the research paper that documents every phase are in ohmychemo/Wren-GGUF.

License

The licenses of the base models apply:

  • Swift Open License v1.0 for the UkisAI contribution, see LICENSE;
  • Qwen Community License 1.0 for Qwen3.8-Flash-Next, see LICENSE-QWEN.

Attribution and the list of changes are in NOTICE. Commercial use above the thresholds stated in the licenses needs a separate license.

Downloads last month
48
Safetensors
Model size
117B params
Tensor type
BF16
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ohmysimo/Wren

Finetuned
(1)
this model
Quantizations
1 model