Darwin-180B-RSI

180B Mixture-of-Experts · vision-language · GPQA Diamond 94.44 % — #1 on the Hugging Face leaderboard · self-improving

reasoning · MoE 512 experts · 262K long context · image + text · Korean + English · self-improvement · ZTC

The newest flagship of the Darwin family — #1 on GPQA Diamond, and a model that gets better by learning from its own verified work.


🧬 The Darwin Family

Darwin is VIDRAFT's measurement-driven reasoning model family — 50+ official models, 400+ community derivatives, and now two places in the GPQA Diamond top 3 (Darwin-180B-RSI #1 · Darwin-397B-ZTC #3).


🧬 Darwin — evolve the parent, keep what works

Darwin treats a strong open model as a parent. It measures where the parent is weak, and strengthens exactly those parts — instead of re-training everything and risking what already works.

  • Diagnose before you change. Every Darwin generation starts from a measured weakness map of the parent.
  • Change little, precisely. Darwin modifies a small, targeted fraction of the network. Knowledge stored in the experts is preserved.
  • Proven capability over new guesses. Earlier Darwin generations grafted the best-performing expert/FFN blocks from other strong models onto a base backbone; Darwin-180B-RSI adds a new ingredient — the model's own verified work.
  • Measured, not claimed. Every change must beat the parent on held-out tests before it ships.
Model Scale GPQA Diamond
Darwin-9B-NEG 9B 84.3
Darwin-27B-Opus 27B dense 86.9
Darwin-36B-Opus 36B MoE 88.4
Darwin-28B-REASON 28B + DELPHI 89.39
Darwin-397B-ZTC 397B MoE (FP8) 93.43
Darwin-180B-RSI 180B MoE 94.44

Lineage

Role
Parent Qwen/Qwen3.8-Flash-Next 180B MoE vision-language backbone · Qwen Community License 1.0
Darwin RSI self-improvement on verified answers the parent's own solutions, checked against verifiable answer keys, fed back as training signal
Preserved 512 routed experts · router · vision encoder untouched — the parent's knowledge stays intact
ZTC zero-token confidence readout see below

📄 Darwin Platform & Research

  • Darwin Family — MRI trust-weighted evolutionary merging for training-free scaling of language-model reasoning (arXiv:2605.14386)
  • Placement Is Free, Composition Is Not — the Latin square as a provably-balanced construction for heterogeneous sequence-mixer stacks (2609.20269) — the AETHER architecture line
  • FINAL Bench — VIDRAFT's measurement-driven evaluation framework (SSRN)
  • Four-layer Pre-AGI roadmap — Darwin → AETHER → PROMETHEUS → HEPHAESTUS
  • Collections: Darwin Family · ZTC Models — JEV ecosystems

🔁 RSI — a model that improves from its own work

Recursive self-improvement (RSI) is the core of this generation. Instead of distilling a bigger teacher, the model improves by learning from itself:

  1. Solve — the model works through practice problems it has never seen in evaluation.
  2. Verify — its answers are checked against verifiable references (answer keys, executable checks). Nothing unverified is learned.
  3. Learn — it is re-trained on the reasoning that turned out to be correct.
  4. Repeat — the improved model becomes the next solver.

What it bought in this release:

Parent (Qwen3.8-Flash-Next) Darwin-180B-RSI
Average reasoning length (MMLU-Pro) 4,320 tokens 3,833 tokens (−11 %)
MMLU-Pro accuracy 88.04 % 88.12 %

Same or better accuracy with shorter reasoning — cheaper and faster to serve. Practice sets are deduplicated against every evaluation set we report (8-gram overlap filter).


🏛️ ZTC — it knows before it answers

Zero-Token Confidence (ZTC) reads the model's own internal state once, before generation, and returns the probability that the answer it is about to give is correct — no extra tokens, no second model.

{"answer": "...", "confidence": 0.93, "ztc_score": 1.84, "truncated": false}

Use it to gate actions: when confidence is low, do not call the tool, escalate, or answer "I don't know". The ZTC readout for this model is being fitted and will ship in ztc/ (same format as Darwin-397B-ZTC).


🏆 Results

Benchmark Score Setting Leaderboard
GPQA Diamond (198) 94.44 majority vote over up to 16 samples · 131,072-token thinking budget #1
MMLU-Pro (12,032) 88.12 single sample · 131,072-token thinking budget #1
MMMU-Pro (vision, 1,730) 79.48 majority vote over 3 samples · 131,072-token thinking budget #1

Sampling for all runs: temperature 1.0 · top_p 0.95 · top_k 20 · bf16. All numbers are self-measured and reproducible with the settings above.

MMLU-Pro by category (single sample) — strongest in math 95.0 · biology 94.6 · physics 92.5; room to grow in law and history.


⚙️ Specifications

Architecture Mixture-of-Experts, hybrid attention (36 linear-attention + 12 full-attention layers)
Layers / hidden 48 / 2,560
Experts 512 routed (10 active per token) + shared expert
Context 262,144 tokens
Vocabulary 248,320
Modalities image + text → text
Precision bf16 (~336 GB)

🚀 Quickstart

Serving with vLLM (8 × B200 or equivalent)

vllm serve FINAL-Bench/Darwin-180B-RSI \
  --tensor-parallel-size 8 --enable-expert-parallel \
  --max-model-len 135168 --trust-remote-code

Chat Completions (OpenAI-compatible)

from openai import OpenAI
c = OpenAI(base_url="http://localhost:8000/v1", api_key="-")
r = c.chat.completions.create(model="FINAL-Bench/Darwin-180B-RSI",
    messages=[{"role": "user", "content": "Explain why the sky is blue in two sentences."}],
    temperature=1.0, top_p=0.95, extra_body={"top_k": 20})
print(r.choices[0].message.content)

Transformers

from transformers import AutoProcessor, AutoModelForImageTextToText
model_id = "FINAL-Bench/Darwin-180B-RSI"
processor = AutoProcessor.from_pretrained(model_id)
model = AutoModelForImageTextToText.from_pretrained(model_id, torch_dtype="auto", device_map="auto")

Tip: this is a thinking model. Give it room — a thinking budget of 32K–131K tokens is recommended for hard reasoning. Short budgets truncate the reasoning and cost accuracy.


⚠️ Limitations and disclosure

  • Scores are self-measured with the settings stated in the Results table; majority-vote numbers use several samples per question.
  • Very long reasoning is normal for hard problems; a short thinking budget will truncate answers and lower accuracy.
  • Like every LLM, the model can be confidently wrong — use the ZTC confidence readout to gate high-stakes actions.

🔗 Related Darwin Models


📚 Citation

@misc{darwin180b_rsi_2026,
  title  = {Darwin-180B-RSI: Recursive Self-Improvement on Verified Answers for a 180B Mixture-of-Experts Reasoning Model},
  author = {FINAL-Bench / Darwin Research Team},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/FINAL-Bench/Darwin-180B-RSI}},
  note   = {GPQA Diamond 94.44 \% · MMLU-Pro 88.12 \%}
}

@misc{darwin_family_2026,
  title  = {Darwin Family: MRI-Trust-Weighted Evolutionary Merging for Training-Free Scaling of Language-Model Reasoning},
  author = {Kim, Taebong and Hong, Youngsik and Kim, Minsik and Choi, Sunyoung and Jang, Jaewon and Shin, Junghoon},
  year   = {2026},
  eprint = {2605.14386},
  archivePrefix = {arXiv}
}

@misc{latin_square_2026,
  title  = {Placement Is Free, Composition Is Not: The Latin Square as a Provably-Balanced Construction for Heterogeneous Sequence-Mixer Stacks},
  author = {Kim, Taebong and Hong, Youngsik and Kim, Minsik and Choi, Sunyoung and Jang, Jaewon and Kim, Minseo},
  year   = {2026},
  eprint = {2609.20269},
  archivePrefix = {arXiv}
}

📜 License

Darwin-180B-RSI is a derivative of Qwen3.8-Flash-Next and is distributed under the Qwen Community License 1.0 (see LICENSE).

🏢 About

Built by VIDRAFT · evaluated with FINAL-Bench.

This model is part of the Darwin Family.

Downloads last month
-
Safetensors
Model size
180B params
Tensor type
BF16
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collections including FINAL-Bench/Darwin-180B-RSI

Papers for FINAL-Bench/Darwin-180B-RSI

Evaluation results

  • Accuracy (majority vote, up to 16 samples, 131K thinking) on GPQA Diamond
    self-reported
    94.440
  • Accuracy (single sample, 131K thinking) on MMLU-Pro
    test set self-reported
    88.120