Darwin-180B-RSI-R3

The second round of model-level self-improvement on top of Darwin-180B-RSI.

Darwin-180B-RSI (R1) holds first place on seven Hugging Face official leaderboards. R3 continues training from R1 with the same recipe: the model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces are used.

R3 is released so anyone can download it, run it, and check the numbers below.

What changed from R1

  • Starting point: R1 weights (not the parent). R3 is a true second round.
  • Practice problems: 3,000 SuperGPQA questions (middle and hard difficulty) that were never used in R1 training. R1 solved each one 8 times.
  • What it learned from: only the "boundary" problems, where R1 was right on 2 to 6 of 8 attempts (462 problems). From those, up to 2 of R1's own correct, untruncated solutions per problem (714 solutions in total).
  • What was trained: the same components as R1 (attention paths and shared experts) via LoRA, then merged. All 512 routed experts, the router and the vision encoder are unchanged.
  • No benchmark data: GPQA Diamond and the held-out set below were never used for training or selection.

Results (same settings for every model)

Held-out SuperGPQA, 1,000 questions never used in training or selection

4 samples per question, 16K thinking budget, temperature 1.0.

Model Single sample Mean of 4 Majority of 4
R1 (Darwin-180B-RSI) 65.30 65.67 68.30
R3 (this model) 66.30 66.70 69.00

Paired per-question difference, R1 → R3 (mean of 4): +1.03 points, 95% CI [+0.05, +2.00]. The gain is small but statistically significant.

GPQA Diamond, 198 questions

8 samples per question, 32K thinking budget, temperature 1.0.

Model Single sample Mean of 8 Majority of 8
R0 (parent, Qwen3.8-Flash-Next) 84.85 85.35 90.91
R1 (Darwin-180B-RSI) 84.85 85.80 89.90
R3 (this model) 85.86 86.05 90.40

Paired differences on GPQA (mean of 8): R0 → R3 +0.69 [−0.71, +2.10], R1 → R3 +0.25 [−1.18, +1.69]. With 198 questions these are within noise; we report them as measured.

How to read this: on a large held-out set, the second round of self-improvement produced a measurable gain over R1. On GPQA the model was already near its ceiling and the differences are not significant.

Leaderboard status

The seven Hugging Face leaderboard #1 results belong to R1 (Darwin-180B-RSI). R3 has not been submitted to any leaderboard. We are asking independent evaluators to measure it directly.

Quickstart

Serving is the same as Darwin-180B-RSI (vLLM, tensor parallel + expert parallel). This is a reasoning model; give it a long generation budget.

vllm serve FINAL-Bench/Darwin-180B-RSI-R3 \
  --tensor-parallel-size 8 --enable-expert-parallel \
  --max-model-len 139264

Recommended sampling: temperature 1.0, top_p 0.95, top_k 20, up to 131,072 generated tokens.

ZTC

The ZTC probe published with Darwin-180B-RSI was fitted on R1's hidden states. A probe fitted on R3 will be added to this repository; until then, use R1's probe only as a rough signal.

Model-level RSI vs. harness-level RSI

Darwin-180B-RSI-R3 is Model-level RSI: the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically. Harness-level RSI (e.g., Google's RRSI) improves the prompts, tools and workflow around a fixed model. The two are complementary.

License

Qwen Community License 1.0 (inherited from Qwen3.8-Flash-Next). See LICENSE.

About

Built by VIDRAFT. Darwin family paper: arXiv 2605.14386.

Downloads last month
-
Safetensors
Model size
180B params
Tensor type
BF16
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FINAL-Bench/Darwin-180B-RSI-R3

Finetuned
(1)
this model

Paper for FINAL-Bench/Darwin-180B-RSI-R3