--- license: other license_name: qwen-community-1.0 license_link: LICENSE language: [en, ko, zh, ja, multilingual] library_name: transformers pipeline_tag: image-text-to-text base_model: FINAL-Bench/Darwin-180B-RSI tags: - darwin - darwin-rsi - model-level-rsi - recursive-self-improvement - self-improvement - reasoning - moe - ztc --- # Darwin-180B-RSI-R3 **The second round of model-level self-improvement on top of [Darwin-180B-RSI](https://huggingface.co/FINAL-Bench/Darwin-180B-RSI).** Darwin-180B-RSI (R1) holds first place on seven Hugging Face official leaderboards. R3 continues training from R1 with the same recipe: the model solves verifiable problems, keeps only its own solutions that check out as correct, and trains on them. No human-written solutions or reasoning traces are used. R3 is released so anyone can download it, run it, and check the numbers below. > ๐Ÿ’ป **Run it on your own machine โ€” [POCKET-Darwin-180B-GGUF](https://huggingface.co/FINAL-Bench/POCKET-Darwin-180B-GGUF)**: the 4-bit GGUF of R3 (111 GB) runs on a **laptop with an 8 GB GPU and 32 GB RAM**, **CPU only at 18โ€“21 tok/s**, a 128 GB mini PC or one DGX Spark โ€” MMLU-Pro **identical to BF16 (87.65%)**. ## What changed from R1 - **Starting point:** R1 weights (not the parent). R3 is a true second round. - **Practice problems:** 3,000 SuperGPQA questions (middle and hard difficulty) that were never used in R1 training. R1 solved each one 8 times. - **What it learned from:** only the "boundary" problems, where R1 was right on 2 to 6 of 8 attempts (462 problems). From those, up to 2 of R1's own correct, untruncated solutions per problem (714 solutions in total). - **What was trained:** the same components as R1 (attention paths and shared experts) via LoRA, then merged. All 512 routed experts, the router and the vision encoder are unchanged. - **No benchmark data:** GPQA Diamond and the held-out set below were never used for training or selection. ## Results (same settings for every model) ### Held-out SuperGPQA, 1,000 questions never used in training or selection 4 samples per question, 16K thinking budget, temperature 1.0. | Model | Single sample | Mean of 4 | Majority of 4 | |---|---:|---:|---:| | R1 (Darwin-180B-RSI) | 65.30 | 65.67 | 68.30 | | **R3 (this model)** | **66.30** | **66.70** | **69.00** | Paired per-question difference, R1 โ†’ R3 (mean of 4): **+1.03 points, 95% CI [+0.05, +2.00]**. The gain is small but statistically significant. ### GPQA Diamond, 198 questions 8 samples per question, 32K thinking budget, temperature 1.0. | Model | Single sample | Mean of 8 | Majority of 8 | |---|---:|---:|---:| | R0 (parent, Qwen3.8-Flash-Next) | 84.85 | 85.35 | 90.91 | | R1 (Darwin-180B-RSI) | 84.85 | 85.80 | 89.90 | | **R3 (this model)** | **85.86** | **86.05** | 90.40 | Paired differences on GPQA (mean of 8): R0 โ†’ R3 +0.69 [โˆ’0.71, +2.10], R1 โ†’ R3 +0.25 [โˆ’1.18, +1.69]. With 198 questions these are within noise; we report them as measured. **How to read this:** on a large held-out set, the second round of self-improvement produced a measurable gain over R1. On GPQA the model was already near its ceiling and the differences are not significant. ## Leaderboard status The seven Hugging Face leaderboard #1 results belong to **R1** ([Darwin-180B-RSI](https://huggingface.co/FINAL-Bench/Darwin-180B-RSI)). R3 has not been submitted to any leaderboard. We are asking independent evaluators to measure it directly. ## Quickstart Serving is the same as Darwin-180B-RSI (vLLM, tensor parallel + expert parallel). This is a reasoning model; give it a long generation budget. ```bash vllm serve FINAL-Bench/Darwin-180B-RSI-R3 \ --tensor-parallel-size 8 --enable-expert-parallel \ --max-model-len 139264 ``` Recommended sampling: temperature 1.0, top_p 0.95, top_k 20, up to 131,072 generated tokens. ## ZTC The ZTC probe published with Darwin-180B-RSI was fitted on R1's hidden states. A probe fitted on R3 will be added to this repository; until then, use R1's probe only as a rough signal. ## Model-level RSI vs. harness-level RSI Darwin-180B-RSI-R3 is **Model-level RSI**: the model itself (its weights) improves by learning only from its own solutions. No human-written solutions or reasoning traces are used; correctness is checked automatically. Harness-level RSI (e.g., Google's RRSI) improves the prompts, tools and workflow around a fixed model. The two are complementary. ## License Qwen Community License 1.0 (inherited from Qwen3.8-Flash-Next). See [LICENSE](LICENSE). ## About Built by [VIDRAFT](https://vidraft.net). Darwin family paper: [arXiv 2605.14386](https://arxiv.org/abs/2605.14386).