๐งฌ Darwin-180B-RSI โ an AI that learns from itself and knows when it's right ๐ FINAL-Bench/Darwin-180B-RSI
๐งฌ Darwin โ crossbreed and evolve the parent Darwin diagnoses strong parent models like an MRI, inherits only their best parts, and evolves the weak spots โ producing a child stronger than its parents. Father model: Qwen3.8-Flash-Next (180B MoE).
๐ RSI ร ๐๏ธ ZTC RSI (recursive self-improvement): solve โ verify against real answers โ learn only the correct reasoning โ repeat. ZTC (Zero-Token Confidence): reads the model's internal state once, before answering, and returns the probability the answer is right โ zero extra tokens. Returns answer + confidence as JSON. {"answer": "...", "confidence": 0.97, "truncated": false}
โจ Synergy: ZTC finds where the model wavers โ RSI learns exactly there โ confidence gets sharper. Low confidence = stop, so agents don't act on wrong answers. โก Same accuracy, 11% shorter reasoning โ faster and cheaper.
๐ The result โ #1 on five Hugging Face official leaderboards ๐ฅ AIME 2026 100% (first perfect score on the board) ๐ฅ HMMT Feb 2026 100% (first perfect score on the board) ๐ฅ GPQA Diamond 94.44% ๐ฅ MMLU-Pro 88.12% ๐ฅ MMMU-Pro 79.48%
๐ 131K-token thinking budget ยท bf16 ยท samples per benchmark listed on the model card. ๐
Run it on defaults and it takes 244 s. Switch to 3 steps and it's 48.6 s. Add VAE tiling and it's 46.4 s.
The biggest culprit was the default. Z-Image Turbo is distilled to paint in few strokes, but the tool's default is 20. We were throwing away 5ร for no reason. So were we, at first.
3 is the floor. Put 4 and 3 side by side and you cannot tell them apart. At 2 it collapses โ water droplets and wood grain vanish, and the surface turns cloth-like.