Post
1441
𧬠Darwin-180B-RSI β an AI that learns from itself and knows when it's right
π FINAL-Bench/Darwin-180B-RSI
𧬠Darwin β crossbreed and evolve the parent
Darwin diagnoses strong parent models like an MRI, inherits only their best parts, and evolves the weak spots β producing a child stronger than its parents.
Father model: Qwen3.8-Flash-Next (180B MoE).
π§ Rewired paths
πΉ 12 full-attention layers Β· πΉ 36 linear-attention layers Β· πΉ 48 shared-expert layers β precision-strengthened
π 512 routed experts Β· router Β· vision encoder β untouched
β Only 0.02% of the weights changed.
π RSI Γ ποΈ ZTC
RSI (recursive self-improvement): solve β verify against real answers β learn only the correct reasoning β repeat.
ZTC (Zero-Token Confidence): reads the model's internal state once, before answering, and returns the probability the answer is right β zero extra tokens. Returns answer + confidence as JSON.
{"answer": "...", "confidence": 0.97, "truncated": false}
β¨ Synergy: ZTC finds where the model wavers β RSI learns exactly there β confidence gets sharper. Low confidence = stop, so agents don't act on wrong answers.
β‘ Same accuracy, 11% shorter reasoning β faster and cheaper.
π https://arxiv.org/abs/2605.14386
π€ FINAL-Bench/Darwin-180B-RSI
ποΈ https://huggingface.co/collections/FINAL-Bench/ztc-models-jev-ecosystems
π The result β #1 on five Hugging Face official leaderboards
π₯ AIME 2026 100% (first perfect score on the board)
π₯ HMMT Feb 2026 100% (first perfect score on the board)
π₯ GPQA Diamond 94.44%
π₯ MMLU-Pro 88.12%
π₯ MMMU-Pro 79.48%
π 131K-token thinking budget Β· bf16 Β· samples per benchmark listed on the model card. π
#Darwin #RSI #ZTC #AIME #HMMT #GPQA #MMLUPro #MMMUPro #OpenSource
π FINAL-Bench/Darwin-180B-RSI
𧬠Darwin β crossbreed and evolve the parent
Darwin diagnoses strong parent models like an MRI, inherits only their best parts, and evolves the weak spots β producing a child stronger than its parents.
Father model: Qwen3.8-Flash-Next (180B MoE).
π§ Rewired paths
πΉ 12 full-attention layers Β· πΉ 36 linear-attention layers Β· πΉ 48 shared-expert layers β precision-strengthened
π 512 routed experts Β· router Β· vision encoder β untouched
β Only 0.02% of the weights changed.
π RSI Γ ποΈ ZTC
RSI (recursive self-improvement): solve β verify against real answers β learn only the correct reasoning β repeat.
ZTC (Zero-Token Confidence): reads the model's internal state once, before answering, and returns the probability the answer is right β zero extra tokens. Returns answer + confidence as JSON.
{"answer": "...", "confidence": 0.97, "truncated": false}
β¨ Synergy: ZTC finds where the model wavers β RSI learns exactly there β confidence gets sharper. Low confidence = stop, so agents don't act on wrong answers.
β‘ Same accuracy, 11% shorter reasoning β faster and cheaper.
π https://arxiv.org/abs/2605.14386
π€ FINAL-Bench/Darwin-180B-RSI
ποΈ https://huggingface.co/collections/FINAL-Bench/ztc-models-jev-ecosystems
π The result β #1 on five Hugging Face official leaderboards
π₯ AIME 2026 100% (first perfect score on the board)
π₯ HMMT Feb 2026 100% (first perfect score on the board)
π₯ GPQA Diamond 94.44%
π₯ MMLU-Pro 88.12%
π₯ MMMU-Pro 79.48%
π 131K-token thinking budget Β· bf16 Β· samples per benchmark listed on the model card. π
#Darwin #RSI #ZTC #AIME #HMMT #GPQA #MMLUPro #MMMUPro #OpenSource