RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents Paper • 2609.22000 • Published 18 days ago • 79
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 22 days ago • 251
Rethinking OPD Collection This collection includes the models used in paper "Rethinking On-Policy Distillation of Large Language Models: Phenomenology, Mechanism, and Recipe" • 5 items • Updated Sep 4 • 2
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published Sep 3 • 104
Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published Aug 31 • 97
Safin-1: Safety from Within through Memory-Native State Evolution Paper • 2609.00092 • Published Aug 31 • 21
Intern-S2-Mobius: Foundation Model with Decoupled Knowledge and Reasoning Paper • 2608.14290 • Published Aug 14 • 35
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published Jul 30 • 188
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published Jul 8 • 144
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published Jul 10 • 83
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 167
NatureBench: Can Coding Agents Match the Published SOTA of Nature-Family Papers? Paper • 2606.24530 • Published Jun 23 • 67