β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation Paper • 2607.28582 • Published 2 days ago • 13
Less is More: Recursive Reasoning with Tiny Networks Paper • 2510.04871 • Published Oct 6, 2025 • 518
IDEA:Enhancing the Rule Learning Ability of Language Agents through Induction, Deduction, and Abduction Paper • 2408.10455 • Published Aug 19, 2024 • 1
Vision-Zero: Scalable VLM Self-Improvement via Strategic Gamified Self-Play Paper • 2509.25541 • Published Sep 29, 2025 • 142