SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video Paper • 2609.37969 • Published 1 day ago • 18
Learning Beyond What You Sample: Off-Policy-Aware Cross-Model Trajectory Exchange for RLVR Paper • 2609.37868 • Published 1 day ago • 27
Recursive Harness Distillation across Agents for Robot Manipulation Paper • 2609.33378 • Published 4 days ago • 35
Surprising Success, Repeated Failure: Entropy-Guided Credit Assignment for Exploration in LLM Reasoning Paper • 2609.33781 • Published 4 days ago • 30
Knowing When Thinking Is Not Enough: Teaching Small Reasoning Models to Reason Beyond Their Parametric Knowledge Paper • 2609.34327 • Published 3 days ago • 34
Agora: Git as Shared Memory for Collective AutoResearch Paper • 2609.18094 • Published 15 days ago • 58
ProgramDistill: From Interactive Web Apps to Verifiable Reference-Guided SWE Tasks Paper • 2609.18805 • Published 15 days ago • 67
ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments Paper • 2609.19134 • Published 15 days ago • 101
Discovery Foundation Models: Toward Open-Ended Discovery Intelligence Paper • 2609.15973 • Published 17 days ago • 33
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 17 days ago • 250
An Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics Paper • 2609.10712 • Published 22 days ago • 45
Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation Paper • 2609.08798 • Published 23 days ago • 84
τ^τ-Bench: An Environment for End-To-End, Realistic Agent Construction Paper • 2609.04611 • Published 27 days ago • 11
To See a World in a Living Context: Unified Indoor-Outdoor Urban World Generation Paper • 2608.05879 • Published Aug 29 • 9
Dr. Claw: An AI Scientist Workspace for Vibe Research Paper • 2609.00365 • Published about 1 month ago • 175