False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents Paper • 2609.39102 • Published 8 days ago • 672
On-Policy or Off-Policy Learning? A Systematic Study of Distillation Dynamics Paper • 2609.35259 • Published 10 days ago • 197
Scaling Trajectories for Complex Tasks through Recursive Self-Rewrite Paper • 2610.02826 • Published 6 days ago • 96
Groupwise Agentic Grading and Advantage Redistribution for Code Agent RL Paper • 2609.32577 • Published 12 days ago • 130
Raven: The Harness of Harnesses for Composable Agentic Intelligence Paper • 2609.33439 • Published 11 days ago • 567
SpeakerMem-R1: Speaker-Centered Dual-Track Memory for Multi-Party Dialogue Paper • 2609.26780 • Published 16 days ago • 103
1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation Paper • 2609.24432 • Published 17 days ago • 16
GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay Paper • 2609.25001 • Published 17 days ago • 132
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses Paper • 2609.24972 • Published 17 days ago • 223
Grounded Skill Synthesis from Code at Scale for Agentic Intelligence Paper • 2609.05571 • Published Sep 4 • 119
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 20 days ago • 138
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents Paper • 2609.22000 • Published 20 days ago • 79