DivOPD: Spread Wide, Look Close for Asynchronous On-Policy Distillation of Multi-turn Agents Paper • 2609.34838 • Published 11 days ago • 1
RRSI: Regularized Recursive Self-Improvement of Agent Harnesses Paper • 2609.24972 • Published 18 days ago • 223
BiPACE: Bisimulation-Guided Policy Optimization with Action Counterfactual Estimation for LLM Agents Paper • 2606.25556 • Published Jun 24 • 1
The Detection--Extraction Gap: Models Know the Answer Before They Can Say It Paper • 2604.06613 • Published Apr 8 • 2
SkillRL: Evolving Agents via Recursive Skill-Augmented Reinforcement Learning Paper • 2602.08234 • Published Feb 9 • 76