Does On-Policy Distillation Really Distill? From Noisy Teacher to Self-Improvement Paper • 2608.31046 • Published Aug 31 • 97
Agentic ESOpt: Fine-Tuning Long-Horizon LLM Agents with Minimal GPU Requirements Paper • 2608.17310 • Published Aug 18 • 110
Stabilizing Reinforcement Learning with LLMs: Formulation and Practices Paper • 2512.01374 • Published Dec 1, 2025 • 109