SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation Paper • 2609.36601 • Published 1 day ago • 24
TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs Paper • 2609.33589 • Published 3 days ago • 4
SEED: Self-Evolving On-Policy Distillation for Agentic Reinforcement Learning Paper • 2607.14777 • Published Jul 16 • 105
OPID: On-Policy Skill Distillation for Agentic Reinforcement Learning Paper • 2606.26790 • Published Jun 25 • 59
Maestro: Reinforcement Learning to Orchestrate Hierarchical Model-Skill Ensembles Paper • 2605.22177 • Published May 21 • 18
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Paper • 2605.00380 • Published May 1 • 7