Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Paper • 2608.00782 • Published 11 days ago • 15
The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170