The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170
Hallucination in World Models is Predictable and Preventable Paper • 2606.27326 • Published Jun 25 • 9