The Mirage of Optimizing Training Policies: Monotonic Inference Policies as the Real Objective for LLM Reinforcement Learning Paper • 2606.29526 • Published Jun 28 • 170
Hallucination in World Models is Predictable and Preventable Paper • 2606.27326 • Published Jun 25 • 9
Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models Paper • 2606.19750 • Published Jun 18 • 3
Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models Paper • 2606.19750 • Published Jun 18 • 3 • 2
Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models Paper • 2606.19750 • Published Jun 18 • 3