Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening Paper • 2609.18708 • Published 24 days ago • 84
T1 Collection CC BY 4.0 RL checkpoints for T1; base-model and third-party terms still apply. • 4 items • Updated 23 days ago • 5
T1: Terminal Agent Reinforcement Learning for Long-Horizon Tasks Paper • 2609.11042 • Published 30 days ago • 64
ADWIN: Adaptive Windows for Horizon-Aware On-Policy Distillation Paper • 2605.28396 • Published May 27 • 2
Achieving Gold-Medal-Level Olympiad Reasoning via Simple and Unified Scaling Paper • 2605.13301 • Published May 13 • 166
Claw-Eval-Live: A Live Agent Benchmark for Evolving Real-World Workflows Paper • 2604.28139 • Published Apr 30 • 42
TEMPO: Scaling Test-time Training for Large Reasoning Models Paper • 2604.19295 • Published Apr 21 • 37
V-Bridge: Bridging Video Generative Priors to Versatile Few-shot Image Restoration Paper • 2603.13089 • Published Mar 13 • 13
P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics Olympiads Paper • 2602.09443 • Published Feb 10 • 59
P1: Mastering Physics Olympiads with Reinforcement Learning Paper • 2511.13612 • Published Nov 17, 2025 • 135
Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning Paper • 2506.01939 • Published Jun 2, 2025 • 190
ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models Paper • 2505.24864 • Published May 30, 2025 • 146
Sherlock: Self-Correcting Reasoning in Vision-Language Models Paper • 2505.22651 • Published May 28, 2025 • 25
The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models Paper • 2505.22617 • Published May 28, 2025 • 132
The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning Paper • 2505.15134 • Published May 21, 2025 • 7