$β$-OPSD: Deriving with Policy Optimization, Training with Self-Distillation Paper • 2607.28582 • Published 4 days ago • 21
Scheduling Thoughts: Learning the Order of Thought in Diffusion Language Models Paper • 2606.23567 • Published Jun 22
Hold Onto That Thought: Assessing KV Cache Compression On Reasoning Paper • 2512.12008 • Published Dec 12, 2025 • 1
MORSE-500: A Programmatically Controllable Video Benchmark to Stress-Test Multimodal Reasoning Paper • 2506.05523 • Published Jun 5, 2025 • 34
β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation Paper • 2607.28582 • Published 4 days ago • 21
β-OPSD: Deriving with Policy Optimization, Training with Self-Distillation Paper • 2607.28582 • Published 4 days ago • 21
SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones? Paper • 2605.30329 • Published May 28 • 8
SoundnessBench: Can Your AI Scientist Really Tell Good Research Ideas from Bad Ones? Paper • 2605.30329 • Published May 28 • 8
ClawEnvKit: Automatic Environment Generation for Claw-Like Agents Paper • 2604.18543 • Published Apr 20 • 30
ROVER: Benchmarking Reciprocal Cross-Modal Reasoning for Omnimodal Generation Paper • 2511.01163 • Published Nov 3, 2025 • 32
LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Paper • 2509.00676 • Published Aug 31, 2025 • 85