PolicyShiftGuard: Benchmarking and Improving Policy-Adaptive Image Guardrails Paper • 2607.05910 • Published 25 days ago • 38
WorldDirector: Building Controllable World Simulators with Persistent Dynamic Memory Paper • 2607.02517 • Published about 1 month ago • 33
EvoPolicyGym: Evaluating Autonomous Policy Evolution in Interactive Environments Paper • 2607.02440 • Published about 1 month ago • 51
MaxProof: Scaling Mathematical Proof with Generative-Verifier RL and Population-Level Test-Time Scaling Paper • 2606.13473 • Published Jun 11 • 94
ComBench: A Benchmark for Rigorous Proof Reasoning and Constructive Realization in Olympiad-Level Combinatorics Paper • 2606.10479 • Published Jun 9 • 20
SubtleMemory: A Benchmark for Fine-Grained Relational Memory Discrimination in Long-Horizon AI Agents Paper • 2606.05761 • Published Jun 4 • 19
SceneCode: Executable World Programs for Editable Indoor Scenes with Articulated Objects Paper • 2605.19587 • Published May 19 • 10