Adaptive Reward Routing: Dynamic Multi-Reward Optimization for Joint Audio-Video Diffusion via Forward-Process RL Paper • 2609.37200 • Published 3 days ago • 81
Post-Training Leaves Behavioral Shadows on Unrelated Decisions Paper • 2609.29233 • Published 8 days ago • 271
WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation Paper • 2608.24479 • Published Aug 25 • 139
Benchmark Radar: A Living Database and Search Engine for AI Benchmarks and Evaluation Paper • 2609.11115 • Published 22 days ago • 173
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper • 2609.00111 • Published Aug 31 • 316
Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models Paper • 2608.25518 • Published Aug 26 • 60
SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification Paper • 2608.08786 • Published Aug 9 • 6