What Makes World Action Models Generalize? An Empirical Study of Test-Time Future Modeling Paper • 2609.34981 • Published 1 day ago • 61
MaxRL Collection Qwen3-Base post-trained checkpoints for our paper, Maximum Likelihood Reinforcement Learning [https://zanette-labs.github.io/MaxRL/] • 5 items • Updated Jul 27 • 3
Co-GRPO: Co-Optimized Group Relative Policy Optimization for Masked Diffusion Model Paper • 2512.22288 • Published Dec 25, 2025 • 3