Beyond Simply Environment Scaling: Designing Effective Environment Distributions for Multimodal Agent Learning Paper • 2608.03571 • Published Aug 6 • 48
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Paper • 2608.03573 • Published Aug 6 • 61
Bottom-up Policy Optimization: Your Language Model Policy Secretly Contains Internal Policies Paper • 2512.19673 • Published Dec 22, 2025 • 66