Verifier-Induced Support Reshaping in On-Policy Optimization Paper • 2608.00220 • Published Jul 31 • 6
Sparse but Critical: A Token-Level Analysis of Distributional Shifts in RLVR Fine-Tuning of LLMs Paper • 2603.22446 • Published Mar 23 • 13
Verifier-Induced Support Reshaping in On-Policy Optimization Paper • 2608.00220 • Published Jul 31 • 6
Verifier-Induced Support Reshaping in On-Policy Optimization Paper • 2608.00220 • Published Jul 31 • 6
SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs Paper • 2608.03573 • Published Aug 6 • 61
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Paper • 2607.28609 • Published Jul 30 • 75
Deferred Exposure of Future Trajectories for Verifiable Reasoning in Autonomous Driving VLMs Paper • 2608.01755 • Published Aug 3 • 58
Video2GUI: Synthesizing Large-Scale Interaction Trajectories for Generalized GUI Agent Pretraining Paper • 2605.14747 • Published May 14 • 57