TMRL: Diffusion Timestep-Modulated Pretraining Enables Exploration for Efficient Policy Finetuning Paper • 2605.12236 • Published May 12
OGPO: Sample Efficient Full-Finetuning of Generative Control Policies Paper • 2605.03065 • Published May 4 • 1
Robometer: Scaling General-Purpose Robotic Reward Models via Trajectory Comparisons Paper • 2603.02115 • Published Mar 2