Post-Training Leaves Behavioral Shadows on Unrelated Decisions Paper • 2609.29233 • Published 7 days ago • 263
WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation Paper • 2608.24479 • Published Aug 25 • 139
ZGCM-1: A Fully Open and Extremely Efficient Foundation Model for Math and Agentic Search Paper • 2609.13356 • Published 20 days ago • 264
OracleZoom: On-Policy Self-Distillation Inspired Reference-Constrained Recursive Image Super Resolution Paper • 2609.06490 • Published 25 days ago • 9