Selecting Diverse SFT Traces Improves Post-RL Generalization Paper • 2609.33780 • Published 3 days ago • 15
Policy Regularized Distributionally Robust Markov Decision Processes with Linear Function Approximation Paper • 2510.14246 • Published Oct 16, 2025 • 1
TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning Paper • 2510.06217 • Published Oct 7, 2025 • 67