Selecting Diverse SFT Traces Improves Post-RL Generalization Paper • 2609.33780 • Published 3 days ago • 15
Dream-RSI: Recursive Self-Improvement through Evolving Worlds Paper • 2609.14858 • Published 16 days ago • 249
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 27 days ago • 114
You Only Need Minimal RLVR Training: Extrapolating LLMs via Rank-1 Trajectories Paper • 2605.21468 • Published May 20 • 51
RubricEM: Meta-RL with Rubric-guided Policy Decomposition beyond Verifiable Rewards Paper • 2605.10899 • Published May 11 • 78
Weak-Driven Learning: How Weak Agents make Strong Agents Stronger Paper • 2602.08222 • Published Feb 9 • 183
Generalizing Test-time Compute-optimal Scaling as an Optimizable Graph Paper • 2511.00086 • Published Oct 29, 2025 • 42
TaTToo: Tool-Grounded Thinking PRM for Test-Time Scaling in Tabular Reasoning Paper • 2510.06217 • Published Oct 7, 2025 • 67
TruthRL: Incentivizing Truthful LLMs via Reinforcement Learning Paper • 2509.25760 • Published Sep 30, 2025 • 56