TGRL: Temperature-Grouped Reinforcement Learning for Efficient Exploration in LLMs Paper • 2609.33589 • Published 6 days ago • 8
UniPrefill: Universal Long-Context Prefill Acceleration via Block-wise Dynamic Sparsification Paper • 2605.06221 • Published May 7 • 21
ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning Paper • 2605.00380 • Published May 1 • 7