Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Paper • 2608.05139 • Published 3 days ago • 25
Transferability for General Reasoning: An Automated Curriculum for Multi-Domain RLVR Paper • 2606.25178 • Published Jun 27 • 7
Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision Paper • 2604.12002 • Published Apr 13 • 12