PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents Paper • 2609.40285 • Published 5 days ago • 21
CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes Paper • 2608.27455 • Published Aug 27 • 12
Toward Skill-Native LLMs: Skill Entropy for Benchmarking and Training Long-Horizon Reasoning Paper • 2608.05139 • Published Aug 5 • 28
Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision Paper • 2604.12002 • Published Apr 13 • 12
AdaptMI: Adaptive Skill-based In-context Math Instruction for Small Language Models Paper • 2505.00147 • Published Apr 30, 2025 • 4