ReOPD On-policy distillation for agents without the environment: replay teacher prefixes instead of live rollouts. Multi-Turn On-Policy Distillation with Prefix Replay Paper • 2607.04763 • Published Jul 16 • 12 baohao/Math_Qwen3-4B-Instruct-2507_SFT-RL 4B • Updated Jul 23 • 17 baohao/Math_Qwen3-4B-Instruct-2507_SFT 4B • Updated Jul 23 • 29 baohao/Math_Qwen3-30B-A3B-Instruct-2507_SFT 31B • Updated Jul 23 • 8
Multi-Turn On-Policy Distillation with Prefix Replay Paper • 2607.04763 • Published Jul 16 • 12
SAGE Self-Hinting Language Models Enhance Reinforcement Learning Self-Hinting Language Models Enhance Reinforcement Learning Paper • 2602.03143 • Published Feb 3 • 31 baohao/aime24 Viewer • Updated Feb 7 • 30 • 314 baohao/aime25 Viewer • Updated Feb 7 • 30 • 302 baohao/amc23 Viewer • Updated Feb 7 • 40 • 298
Self-Hinting Language Models Enhance Reinforcement Learning Paper • 2602.03143 • Published Feb 3 • 31
ReOPD On-policy distillation for agents without the environment: replay teacher prefixes instead of live rollouts. Multi-Turn On-Policy Distillation with Prefix Replay Paper • 2607.04763 • Published Jul 16 • 12 baohao/Math_Qwen3-4B-Instruct-2507_SFT-RL 4B • Updated Jul 23 • 17 baohao/Math_Qwen3-4B-Instruct-2507_SFT 4B • Updated Jul 23 • 29 baohao/Math_Qwen3-30B-A3B-Instruct-2507_SFT 31B • Updated Jul 23 • 8
Multi-Turn On-Policy Distillation with Prefix Replay Paper • 2607.04763 • Published Jul 16 • 12
SAGE Self-Hinting Language Models Enhance Reinforcement Learning Self-Hinting Language Models Enhance Reinforcement Learning Paper • 2602.03143 • Published Feb 3 • 31 baohao/aime24 Viewer • Updated Feb 7 • 30 • 314 baohao/aime25 Viewer • Updated Feb 7 • 30 • 302 baohao/amc23 Viewer • Updated Feb 7 • 40 • 298
Self-Hinting Language Models Enhance Reinforcement Learning Paper • 2602.03143 • Published Feb 3 • 31