Look Before You Select: Rethinking Vocabulary Sparsification in On-Policy Distillation Paper • 2609.34386 • Published 5 days ago • 7
WTF GENIUS PAPERS Collection Papers that made me appreciate my major and my life a little more. obs=Observation, innov=Innovation. Most papers are abt improving tiny models. • 414 items • Updated about 17 hours ago • 93
Can LLMs Learn to Reason Robustly under Noisy Supervision? Paper • 2604.03993 • Published Apr 5 • 39 • 6
Can LLMs Learn to Reason Robustly under Noisy Supervision? Paper • 2604.03993 • Published Apr 5 • 39 • 6
Can LLMs Learn to Reason Robustly under Noisy Supervision? Paper • 2604.03993 • Published Apr 5 • 39 • 6
Seeing What Matters: Visual Preference Policy Optimization for Visual Generation Paper • 2511.18719 • Published Nov 24, 2025 • 1 • 1
TraPO: A Semi-Supervised Reinforcement Learning Framework for Boosting LLM Reasoning Paper • 2512.13106 • Published Dec 15, 2025 • 4
TraPO: A Semi-Supervised Reinforcement Learning Framework for Boosting LLM Reasoning Paper • 2512.13106 • Published Dec 15, 2025 • 4
TraPO: A Semi-Supervised Reinforcement Learning Framework for Boosting LLM Reasoning Paper • 2512.13106 • Published Dec 15, 2025 • 4