RenderRank: Learning to Rerank Text with Compressed Visual Tokens Paper • 2609.35069 • Published 11 days ago • 7
Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published Sep 3 • 104
ProcessBench: Identifying Process Errors in Mathematical Reasoning Paper • 2412.06559 • Published Dec 9, 2024 • 87
I Have Covered All the Bases Here: Interpreting Reasoning Features in Large Language Models via Sparse Autoencoders Paper • 2503.18878 • Published Mar 24, 2025 • 121
FIPO: Eliciting Deep Reasoning with Future-KL Influenced Policy Optimization Paper • 2603.19835 • Published Mar 20 • 121
Group-in-Group Policy Optimization for LLM Agent Training Paper • 2505.10978 • Published May 16, 2025 • 24