An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning Paper • 2609.35505 • Published 5 days ago • 21
Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It Paper • 2609.32444 • Published 7 days ago • 30
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 16 days ago • 110
Self-Correction Can Amplify Hallucinations: Fact-Level Repair with Graph-Based Evidence Routing in Multimodal Generation Paper • 2606.00232 • Published Aug 30 • 1
A Survey on Model Extraction Attacks and Defenses for Large Language Models Paper • 2506.22521 • Published Jun 26, 2025 • 1