Rethinking Training-Inference Mismatch in LLM Reinforcement Learning: Where It Arises and How to Correct It Paper • 2609.32444 • Published 15 days ago • 35
When EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation Paper • 2609.20511 • Published 24 days ago • 115
Youthquake123/gemma3-it_FINAL_TRAIN_EPOCH5_MERGED_int4_ao_quantized Image-Text-to-Text • Updated Aug 6, 2025 • 10
Youthquake123/gemma3-it_FINAL_TRAIN_EPOCH5_MERGED_int4_ao_quantized Image-Text-to-Text • Updated Aug 6, 2025 • 10