The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation Paper • 2609.36484 • Published 8 days ago • 578
Edge0/Audio8-ASR-Infinite Automatic Speech Recognition • 4B • Updated 13 days ago • 40.4k • 2.43k
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents Paper • 2609.39102 • Published 7 days ago • 672
openai/clip-vit-base-patch32 Zero-Shot Image Classification • Updated Feb 29, 2024 • 20.5M • 1.58k
sentence-transformers/all-MiniLM-L6-v2 Sentence Similarity • 22.7M • Updated Jun 1 • 235M • • 6.2k
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs Paper • 2608.20492 • Published Aug 20 • 87