LoGRA: Scaling LLM Reinforcement Learning with Low-Rank Gradient Sketches Paper • 2610.06647 • Published 4 days ago • 121
Omni-Embed-Mini: Binding Modalities Without Forgetting via Dense Distillation Paper • 2610.02148 • Published 8 days ago • 20
ScienceIDE: Turning World's Scientific Codebase into Agent Learnable Environments Paper • 2609.19134 • Published 23 days ago • 102
Beyond Solver Verdicts: Generative Reward Models for Autoformalization Paper • 2609.11085 • Published 29 days ago • 35
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published Sep 8 • 328