NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published 23 days ago • 327
What Does an Evaluation License? A Commit-Bound Census of Claim Replay in Inspect Evals Paper • 2608.19269 • Published Aug 29 • 6
StudentSim: Training LLM-based Student Simulators Paper • 2609.01591 • Published about 1 month ago • 494
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published Aug 17 • 122
WarpSAC: Towards the Pinnacle of Scalable Off-policy RL by Rethinking Exploration and Exploitation Paper • 2608.24479 • Published Aug 25 • 139
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill Paper • 2608.11924 • Published Aug 12 • 110
When Activation Oracles Learn Not to Read: Concept-Specific Blind Spots in Fine-Tuned Oracles Paper • 2607.23379 • Published Jul 25 • 18
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA Paper • 2608.09819 • Published Aug 10 • 179
Progress Reward Modeling for Robotic Learning: A Comprehensive Survey Paper • 2607.21655 • Published Jul 22 • 112
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Paper • 2607.25895 • Published Jul 28 • 95