ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds Paper • 2609.30199 • Published 10 days ago • 29
RecreationWorld: Scalable and Verifiable Environments for Hybrid Computer-Use Agents Paper • 2609.22000 • Published 16 days ago • 79
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 16 days ago • 138
ModularRSI: Modular and Generalizable Recursive Harness Self-Improvement Paper • 2609.14857 • Published 20 days ago • 215
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 19 days ago • 78
SAS: Simple Attention Sparsification via End-to-End Optimization of Context Ranking Paper • 2609.13141 • Published 23 days ago • 68
FrontierChallenge: Evaluating Scientific Workflow Completion Paper • 2608.24979 • Published Aug 25 • 152
OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models Paper • 2607.28609 • Published Jul 30 • 75
Novel Claim or Déjà Vu? Rethinking "Contamination-Free'' Dynamic Evaluation for Multimodal Automated Fact-Checking Paper • 2607.23514 • Published Jul 26 • 14
Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence Paper • 2606.15932 • Published Jun 16 • 39
WBench: A Comprehensive Multi-turn Benchmark for Interactive Video World Model Evaluation Paper • 2605.25874 • Published May 25 • 82
EnvFactory: Scaling Tool-Use Agents via Executable Environments Synthesis and Robust RL Paper • 2605.18703 • Published May 18 • 50
OpenMobile: Building Open Mobile Agents with Task and Trajectory Synthesis Paper • 2604.15093 • Published Apr 16 • 30
How to Fine-Tune a Reasoning Model? A Teacher-Student Cooperation Framework to Synthesize Student-Consistent SFT Data Paper • 2604.14164 • Published Mar 23 • 35
The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Paper • 2604.11297 • Published Apr 13 • 144
Project Imaging-X: A Survey of 1000+ Open-Access Medical Imaging Datasets for Foundation Model Development Paper • 2603.27460 • Published Mar 29 • 72
MetaClaw: Just Talk -- An Agent That Meta-Learns and Evolves in the Wild Paper • 2603.17187 • Published Mar 17 • 141