Workflow-GYM: Towards Long-Horizon Evaluation of Computer-use Agentic tasks in Real-World Professional Fields Paper • 2606.11042 • Published Jun 9 • 221
NeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness Paper • 2609.08183 • Published Sep 8 • 321
Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization Paper • 2609.05258 • Published Sep 4 • 20
MNIST-PRO: MNIST is Back as a Partially Observable World for AI Agents Paper • 2608.31022 • Published Aug 31 • 9
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published Aug 17 • 122