CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks Paper • 2608.06352 • Published 2 days ago • 11
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published 16 days ago • 151
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published 19 days ago • 61
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Paper • 2607.17250 • Published 20 days ago • 92
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published about 1 month ago • 77
BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding Paper • 2606.31315 • Published Jun 30 • 77
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents Paper • 2606.28480 • Published Jun 26 • 48
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks Paper • 2606.29537 • Published Jun 28 • 24
The Verification Horizon: No Silver Bullet for Coding Agent Rewards Paper • 2606.26300 • Published Jun 24 • 53
Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence Paper • 2606.15932 • Published Jun 16 • 38
Autodata: An agentic data scientist to create high quality synthetic data Paper • 2606.25996 • Published Jun 24 • 18
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 155
EnterpriseClawBench: Benchmarking Agents from Real Workplace Sessions Paper • 2606.23654 • Published Jun 22 • 80
PlanBench-XL: Evaluating Long-Horizon Planning of LLM Tool-Use Agents in Large-Scale Tool Ecosystems Paper • 2606.22388 • Published Jun 21 • 96