Agent-Editing World Model: Rethinking World Modeling for LLM Agents Paper • 2609.28416 • Published 12 days ago • 43
SPADE: Self-Play in Adaptive Synthetic Executable Environments Paper • 2608.19197 • Published Aug 19 • 54
Improving the matrix multiplication exponent with modern optimization and AlphaEvolve Paper • 2608.16884 • Published Aug 17 • 20
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks Paper • 2608.06352 • Published Aug 6 • 24
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published Jul 23 • 108
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published Jul 20 • 55
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Paper • 2607.17250 • Published Jul 19 • 93
Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading Paper • 2607.08964 • Published Jul 9 • 77
BlockPilot: Instance-Adaptive Policy Learning for Diffusion-based Speculative Decoding Paper • 2606.31315 • Published Jun 30 • 76
TUA-Bench: A Benchmark for General-Purpose Terminal-Use Agents Paper • 2606.28480 • Published Jun 26 • 49
OSWorld2.0: Benchmarking Computer Use Agents on Long-Horizon Real-World Tasks Paper • 2606.29537 • Published Jun 28 • 25
The Verification Horizon: No Silver Bullet for Coding Agent Rewards Paper • 2606.26300 • Published Jun 24 • 53
Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence Paper • 2606.15932 • Published Jun 16 • 39
Autodata: An agentic data scientist to create high quality synthetic data Paper • 2606.25996 • Published Jun 24 • 21
Qwen-AgentWorld: Language World Models for General Agents Paper • 2606.24597 • Published Jun 23 • 167