Agent-Editing World Model: Rethinking World Modeling for LLM Agents Paper • 2609.28416 • Published 13 days ago • 43
Agent-Editing World Model Collection Resources for Agent-Editing World Model (AEWM), a world model for improving long-horizon LLM agents through decision-effect modeling and state rev • 3 items • Updated 11 days ago • 1
CalibForge: Adversarial Solver Calibration for Scaling Learnable Terminal Tasks Paper • 2608.06352 • Published Aug 6 • 24
AREX: Towards a Recursively Self-Improving Agent for Deep Research Paper • 2607.21461 • Published Jul 23 • 108
FORT-Searcher: Synthesizing Shortcut-Resistant Search Tasks for Training Deep Search Agents Paper • 2606.12087 • Published Jun 10 • 51
DeNovoSWE: Scaling Long-Horizon Environments for Generating Entire Repositories from Scratch Paper • 2606.10728 • Published Jun 9 • 37
π-Bench: Evaluating Proactive Personal Assistant Agents in Long-Horizon Workflows Paper • 2605.14678 • Published May 19 • 96
ClawGym: A Scalable Framework for Building Effective Claw Agents Paper • 2604.26904 • Published Apr 29 • 55
Toward Autonomous Long-Horizon Engineering for ML Research Paper • 2604.13018 • Published Apr 14 • 34
Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs Paper • 2604.10480 • Published Apr 12 • 20
QuanBench+: A Unified Multi-Framework Benchmark for LLM-Based Quantum Code Generation Paper • 2604.08570 • Published Mar 25 • 126
Pseudo-Unification: Entropy Probing Reveals Divergent Information Patterns in Unified Multimodal Models Paper • 2604.10949 • Published Apr 13 • 15
Cheers: Decoupling Patch Details from Semantic Representations Enables Unified Multimodal Comprehension and Generation Paper • 2603.12793 • Published Mar 13 • 38
OmniForcing: Unleashing Real-time Joint Audio-Visual Generation Paper • 2603.11647 • Published Mar 12 • 31