YuE2: Unifying Symbolic and Audio Music Generation at Frontier Quality Paper • 2609.33757 • Published 3 days ago • 199
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 12 days ago • 138
Terminal-Universe: Turning Agent Trajectories into Scalable Terminal Environments Paper • 2609.04148 • Published 27 days ago • 248
Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving Paper • 2609.00111 • Published about 1 month ago • 315
S3Gym: Can LLMs Turn Self-Testing and Self-Judging into Self-Improvement? Paper • 2608.31100 • Published about 1 month ago • 42
Aspire: Can Models Self-Evolve from Vague Goals? Paper • 2608.31111 • Published about 1 month ago • 169
HarnessDev: Can LLMs Create and Evolve Their Own Agent Harness? Paper • 2609.01437 • Published 29 days ago • 222
SMELT: Scaling Laws for Compute-Matched MoE Looped Transformers Paper • 2609.01343 • Published 29 days ago • 91
StartupBench: Benchmarking General-Purpose Agents on Market-Validated End-to-End Workflows Paper • 2608.17800 • Published Aug 18 • 12
LoopsBench: From Harness Engineering to Loop Engineering in Coding Agent Evaluation Paper • 2608.00267 • Published Jul 31 • 3
AVE-Compass: Towards Holistic Evaluation for Audio-Video Editing Abilities Paper • 2607.24821 • Published Jul 17 • 18
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published Aug 3 • 188
Beacon: Knowing When and How to Perform Agentic Visual Reasoning Paper • 2607.28595 • Published Jul 30 • 56
RefCaptioner: Multi-Reference Image-Grounded Video Captioning Paper • 2607.28509 • Published Jul 30 • 30
ICAE-Bench: Evaluating Coding Agents as Interactive Project Builders Paper • 2607.21217 • Published Jul 23 • 9
ABot-AgentOS: A General Robotic Agent OS with Lifelong Multi-modal Memory Paper • 2607.10350 • Published Jul 15 • 80
ABot-N1: Toward a General Visual Language Navigation Foundation Model Paper • 2607.10383 • Published Jul 14 • 97