MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published 6 days ago • 82
CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Paper • 2608.02589 • Published 3 days ago • 20
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published 7 days ago • 180
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 7 days ago • 300