MerchantBench: Benchmarking LLM Agents for Long-Term Coherence in E-Commerce Operations Paper • 2607.28956 • Published 12 days ago • 96
RL^2-VLA: Adaptive RL Latent Compositional Steering with Test-Time Scaling for Vision-Language-Action Models Paper • 2607.26991 • Published 13 days ago • 9
N_0-VTLA: Scaling Vision-Tactile-Language-Action Model with Latent Tactile Tokens Paper • 2607.23782 • Published 17 days ago • 77
LongHorizon-Harness: Advancing Long-Horizon Agents for Real-World Tasks Paper • 2608.01964 • Published 9 days ago • 166
Semi-off-Policy Reinforcement Learning for Vision-Language Slow-thinking Reasoning Paper • 2507.16814 • Published Jul 22, 2025 • 22
Progressive Agent Skill Generation via Reinforcement Learning Paper • 2608.01678 • Published 9 days ago • 58
INTACT: Isomorphic Intent-to-Action Learning for Search-Free World Models Paper • 2607.26056 • Published 15 days ago • 17
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published 13 days ago • 302
VideoCoCo: Code-as-CoT for Physically-Consistent Video Generation via an Agentic Dual-Engine System Paper • 2607.27380 • Published 14 days ago • 71
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published 13 days ago • 182
Can AI agents conduct open-ended AI research? Early evidence from two case studies Paper • 2607.27191 • Published 14 days ago • 18
TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 14 days ago • 139
Molt: A Scalable PyTorch-Native Training Framework for Agentic Reinforcement Learning Paper • 2607.21653 • Published 21 days ago • 32
Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers Paper • 2607.21594 • Published 20 days ago • 16
G-MAD: A Game-Based Data Generation Framework for Multi-View RGB-T Aerial Object Detection Paper • 2607.19942 • Published 21 days ago • 4
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Paper • 2607.13429 • Published 28 days ago • 16
Trajectory-aware Cross-view Geo-localization with Sequential Observations Paper • 2607.15491 • Published 27 days ago • 8
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 21 days ago • 311