Multi-module GRPO: Composing Policy Gradients and Prompt Optimization for Language Model Programs Paper • 2508.04660 • Published Aug 6, 2025 • 7
SmolVLA: A Vision-Language-Action Model for Affordable and Efficient Robotics Paper • 2506.01844 • Published Jun 2, 2025 • 164
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence Paper • 2606.14777 • Published Jun 10 • 216
Xiaomi-Robotics-1: Scaling Vision-Language-Action Models with over 100K Hours of Real-World Trajectories Paper • 2607.15330 • Published 13 days ago • 70
GEPA: Reflective Prompt Evolution Can Outperform Reinforcement Learning Paper • 2507.19457 • Published Jul 25, 2025 • 35
GigaWorld-1: A Roadmap to Build World Models for Robot Policy Evaluation Paper • 2607.02642 • Published 27 days ago • 40
IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Paper • 2502.05512 • Published Feb 8, 2025 • 8
MediaPipe: A Framework for Building Perception Pipelines Paper • 1906.08172 • Published Jun 14, 2019 • 3
PyTorch Distributed: Experiences on Accelerating Data Parallel Training Paper • 2006.15704 • Published Jun 28, 2020 • 12
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published 9 days ago • 59
RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources Paper • 2606.29538 • Published 13 days ago • 141
Very Large-Scale Multi-Agent Simulation in AgentScope Paper • 2407.17789 • Published Jul 25, 2024 • 46
AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications Paper • 2508.16279 • Published Aug 22, 2025 • 68
SearchOS-V1: Towards Robust Open-Domain Information-Seeking Agent Collaboration Paper • 2607.15257 • Published 13 days ago • 71
EverMemOS: A Self-Organizing Memory Operating System for Structured Long-Horizon Reasoning Paper • 2601.02163 • Published Jan 5 • 15
Ultralytics YOLO26: Unified Real-Time End-to-End Vision Models Paper • 2606.03748 • Published Jun 2 • 21