TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 13 days ago • 139
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model Paper • 2607.17977 • Published 22 days ago • 198
MAGIC: Transition-Aware Generation of Navigable Multi-Scene Game Worlds with Large Language Models Paper • 2607.11594 • Published 29 days ago • 7
LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies Paper • 2606.15768 • Published Jun 14 • 6
OmniDirector: General Multi-Shot Camera Cloning without Cross-Paired Data Paper • 2606.13432 • Published Jun 11 • 113
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 253
Lost in Sampling: Assessing Lexical Reachability in LLMs via the Word Coverage Score (WCS) Paper • 2605.27268 • Published May 26 • 13
EverAnimate: Minute-Scale Human Animation via Latent Flow Restoration Paper • 2605.15042 • Published May 14 • 5
ClawKeeper: Comprehensive Safety Protection for OpenClaw Agents Through Skills, Plugins, and Watchers Paper • 2603.24414 • Published Mar 25 • 183
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models Paper • 2603.16859 • Published Mar 17 • 248