HumanCLAW: Can Vision-Language Models Act Through a Body? Paper • 2607.27180 • Published 2 days ago • 64
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU Paper • 2607.19191 • Published 9 days ago • 304
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing Paper • 2606.26740 • Published Jun 25 • 82
S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence Paper • 2606.20515 • Published Jun 18 • 41
TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio Motion Embedding and Diffusion Interpolation Paper • 2410.04221 • Published Oct 5, 2024 • 4
InfiniteDance: Scalable 3D Dance Generation Towards in-the-wild Generalization Paper • 2603.13375 • Published Mar 10 • 3
HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions Paper • 2603.15612 • Published Mar 16 • 153
Klear: Unified Multi-Task Audio-Video Joint Generation Paper • 2601.04151 • Published Jan 7 • 17
Simulating the Visual World with Artificial Intelligence: A Roadmap Paper • 2511.08585 • Published Nov 11, 2025 • 30
PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image Paper • 2511.13648 • Published Nov 17, 2025 • 53
The Quest for Generalizable Motion Generation: Data, Model, and Evaluation Paper • 2510.26794 • Published Oct 30, 2025 • 27
SpeakerVid-5M: A Large-Scale High-Quality Dataset for Audio-Visual Dyadic Interactive Human Generation Paper • 2507.09862 • Published Jul 14, 2025 • 52