HumanCLAW: Can Vision-Language Models Act Through a Body? Paper • 2607.27180 • Published 15 days ago • 76
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published about 1 month ago • 85
Function-Aware Fill-in-the-Middle as Mid-Training for Coding Agent Foundation Models Paper • 2607.12463 • Published 30 days ago • 108
Drop-Then-Recovery: How Redundant Are Vision-Language-Action Models? Paper • 2606.27755 • Published Jun 26 • 6
DragMesh-2: Physically Plausible Dexterous Hand-Object Interaction with Articulated Objects Paper • 2606.15133 • Published Jun 13 • 74
MotiMotion: Motion-Controlled Video Generation with Visual Reasoning Paper • 2605.22818 • Published May 21 • 5
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality? Paper • 2605.22109 • Published May 21 • 171
SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise Paper • 2602.12783 • Published Feb 13 • 246
SEIF: Self-Evolving Reinforcement Learning for Instruction Following Paper • 2605.07465 • Published May 8 • 30
ShotStream: Streaming Multi-Shot Video Generation for Interactive Storytelling Paper • 2603.25746 • Published Mar 26 • 155
SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models Paper • 2603.16859 • Published Mar 17 • 248