PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models Paper • 2609.14973 • Published 13 days ago • 174
Breaking the Vision-Action Shortcut: Latent Interface Training for Generalizable Robotics Foundation Models Paper • 2609.12641 • Published 16 days ago • 71
τ_0-VLA: a Hierarchical Robot Foundation Model with World-Model-Guided Test-Time Computation Paper • 2608.16885 • Published Aug 17 • 17
Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders Paper • 2601.16208 • Published Jan 22 • 55
Too Good to be Bad: On the Failure of LLMs to Role-Play Villains Paper • 2511.04962 • Published Nov 7, 2025 • 57
OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language Models Paper • 2506.03135 • Published Jun 3, 2025 • 41
Revisiting Modeling and Evaluation Approaches in Speech Emotion Recognition: Considering Subjectivity of Annotators and Ambiguity of Emotions Paper • 2510.05934 • Published Oct 7, 2025 • 3
VCRL: Variance-based Curriculum Reinforcement Learning for Large Language Models Paper • 2509.19803 • Published Sep 24, 2025 • 122
Video-MTR: Reinforced Multi-Turn Reasoning for Long Video Understanding Paper • 2508.20478 • Published Aug 28, 2025 • 18
nablaNABLA: Neighborhood Adaptive Block-Level Attention Paper • 2507.13546 • Published Jul 17, 2025 • 126
Thinking with Images for Multimodal Reasoning: Foundations, Methods, and Future Frontiers Paper • 2506.23918 • Published Jun 30, 2025 • 90
Geometry-Editable and Appearance-Preserving Object Compositon Paper • 2505.20914 • Published May 27, 2025 • 6