ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models Paper • 2609.18487 • Published 19 days ago • 48
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models Paper • 2609.14973 • Published 21 days ago • 175
AI-generated Images Challenge Visual Trust in High-risk Scenarios Paper • 2607.22745 • Published Jul 23
UMEM: Unified Memory Extraction and Management Framework for Generalizable Memory Paper • 2602.10652 • Published Feb 11 • 4
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models Paper • 2609.14973 • Published 21 days ago • 175
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models Paper • 2609.14973 • Published 21 days ago • 175
MR-Align: Meta-Reasoning Informed Factuality Alignment for Large Reasoning Models Paper • 2510.24794 • Published Oct 27, 2025 • 32
ScalSelect: Scalable Training-Free Multimodal Data Selection for Efficient Visual Instruction Tuning Paper • 2602.11636 • Published Feb 12 • 2
IntentVLA: Short-Horizon Intent Modeling for Aliased Robot Manipulation Paper • 2605.14712 • Published May 14 • 16
RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models Paper • 2606.02277 • Published Jun 1 • 4
SIEVE: Structure-Aware Data Selection for Imitation Learning with VLA Models Paper • 2607.06442 • Published Jul 7 • 5
Human-as-Humanoid: Enabling Zero-Shot Humanoid Learning from Ego-Exo Human Videos with Human-Aligned Embodiments Paper • 2606.32009 • Published Jun 30
RoboSemanticBench: Diagnosing Semantic Grounding in Action Prediction for VLA Models Paper • 2606.02277 • Published Jun 1 • 4