ActiveVLA: Injecting Active Perception into Vision-Language-Action Models for Precise 3D Robotic Manipulation Paper • 2601.08325 • Published Jan 13 • 1
Spatial-Temporal Aware Visuomotor Diffusion Policy Learning Paper • 2507.06710 • Published Jul 9, 2025
In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion Paper • 2609.32540 • Published 9 days ago • 33
WorldPlay2: Extending Real-Time Interactive World Models in Control and Horizon Paper • 2609.35560 • Published 7 days ago • 31
In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion Paper • 2609.32540 • Published 9 days ago • 33
In-Flight KV Cache with Clean Anchors for Faster Autoregressive Video Diffusion Paper • 2609.32540 • Published 9 days ago • 33
Direct 3D-Aware Object Insertion via Decomposed Visual Proxies Paper • 2606.06601 • Published Jun 4 • 26
Direct 3D-Aware Object Insertion via Decomposed Visual Proxies Paper • 2606.06601 • Published Jun 4 • 26
From Pixels to Words -- Towards Native One-Vision Models at Scale Paper • 2605.28820 • Published May 27 • 75
Tuna-2: Pixel Embeddings Beat Vision Encoders for Multimodal Understanding and Generation Paper • 2604.24763 • Published Apr 27 • 71
From Pixels to Words -- Towards Native Vision-Language Primitives at Scale Paper • 2510.14979 • Published Oct 16, 2025 • 70
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation Paper • 2510.08673 • Published Oct 9, 2025 • 128
Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and Generation Paper • 2510.08673 • Published Oct 9, 2025 • 128