TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining Paper • 2609.33419 • Published 10 days ago • 19
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding Paper • 2609.07941 • Published 30 days ago • 19
Scal3R: Learning Efficient Multi-Relative Pose Query for Scalable Online 3D Reconstruction Paper • 2609.04201 • Published Sep 3 • 50
ReactVAU: A Slow-Fast Decoupled Framework for Streaming Video Anomaly Understanding Paper • 2609.07941 • Published 30 days ago • 19
PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration Paper • 2608.21031 • Published Aug 21 • 4
PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration Paper • 2608.21031 • Published Aug 21 • 4
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Paper • 2606.13673 • Published Jun 11 • 115
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Paper • 2606.13673 • Published Jun 11 • 115
SpatialClaw: Rethinking Action Interface for Agentic Spatial Reasoning Paper • 2606.13673 • Published Jun 11 • 115
FineBench: Benchmarking and Enhancing Vision-Language Models for Fine-grained Human Activity Understanding Paper • 2605.19846 • Published May 20 • 3
Physics in 2-Steps: Locking Motion Priors Before Visual Refinement Erases Them Paper • 2606.06361 • Published Jun 4 • 16
Eagle 2: Building Post-Training Data Strategies from Scratch for Frontier Vision-Language Models Paper • 2501.14818 • Published Jan 20, 2025 • 10
QLIP: Text-Aligned Visual Tokenization Unifies Auto-Regressive Multimodal Understanding and Generation Paper • 2502.05178 • Published Feb 7, 2025 • 10
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots Paper • 2503.14734 • Published Mar 18, 2025 • 9
Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models Paper • 2504.03624 • Published Apr 4, 2025 • 20
OpenVision 3: A Family of Unified Visual Encoder for Both Understanding and Generation Paper • 2601.15369 • Published Jan 21 • 22
Fully Attentional Networks with Self-emerging Token Labeling Paper • 2401.03844 • Published Jan 8, 2024
Argus: Vision-Centric Reasoning with Grounded Chain-of-Thought Paper • 2505.23766 • Published May 29, 2025