Active Video Perception: Iterative Evidence Seeking for Agentic Long Video Understanding Paper • 2512.05774 • Published Dec 5, 2025 • 7
Learning Visual Grounding from Generative Vision and Language Model Paper • 2407.14563 • Published Jul 18, 2024
Deform360: A Massive Multi-view Visuotactile Dataset for Deformable World Models Paper • 2607.05390 • Published Jul 6 • 11
A Cookbook of 3D Vision: Data, Learning Paradigms, and Application Paper • 2606.04291 • Published Jun 2 • 6
ParBalans: Parallel Multi-Armed Bandits-based Adaptive Large Neighborhood Search Paper • 2508.06736 • Published Aug 8, 2025
FinToolBench: Evaluating LLM Agents for Real-World Financial Tool Use Paper • 2603.08262 • Published Mar 9 • 42
UMO: Unified In-Context Learning Unlocks Motion Foundation Model Priors Paper • 2603.15975 • Published Mar 16 • 3
Causal-JEPA: Learning World Models through Object-Level Latent Interventions Paper • 2602.11389 • Published Feb 11 • 13