BIGS: Bimanual Category-agnostic Interaction Reconstruction from Monocular Videos via 3D Gaussian Splatting Paper • 2504.09097 • Published Apr 12, 2025
Cross-Modal Emotion Transfer for Emotion Editing in Talking Face Video Paper • 2604.07786 • Published Apr 9 • 7
HandVQA: Diagnosing and Improving Fine-Grained Spatial Reasoning about Hands in Vision-Language Models Paper • 2603.26362 • Published Mar 27
LighthouseGS: Indoor Structure-aware 3D Gaussian Splatting for Panorama-Style Mobile Captures Paper • 2507.06109 • Published Jul 8, 2025
Enabling Chatbots with Eyes and Ears: An Immersive Multimodal Conversation System for Dynamic Interactions Paper • 2506.00421 • Published May 31, 2025 • 5
VPOcc: Exploiting Vanishing Point for 3D Semantic Occupancy Prediction Paper • 2408.03551 • Published Aug 7, 2024