Circuit Hypernetworks for Quantum-Augmented Diffusion Language Models Paper • 2609.24657 • Published 13 days ago • 40
NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video Paper • 2608.13210 • Published Aug 13 • 10
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos Paper • 2607.11523 • Published Jul 13 • 9
Vinci2: Providing Proactive Assistance in Continuous Egocentric Videos Paper • 2607.11523 • Published Jul 13 • 9
Jailbreaking Multimodal Large Language Models via Shuffle Inconsistency Paper • 2501.04931 • Published Jan 9, 2025
From reactive to cognitive: brain-inspired spatial intelligence for embodied agents Paper • 2508.17198 • Published Aug 24, 2025 • 10
Towards Interactive Intelligence for Digital Humans Paper • 2512.13674 • Published Dec 15, 2025 • 12
SFHand: A Streaming Framework for Language-guided 3D Hand Forecasting and Embodied Manipulation Paper • 2511.18127 • Published Nov 22, 2025 • 1
Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions Paper • 2510.27195 • Published Oct 31, 2025 • 1
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality? Paper • 2605.22109 • Published May 21 • 50
Perception or Prejudice: Can MLLMs Go Beyond First Impressions of Personality? Paper • 2605.22109 • Published May 21 • 50
Can MLLMs Read the Room? A Multimodal Benchmark for Verifying Truthfulness in Multi-Party Social Interactions Paper • 2510.27195 • Published Oct 31, 2025 • 1
SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise Paper • 2602.12783 • Published Feb 13 • 25
SQuTR: A Robustness Benchmark for Spoken Query to Text Retrieval under Acoustic Noise Paper • 2602.12783 • Published Feb 13 • 25