MM-SpuBench: Towards Better Understanding of Spurious Biases in Multimodal LLMs Paper • 2406.17126 • Published Jun 24, 2024
PALM: Progress-Aware Policy Learning via Affordance Reasoning for Long-Horizon Robotic Manipulation Paper • 2601.07060 • Published Jan 11 • 1
TrialBench: Multi-Modal Artificial Intelligence-Ready Clinical Trial Datasets Paper • 2407.00631 • Published Jun 15, 2025
GRASP: Learning to Ground Social Reasoning in Multi-Person Non-Verbal Interactions Paper • 2605.15764 • Published May 15 • 4
Evaluating Cognitive Age Alignment in Interactive AI Agents Paper • 2605.17894 • Published May 18 • 5
CogniRoute: Learning to Route Social Evidence in Omni-Modal Models Paper • 2606.20970 • Published Jun 18 • 4
Toward Cognitive Supersensing in Multimodal Large Language Model Paper • 2602.01541 • Published Feb 2 • 16
Drive as You Speak: Enabling Human-Like Interaction with Large Language Models in Autonomous Vehicles Paper • 2309.10228 • Published Sep 19, 2023
On-Board Vision-Language Models for Personalized Autonomous Vehicle Motion Control: System Design and Real-World Validation Paper • 2411.11913 • Published Nov 17, 2024
MedSAM3: Delving into Segment Anything with Medical Concepts Paper • 2511.19046 • Published Nov 24, 2025 • 56
SocialGesture: Delving into Multi-person Gesture Understanding Paper • 2504.02244 • Published Apr 3, 2025
Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMs Paper • 2506.21656 • Published Jun 26, 2025 • 16
If LLM Is the Wizard, Then Code Is the Wand: A Survey on How Code Empowers Large Language Models to Serve as Intelligent Agents Paper • 2401.00812 • Published Jan 1, 2024 • 12
What is the Visual Cognition Gap between Humans and Multimodal LLMs? Paper • 2406.10424 • Published Jun 14, 2024
Mitigating Transformer Overconfidence via Lipschitz Regularization Paper • 2306.06849 • Published Jun 12, 2023