Scaling and Distilling Text Embeddings for Better Diffusibility Paper • 2610.01016 • Published 8 days ago • 57
CoLT-Drive: Counterfactual Long-Tail Benchmarking and Knowledge-Preserving Adaptation for Driving Affordance Prediction Paper • 2609.00242 • Published Aug 31 • 1
Beyond Textual Chain-of-Thought: A Survey on Action-Grounded Reasoning in Autonomous Driving Paper • 2609.01659 • Published Aug 31 • 2
CausalDS: Benchmarking Causal Reasoning in Data-Science Agents Paper • 2607.08093 • Published Jul 9 • 6
One Scene, Two Depths: Probing Geometric Ambiguity in Monocular Foundation Models Paper • 2606.29600 • Published Jun 28 • 6
Foresight: Failure Detection for Long-Horizon Robotic Manipulation with Action-Conditioned World Model Latents Paper • 2606.23085 • Published Jun 22 • 15
See What I See, Know What I Think: Dense Latent Communication Across Heterogeneous Agents Paper • 2606.13594 • Published Jun 11 • 6
AFUN: Towards an Affordance Foundation Model for Functionality Understanding Paper • 2606.02551 • Published Jun 1 • 7
TactAlign: Human-to-Robot Policy Transfer via Tactile Alignment Paper • 2602.13579 • Published Feb 14 • 11
Next-Embedding Prediction Makes Strong Vision Learners Paper • 2512.16922 • Published Dec 18, 2025 • 91
Towards Scalable Language-Image Pre-training for 3D Medical Imaging Paper • 2505.21862 • Published May 28, 2025 • 1
DANLI: Deliberative Agent for Following Natural Language Instructions Paper • 2210.12485 • Published Oct 22, 2022
What Gives the Answer Away? Question Answering Bias Analysis on Video QA Datasets Paper • 2007.03626 • Published Jul 7, 2020
3D-GRAND: A Million-Scale Dataset for 3D-LLMs with Better Grounding and Less Hallucination Paper • 2406.05132 • Published Jun 7, 2024 • 30
LLM-Grounder: Open-Vocabulary 3D Visual Grounding with Large Language Model as an Agent Paper • 2309.12311 • Published Sep 21, 2023 • 18