MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs Paper • 2606.30026 • Published Jun 29 • 6
Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents Paper • 2610.07785 • Published 2 days ago • 4
Attacca: Goal-Directed Control under State Continuity for Long-Horizon Embodied Agents Paper • 2610.07785 • Published 2 days ago • 4
MetaRubric: Learning to Reward for Rubric-Based Reinforcement Learning Paper • 2610.02824 • Published 6 days ago • 32
MuseBench: Benchmarking Intent-Level Audiovisual Arts Understanding in MLLMs Paper • 2606.30026 • Published Jun 29 • 6
Confidence-Aware Tool Orchestration for Robust Video Understanding Paper • 2606.26904 • Published Jun 25 • 12
Judging What We Cannot Solve: A Consequence-Based Approach for Oracle-Free Evaluation of Research-Level Math Paper • 2602.06291 • Published Feb 6 • 24
D2E: Scaling Vision-Action Pretraining on Desktop Data for Transfer to Embodied AI Paper • 2510.05684 • Published Oct 7, 2025 • 148