Human Cognition in Machines: A Unified Perspective of World Models Paper • 2604.16592 • Published Apr 17
LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception Paper • 2609.39938 • Published 7 days ago • 2
LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception Paper • 2609.39938 • Published 7 days ago • 2
Human Cognition in Machines: A Unified Perspective of World Models Paper • 2604.16592 • Published Apr 17
Flash-WAM: Modality-Aware Distillation for World Action Models Paper • 2606.05254 • Published Jun 3 • 7
PhyGround: Benchmarking Physical Reasoning in Generative World Models Paper • 2605.10806 • Published May 11 • 4
PhyGround: Benchmarking Physical Reasoning in Generative World Models Paper • 2605.10806 • Published May 11 • 4
ActQuant: Sub-4-bit Action-Guided Quantization for Vision-Language-Action Models Paper • 2605.24011 • Published May 19 • 2
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Paper • 2507.05116 • Published Jul 7, 2025
CircuitSense: A Hierarchical Circuit System Benchmark Bridging Visual Comprehension and Symbolic Reasoning in Engineering Design Process Paper • 2509.22339 • Published Sep 26, 2025 • 1
Dream-VL & Dream-VLA: Open Vision-Language and Vision-Language-Action Models with Diffusion Language Model Backbone Paper • 2512.22615 • Published Dec 27, 2025 • 51
VOTE: Vision-Language-Action Optimization with Trajectory Ensemble Voting Paper • 2507.05116 • Published Jul 7, 2025
DiffuCoder: Understanding and Improving Masked Diffusion Models for Code Generation Paper • 2506.20639 • Published Jun 25, 2025 • 32
L-Eval: Instituting Standardized Evaluation for Long Context Language Models Paper • 2307.11088 • Published Jul 20, 2023 • 5
BBA: Bi-Modal Behavioral Alignment for Reasoning with Large Vision-Language Models Paper • 2402.13577 • Published Feb 21, 2024 • 9