EgoTools: Towards Tool-Centric Reasoning in Real-World Egocentric Videos Paper • 2609.39378 • Published 4 days ago • 44
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning Paper • 2609.35767 • Published 6 days ago • 48
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning Paper • 2609.35767 • Published 6 days ago • 48
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Paper • 2608.26105 • Published Aug 26 • 192
HarnessEval-W: Agentifying the Evaluation of Visual Worlds Paper • 2608.16859 • Published Aug 17 • 122
Apple-$π$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Paper • 2607.16401 • Published Jul 17 • 45
Learning Native Reflection in Unified Models with Interleaved Reinforcement Learning Paper • 2609.35767 • Published 6 days ago • 48
RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents Paper • 2609.25636 • Published 12 days ago • 13
RoboFollow: Unveiling the Instruction Following Mirage in Embodied Agents Paper • 2609.25636 • Published 12 days ago • 13
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training Paper • 2609.15051 • Published 20 days ago • 14
Not All Prompts Are Equal: Exploration-Guided Prompt Scaffolding for Multimodal Reinforcement Post-Training Paper • 2609.15051 • Published 20 days ago • 14
Why Struggle with Continuous Latents? Interpretable Discrete Latent Reasoning via Rendered Compression Paper • 2606.29712 • Published Jun 29 • 1
OmniPresent: Generating Coherent Presentation Suites from Scientific Papers Paper • 2607.02590 • Published Jul 1 • 1
Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System Paper • 2609.01607 • Published Sep 1 • 24
Apple-$π$: Benchmarking Thinking with Video Towards Law-Grounded Physical Intelligence Paper • 2607.16401 • Published Jul 17 • 45
DistillAlign: Coordinating Mode Covering and Mode Seeking in Autoregressive Video Distillation Paper • 2607.26811 • Published Jul 29 • 18
ShotBench: Expert-Level Cinematic Understanding in Vision-Language Models Paper • 2506.21356 • Published Jun 26, 2025 • 22
Cut2Next: Generating Next Shot via In-Context Tuning Paper • 2508.08244 • Published Aug 11, 2025 • 13