OneSearch-VL: Unified Multimodal Deep Research Agent for Image and Video Paper • 2610.12419 • Published 2 days ago • 16
TACD: Distilling Efficient Text-to-Motion Models via Terminal Amplification Control Paper • 2610.02867 • Published 8 days ago • 4
Puffin-World: Scaling a Unified Multimodal Model with Native 3D World States Paper • 2609.04196 • Published Sep 3 • 71
Video-IFBench: Evaluating Instruction Following of Multimodal LLMs in Video Understanding Scenarios Paper • 2608.25529 • Published Aug 26 • 17
Video-MME-Logical: A Controlled Diagnostic Benchmark for Video Temporal-Logical Reasoning Paper • 2606.27828 • Published Jun 26 • 26
Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning Paper • 2605.21487 • Published May 20 • 21
InterleaveThinker: Reinforcing Agentic Interleaved Generation Paper • 2606.13679 • Published Jun 11 • 38
InterleaveThinker: Reinforcing Agentic Interleaved Generation Paper • 2606.13679 • Published Jun 11 • 38
InterleaveThinker: Reinforcing Agentic Interleaved Generation Paper • 2606.13679 • Published Jun 11 • 38
InterleaveThinker/InterleaveThinker-Planner-8B Image-Text-to-Text • 770k • Updated Jun 12 • 25 • 3
InterleaveThinker/InterleaveThinker-Critic-8B Image-Text-to-Text • 9B • Updated Jun 12 • 22 • 2
Uni-Edit: Intelligent Editing Is A General Task For Unified Model Tuning Paper • 2605.21487 • Published May 20 • 21