Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation Paper • 2610.05608 • Published 5 days ago • 152
In-Distribution Forcing for Long Video Generation at Test Time Paper • 2610.03120 • Published 7 days ago • 44
World Action Modeling with Progressive Visual Planning Paper • 2610.02508 • Published 8 days ago • 95
DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation Paper • 2609.33711 • Published 12 days ago • 19
SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video Paper • 2609.37969 • Published 10 days ago • 42
LEGO-Anything: Coding Agents for 3D Scene Reconstruction Paper • 2609.36380 • Published 11 days ago • 139
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation Paper • 2609.30221 • Published 15 days ago • 47
VideoPhysEdit: Physical Counterfactual Video Editing via Rigid-Body Physical Scene Reconstruction Paper • 2609.35134 • Published 11 days ago • 19
Precise Editing and Flexible Referencing for Interactable Worlds Paper • 2609.34470 • Published 11 days ago • 20
TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding Paper • 2609.30670 • Published 14 days ago • 11
Schrödinger's Code Repository: Have LLMs Learned SWE-bench or Memorized It? Paper • 2609.27891 • Published Aug 21 • 32
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 18 days ago • 157
Flash-dLLM: IO-Aware KV Caching and Parallel Decoding for Fast, Memory-Efficient Diffusion LLMs Paper • 2609.26796 • Published 17 days ago • 37
Grounded Action Model: 3D Grounding as a Foundation for Robotics Paper • 2609.23863 • Published 19 days ago • 91
Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms Paper • 2609.23658 • Published 19 days ago • 30