Motion Beyond Morphology: Bootstrapping Cross-Category Motion Transfer from Abstract Motion Representations Paper • 2608.01628 • Published 4 days ago • 22
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published 7 days ago • 36
PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models Paper • 2607.24957 • Published 11 days ago • 18
Self Gradient Forcing: Native Long Video Extrapolation Paper • 2607.20368 • Published 16 days ago • 35
Weak-to-Strong Generalization via Direct On-Policy Distillation Paper • 2607.05394 • Published about 1 month ago • 140
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published 28 days ago • 87
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published about 1 month ago • 64
PixWorld: Unifying 3D Scene Generation and Reconstruction in Pixel Space Paper • 2607.05373 • Published Jul 6 • 66
MVTrack4Gen: Multi-View Point Tracking as Geometric Supervision for 4D Video Generation Paper • 2606.26087 • Published Jun 24 • 35
World Tracing: Generative Pixel-Aligned Geometry Beyond the Visible Paper • 2606.13652 • Published Jun 11 • 16
RepFusion: Leveraging Multimodal Priors for Denoising in Representation Space Paper • 2606.14700 • Published Jun 12 • 18
Gamma-World: Generative Multi-Agent World Modeling Beyond Two Players Paper • 2605.28816 • Published May 27 • 433
SenseNova-U1: Unifying Multimodal Understanding and Generation with NEO-unify Architecture Paper • 2605.12500 • Published May 12 • 195