Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published 25 days ago • 706
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Paper • 2609.03796 • Published Sep 3 • 185
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published Aug 4 • 107
ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Paper • 2608.04436 • Published Aug 5 • 60
Unlocking the Potential of Image Editing via Concept Scaling and Dense Supervision Paper • 2608.16812 • Published Aug 17 • 50
Read It Back: Pretrained MLLMs Are Zero-Shot Reward Models for Text-to-Image Generation Paper • 2607.11886 • Published Jul 13 • 60
Mage-Flow: An Efficient Native-Resolution Foundation Model for Image Generation and Editing Paper • 2607.19064 • Published Jul 21 • 77
Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers Paper • 2607.19139 • Published Jul 21 • 74
Imagine Before You Draw: Visual Prompt Engineering for Image Generation Paper • 2606.04457 • Published Jun 3
ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition Paper • 2607.25565 • Published Jul 28 • 64
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published Jul 20 • 54