The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation Paper • 2609.02367 • Published Sep 2 • 38
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Paper • 2609.03796 • Published Sep 3 • 182
EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing Paper • 2608.18063 • Published Aug 18 • 23
AVA-Encoder: Towards Agent-Native Video Representation Learning Paper • 2608.12313 • Published Aug 12 • 43
LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time Paper • 2608.11745 • Published Aug 13 • 25
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Paper • 2608.11752 • Published Aug 13 • 23
What to Edit Next: Visually Aligned Image-Editing Follow-Up Suggestions in Conversational Systems Paper • 2608.07565 • Published Aug 3 • 29
Rethinking Classifier-Free Guidance in On-Policy Diffusion Distillation Paper • 2607.24731 • Published Jul 27 • 47
HOMIE: Human-object Centric Video Personalization via Multimodal Intelligent Enchancement Paper • 2607.18217 • Published Jul 20 • 55
CollectionLoRA: Collecting 50 Effects in 1 LoRA via Multi-Teacher On-Policy Distillation Paper • 2605.25378 • Published May 25 • 38
Soap2Soap: Long Cinematic Video Remaking via Multi-Agent Collaboration Paper • 2605.17423 • Published May 17 • 31
RationalRewards: Reasoning Rewards Scale Visual Generation Both Training and Test Time Paper • 2604.11626 • Published Apr 13 • 28
DataFlow: An LLM-Driven Framework for Unified Data Preparation and Workflow Automation in the Era of Data-Centric AI Paper • 2512.16676 • Published Dec 18, 2025 • 225
End-to-End Training for Autoregressive Video Diffusion via Self-Resampling Paper • 2512.15702 • Published Dec 17, 2025 • 16
Z-Image: An Efficient Image Generation Foundation Model with Single-Stream Diffusion Transformer Paper • 2511.22699 • Published Nov 27, 2025 • 249
PosterCopilot: Toward Layout Reasoning and Controllable Editing for Professional Graphic Design Paper • 2512.04082 • Published Dec 3, 2025 • 14
Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length Paper • 2512.04677 • Published Dec 4, 2025 • 134
UI-Ins: Enhancing GUI Grounding with Multi-Perspective Instruction-as-Reasoning Paper • 2510.20286 • Published Oct 23, 2025 • 24