The Past Frames the Future: Memory for Autoregressive Video Generation Paper • 2609.28466 • Published 15 days ago • 64
GAE: Learning a Geometry-Native Latent Space for 3D-Consistent World Generation Paper • 2609.24981 • Published 17 days ago • 76
Paint-Anything: Unified Any-Color Control for Image Generation and Editing Paper • 2609.20816 • Published 21 days ago • 57
OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation Paper • 2609.22069 • Published 20 days ago • 37
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 21 days ago • 224
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation Paper • 2609.20744 • Published 21 days ago • 53
PhysBrain 1.5: From Vision-Language Models to Physical Foundation Models Paper • 2609.14973 • Published 24 days ago • 175
Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation Paper • 2609.11638 • Published 28 days ago • 707
The Missing Temporal Link: Temporal Context Routing for Script-Driven Audio-Video Generation Paper • 2609.02367 • Published Sep 2 • 38
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Paper • 2609.03796 • Published Sep 3 • 186
H3-World: Turning Language Understanding into World Control Paper • 2609.01560 • Published Sep 1 • 53
Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion Paper • 2608.26794 • Published Aug 27 • 16
LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation Paper • 2608.28460 • Published Aug 28 • 29
WithEveryone: Unified Planning and Identity Grounding for Group Image Generation Paper • 2608.20336 • Published Aug 20 • 43
Alaya-EVOKE: From Linear-Scaling Supervision to Endless World Paper • 2608.13546 • Published Aug 13 • 143
ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Paper • 2608.04436 • Published Aug 5 • 61