DC-SAE: Deep Compression Semantic Autoencoder for Faster Diffusion Convergence Paper • 2609.39222 • Published 2 days ago • 35
Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model Paper • 2609.18323 • Published 16 days ago • 133
VBVR-Pro: A Scalable and Verifiable Suite for Native Visual Reasoning Paper • 2608.26105 • Published Aug 26 • 191
ReWorld: An Interactive World Model with Long-Horizon Memory Paper • 2608.23565 • Published Aug 24 • 25
view article Article NEO-unify: Building Native Multimodal Unified Models End to End sensenova • Mar 5 • 182
WithEveryone: Unified Planning and Identity Grounding for Group Image Generation Paper • 2608.20336 • Published Aug 20 • 43
EDITBRIDGE: Towards Faithful and Efficient Ultra-High-Resolution Image Editing Paper • 2608.18063 • Published Aug 18 • 23
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models Paper • 2608.16887 • Published Aug 17 • 36
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published Aug 4 • 107
ReDesign: Recovering Editable Design Structures from Images via Agentic Decomposition Paper • 2607.25565 • Published Jul 28 • 65
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification Paper • 2607.24027 • Published Jul 27 • 40
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published Jul 8 • 64
Multi-Resolution Flow Matching: Training-Free Diffusion Acceleration via Staged Sampling Paper • 2607.01642 • Published Jul 2 • 38
GEAR: Guided End-to-End AutoRegression for Image Synthesis Paper • 2606.32039 • Published Jun 30 • 34
LiveEdit: Towards Real-Time Diffusion-Based Streaming Video Editing Paper • 2606.26740 • Published Jun 25 • 82
Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation Paper • 2606.26907 • Published Jun 25 • 55