LOCI: Spatial Linear Memory for Streaming World Models Paper • 2609.40222 • Published 6 days ago • 12
Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces Paper • 2609.40362 • Published 7 days ago • 37
StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training Paper • 2609.26774 • Published 15 days ago • 56
Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation Paper • 2609.20744 • Published 20 days ago • 53
DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 20 days ago • 224
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance Paper • 2606.19195 • Published Jun 17 • 80
RAD-2: Scaling Reinforcement Learning in a Generator-Discriminator Framework Paper • 2604.15308 • Published Apr 16 • 29
Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding Paper • 2604.05015 • Published Apr 6 • 232
UniDriveVLA: Unifying Understanding, Perception, and Action Planning for Autonomous Driving Paper • 2604.02190 • Published Apr 2 • 26
Spa3R: Predictive Spatial Field Modeling for 3D Visual Reasoning Paper • 2602.21186 • Published Feb 24 • 5
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models Paper • 2512.15713 • Published Dec 17, 2025 • 20
DiffusionVL: Translating Any Autoregressive Models into Diffusion Vision Language Models Paper • 2512.15713 • Published Dec 17, 2025 • 20
Towards Scalable Pre-training of Visual Tokenizers for Generation Paper • 2512.13687 • Published Dec 15, 2025 • 108