SANA-Video 2.0: Hybrid Linear Attention with Attention Residuals for Efficient Video Generation Paper • 2607.21553 • Published 8 days ago • 38
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published 21 days ago • 85
InternVLA-A1.5: Unifying Understanding, Latent Foresight, and Action for Compositional Generalization Paper • 2607.04988 • Published 25 days ago • 28
GLM-5V-Turbo: Toward a Native Foundation Model for Multimodal Agents Paper • 2604.26752 • Published Apr 29 • 113
VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric Control Paper • 2601.05138 • Published Jan 8 • 19
NitroGen: An Open Foundation Model for Generalist Gaming Agents Paper • 2601.02427 • Published Jan 4 • 46
DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models Paper • 2512.02556 • Published Dec 2, 2025 • 271
VIST3A: Text-to-3D by Stitching a Multi-view Reconstruction Network to a Video Generator Paper • 2510.13454 • Published Oct 15, 2025 • 10
Learning an Image Editing Model without Image Editing Pairs Paper • 2510.14978 • Published Oct 16, 2025 • 9
SRDiffusion: Accelerate Video Diffusion Inference via Sketching-Rendering Cooperation Paper • 2505.19151 • Published May 25, 2025 • 2
Kimi-VL-A3B Collection Moonshot's efficient MoE VLMs, exceptional on agent, long-context, and thinking • 6 items • Updated Mar 2 • 84
SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features Paper • 2502.14786 • Published Feb 20, 2025 • 168
SuperGPQA: Scaling LLM Evaluation across 285 Graduate Disciplines Paper • 2502.14739 • Published Feb 20, 2025 • 110