TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 3 days ago • 122
Self Gradient Forcing: Native Long Video Extrapolation Paper • 2607.20368 • Published 10 days ago • 34
OpenCoF: Learning to Reason Through Video Generation Paper • 2607.08763 • Published 23 days ago • 29
DreamForge-World 0.1 Preview: A Low-Compute Real-Time Controllable World Model Paper • 2606.30292 • Published Jun 29 • 16
JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence Paper • 2606.14777 • Published Jun 10 • 216
Youtu-LLM: Unlocking the Native Agentic Potential for Lightweight Large Language Models Paper • 2512.24618 • Published Dec 31, 2025 • 155
TurboDiffusion: Accelerating Video Diffusion Models by 100-200 Times Paper • 2512.16093 • Published Dec 18, 2025 • 96
DreaMontage: Arbitrary Frame-Guided One-Shot Video Generation Paper • 2512.21252 • Published Dec 24, 2025 • 35
Details Collection A gated collection of datasets containing evaluation details • 4500 items • Updated Mar 2 • 6
Running on Zero Agents Featured 1.64k Stable Diffusion 3 Medium 🎨 1.64k Generate custom images from text prompts with Stable Diffusion 3