TurboVLA: Real-Time Vision-Language-Action Model at 32 Hz on an RTX 4090 with <1 GB VRAM Paper • 2607.27205 • Published 4 days ago • 126
Remember When It Matters: Proactive Memory Agent for Long-Horizon Agents Paper • 2607.08716 • Published 24 days ago • 15
lite-infer/Qwen-Image-nunchaku-lite-int4_r32-bnb4-text-encoder Text-to-Image • 8B • Updated 25 days ago • 32 • 1
Trimming the Long-Tail of Visual World Modeling Evaluation Paper • 2606.24256 • Published Jun 23 • 43
Imaginative Perception Tokens Enhance Spatial Reasoning in Multimodal Language Models Paper • 2606.03988 • Published Jun 3 • 126
Crafter: A Multi-Agent Harness for Editable Scientific Figure Generation from Diverse Inputs Paper • 2605.30611 • Published May 28 • 253
Speculative Pipeline Decoding: Higher-Accruacy and Zero-Bubble Speculation via Pipeline Parallelism Paper • 2605.30852 • Published May 29 • 10
VLMs are Good Teachers for Video Reasoning via Adaptive Test-Time Optimization Paper • 2606.02564 • Published Jun 1 • 29