Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies Paper • 2609.38155 • Published 5 days ago • 106
Running on Zero Agents Featured 1.38k OmniVoice 🌍 1.38k High-quality voice cloning TTS for 600+ languages
LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding Paper • 2605.27365 • Published May 26 • 143
Running on Zero Agents Featured 998 MMAudio — generating synchronized audio from video/text 🔊 998 Generate audio from video and text prompts
Video Analysis and Generation via a Semantic Progress Function Paper • 2604.22554 • Published Apr 24 • 59
ImagenWorld: Stress-Testing Image Generation Models with Explainable Human Evaluation on Open-ended Real-World Tasks Paper • 2603.27862 • Published Mar 29 • 32
UniGRPO: Unified Policy Optimization for Reasoning-Driven Visual Generation Paper • 2603.23500 • Published Mar 24 • 37
Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model Paper • 2603.21986 • Published Mar 23 • 125