Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation Paper • 2211.06687 • Published Nov 12, 2022 • 7
Video Generation Models are General-Purpose Vision Learners Paper • 2607.09024 • Published 18 days ago • 85
Vidu S1: A Real-Time Interactive Video Generation Model Paper • 2607.03118 • Published 25 days ago • 141
CohereLabs/cohere-transcribe-arabic-07-2026 Automatic Speech Recognition • 2B • Updated 14 days ago • 39.4k • 137
Program-as-Weights: A Programming Paradigm for Fuzzy Functions Paper • 2607.02512 • Published 26 days ago • 236
Towards Automating Scientific Review with Google's Paper Assistant Tool Paper • 2606.28277 • Published Jun 26 • 11
Autodata: An agentic data scientist to create high quality synthetic data Paper • 2606.25996 • Published Jun 24 • 18
Moebius: 0.2B Lightweight Image Inpainting Framework with 10B-Level Performance Paper • 2606.19195 • Published Jun 17 • 141