Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces Paper • 2609.40362 • Published 5 days ago • 23
Qwen-VL: A Frontier Large Vision-Language Model with Versatile Abilities Paper • 2308.12966 • Published Aug 24, 2023 • 12
view article Article Introducing Idefics2: A Powerful 8B Vision-Language Model for the community +1 Leyo, HugoLaurencon, VictorSanh • Apr 15, 2024 • 192