thomas-mayne 's Collections Vision Language Models
updated
Mini-Gemini: Mining the Potential of Multi-modality Vision Language
Models
Paper
• 2403.18814
• Published • 49
meta-llama/Llama-3.2-11B-Vision
Image-Text-to-Text
• 11B • Updated • 11.4k
• 611
google/paligemma-3b-pt-224
Image-Text-to-Text
• 3B • Updated • 268k
• 613
Qwen/Qwen2-VL-2B-Instruct
Image-Text-to-Text
• 2B • Updated • 1.42M
• 520
Qwen/Qwen2-VL-7B-Instruct
Image-Text-to-Text
• 8B • Updated • 623k
• 1.29k
Image-Text-to-Text
• 0.7B • Updated • 604k
• 1.56k
Image-Text-to-Text
• 25B • Updated • 43.1k
• 639
Salesforce/blip-image-captioning-large
Image-to-Text
• 0.5B • Updated • 592k
• 1.49k
black-forest-labs/FLUX.1-dev
Text-to-Image
• 12B • Updated • 713k
• • 15.3k
black-forest-labs/FLUX.1-schnell
Text-to-Image
• 12B • Updated • 626k
• • 6.11k
stabilityai/stable-diffusion-3.5-large
Text-to-Image
• 8B • Updated • 100k
• • 4.09k
Image-Text-to-Text
• 8B • Updated • 36.1k
• 1.06k
stabilityai/stable-diffusion-3.5-medium
Text-to-Image
• 2B • Updated • 91.1k
• • 1.17k
Image-Text-to-Text
• Updated • 611
• 1.72k