unsloth/Qwen3-VL-8B-Instruct-unsloth-bnb-4bit Image-Text-to-Text • 9B • Updated Oct 31, 2025 • 24.5k • 22
google/siglip2-base-patch16-512 Zero-Shot Image Classification • 0.4B • Updated Feb 21, 2025 • 126k • 50
microsoft/Phi-4-multimodal-instruct Automatic Speech Recognition • 6B • Updated Dec 10, 2025 • 270k • 1.61k
Qwen-VLA: Unifying Vision-Language-Action Modeling across Tasks, Environments, and Robot Embodiments Paper • 2605.30280 • Published May 28 • 146