mlx-community/Llama-3.2-90B-Vision-Instruct-4bit Image-Text-to-Text • 89B • Updated Dec 21, 2024 • 402 • 5
Large-scale Contrastive Language-Audio Pretraining with Feature Fusion and Keyword-to-Caption Augmentation Paper • 2211.06687 • Published Nov 12, 2022 • 7
Voost: A Unified and Scalable Diffusion Transformer for Bidirectional Virtual Try-On and Try-Off Paper • 2508.04825 • Published Aug 6, 2025 • 60