HuggingFaceTB/SmolVLM-256M-Instruct Image-Text-to-Text ⢠0.3B ⢠Updated Apr 8, 2025 ⢠1.07M ⢠393
Running on Zero Agents Featured 1.78k Dia 1.6B đŻ 1.78k Generate realistic dialogue from a script, using Dia!
Paused Agents Featured 229 Spark TTS đ 229 A text-to-speech model powered by SparkAudio and Mobvoi.
HuggingFaceTB/SmolVLM2-500M-Video-Instruct Image-Text-to-Text ⢠0.5B ⢠Updated Apr 8, 2025 ⢠995k ⢠164
microsoft/Phi-4-multimodal-instruct Automatic Speech Recognition ⢠6B ⢠Updated Dec 10, 2025 ⢠554k ⢠1.61k
Running Featured 363 Kokoro Text-to-Speech (WebGPU) đŁ 363 High-quality speech synthesis powered by Kokoro TTS
mlx-community/SmolVLM2-500M-Video-Instruct-mlx Video-Text-to-Text ⢠0.5B ⢠Updated Feb 20, 2025 ⢠2.28k ⢠18
Running on Zero Agents Featured 3.63k InstantID đť 3.63k Generate personalized images preserving your face identity
Runtime error Agents 48 InstructBLIP đ 48 Instruction-tuned model for a range of vision-language tasks
Runtime error Agents Featured 32 CLIPnCROP đ 32 Extract and crop image sections based on text description