🏗️ Building on HF
Adarsh Zolekar
adarshzolekar
AI & ML interests
Exploring AI, ML, Deep Learning, models and datasets while building and contributing to the Hugging Face community.
Recent Activity
liked a model 8 days ago
Lightricks/LTX-2.5 liked a model 8 days ago
Qwen/Qwen-Image-2.1 upvoted a paper 8 days ago
Grounded Skill Synthesis from Code at Scale for Agentic IntelligenceOrganizations
Multimodal AI Models
Purpose: Models that understand text + image + audio together.
Vision Models (Image & Video)
Purpose: Text-to-image, image classification, detection, segmentation.
-
openai/clip-vit-base-patch32
Zero-Shot Image Classification • Updated • 21.6M • 1.57k -
facebook/detr-resnet-50
Object Detection • 41.6M • Updated • 319k • • 980 -
Tongyi-MAI/Z-Image-Turbo
Text-to-Image • 6B • Updated • 585k • • 5.39k -
black-forest-labs/FLUX.1-dev
Text-to-Image • 12B • Updated • 690k • • 15.2k
Embeddings & Retrieval Models (RAG)
-
sentence-transformers/all-MiniLM-L6-v2
Sentence Similarity • 22.7M • Updated • 243M • • 6.15k -
BAAI/bge-m3
Sentence Similarity • Updated • 35.9M • • 3.74k -
nomic-ai/nomic-embed-text-v1.5
Sentence Similarity • 0.1B • Updated • 13.5M • 943 -
BAAI/bge-reranker-v2-m3
Text Classification • 0.6B • Updated • 17.1M • • 1.22k
Audio & Speech Models
Purpose: Speech recognition, text-to-speech, music, audio analysis.
-
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 4.32M • • 6.47k -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 6.4M • • 3.41k -
hexgrad/Kokoro-82M
Text-to-Speech • Updated • 11.5M • • 7.06k -
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Text-to-Speech • 2B • Updated • 2.4M • 2.02k
Text & Code Models (NLP)
Purpose: Text generation, summarization, translation, embeddings, coding.
-
mistralai/Mistral-7B-Instruct-v0.3
7B • Updated • 2.18M • 3.65k -
Qwen/Qwen3-8B
Text Generation • 8B • Updated • 11.7M • • 2.06k -
deepseek-ai/DeepSeek-V4-Flash-0731
Text Generation • 304B • Updated • 4.52M • • 4.01k -
unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF
Text Generation • 31B • Updated • 9.9M • 1.09k
Reasoning & Agentic Models
Embeddings & Retrieval Models (RAG)
-
sentence-transformers/all-MiniLM-L6-v2
Sentence Similarity • 22.7M • Updated • 243M • • 6.15k -
BAAI/bge-m3
Sentence Similarity • Updated • 35.9M • • 3.74k -
nomic-ai/nomic-embed-text-v1.5
Sentence Similarity • 0.1B • Updated • 13.5M • 943 -
BAAI/bge-reranker-v2-m3
Text Classification • 0.6B • Updated • 17.1M • • 1.22k
Multimodal AI Models
Purpose: Models that understand text + image + audio together.
Audio & Speech Models
Purpose: Speech recognition, text-to-speech, music, audio analysis.
-
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 4.32M • • 6.47k -
openai/whisper-large-v3-turbo
Automatic Speech Recognition • 0.8B • Updated • 6.4M • • 3.41k -
hexgrad/Kokoro-82M
Text-to-Speech • Updated • 11.5M • • 7.06k -
Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice
Text-to-Speech • 2B • Updated • 2.4M • 2.02k
Vision Models (Image & Video)
Purpose: Text-to-image, image classification, detection, segmentation.
-
openai/clip-vit-base-patch32
Zero-Shot Image Classification • Updated • 21.6M • 1.57k -
facebook/detr-resnet-50
Object Detection • 41.6M • Updated • 319k • • 980 -
Tongyi-MAI/Z-Image-Turbo
Text-to-Image • 6B • Updated • 585k • • 5.39k -
black-forest-labs/FLUX.1-dev
Text-to-Image • 12B • Updated • 690k • • 15.2k
Text & Code Models (NLP)
Purpose: Text generation, summarization, translation, embeddings, coding.
-
mistralai/Mistral-7B-Instruct-v0.3
7B • Updated • 2.18M • 3.65k -
Qwen/Qwen3-8B
Text Generation • 8B • Updated • 11.7M • • 2.06k -
deepseek-ai/DeepSeek-V4-Flash-0731
Text Generation • 304B • Updated • 4.52M • • 4.01k -
unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF
Text Generation • 31B • Updated • 9.9M • 1.09k