Susant Achary PRO
Susant-Achary
AI & ML interests
Building from India. Post Training ,Quantisation(GGUF,ONNX,MLX) Tiny to Small Language Models(Vision, Text, Audio),
Specialisation in Computer vision, Information Retrieval & Representation, Personalisation.(Healthcare Imaging, Geospatial, DocAI,Ecommerce)
Recent Activity
liked a model 11 days ago
unsloth/Kimi-K3-GGUF liked a model 11 days ago
moonshotai/Kimi-K3 liked a model about 1 month ago
baidu/Unlimited-OCROrganizations
🛩️Qwen3-VL
the most powerful vision-language model in the Qwen series to date. Available in Dense and MoE architectures
-
Qwen/Qwen3-VL-30B-A3B-Thinking
Image-Text-to-Text • 31B • Updated • 66.5k • • 200 -
mlx-community/Qwen3-VL-30B-A3B-Instruct-4bit
Image-Text-to-Text • Updated • 919 • 8 -
mlx-community/Qwen3-VL-30B-A3B-Instruct-8bit
Image-Text-to-Text • Updated • 233 • 3 -
mlx-community/Qwen3-VL-8B-Instruct-4bit
Image-Text-to-Text • Updated • 2.38k • 6
🍎 MLX-Quantized Models (3/4/5/6-bit) Mac & iOS
Curated MLX-ready quantized LLMs that run fast on Apple Silicon (and some on iOS). Every card lists Bits · Group size · Peak UM (GB) · Stable context.
-
mlx-community/Apriel-1.5-15b-Thinker-3bit-MLX
Image-Text-to-Text • Updated • 11 -
mlx-community/Apriel-1.5-15b-Thinker-6bit-MLX
Image-Text-to-Text • Updated • 35 • 1 -
mlx-community/granite-4.0-h-tiny-3bit-MLX
Text Generation • 0.9B • Updated • 76 • 2 -
mlx-community/granite-4.0-tiny-preview-4bit
Text Generation • 1B • Updated • 86
🖼️ Vision Backbones & Image Embeddings
-
facebook/dinov2-base
Image Feature Extraction • 86.6M • Updated • 2.4M • 188 -
openai/clip-vit-large-patch14-336
Zero-Shot Image Classification • Updated • 4.12M • 308 -
google/siglip-so400m-patch14-384
Zero-Shot Image Classification • 0.9B • Updated • 2.28M • 682 -
BAAI/EVA-CLIP-8B
Feature Extraction • Updated • 1.06k • 51
🧊Sept 25 <Image-to-3D> [Top Releases]
Models that turn a single image (or image+prompt) into 3D assets meshes, Gaussians, or point clouds suited for AR/VR, product turntables, game props.
🎬 ✍️ Sept 25 <Video & Text2Video> (Top Releases)
open T2V & animation models emphasizing temporal coherence, controllability, and real-time playback. Great starting point for creative tools, Ads.
Top Apache 2.0 License
Free and Open Source provided you don't source model and claim right
-
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 5.25M • • 6.11k -
facebook/wav2vec2-base-960h
Automatic Speech Recognition • 94.4M • Updated • 1.73M • 402 -
openai/whisper-small
Automatic Speech Recognition • 0.2B • Updated • 2.22M • 578 -
openai/whisper-tiny
Automatic Speech Recognition • 37.8M • Updated • 1.83M • 437
✍️➡️🎬 Text-to-Video
Models that create short videos from written prompts. Perfect for experimentation in generative video and creative storytelling.
🖌️ Image-to-Image
Image editing and transformation models :- from style transfer to super-resolution, inpainting, and diffusion-based edits.
-
stabilityai/stable-diffusion-xl-refiner-1.0
Image-to-Image • 2B • Updated • 119k • 2.06k -
black-forest-labs/FLUX.1-Kontext-dev
Image-to-Image • 12B • Updated • 224k • • 2.76k -
Qwen/Qwen-Image-Edit
Image-to-Image • 20B • Updated • 109k • • 2.47k -
lllyasviel/sd-controlnet-canny
Image-to-Image • 0.4B • Updated • 31.8k • 249
🖼️➡️📚 Image-Text-to-Text
Multimodal models that take image + text as input and produce natural language output. Use cases: chart QA, visual document reasoning, VQA.
-
Qwen/Qwen2.5-VL-7B-Instruct
Image-Text-to-Text • 8B • Updated • 8.98M • • 1.67k -
Qwen/Qwen2.5-VL-3B-Instruct
Image-Text-to-Text • 4B • Updated • 7.55M • • 683 -
google/gemma-3-4b-it
Image-Text-to-Text • 4B • Updated • 1.67M • • 1.44k -
nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1
Image-Text-to-Text • 9B • Updated • 1.29M • 181
✍️ Text Generation
Collection of top open LLMs for writing, summarization, chat, reasoning, and document drafting. Includes small SLMs for devices and large models .
🧠General Purpose Dataset < 10M samples
Dataset that can 🌐chat, ⚡code and 🧮reasoning
🍎 MLX-Ready LLMs
MLX weights and proven for MLX inference
-
mlx-community/gpt-oss-20b-MXFP4-Q8
Text Generation • 21B • Updated • 339k • 84 -
lmstudio-community/Seed-OSS-36B-Instruct-MLX-4bit
Text Generation • 36B • Updated • 37.3k • 1 -
lmstudio-community/Qwen3-4B-Thinking-2507-MLX-4bit
Text Generation • 0.6B • Updated • 53.1k • 14 -
mlx-community/parakeet-tdt-0.6b-v2
Automatic Speech Recognition • 0.6B • Updated • 1.85M • 45
📱 OnDevice -Ready SLMs (≤4B)
Tiny, fast models that run on iPhone/iPad or Mac with very low memory. Great for quick replies, offline note-assist, and routing
-
lmstudio-community/Qwen3-4B-Thinking-2507-MLX-8bit
Text Generation • 1B • Updated • 52.6k • 8 -
lmstudio-community/DeepSeek-R1-0528-Qwen3-8B-MLX-4bit
Text Generation • 1B • Updated • 274k • 14 -
lmstudio-community/gemma-3n-E4B-it-MLX-4bit
Image-Text-to-Text • Updated • 37.4k • 3 -
mlx-community/gemma-3-4b-it-qat-4bit
Image-Text-to-Text • 0.9B • Updated • 60.1k • 9
GPT2-JungleBook-from-Scratch-Models
The primary objective of project is to explore & analyze the impact of model size on text generation quality with GPT-2 arch trained from scratch.
Vision-LM
<7B Best of MoE 🧠
Collection of Small size big impact MoE.
-
LiquidAI/LFM2-8B-A1B
Text Generation • 8B • Updated • 30.3k • 371 -
ibm-granite/granite-4.0-h-tiny
Text Generation • 7B • Updated • 270k • 206 -
microsoft/Phi-4-multimodal-instruct
Automatic Speech Recognition • 6B • Updated • 566k • 1.61k -
google/gemma-3n-E4B-it
Image-Text-to-Text • 8B • Updated • 36.3k • • 919
Audio Features
Feature Extraction with 🧠 Text Embeddings
models for turning text, images, audio (and combos) into useful vectors or feature maps. Ideal for search/RAG, clustering, recommendation, retrieval.
🪶 Sept’25 <Text Generation Language Models >(Top Releases)
coding models and pipelines released this month that boost repo-level reasoning, GUI automation, and tool use. Focused on practical editing.
🖼️ **Text2Image, i2i ** September ’25 (Top Releases)
Cutting-edge image generation & VLM updates from September ’25. This collection spotlights models that improved text rendering, layout control & more.
📄➡️🔊 Text-to-Speech (TTS)
Speech synthesis models that turn text into natural audio. Includes multilingual TTS, low-latency real-time models, and voice-cloning variants.
📚➡️🎨Text-to-Image
State-of-the-art diffusion and generative models that turn text prompts into detailed images. Includes lightweight CPU-friendly and photorealistic mdl
-
stable-diffusion-v1-5/stable-diffusion-v1-5
Text-to-Image • 0.9B • Updated • 1.51M • 1.23k -
stabilityai/stable-diffusion-xl-base-1.0
Text-to-Image • 3B • Updated • 1.4M • • 8.02k -
stabilityai/sd-turbo
Text-to-Image • 0.9B • Updated • 693k • 457 -
black-forest-labs/FLUX.1-dev
Text-to-Image • 12B • Updated • 513k • • 14k
🎨➡️✍️ Image-to-Text
OCR, captioning, and visual QA models that turn pure images into descriptive or structured text.
-
Salesforce/blip-image-captioning-base
Image-to-Text • Updated • 1.99M • 875 -
Salesforce/blip-image-captioning-large
Image-to-Text • 0.5B • Updated • 535k • 1.48k -
nlpconnect/vit-gpt2-image-captioning
Image-to-Text • Updated • 90.9k • 932 -
microsoft/trocr-base-handwritten
Image-to-Text • 0.3B • Updated • 215k • 505
🌀 Any-to-Any Multimodal Models
Models that can flexibly convert across modalities (text, image, audio, video). Ideal for researchers exploring unified multimodal-AI.
👨💻Mathematical Reasoning 🧮
Datasets tackling AI Toughest Challenges
🧩 Long-Context Models (≥128k) CODING
10 CODING models that support ≥128k context (native or via officially documented scaling)
-
meta-llama/Llama-3.1-8B-Instruct
Text Generation • 8B • Updated • 7.67M • • 6.54k -
google/gemma-3-4b-it
Image-Text-to-Text • 4B • Updated • 1.67M • • 1.44k -
Qwen/Qwen3-Coder-30B-A3B-Instruct
Text Generation • 31B • Updated • 1.36M • • 1.19k -
deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct
Text Generation • 16B • Updated • 542k • 628
🧩 Long-Context Models (≥128k) under 8B
Qwen3
Best of Qwen3 Series of Models
-
Qwen/Qwen3-30B-A3B-Instruct-2507
Text Generation • 31B • Updated • 1.79M • • 824 -
Qwen/Qwen3-Next-80B-A3B-Thinking
Text Generation • 81B • Updated • 34.8k • • 492 -
Qwen/Qwen3-Coder-30B-A3B-Instruct
Text Generation • 31B • Updated • 1.36M • • 1.19k -
Qwen/Qwen3-Omni-30B-A3B-Instruct
Any-to-Any • 35B • Updated • 1.26M • 971
Geospatial
Vision-LM
🛩️Qwen3-VL
the most powerful vision-language model in the Qwen series to date. Available in Dense and MoE architectures
-
Qwen/Qwen3-VL-30B-A3B-Thinking
Image-Text-to-Text • 31B • Updated • 66.5k • • 200 -
mlx-community/Qwen3-VL-30B-A3B-Instruct-4bit
Image-Text-to-Text • Updated • 919 • 8 -
mlx-community/Qwen3-VL-30B-A3B-Instruct-8bit
Image-Text-to-Text • Updated • 233 • 3 -
mlx-community/Qwen3-VL-8B-Instruct-4bit
Image-Text-to-Text • Updated • 2.38k • 6
<7B Best of MoE 🧠
Collection of Small size big impact MoE.
-
LiquidAI/LFM2-8B-A1B
Text Generation • 8B • Updated • 30.3k • 371 -
ibm-granite/granite-4.0-h-tiny
Text Generation • 7B • Updated • 270k • 206 -
microsoft/Phi-4-multimodal-instruct
Automatic Speech Recognition • 6B • Updated • 566k • 1.61k -
google/gemma-3n-E4B-it
Image-Text-to-Text • 8B • Updated • 36.3k • • 919
🍎 MLX-Quantized Models (3/4/5/6-bit) Mac & iOS
Curated MLX-ready quantized LLMs that run fast on Apple Silicon (and some on iOS). Every card lists Bits · Group size · Peak UM (GB) · Stable context.
-
mlx-community/Apriel-1.5-15b-Thinker-3bit-MLX
Image-Text-to-Text • Updated • 11 -
mlx-community/Apriel-1.5-15b-Thinker-6bit-MLX
Image-Text-to-Text • Updated • 35 • 1 -
mlx-community/granite-4.0-h-tiny-3bit-MLX
Text Generation • 0.9B • Updated • 76 • 2 -
mlx-community/granite-4.0-tiny-preview-4bit
Text Generation • 1B • Updated • 86
Audio Features
🖼️ Vision Backbones & Image Embeddings
-
facebook/dinov2-base
Image Feature Extraction • 86.6M • Updated • 2.4M • 188 -
openai/clip-vit-large-patch14-336
Zero-Shot Image Classification • Updated • 4.12M • 308 -
google/siglip-so400m-patch14-384
Zero-Shot Image Classification • 0.9B • Updated • 2.28M • 682 -
BAAI/EVA-CLIP-8B
Feature Extraction • Updated • 1.06k • 51
Feature Extraction with 🧠 Text Embeddings
models for turning text, images, audio (and combos) into useful vectors or feature maps. Ideal for search/RAG, clustering, recommendation, retrieval.
🧊Sept 25 <Image-to-3D> [Top Releases]
Models that turn a single image (or image+prompt) into 3D assets meshes, Gaussians, or point clouds suited for AR/VR, product turntables, game props.
🪶 Sept’25 <Text Generation Language Models >(Top Releases)
coding models and pipelines released this month that boost repo-level reasoning, GUI automation, and tool use. Focused on practical editing.
🎬 ✍️ Sept 25 <Video & Text2Video> (Top Releases)
open T2V & animation models emphasizing temporal coherence, controllability, and real-time playback. Great starting point for creative tools, Ads.
🖼️ **Text2Image, i2i ** September ’25 (Top Releases)
Cutting-edge image generation & VLM updates from September ’25. This collection spotlights models that improved text rendering, layout control & more.
Top Apache 2.0 License
Free and Open Source provided you don't source model and claim right
-
openai/whisper-large-v3
Automatic Speech Recognition • 2B • Updated • 5.25M • • 6.11k -
facebook/wav2vec2-base-960h
Automatic Speech Recognition • 94.4M • Updated • 1.73M • 402 -
openai/whisper-small
Automatic Speech Recognition • 0.2B • Updated • 2.22M • 578 -
openai/whisper-tiny
Automatic Speech Recognition • 37.8M • Updated • 1.83M • 437
📄➡️🔊 Text-to-Speech (TTS)
Speech synthesis models that turn text into natural audio. Includes multilingual TTS, low-latency real-time models, and voice-cloning variants.
✍️➡️🎬 Text-to-Video
Models that create short videos from written prompts. Perfect for experimentation in generative video and creative storytelling.
📚➡️🎨Text-to-Image
State-of-the-art diffusion and generative models that turn text prompts into detailed images. Includes lightweight CPU-friendly and photorealistic mdl
-
stable-diffusion-v1-5/stable-diffusion-v1-5
Text-to-Image • 0.9B • Updated • 1.51M • 1.23k -
stabilityai/stable-diffusion-xl-base-1.0
Text-to-Image • 3B • Updated • 1.4M • • 8.02k -
stabilityai/sd-turbo
Text-to-Image • 0.9B • Updated • 693k • 457 -
black-forest-labs/FLUX.1-dev
Text-to-Image • 12B • Updated • 513k • • 14k
🖌️ Image-to-Image
Image editing and transformation models :- from style transfer to super-resolution, inpainting, and diffusion-based edits.
-
stabilityai/stable-diffusion-xl-refiner-1.0
Image-to-Image • 2B • Updated • 119k • 2.06k -
black-forest-labs/FLUX.1-Kontext-dev
Image-to-Image • 12B • Updated • 224k • • 2.76k -
Qwen/Qwen-Image-Edit
Image-to-Image • 20B • Updated • 109k • • 2.47k -
lllyasviel/sd-controlnet-canny
Image-to-Image • 0.4B • Updated • 31.8k • 249
🎨➡️✍️ Image-to-Text
OCR, captioning, and visual QA models that turn pure images into descriptive or structured text.
-
Salesforce/blip-image-captioning-base
Image-to-Text • Updated • 1.99M • 875 -
Salesforce/blip-image-captioning-large
Image-to-Text • 0.5B • Updated • 535k • 1.48k -
nlpconnect/vit-gpt2-image-captioning
Image-to-Text • Updated • 90.9k • 932 -
microsoft/trocr-base-handwritten
Image-to-Text • 0.3B • Updated • 215k • 505
🖼️➡️📚 Image-Text-to-Text
Multimodal models that take image + text as input and produce natural language output. Use cases: chart QA, visual document reasoning, VQA.
-
Qwen/Qwen2.5-VL-7B-Instruct
Image-Text-to-Text • 8B • Updated • 8.98M • • 1.67k -
Qwen/Qwen2.5-VL-3B-Instruct
Image-Text-to-Text • 4B • Updated • 7.55M • • 683 -
google/gemma-3-4b-it
Image-Text-to-Text • 4B • Updated • 1.67M • • 1.44k -
nvidia/Llama-3.1-Nemotron-Nano-VL-8B-V1
Image-Text-to-Text • 9B • Updated • 1.29M • 181
🌀 Any-to-Any Multimodal Models
Models that can flexibly convert across modalities (text, image, audio, video). Ideal for researchers exploring unified multimodal-AI.
✍️ Text Generation
Collection of top open LLMs for writing, summarization, chat, reasoning, and document drafting. Includes small SLMs for devices and large models .
👨💻Mathematical Reasoning 🧮
Datasets tackling AI Toughest Challenges
🧠General Purpose Dataset < 10M samples
Dataset that can 🌐chat, ⚡code and 🧮reasoning
🧩 Long-Context Models (≥128k) CODING
10 CODING models that support ≥128k context (native or via officially documented scaling)
-
meta-llama/Llama-3.1-8B-Instruct
Text Generation • 8B • Updated • 7.67M • • 6.54k -
google/gemma-3-4b-it
Image-Text-to-Text • 4B • Updated • 1.67M • • 1.44k -
Qwen/Qwen3-Coder-30B-A3B-Instruct
Text Generation • 31B • Updated • 1.36M • • 1.19k -
deepseek-ai/DeepSeek-Coder-V2-Lite-Instruct
Text Generation • 16B • Updated • 542k • 628
🍎 MLX-Ready LLMs
MLX weights and proven for MLX inference
-
mlx-community/gpt-oss-20b-MXFP4-Q8
Text Generation • 21B • Updated • 339k • 84 -
lmstudio-community/Seed-OSS-36B-Instruct-MLX-4bit
Text Generation • 36B • Updated • 37.3k • 1 -
lmstudio-community/Qwen3-4B-Thinking-2507-MLX-4bit
Text Generation • 0.6B • Updated • 53.1k • 14 -
mlx-community/parakeet-tdt-0.6b-v2
Automatic Speech Recognition • 0.6B • Updated • 1.85M • 45
🧩 Long-Context Models (≥128k) under 8B
📱 OnDevice -Ready SLMs (≤4B)
Tiny, fast models that run on iPhone/iPad or Mac with very low memory. Great for quick replies, offline note-assist, and routing
-
lmstudio-community/Qwen3-4B-Thinking-2507-MLX-8bit
Text Generation • 1B • Updated • 52.6k • 8 -
lmstudio-community/DeepSeek-R1-0528-Qwen3-8B-MLX-4bit
Text Generation • 1B • Updated • 274k • 14 -
lmstudio-community/gemma-3n-E4B-it-MLX-4bit
Image-Text-to-Text • Updated • 37.4k • 3 -
mlx-community/gemma-3-4b-it-qat-4bit
Image-Text-to-Text • 0.9B • Updated • 60.1k • 9
Qwen3
Best of Qwen3 Series of Models
-
Qwen/Qwen3-30B-A3B-Instruct-2507
Text Generation • 31B • Updated • 1.79M • • 824 -
Qwen/Qwen3-Next-80B-A3B-Thinking
Text Generation • 81B • Updated • 34.8k • • 492 -
Qwen/Qwen3-Coder-30B-A3B-Instruct
Text Generation • 31B • Updated • 1.36M • • 1.19k -
Qwen/Qwen3-Omni-30B-A3B-Instruct
Any-to-Any • 35B • Updated • 1.26M • 971
GPT2-JungleBook-from-Scratch-Models
The primary objective of project is to explore & analyze the impact of model size on text generation quality with GPT-2 arch trained from scratch.