Qwen/Qwen3.6-35B-A3B
Image-Text-to-Text • 36B • Updated • 6.12M • • 2.59k
Small, capable models I run locally on a single RTX 3090 (Ollama / llama.cpp / transformers) — the backbone of self-hosted, sovereign AI.
Note Daily driver on the RTX 3090 (ollama qwen3.6:latest, Q4_K_M) - Qwen3.6 MoE, 35B total / ~3B active per token.
Note Qwen3.6-27B in ternary 2-bit — a 27B in 7 GB. Installed on the 3090 (Q2_0) next to its Q4_K_M sibling for quality A/B.
Note Official QAT GGUF of Gemma 4 26B MoE (4B active) — a 3090-friendly middleweight agent
Note Local vision on the RTX 3090 (runs as qwen3-vl:8b via ollama) — screenshots, docs, UI grounding