ShinMK
ShinMK3
ยท
AI & ML interests
None yet
Recent Activity
reacted to ManniX-ITA's post with ๐ 11 days ago
๐ New release: Qwen3.6-27B-A3B-Coder
A code-specialized MoE carved out of Qwen3.6-35B-A3B by pure expert pruning โ no fine-tuning, no distillation. I profiled all 256 experts on balanced corpora plus targeted code benchmarks (LiveCodeBench +
MultiPL-E), built a competence map with the code classes up-weighted 1.5ร, and dropped the 72 weakest experts per layer (256โ184, ~35Bโ27B). Router, attention, norms, the MTP head and the vision tower are all
preserved; active params stay at A3B and routing is baked to top-10 (revert to top-8 anytime).
๐ Benchmarks (Q6_K, temp 0.6):
โข MultiPL-E 0.840
โข HumanEval 0.970
โข LiveCodeBench 0.688
โข GSM8K 0.970 ยท ARC-C 0.944 ยท AIME 0.733
โข GPQA-Diamond 0.773 ยท MATH-500 0.620 ยท IFEval 0.730
โข Average 0.808
27B footprint, A3B speed, coding that punches well above its size โ and the preserved MTP head gives you speculative decoding out of the box (text + vision).
๐ Model: https://huggingface.co/ManniX-ITA/Qwen3.6-27B-A3B-Coder
๐ฆ GGUF (+MTP): https://huggingface.co/ManniX-ITA/Qwen3.6-27B-A3B-Coder-MTP-GGUF
๐ฆ Ollama: https://ollama.com/mannix/qwen3.6-27b-a3b-coder updated a dataset 19 days ago
ShinMK3/Mega-Brain-Distill liked a dataset 20 days ago
ShinMK3/Mega-Brain-Distill