A code-specialized MoE carved out of Qwen3.6-35B-A3B by pure expert pruning β no fine-tuning, no distillation. I profiled all 256 experts on balanced corpora plus targeted code benchmarks (LiveCodeBench + MultiPL-E), built a competence map with the code classes up-weighted 1.5Γ, and dropped the 72 weakest experts per layer (256β184, ~35Bβ27B). Router, attention, norms, the MTP head and the vision tower are all preserved; active params stay at A3B and routing is baked to top-10 (revert to top-8 anytime).
27B footprint, A3B speed, coding that punches well above its size β and the preserved MTP head gives you speculative decoding out of the box (text + vision).
πππ The largest ever dataset of co-folded 3D protein-ligand structures just dropped on HF!!
Meet SAIR (Structurally Augmented ICβ β Repository): 5M+ AI-generated complexes with experimentally measured drug potency data from SandboxAQ. πππ