view article Article Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem MultiverseComputingCAI • 3 days ago • 27
LLM Compression by Block Removal with Constrained Binary Optimization Paper • 2602.00161 • Published Jun 17 • 9
Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal Paper • 2609.04482 • Published 21 days ago • 11
Quantization-Aware Healing: A Practical Recipe for Recovering Compressed, 4-Bit LLMs Paper • 2608.20953 • Published Aug 21 • 14
view article Article Quantization-Aware Healing: a compressed, 4-bit model that outperforms its full-precision original MultiverseComputingCAI • 30 days ago • 69
view article Article Making Knowledge Distillation Cheap Enough to Run at Scale MultiverseComputingCAI • Aug 10 • 42
Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss Paper • 2608.03796 • Published Aug 4 • 20
MultiverseComputingCAI/Qwen3-Next-80B-A3B-Thinking-Uncensored Text Generation • 80B • Updated Jun 2 • 175 • 17