view article Article Kimi K3 Model Overview: 2.8T Parameters, MXFP4 Quantization, and What the Open Weights Mean for the Community ResterChed • 11 days ago • 172
view article Article Native-speed vLLM transformers modeling backend hmellor, lysandre • 21 days ago • 60
Running on CPU Upgrade Featured 3.25k The Smol Training Playbook 📚 3.25k The secrets to building world-class LLMs
DFlash: Block Diffusion for Flash Speculative Decoding Paper • 2602.06036 • Published Feb 5 • 89
CohereLabs/command-a-plus-05-2026-bf16 Image-Text-to-Text • 219B • Updated Jun 15 • 67k • • 141
CohereLabs/command-a-plus-05-2026-w4a4 Image-Text-to-Text • 126B • Updated Jun 16 • 6.65k • • 235