mlx-community/Kimi-Linear-48B-A3B-Instruct-5bit
Text Generation • 49B • Updated • 76 • 1
local, open-source models will play major part in future progress, but we can`t match "1.1M tokens/sec on just one rack of GB300 GPUs in our Azure fleet"