view post Post 799 We compared 1-bit Kimi K3 to Claude Opus 5 and GPT 5.6. 🤯We gave 4 models the same prompt: Create a glass aquarium whose side panel develops a visible crack and then bursts...1-bit Kimi K3 GGUF ran locally on 4x B200s at 36 tok/s.GGUF: unsloth/Kimi-K3-GGUFGitHub repo: https://github.com/unslothai/unsloth See translation 1 reply · 👍 1 1 🔥 1 1 🤗 1 1 + Reply
view post Post 3636 Kimi K3 can now be run locally! ✨The 1-bit model retains ~78.9% accuracy after we shrunk it from 1.56TB to 594GB (-62% size).Run on a Mac Studio connected with 128GB RAM device. Kimi K3 is the strongest open model to date.GGUF: unsloth/Kimi-K3-GGUFGuide: https://unsloth.ai/docs/models/kimi-k3 See translation 5 replies · ❤️ 11 11 🔥 9 9 👍 3 3 🤗 1 1 + Reply
view post Post 4723 Introducing Unsloth for AMD 🚀You can now train & run LLMs on your AMD hardware• We collaborated with AMD to enable you to train & run 500+ models on AMD GPUs• Works on Windows, WSL, Linux• Train Qwen, Gemma on just 3GB VRAMGitHub: https://github.com/unslothai/unslothBlog + Guide: https://unsloth.ai/docs/basics/amd See translation 3 replies · 🔥 27 27 ❤️ 10 10 🚀 6 6 🤗 6 6 👀 5 5 👍 5 5 🧠 3 3 ➕ 3 3 🤯 3 3 🤝 2 2 😎 2 2 + Reply
view post Post 5790 Gemma 4 is now faster and much more accurate! 🚀Google made huge improvements to tool-calling and chat accuracy, reliability + speed.To get fixes, re-download our updated GGUF, MLX, NVFP4 quants!Unsloth quants: https://huggingface.co/collections/unsloth/gemma-4Gemma 4 Guide: https://unsloth.ai/docs/models/gemma-4 See translation 8 replies · 🚀 30 30 👍 15 15 😎 6 6 🤗 5 5 🤝 1 1 + Reply
view post Post 4674 We’re releasing Gemma 4 NVFP4 quants that run 1.5× faster on your GPU.Gemma-4-12B NVFP4 works on 11GB VRAM.26B-A4B hits 13K tok/s (B200).Unsloth NVFP4 enables faster, more accurate 4-bit Blackwell inference.Blog: https://unsloth.ai/docs/basics/nvfp4Gemma NVFP4: https://huggingface.co/collections/unsloth/nvfp4 See translation 3 replies · 🔥 12 12 🤗 3 3 🚀 2 2 👍 2 2 + Reply
CohereLabs/cohere-transcribe-arabic-07-2026 Automatic Speech Recognition • 2B • Updated 21 days ago • 49.1k • 142
view post Post 4258 We’re releasing new Qwen3.6 quants that run 2.5× faster on your GPU. ⚡Qwen3.6-27B NVFP4 runs on 24GB VRAM.35B-A3B can hit 17,561 tok/s (B200).We also improved accuracy, tool calling, agent use, and looping.Qwen3.6 NVFP4: https://huggingface.co/collections/unsloth/nvfp4Guide: https://unsloth.ai/docs/models/qwen3.6#nvfp4 See translation 1 reply · 🚀 15 15 🔥 11 11 🤗 1 1 + Reply
view post Post 6201 DeepSeek-V4 can now run locally with Unsloth GGUFs! 🐳Run lossless DeepSeek-V4-Flash on 168GB RAM or3-bit works on 110GB Mac, RAM, VRAM setups.Run via Unsloth Studio or llama.cpp.GGUF: unsloth/DeepSeek-V4-Flash-GGUFGuide: https://unsloth.ai/docs/models/deepseek-v4 See translation 🔥 20 20 🚀 5 5 👍 3 3 🤗 2 2 + Reply
Running on CPU Upgrade Agents 21 Cohere Transcribe Arabic ASR 🎙 21 Transcribe Arabic or English audio into text
CohereLabs/cohere-transcribe-arabic-07-2026 Automatic Speech Recognition • 2B • Updated 21 days ago • 49.1k • 142
Running on CPU Upgrade Agents 21 Cohere Transcribe Arabic ASR 🎙 21 Transcribe Arabic or English audio into text
view post Post 3367 1-bit GLM-5.2 GGUF vs. Claude 4.8 Opus vs. GPT-5.5We gave 3 models the same prompt and compared one-shot outputs.The 1-bit GLM-5.2 GGUF ran locally on a Mac Studio M3 Ultra with 256GB RAM at ~21.6 tok/s.Which output do you like best?GGUF: unsloth/GLM-5.2-GGUF See translation 3 replies · 🤗 11 11 👍 4 4 🔥 1 1 + Reply
view post Post 4611 Google's new DiffusionGemma can now run at 2000+ tokens/sec! ⚡We made local DiffusionGemma inference 1.8× faster.Run it on 18GB RAM via Unsloth Studio.GitHub: https://github.com/unslothai/unslothGuide: https://unsloth.ai/docs/models/diffusiongemma See translation 4 replies · 🔥 8 8 🤗 2 2 😔 1 1 + Reply