432 GB of ultra-fast HBM4 and up to 23.3 TB/s of memory bandwidth on a single GPU š¤Æ.
Two weeks ago, we got early access to AMD's new Instinct MI455X, and our first goal was simple: make sure š¤ Transformers works on day one.
Over the past few weeks, we worked closely with the AMD team to validate the platform, enable Flash Attention, add torchcodec support for multimodal models, and resolve issues uncovered during testing.
The result: ā 99.5% success rate across our 24 core Transformers model architectures - already on par with our daily CI on previous AMD and NVIDIA platforms.
The hardware is just as exciting. With 432 GB of HBM per GPU, our early capacity experiments showed more than 3Ć the concurrent long-context requests compared to MI300, thanks to the much larger KV cache capacity.
A huge thanks to the AMD team for the early access and the great collaboration!
Huge news from MiniMax: weāve secured a $2B funding round, paired with a formal long-term commitment from our CEO IO to allocate 1% of total company equity from his personal holdings to support the global open-source AI community over the next four years.
This capital backs our continuous open model releases, community tooling and transparent frontier AI research. Weāre just getting started on our open-source roadmap toward accessible AGI.
If you build with open foundation models and want to push frontier AI together, come join us. Intelligence with Everyone. š
š®š³ New in my Hindi LLM Series: Gemma-4 E4B, fine-tuned for Hindi ā and it runs on your laptop's CPU. I fine-tuned Google's new Gemma-4 E4B on ~10k Hindi instruction pairs (AI4Bharat: anudesh + dolly) using Unsloth + LoRA, on a single L4 GPU. Then I ran an honest side-by-side eval: base Gemma-4 vs my fine-tune, across 25 Hindi prompts. The results were interesting š ā My fine-tune is more concise ā ask for "3 tips" and it gives exactly 3. Base writes a 1,200-character essay.
ā Pure native Hindi ā base keeps slipping into English ("ą¤øą¤ą¤¤ą„लित ą¤ą¤¹ą¤¾ą¤° (Eat a Balanced Diet)", "तारा (Star)"). My fine-tune stays in clean Hindi.
ā Tighter instruction-following ā ask for a "short message" and it gives one, not a menu of options. āļø And to be honest: base Gemma-4 is more detailed and comprehensive. I didn't build a "smarter" model ā I built a focused, Hindi-native, edge-friendly one that runs as a 5GB GGUF (Q4) on CPU. š Try it: