Instella-MoE ✨ Collection Family of fully open 16B MoE LLM with 2.8B active params per token, trained on AMD Instinct™ MI300 & MI325 GPUs. https://arxiv.org/abs/2609.00791 • 6 items • Updated 10 days ago • 19
Nemotron-Pre-Training-Datasets Collection Large scale pre-training datasets used in the Nemotron family of models. • 15 items • Updated Aug 11 • 195