deepseek-ai/DeepSeek-R1-Distill-Qwen-32B Text Generation • 33B • Updated Feb 24, 2025 • 434k • • 1.63k
princeton-nlp/Llama-3-8B-ProLong-64k-Instruct Text Generation • 8B • Updated Oct 31, 2024 • 8.64k • • 13
Running on CPU Upgrade Agents Featured 1.02k Model Memory Utility 🚀 1.02k Calculate GPU memory needed for training Hugging Face models