view article Article Scaling Pedagogical Pre-training: From Optimal Mixing to 10 Billion Tokens codelion • Mar 6 • 8
view article Article The 1 Billion Token Challenge: Finding the Perfect Pre-training Mix codelion • Nov 3, 2025 • 67
view article Article The Optimal Architecture for Small Language Models codelion • Dec 26, 2025 • 123
Running on CPU Upgrade Featured 3.31k The Smol Training Playbook 📚 3.31k The secrets to building world-class LLMs
Running Agents Featured 117 Qwen TTS Demo 💻 117 Generate spoken audio from text using selectable voices
Qwen2.5 Collection Qwen2.5 language models, including pretrained and instruction-tuned models of 7 sizes, including 0.5B, 1.5B, 3B, 7B, 14B, 32B, and 72B. • 43 items • Updated Mar 2 • 737
Build error Agents Featured 1.43k SadTalker 😭 1.43k Generate a talking face video from an image and audio
Running on CPU Upgrade 14.1k Open LLM Leaderboard 🏆 14.1k Track, rank and evaluate open LLMs and chatbots