view article Article Harness, Scaffold, and the AI Agent Terms Worth Getting Right sergiopaniego, ariG23498 • May 25 • 135
Running Featured 81 QED-Nano: Teaching a Tiny Model to Prove Hard Theorems 📝 81 Who needs 1T parameters? Olympiad proofs with a 4B model
Running 84 Maintain the unmaintainable 📚 84 Explore the complex relationships between 400+ machine learning models
Running 95 Scaling FineWeb to 1000+ languages: Step 1: finding signal in 100s of evaluation tasks 📝 95 Evaluate multilingual models using FineTasks
Running on CPU Upgrade 268 The Synthetic Data Playbook: Generating Trillions of the Finest Tokens 📝 268 Visualize synthetic‑data experiments as an interactive bookshelf
Running on CPU Upgrade 14.1k Open LLM Leaderboard 🏆 14.1k Track, rank and evaluate open LLMs and chatbots
Running on CPU Upgrade Featured 3.25k The Smol Training Playbook 📚 3.25k The secrets to building world-class LLMs
view article Article Open-R1: a fully open reproduction of DeepSeek-R1 +1 eliebak, lvwerra, lewtun • Jan 28, 2025 • 891
Running 3.96k The Ultra-Scale Playbook 🌌 3.96k The ultimate guide to training LLM on large GPU Clusters
view article Article Open-source DeepResearch – Freeing our search agents +3 m-ric, albertvillanova, merve, thomwolf, clefourrier • Feb 4, 2025 • 1.32k
Scaling Test-Time Compute with Open Models Collection Models and datasets used in our blog post: https://huggingface.co/spaces/HuggingFaceH4/blogpost-scaling-test-time-compute • 10 items • Updated Jan 6, 2025 • 31