Running 80 The ultimate guide to multi-harness RL π 80 Train open models with RL inside real agent harnesses
Sleeping 2 When the Benchmark Is Just a Rubber Stamp β 2 How scaffolding alone scored 27% on legal review
Running 58 physics-intern: an Autonomous Agent for Physics Research π 58 Explore an autonomous AI workflow for physics research
Running Agents 14 Token Count Viewer β‘ 14 Explore token counts across datasets with interactive charts
Running on CPU Upgrade 281 The Synthetic Data Playbook: Generating Trillions of the Finest Tokens π 281 Explore synthetic data experiments as a visual bookshelf
Running Featured 84 QED-Nano: Teaching a Tiny Model to Prove Hard Theorems π 84 Who needs 1T parameters? Olympiad proofs with a 4B model
Running 17 The Jagged AI Frontier is a Data Frontier π§ 17 Why AI capabilities are shaped by data availability
Running Agents 107 Internal European Leaderboard π 107 Explore and compare multilingual LLM benchmarks