Sleeping 2 When the Benchmark Is Just a Rubber Stamp ⚖ 2 How scaffolding alone scored 27% on legal review
Running 57 physics-intern: an Autonomous Agent for Physics Research 📝 57 Explore an autonomous AI workflow for physics research
Running on CPU Upgrade 270 The Synthetic Data Playbook: Generating Trillions of the Finest Tokens 📝 270 Visualize synthetic‑data experiments as an interactive bookshelf
Running Featured 82 QED-Nano: Teaching a Tiny Model to Prove Hard Theorems 📝 82 Who needs 1T parameters? Olympiad proofs with a 4B model
Running 17 The Jagged AI Frontier is a Data Frontier 🧭 17 Why AI capabilities are shaped by data availability
Running Agents 105 Internal European Leaderboard 🌍 105 Explore and compare multilingual LLM benchmarks