Running 1 Combining LLMs Rarely Beats the Single Best Model ๐ฒ beta=P(all wrong): the co-failure ceiling on LLM ensembles
Running Agents 1 FlavourBench ๐ฒ Explore and compare LLM performance on the FlavourBench leaderboard
Sleeping Agents 1 Ready Cohorts Explorer ๐ Explore GPU-ready cohorts for deterministic agent control.
Running 1 The Physical AI Inference Gap in Batch-1 LLM Decode ๐ช Interactive companion to the batch-1 LLM decode paper