Running 218 The ultimate guide to multi-harness RL π 218 Train open models with RL inside real agent harnesses
Running Featured 94 Flow Matching for Image Generation π 94 Visualize how noise transforms into images in your browser
Running Featured 71 How to turn a game into an RL environment π 71 From an idea to a trained 4B, with the dead ends left in
Running on CPU Upgrade 282 The Synthetic Data Playbook: Generating Trillions of the Finest Tokens π 282 Explore synthetic data benchmarks with an interactive bookshelf
Running Featured 95 Distilling 100B+ Models 40x Faster with TRL π 95 TRL distillation for 100B+ teachers, 40x faster
Running 259 The ultimate guide to RL environments: building and scaling them in the LLM era π 259 Building and scaling RL environments for LLM training