The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Paper • 2606.24937 • Published Jun 22 • 22
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • Sep 3 • 147
Running on Zero MCP 1 Chess Reasoning LLM ♟ 1 Play chess against a tiny reasoning LLM that thinks first
Pre2Post-Chess Collection Open-sourced models and datasets for training the chess reasoning models. • 8 items • Updated Aug 8 • 7
Running 255 The ultimate guide to RL environments: building and scaling them in the LLM era 📝 255 Building and scaling RL environments for LLM training
OpenEnv Environment Hub Collection Validated first-party canonical OpenEnv environments on Hugging Face Hub • 10 items • Updated Jun 24 • 12
Sleeping Agents 2 Qwen3-1.7B Wordle GRPO Training Dashboard 📈 2 Live training metrics for Qwen3-1.7B GRPO run on Wordle