view article Article Training a coding agent using the OpenCode harness in remote HF sandboxes with TRL and OpenEnv sergiopaniego • Aug 5 • 30
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • Sep 3 • 147
MathSmith: Towards Extremely Hard Mathematical Reasoning by Forging Synthetic Problems with a Reinforced Policy Paper • 2508.05592 • Published Aug 7, 2025 • 6