PPO agent playing SnowballTarget

Trained from scratch for 91,608 steps in Google Colab using Unity ML-Agents, with coding and execution assistance from Codex. This is an introductory course model, not a fully converged policy.

Evaluation

Evaluated using the exported ONNX policy with deterministic actions over 20 completed agent episodes, seed 12345. Mean reward: 24.75; standard deviation: 2.384848003542364. Soccer rewards include the team reward. Full episode returns are in evaluation.json.

Files

The ONNX file is the evaluated policy. configuration.yaml and config.json contain training settings. checkpoint.pt permits continued training; training.log records the actual run.

References

Downloads last month
6
Video Preview
loading

Evaluation results

  • mean_reward on ML-Agents-SnowballTarget
    self-reported
    24.75 +/- 2.384848003542364