File size: 1,307 Bytes
9e3ffe2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
---
library_name: ml-agents
tags:
- reinforcement-learning
- deep-reinforcement-learning
- Pyramids
- ML-Agents-Pyramids
model-index:
- name: ppo-Pyramids
  results:
  - task:
      type: reinforcement-learning
      name: Reinforcement Learning
    dataset:
      name: ML-Agents-Pyramids
      type: ML-Agents-Pyramids
    metrics:
    - name: mean_reward
      type: mean_reward
      value: -0.9999999310821295 +/- 0.0
---
# PPO agent playing Pyramids

Trained from scratch for 63,991 steps in Google Colab using Unity ML-Agents, with coding and execution assistance from Codex. This is an introductory course model, not a fully converged policy.

## Evaluation
Evaluated using the exported ONNX policy with deterministic actions over 20 completed agent episodes, seed 12345. Mean reward: -0.9999999310821295; standard deviation: 0.0. Soccer rewards include the team reward. Full episode returns are in evaluation.json.

## Files
The ONNX file is the evaluated policy. configuration.yaml and config.json contain training settings. checkpoint.pt permits continued training; training.log records the actual run.

## References
- https://huggingface.co/learn/deep-rl-course/en/unit5/hands-on
- https://huggingface.co/learn/deep-rl-course/en/unit7/hands-on
- https://github.com/Unity-Technologies/ml-agents