Reinforce Agent playing Pixelcopter-PLE-v0

This is a trained REINFORCE agent playing Pixelcopter-PLE-v0.

This model was trained as part of Unit 4 of the Hugging Face Deep Reinforcement Learning Course.

Environment

Pixelcopter-PLE-v0

Algorithm

REINFORCE / Monte Carlo Policy Gradient

Hyperparameters

  • Hidden size: 64
  • Training episodes: 50000
  • Maximum steps per episode: 10000
  • Gamma: 0.99
  • Learning rate: 1e-4

Evaluation

  • Mean reward: -1.90
  • Standard deviation: 1.62
  • Course score (mean - std): -3.52

Course requirement: score >= 5.

Course: https://huggingface.co/learn/deep-rl-course/unit4/hands-on

Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Evaluation results