PPO Agent playing LunarLander-v2 from scratch
This is a trained model of a custom PPO agent implemented in PyTorch from scratch for Unit 8 Part 1 of the Hugging Face Deep Reinforcement Learning Course.
Evaluation Results
- Mean Reward: 215.30 +/- 22.10
- Threshold Required: >= -500
- Pass Status: Passed ✅
Evaluation results
- mean_reward on LunarLander-v2self-reported215.30 +/- 22.10