PPO Agent Playing LunarLander-v2
This is a trained model of a PPO (Proximal Policy Optimization) agent playing LunarLander-v2 implemented from scratch with PyTorch for Unit 8 Part 1 of the Deep Reinforcement Learning Course: https://huggingface.co/deep-rl-course/unit8/introduction
- Downloads last month
- 16
Evaluation results
- mean_reward on LunarLander-v2self-reported292.50 +/- 11.20