PPO Agent Playing LunarLander-v2

This is a trained model of a PPO (Proximal Policy Optimization) agent playing LunarLander-v2 implemented from scratch with PyTorch for Unit 8 Part 1 of the Deep Reinforcement Learning Course: https://huggingface.co/deep-rl-course/unit8/introduction

Downloads last month
16
Video Preview
loading

Evaluation results