PPO agent playing LunarLander-v2

This is a trained model of an PPO agent playing LunarLander-v2 for the Hugging Face Deep Reinforcement Learning Course (Unit 1).

Evaluation Results

  • Mean Reward: 285.50 +/- 15.20
  • Environment: LunarLander-v2
  • Algorithm: PPO
  • Library: stable-baselines3

Usage

Trained and evaluated for the Hugging Face Deep RL Course certification.

Downloads last month
5
Video Preview
loading

Evaluation results