ppo-LunarLander-v2 / README.md
umesh251's picture
Update solved score: 292.50 +/- 11.20 (Result: 281.30)
b84ffe5 verified
|
Raw History Blame Contribute Delete
1.58 kB
metadata
tags:
  - LunarLander-v2
  - ppo
  - deep-reinforcement-learning
  - reinforcement-learning
  - custom-implementation
  - deep-rl-course
model-index:
  - name: PPO
    results:
      - task:
          type: reinforcement-learning
          name: reinforcement-learning
        dataset:
          name: LunarLander-v2
          type: LunarLander-v2
        metrics:
          - type: mean_reward
            value: 292.50 +/- 11.20
            name: mean_reward
            verified: false

PPO Agent Playing LunarLander-v2

This is a trained model of a PPO (Proximal Policy Optimization) agent playing LunarLander-v2 implemented from scratch with PyTorch using the CleanRL framework.

๐Ÿ† Evaluation Results

  • Mean Reward: 292.50 +/- 11.20
  • Lower Bound Result (Mean - Std): 281.30
  • Episodes Evaluated: 10
  • Status: Solved (Environment threshold: 200.0)

Hyperparameters

{
  'anneal_lr': True,
  'batch_size': 2048,
  'capture_video': False,
  'clip_coef': 0.2,
  'clip_vloss': True,
  'cuda': True,
  'ent_coef': 0.01,
  'env_id': 'LunarLander-v2',
  'exp_name': 'ppo',
  'gae': True,
  'gae_lambda': 0.95,
  'gamma': 0.99,
  'learning_rate': 0.00025,
  'max_grad_norm': 0.5,
  'num_envs': 16,
  'num_minibatches': 4,
  'num_steps': 128,
  'repo_id': 'umesh251/ppo-LunarLander-v2',
  'seed': 1,
  'total_timesteps': 500000,
  'update_epochs': 4,
  'vf_coef': 0.5
}

How to use

import gym
import torch
from test_agent import Agent

device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
env = gym.make("LunarLander-v2")
agent = Agent(env).to(device)