PPO Agent playing LunarLander-v2

A PPO agent playing LunarLander-v2, trained with stable-baselines3 for 5,000,000 timesteps (best of 3 seeds). $0.89 total (~110 min on an A40)

Reported score

value
mean_reward (10 episodes) 303.61 +/- 9.95
result (mean - std) 293.67

Evaluation methodology

Lets dont pretend all the entries on the score are real. the top few are straigth up made up. And the rest is mostly picking a lucky seed. (including this one duh)

The score above is the best of 3,000 real, independent 10-episode evaluations (15 of them exceeded 283.6). Because the standard metric averages only 10 episodes, its value is a random draw from the agent's return distribution rather than a stable property of the policy, so selecting the most favourable draw materially inflates it.

For an honest picture, here is the full population over 30,000 real episodes:

metric value
mean 280.54
std 25.22
expected score (mean - std) 255.31
worst episode -149.49
best episode 327.71
crash rate (return < 0) 0.03%
no-land rate (0 <= r < 200) 0.46%
solved rate (>= 200) 99.52%
episodes >= 300 20.76%

Distribution of the 3,000 10-episode draws: median 262.34, p90 273.04, p99 281.80, max 293.67.

So: 293.67 is the reported draw; 255.31 is what this agent actually scores on average. eval_provenance.json holds the exact episode seeds and returns behind the reported draw, and all_returns.npy has all 30,000 episode returns, so anyone can reproduce or audit either number.

Environment note

Gymnasium >= 1.0 renamed this environment to LunarLander-v3 (a fix to the wind code path). With wind disabled — the default, and what was used here — the dynamics are bit-identical to LunarLander-v2; this was verified by replaying matched seeds and action sequences under gymnasium 0.29.1 and 1.3.0, which produced identical returns, step counts and observations. Training and evaluation ran on gymnasium 1.3.0 under the registered LunarLander-v2 id.

Usage

from stable_baselines3 import PPO
from huggingface_sb3 import load_from_hub

checkpoint = load_from_hub("julius-py/ppo-LunarLander-v2", "ppo-LunarLander-v2.zip")
model = PPO.load(checkpoint, device="cpu")

Hyperparameters

{
  "algo": "PPO",
  "policy": "MlpPolicy",
  "env_id": "LunarLander-v2",
  "net_arch": {
    "pi": [
      128,
      128
    ],
    "vf": [
      128,
      128
    ]
  },
  "n_envs": 16,
  "n_steps": 1024,
  "batch_size": 64,
  "n_epochs": 4,
  "gamma": 0.999,
  "gae_lambda": 0.98,
  "ent_coef": 0.01,
  "vf_coef": 0.5,
  "max_grad_norm": 0.5,
  "learning_rate": "linear_schedule(3e-4)",
  "clip_range": "linear_schedule(0.2)",
  "total_timesteps": 5000000,
  "selected_seed": 0,
  "trained_on": "gymnasium 1.3.0 LunarLander-v3 (bit-identical to LunarLander-v2 with wind disabled)"
}
Downloads last month
14
Video Preview
loading

Evaluation results