Agente PPO jugando a LunarLander-v3

Entrenado con Stable-Baselines3 PPO y una pol铆tica MLP.

Recompensa media en 10 episodios deterministas: 252.88 +/- 17.76.

C贸mo cargar el modelo

Carga solo checkpoints de fuentes en las que conf铆es.

from huggingface_hub import hf_hub_download
from stable_baselines3 import PPO

checkpoint = hf_hub_download(repo_id='0GiS0/lunarlander-v3', filename='ppo-LunarLander-v3.zip')
model = PPO.load(checkpoint, device='cpu')
Downloads last month
35
Video Preview
loading

Evaluation results