|
Download README.md from SnEhAh018/ppo-agent: direct link, hf CLI and curl.
- Browser
- Download file 1.11 kB
-
https://huggingface.co/SnEhAh018/ppo-agent/resolve/main/README.md
- Command line
-
hf download hf://SnEhAh018/ppo-agent/README.md
-
curl -L -o README.md https://huggingface.co/SnEhAh018/ppo-agent/resolve/main/README.md
1.11 kB
| tags: | |
| - LunarLander-v2 | |
| - ppo | |
| - deep-reinforcement-learning | |
| - reinforcement-learning | |
| - custom-implementation | |
| - deep-rl-course | |
| model-index: | |
| - name: PPO | |
| results: | |
| - task: | |
| type: reinforcement-learning | |
| name: reinforcement-learning | |
| dataset: | |
| name: LunarLander-v2 | |
| type: LunarLander-v2 | |
| metrics: | |
| - type: mean_reward | |
| value: -112.34 +/- 57.15 | |
| name: mean_reward | |
| verified: false | |
| # PPO Agent Playing LunarLander-v2 | |
| ## Evaluation | |
| - Mean reward: -112.34 +/- 57.15 | |
| - Evaluation episodes: 10 | |
| - Environment: `LunarLander-v2` | |
| ## Hyperparameters | |
| ```python | |
| {'seed': 1, 'torch_deterministic': True, 'cuda': True, 'env_id': 'LunarLander-v2', 'total_timesteps': 50000, 'learning_rate': 0.00025, 'num_envs': 8, 'num_steps': 128, 'anneal_lr': True, 'gae': True, 'gamma': 0.99, 'gae_lambda': 0.95, 'num_minibatches': 4, 'update_epochs': 4, 'norm_adv': True, 'clip_coef': 0.2, 'clip_vloss': True, 'ent_coef': 0.01, 'vf_coef': 0.5, 'max_grad_norm': 0.5, 'target_kl': None, 'repo_id': 'YOUR_HF_USERNAME/ppo-LunarLander-v2', 'batch_size': 1024, 'minibatch_size': 256} | |
| ``` | |