| --- |
| library_name: hd-ppo |
| tags: |
| - InvertedPendulum-v5 |
| - deep-reinforcement-learning |
| - reinforcement-learning |
| - hyperdimensional-computing |
| - fractional-power-encoding |
| - LTU-AI |
| model-index: |
| - name: Hybrid-HD-PPO |
| results: |
| - task: |
| type: reinforcement-learning |
| name: reinforcement-learning |
| dataset: |
| name: InvertedPendulum-v5 |
| type: InvertedPendulum-v5 |
| metrics: |
| - type: mean_reward |
| value: 1000.00 +/- 0.00 |
| name: mean_reward |
| verified: false |
| --- |
| |
| # **Hybrid-HD-PPO** Agent playing **InvertedPendulum-v5** |
|
|
| This is a trained **Hybrid-HD-PPO** (Hyperdimensional Proximal Policy Optimization) agent |
| playing **InvertedPendulum-v5** using **gradient-adaptive Fractional Power Encoding (FPE)** |
| with a prune-and-fine-tune pipeline. |
|
|
| Published by [LTU-AI](https://huggingface.co/LTU-AI). |
|
|
| ## Pipeline |
|
|
| 1. Train a teacher at **D=512** with gradient-adaptive single-beta FPE. |
| 2. Prune by actor-weight importance through **D=512 → 128 → 64**. |
| 3. Fine-tune each pruned checkpoint with PPO. |
|
|
| Published checkpoint: seed **123**, compact **D=64** model |
| (held-out eval mean reward **1000.00 ± 0.00**). |
|
|
| ## Usage |
|
|
| Install dependencies: |
|
|
| ```bash |
| pip install -r requirements.txt |
| ``` |
|
|
| Evaluate the local checkpoint: |
|
|
| ```bash |
| python enjoy.py --weights hdppo-InvertedPendulum-v5/weights.npz --episodes 10 |
| ``` |
|
|
| Render episodes: |
|
|
| ```bash |
| python enjoy.py --weights hdppo-InvertedPendulum-v5/weights.npz --render --episodes 3 |
| ``` |
|
|
| Record a replay video: |
|
|
| ```bash |
| python record_video.py --weights hdppo-InvertedPendulum-v5/weights.npz --output replay.mp4 |
| ``` |
|
|
| Load from Hugging Face Hub: |
|
|
| ```bash |
| python enjoy.py --weights LTU-AI/hdppo-InvertedPendulum-v5 --episodes 10 |
| ``` |
|
|
| ## Training pipeline |
|
|
| Reproduce the teacher → prune → fine-tune workflow: |
|
|
| ```bash |
| python train_hdppo.py |
| ``` |
|
|
| ## Hyperparameters |
|
|
| ```python |
| { |
| "env": "InvertedPendulum-v5", |
| "algo": "Hybrid-HD-PPO (HD actor + MLP critic, gradient-adaptive FPE)", |
| "teacher_D": 512, |
| "pruned_D": 64, |
| "beta_base": 0.7071, |
| "timesteps_per_stage": 1000000, |
| "seed": 123 |
| } |
| ``` |
|
|
| ## Environment Arguments |
|
|
| ```python |
| { |
| "render_mode": "rgb_array" |
| } |
| ``` |
|
|
| ## Model files |
|
|
| | File | Description | |
| |------|-------------| |
| | `hdppo-InvertedPendulum-v5/weights.npz` | Published actor (+ critic if HD) and FPE encoder (D=64) | |
| | `hdppo-InvertedPendulum-v5/weights_D512_teacher.npz` | Teacher checkpoint (D=512) | |
| | `replay.mp4` | Sample rollout video from the published min-D checkpoint | |
| | `results.json` | Evaluation summary for the published checkpoint | |
| | `results_D512_teacher.json` | Evaluation summary for the teacher | |
| | `config.yml` | Training hyperparameters | |
| | `train_hdppo.py` / training modules | Self-contained training code | |
|
|
| ## Citation |
|
|
| If you use this model, please cite the HD-PPO / Hybrid-HD-PPO work. |
|
|