DexPolicy / README.md
wwhyyyyyy's picture
Update representative checkpoint model card
575645a verified
|
Raw History Blame Contribute Delete
2.19 kB
---
tags:
- reinforcement-learning
- robotics
- dexterous-manipulation
---
# DexPolicy representative checkpoints
This repository contains a representative subset of the object-specific policies reported in the DexPolicy paper. The released pairs use the YCB mustard-bottle relocation task and isolate the paper's exploration-scheduling intervention under PPO, GRPO, and FPO.
## Checkpoints
| File | Method | Object | Seed | Exploration during training |
|---|---|---|---:|---|
| `models/ppo_mustard_baseline_seed0.zip` | PPO baseline | YCB 006 mustard bottle | 0 | Learned policy standard deviation |
| `models/ppo_mustard_dexpolicy_seed0.zip` | PPO + DexPolicy | YCB 006 mustard bottle | 0 | Piecewise schedule from `0.20` to a final `0.04` plateau |
| `models/grpo_mustard_baseline_seed0.zip` | GRPO baseline | YCB 006 mustard bottle | 0 | Inherited, learned policy standard deviation |
| `models/grpo_mustard_dexpolicy_seed0.zip` | GRPO + DexPolicy | YCB 006 mustard bottle | 0 | Scheduled standard deviation, `0.10 -> 0.03` |
| `models/fpo_mustard_baseline_seed0.pt` | FPO baseline | YCB 006 mustard bottle | 0 | Fixed standard deviation, `0.10` |
| `models/fpo_mustard_dexpolicy_seed0.pt` | FPO + DexPolicy | YCB 006 mustard bottle | 0 | Linear schedule, `0.20 -> 0.05` |
The PPO and FPO pairs train for approximately 5M environment steps with matched settings within each pair. PPO uses an MLP actor mean; FPO uses a conditional-flow actor mean. The GRPO pair continues from matched object-and-seed checkpoints at 507,904 environment steps for another 250,000 steps. It uses the same reward, terminal success definition, curriculum stage, and return-to-go credit module; the exploration schedule is the controlled difference.
The PPO and GRPO `.zip` archives follow the Stable-Baselines3 checkpoint layout and include the policy, PyTorch variables, optimizer state, and library-version metadata. The FPO `.pt` files are PyTorch training checkpoints for the conditional-flow actor and critic. These policies are object-specific and are not presented as a single cross-object policy.
This is a compact representative release rather than the complete multi-object checkpoint suite.