Add pipeline tag and links to paper, project page, and code
#1
by nielsr HF Staff - opened
README.md
CHANGED
|
@@ -3,11 +3,16 @@ tags:
|
|
| 3 |
- reinforcement-learning
|
| 4 |
- robotics
|
| 5 |
- dexterous-manipulation
|
|
|
|
| 6 |
---
|
| 7 |
|
| 8 |
# DexPolicy representative checkpoints
|
| 9 |
|
| 10 |
-
This repository contains a representative subset of the object-specific policies reported in the
|
|
|
|
|
|
|
|
|
|
|
|
|
| 11 |
|
| 12 |
## Checkpoints
|
| 13 |
|
|
@@ -24,4 +29,4 @@ The PPO and FPO pairs train for approximately 5M environment steps with matched
|
|
| 24 |
|
| 25 |
The PPO and GRPO `.zip` archives follow the Stable-Baselines3 checkpoint layout and include the policy, PyTorch variables, optimizer state, and library-version metadata. The FPO `.pt` files are PyTorch training checkpoints for the conditional-flow actor and critic. These policies are object-specific and are not presented as a single cross-object policy.
|
| 26 |
|
| 27 |
-
This is a compact representative release rather than the complete multi-object checkpoint suite.
|
|
|
|
| 3 |
- reinforcement-learning
|
| 4 |
- robotics
|
| 5 |
- dexterous-manipulation
|
| 6 |
+
pipeline_tag: robotics
|
| 7 |
---
|
| 8 |
|
| 9 |
# DexPolicy representative checkpoints
|
| 10 |
|
| 11 |
+
This repository contains a representative subset of the object-specific policies reported in the paper [DexPolicy: Scheduled Exploration for Trajectory-Guided Dexterous Manipulation](https://huggingface.co/papers/2610.00360).
|
| 12 |
+
Project page: [https://aigeeksgroup.github.io/DexPolicy/](https://aigeeksgroup.github.io/DexPolicy/)
|
| 13 |
+
Code: [https://github.com/AIGeeksGroup/DexPolicy](https://github.com/AIGeeksGroup/DexPolicy)
|
| 14 |
+
|
| 15 |
+
The released pairs use the YCB mustard-bottle relocation task and isolate the paper's exploration-scheduling intervention under PPO, GRPO, and FPO.
|
| 16 |
|
| 17 |
## Checkpoints
|
| 18 |
|
|
|
|
| 29 |
|
| 30 |
The PPO and GRPO `.zip` archives follow the Stable-Baselines3 checkpoint layout and include the policy, PyTorch variables, optimizer state, and library-version metadata. The FPO `.pt` files are PyTorch training checkpoints for the conditional-flow actor and critic. These policies are object-specific and are not presented as a single cross-object policy.
|
| 31 |
|
| 32 |
+
This is a compact representative release rather than the complete multi-object checkpoint suite.
|