georginio2000 commited on
Commit
ee2bf0b
·
verified ·
1 Parent(s): a8b406f

Add model card

Browse files
Files changed (1) hide show
  1. README.md +40 -0
README.md ADDED
@@ -0,0 +1,40 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ tags:
4
+ - robomimic
5
+ - diffusion-policy
6
+ - imitation-learning
7
+ - robosuite
8
+ ---
9
+
10
+ # Diffusion Policy — NutAssemblySquare (Square/PH/low-dim)
11
+
12
+ A robomimic `DiffusionPolicyUNet` checkpoint (UNet + DDPM noise scheduler,
13
+ action-chunked receding-horizon control: `observation_horizon=2`,
14
+ `action_horizon=8`, `prediction_horizon=16`) trained for 2000 epochs on
15
+ robomimic's public 200-demo Square/PH/low-dim dataset.
16
+
17
+ Rollout success rate (20 episodes, evaluated every 200 epochs):
18
+
19
+ | Epoch | 200 | 400 | 600 | 800 | 1000 | 1200 | 1400 | 1600 | **1800** | 2000 |
20
+ |---|---|---|---|---|---|---|---|---|---|---|
21
+ | Success | 85% | 85% | 85% | 70% | 75% | 90% | 85% | 85% | **95%** | 90% |
22
+
23
+ `checkpoints/model_epoch_1800_low_dim_success_0.95.pth` (95% success, first
24
+ epoch to hit the run's peak) is the checkpoint used as the frozen base
25
+ policy for downstream residual-RL experiments.
26
+
27
+ This is a second, parallel base-policy lineage alongside a BC-RNN baseline
28
+ trained on the same task/dataset, for a residual-RL project comparing how a
29
+ heuristic-triggered SAC residual correction interacts with each. Full
30
+ writeup, code, and results: https://github.com (see the project's own
31
+ README for the actual repo link — not filled in here since it isn't public).
32
+
33
+ Downloading a specific checkpoint:
34
+ ```python
35
+ from huggingface_hub import hf_hub_download
36
+ hf_hub_download(
37
+ "georginio2000/diffusion-square-nutassembly",
38
+ "checkpoints/model_epoch_1800_low_dim_success_0.95.pth",
39
+ )
40
+ ```