File size: 9,229 Bytes
5d08972
 
 
 
 
 
 
 
 
 
28ed187
5d08972
 
28ed187
 
5d08972
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
28ed187
 
 
5d08972
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
28ed187
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
5d08972
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
# pose6daug β€” augmentation, dataset build, training and evaluation scripts

The scripts behind the MimicGen and VACE augmentation baselines for RoboCasa
PickPlaceCounterToCabinet, and the GR00T 1.5 fine-tuning and exact-replay evaluation run on
top of them. Archived as-run: paths are absolute and point at one particular machine, so treat
this as a record of the procedure rather than a turnkey package. See **Paths to change** below.

```
augment/mimicgen/   generate episodes with MimicGen, render preview videos
augment/vace/       VACE object-swap augmentation (assignment builder, runner, GT masks)
augment/actaug/     actaug episode folders -> gr00t_views (no simulator needed)
dataset/            convert generated episodes into a gr00t_views (LeRobot v2.1) dataset
train/              GR00T 1.5 fine-tuning launcher
eval/               exact-state replay evaluation -- read eval/EVALUATION.md for the protocol
ops/                checkpoint retention, optimizer pruning, eval-on-checkpoint triggers
tools/              dataset sanity checks
```

## Pipeline

```
MimicGen generation ──► demo.hdf5 (MuJoCo states + actions, no pixels)
                          β”‚
                β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
                β–Ό                    β–Ό
       preview mp4 (3 views)   gr00t_views dataset  ──► fine-tune ──► exact-replay eval
```

VACE skips the simulator: it repaints pixels and keeps the source trajectory, so its builder
copies the source parquet instead of replaying.

### 1. Augment

`augment/mimicgen/run_generation.sh` launches one MimicGen process per worker from a JSON
config per (object, worker); `config_template.json` is the template. Two knobs matter:

- `guarantee=false` makes `num_trials` mean **attempts**, not successes. Giving every object
  the same attempt budget is what produces the generator's natural yield instead of a
  per-object quota.
- `obj_registries` must include `aigen` for objects under `aigen_objs/` (e.g. `wine_5`), or
  `sample_kitchen_object_helper` raises a bare `ValueError`.

`snapshot_episode_times.sh` records the per-episode temporary filenames while they exist:
MimicGen writes each success to `tmp/date_..._time_HH_MM_SS.hdf5`, then `merge_all_hdf5`
sorts by timestamp and deletes the folder. That snapshot is the only record of when each
episode was produced, and `select_by_generation_time.py` joins it back to order episodes
globally across workers. The poller can only ever miss a worker's **last** file, which the
selector pads with the merged file's mtime.

### 2. Build the dataset

`dataset/render_to_gr00t.py` replays each episode in MuJoCo, renders three cameras and writes
parquet + videos + meta. `dataset/run_convert.sh` shards it (each shard needs its own
`GEN_DIR` and `DATASET_OUT`, or the glob picks up the others), then
`merge_gr00t_view_datasets.py` merges and `repair_mimicgen_task_ids.py` fixes task ids.

`repair_mimicgen_task_ids.py` is not optional: the writer stores every parquet task column as
0 while `episodes.jsonl` holds the intended language, so without it every episode trains as
task 0 and the language conditioning silently collapses.

### 3. Train

```bash
DATASET_PATH=<gr00t_views dataset> DATASETNAME=<name> \
GPUS=2,3 PER_GPU_BATCH=32 MAX_STEPS=30000 SAVE_STEPS=5000 \
bash train/train_groot15_single_dataset.sh
```

Fine-tunes the action-head projector and diffusion head from a base checkpoint; the backbone
stays frozen. `RESUME=1` picks up the newest checkpoint in the output directory.

### 4. Evaluate

The full protocol -- episode set, seeds, which checkpoints to compare, how to read the
stage flags -- is in [`eval/EVALUATION.md`](eval/EVALUATION.md). The short version:

```bash
MODEL_PATH=<checkpoint> MYGROOT_ROOT=<myGR00T tree> \
EVAL_CLIENT=<eval/eval_robocasa_replay_state_grasp.py> \
REPLAY_STATE_ROOT=<replay set> N_EPISODES=160 GPUS=4,5,6,7 POLICY_SEED=12345 SEED_BASE=42 \
bash eval/eval_groot15_exact_replay.sh
```

Each episode restores a saved scene XML and flattened MuJoCo state, so runs are comparable
across checkpoints. Two independent seeds:

- `POLICY_SEED` β€” the policy server's global RNG **and** the per-step action seed
  (`policy_seed + episode_index Γ— stride + step`). This is the one that changes the rollout.
- `SEED_BASE` β€” the env seed, `SEED_BASE + worker_id` per worker. Exact replay overwrites the
  scene immediately after, so it should not affect the initial state.

Worker count follows the GPU list, and episodes are split evenly across workers β€” so the same
`SEED_BASE` with a different GPU count gives each episode a different worker seed.

Client variants:

| file | adds |
|---|---|
| `eval_robocasa_replay_state_3view.py` | saves the wrist (ego) view; `RecordVideo` only captures `robot0_agentview_center` |
| `eval_robocasa_replay_state_grasp.py` | the above, plus `--episode_indices` for an arbitrary subset, and per-episode `grasped` / `lifted` / `in_cab` stage flags |

Stock RoboCasa success is `obj_inside_of(cab) and gripper_obj_far` β€” a single boolean, which
says nothing about where a failed episode broke down. The stage flags come from
`_check_grasp`, a 3 cm rise in the object's body z, and `OU.obj_inside_of`.

`organize_videos.py` regroups the output into one folder per global episode
(`rollouts/episode_NNNNNN/{center.mp4, wrist.mp4, info.txt}`); `RecordVideo` names files by
worker-local index, which does not match the global index.

## The action-order bug β€” check any dataset before training on it

RoboCasa's simulator exports actions **arm-first**:

```
[eef_pos(3), eef_rot(3), gripper, base(3), torso, base_mode]
```

A gr00t_views dataset declares them **base-first**:

```
base_motion[0:4]  control_mode[4:5]  eef_pos[5:8]  eef_rot[8:11]  gripper_close[11:12]
```

A builder that re-derives actions from a simulator rollout has to reorder; one that copies
rows from an existing LeRobot dataset does not. Copying them through unchanged puts arm motion
on the base-motion channel and `base_mode` on the gripper. Nothing errors, the loss converges
to a small value, and the policy drives the base away from the counter and never closes the
gripper β€” **0/160** on an exact-replay eval whose base checkpoint scored 11/160.

```python
actions = raw_actions[:, [7, 8, 9, 10, 11, 0, 1, 2, 3, 4, 5, 6]]
```

`tools/check_action_layout.py` tells the two layouts apart from the data alone (the mobile
base never moves in this task, so the constant dimensions give it away), and
`dataset/repair_action_order.py` fixes an already-built dataset in place β€” parquet plus the
action entry of `meta/stats.json`, no re-render. It refuses to run on a dataset that does not
look like simulator order, so it cannot be applied twice. Read `tools/ACTION_LAYOUT.md` first.

## Paths to change

Every script hard-codes absolute paths from the machine this was run on. At minimum:

| what | appears as |
|---|---|
| RoboCasa / robosuite checkouts | `/lp-dev/jonghoon/robocasa_calib/repos/...` |
| MimicGen env + augmentation code | `/lp-dev/jonghoon/mimicgen_augment/...` |
| myGR00T tree and conda envs | `/data/minha/pose6daug/train_robocasa/myGR00T`, `/data/nvidia/gripper_augmentator/conda-envs/...` |
| base checkpoint | `/lp-dev/jonghoon/myGR00T_outputs/pnpcountertocab_all502_.../checkpoint-60000` |
| replay sets and eval output | `/lp-dev/jonghoon/isaac-gr00t/eval_results/...` |
| generation scratch | `/tmp/claude-.../scratchpad/...` |

`eval/eval_groot15_exact_replay.sh` defaults `MYGROOT_ROOT` to a path that no longer exists;
pass it explicitly. No credentials are embedded β€” the training launcher reads
`WANDB_API_KEY` from an env var or a file path you supply.

## Attribution

`eval/eval_groot15_exact_replay.sh` and the upstream of `dataset/render_to_gr00t.py` and the
eval clients come from the pose6daug project's shared tree; the copies here carry the fixes
described above (action reorder, mesh-swap bbox fix, target sharding, wrist-view capture,
stage logging).

## actaug

`augment/actaug/build_actaug_gr00t.py` converts actaug's per-episode folders
(`actions.npy`, `states.npz`, `left/right/wrist.mp4`, `meta.json`) into a gr00t_views dataset.
Like the VACE builder it needs no simulator -- the videos and states already exist.

Two things it handles:

- **Instruction rewrite.** actaug keeps the *source* episode's language (65 distinct strings,
  e.g. "Pick the wine ..." on a SoapDispenser010 episode). The augmented object is the one in
  the folder name, so the instruction is regenerated from that, giving the same 7 task strings
  the other datasets use.
- **Action order.** actaug already writes base-first, matching `modality.json`, so no reorder
  is needed -- unlike the MimicGen path. Its actions do exercise the mobile base in some
  episodes, which the other datasets never do, so `tools/check_action_layout.py` cannot
  classify it (that check assumes a static base). Verify by hand there instead.

## Related artifacts

- `mlnha/mimicgen-batch64-30k-ckpts` β€” checkpoints from the fixed-action-order run
- `mlnha/vace-batch64-30k-ckpts` β€” VACE run checkpoints
- `mlnha/mimicgen-pi05-aug256` β€” MimicGen augmentation for the pi0.5 hard-object set